Showing posts with label mistakes. Show all posts
Showing posts with label mistakes. Show all posts

Tuesday, July 7, 2026

Can the AI learn from its mistakes?



That is a good question. The answer is simple. That depends on how the error or mistake is determined. Determining what the mistake or error is is one of the bottlenecks to creating artificial general intelligence (AGI): how to determine right or wrong? How to describe favorable cases. That the AI should use. And. How to determine the non-favorable cases? The latter are cases that the AI should avoid. In the teaching process. The operator determines and describes the case. And then gives it. Positive (favorable) or negative (non-favorable) values. 

In this text. The main topic is reinforcement learning. There, the system makes something, and the environment. It gives feedback. The feedback. Or. The actor who gives feedback. Determines. If the AI acts right or wrong. And how to determine the values “yes” (Positive)(+)  and “no” (Negative)(-)? 

The system could partially follow Boolean algebra. The chain of positive (+) solutions. It can be conjucted by using “AND”. The “OR” changes the model. And if the model gives a negative (-) value. The system turns to using “NOT,” and then the system. It must change the model. The disjunction happens when there are too many negative values in the series of cases. The negation operation makes the system retake the algorithm. And then try another way, or algorithm, to solve the problem. The problem. It must always be solved by following the rules. 

When we think about the trial-and-error model. That model is effective. But not in all cases. This model is also known as the reinforcement model. Trial-and-error model. It is a good tool for virtual cases. But in cases where the AI must drive a car. That kind of learning solution. That can turn very expensive. There are not many ways. How to react to things the right way. Wrong reaction. It can turn fatal. If the AI driver reacts the wrong way. That can be a very big risk.  

When the car stops at a red light. That is the rule. This instruction is for public safety. But what if somebody tries to rob the car? What if a street gang member  shows a red light or “stop sign” to the car? Trying to rob it? That case is not very common. But those special cases show. How difficult. It is to program the AI. The AI is like a student. That system requires intensive training. The AI trainer must give instructions on what to do. And what not to do. 

In simple cases, the AI uses a limited data type. The AI is very easy to teach. The system requires a description of the favorable case. That case is determined as plus. But then the AI requires determination. About the non-favorable cases. The thing that the AI should not do. That is as important as what the AI should do. The AI should also have value. 

What to do if it doesn’t recognize the case? In a virtual world. The AI. It can make as many mistakes as the user allows. But in real life. When AI controls physical things. There is no room for errors. If the AI controls robot forklifts. Those systems can break lots of merchandise. If they work wrong. If the AI controls vehicles. like cars. And it reacts the wrong way. Results can be devastating. In real traffic, the vehicle has no time to wait and analyze opportunities. 

If we want to use virtual environments. The AI can wait. More information for the entire day. The virtual system. It can have endless time to try again. Or wait for more information. 

The world in the virtual environment. There, the system handles things like numbers. There are only two possible cases. Right (+) or wrong (-). But in cases like traffic, there are also plus-minus (±) cases. When AI controls a car. It can face a situation. That there is an emergency vehicle behind it. The AI can be ordered to drive to the sidewalk. The AI must also have orders that it must not impact people. And those cases. That don’t happen very often. They are the most challenging things for the AI. The AI must be prepared. That somebody tries to rob the car. Or there is an emergency vehicle behind it. In a tight avenue. 



Boolean algebra. 


The AI can learn in three main ways. 


1) Reinforcement learning


“In machine learning and optimal control, reinforcement learning (RL) is concerned with how an intelligent agent should take actions in a dynamic environment in order to maximize a reward signal. Reinforcement learning is one of the three basic machine learning paradigms, alongside supervised learning and unsupervised learning.” (Wikipedia, Reinforcement learning)


2) Supervised learning


“In machine learning, supervised learning (SL) is a paradigm in which an algorithm learns to map input data to a specific output based on example input-output pairs. This process involves training a statistical model on labeled data. Each input is paired with the correct output. The term "supervised" refers to the role of a teacher, or supervisor. Who provides. This training data guides the algorithm. Towards correct predictions. For instance, if you want a model to identify cats in images, supervised learning would involve feeding it many images of cats (inputs) that are explicitly labeled "cat" (outputs).” (Wikipedia, Supervised learning)


3) Unsupervised learning

“Unsupervised learning is a framework in machine learning where, in contrast to supervised learning, algorithms learn patterns exclusively from unlabeled data. Other frameworks. In the spectrum of supervision. Including weak- or semi-supervision, where a small portion of the data is tagged, and self-supervision. Some researchers consider self-supervised learning a form of unsupervised learning” (Wikipedia, Unsupervised learning)


“In supervised learning, the training data is labeled with the expected answers, while in unsupervised learning, the model identifies patterns or structures in unlabeled data.” (Wikipedia, Supervised learning)



“The typical framing of a reinforcement learning (RL) scenario: an agent takes actions in an environment, which is interpreted into a reward and a state representation, which are fed back to the agent.” (Wikipedia, Reinforcement learning) The thing that gives feedback. Like determining whether the case is favorable. Or non-favorable. It can be the human. 

The main problem with the AI and learning system is. How to determine whether the solution is good or bad. The simplest way is to use a human as a controller. When the algorithm ends its operation. Human operators. They select whether the solution is right or wrong. Determination of the desired solutions. It can also be programmed into the algorithm. In the case of stock marketing, desired. Or. A favorable solution could be maximized income. In the series of actions, the algorithm repeats the action. Time after time. The solution that it pursues. That is, maximizing income. 

Stock market analysis is a simple solution for modeling. The rising line, or rising income. It is the positive solution. The decreasing line is the negative thing. 

This type of machine learning is not hard to make. The user must only determine the highest number. That is, in a certain column. That is what the AI should pursue. In a series of cases, the user marks the wanted solutions, or actions. As positive (+) and negative (-). The thing. The algorithm must pursue. It is the highest possible number of positive solutions. 

The user determines the plus and the minus. And the AI tries to take as many points in the plus column. As possible. The AI, or its teacher, just selects the answer. That is marked as plus. The process requires more than one point. And then the AI follows the line. When the line is rising. The AI makes the right (+) solution. When the line decreases, the solution is wrong (-). 



https://vertexestechnology.com/levels-of-ai/


https://en.wikipedia.org/wiki/Boolean_algebra


https://en.wikipedia.org/wiki/Supervised_learning


https://en.wikipedia.org/wiki/Reinforcement_learning


https://en.wikipedia.org/wiki/Unsupervised_learning


Thursday, July 2, 2026

Why does the AI change answers all the time?



“ChatGPT may sound confident, but when tested on complex scientific claims, it often guesses and even contradicts itself. Researchers found it struggles especially with spotting false information. Credit: Shutterstock”(ScitechDaily, ChatGPT Was Asked the Same Question 10 Times. The Answers Kept Changing)

“In the initial 2024 experiment, ChatGPT answered correctly 76.5% of the time. When the study was repeated in 2025, accuracy rose slightly to 80%. However, once the results were adjusted for random guessing, the performance looked far less reliable. The AI was only about 60% better than chance, which the researchers described as closer to a low D than strong performance.”(ScitechDaily, ChatGPT Was Asked the Same Question 10 Times. The Answers Kept Changing)

“The system had particular difficulty identifying false statements, correctly labeling them only 16.4% of the time. It also showed inconsistency. When given the exact same prompt 10 times, ChatGPT produced consistent results for only about 73% of the cases.”(ScitechDaily, ChatGPT Was Asked the Same Question 10 Times. The Answers Kept Changing)

Researchers talk about those cases like this. “We used 10 prompts with the same exact question. Everything was identical. It would answer true. Next, it says it’s false. It’s true, it’s false, false, true. There were several cases where there were five true, five false.” (ScitechDaily, ChatGPT Was Asked the Same Question 10 Times. The Answers Kept Changing)

The reason for that is the accumulation of information. When the researchers ask exactly the same questions hundreds or thousands of times. That causes data accumulation. Another thing is that. There is probably a pointer in the algorithm. That tells whether the answer satisfies the user. If the user does not stop the algorithm. While bombarding it with the same question. 

The algorithm interprets. That situation. That there is something wrong in the answer. If the user drives the same algorithm.  Again and again.  There is a possibility. That some router. Or a switch is stuck. This can cause a situation. That the system gives different answers. When the AI gives an answer that remains. It's RAM memory. If the algorithm runs in a loop, the question repeats. 

Time after time. That can also fill the memory of the central servers. So, before the next run. The user must remove the garbage or the data from the RAM. The system must have time to stop the algorithm and then clean its memory. If that is not done. The memory. It can be filled. And data. It can be polluted. 


The user should finish the job first. That happens by telling the AI that it’s time to begin the new operation. 


We all know that 1+1=2.But sometimes a stuck switch, or gate. Causes. That microchip fails in the simplest possible calculation.  And tells researchers that 1+1=3. The failure in a processor’s internal structure. It can make it break mathematical rules. 

The situation is similar. To cases. There are binary microchips that calculate 1+1=3. This happens sometimes. When the binary processors are tested by calculating 1+1 thousands or billions of times. Suddenly. A microchip. It can give an answer that can surprise us. The reason why the computer gives 1+1=3 is in Boolean algebra and switches. When the system runs 1+1 many times. That simple calculation jams the switch. And that causes the wrong answer. 

The biggest problem with AI is that. The same AI should handle all types of things. There is no human on Earth. Who knows everything. The limit. Of the sector. Of AI use. It can make it more trustworthy. If the same AI is used for fun and for things like scientific work. That causes the data that it involves. Scientific data and other types of data. Like poems. The biggest problem is that. The AI will not make. A difference between the sources. The AI handles all homepages the same way. The thing that makes AI more trustworthy is that. Its data source. It can use. They are limited to trusted scientific homepages. In cases like entertainment purposes. The AI must not use data in a similar way. 

Did you ask something? From the AI? Then you might notice that the answer changes all the time. The reason for that is in the AI, or LLM (Large Language Model). And especially in its structure. The AI is a cloud-based solution. The giant entirety. People who use AI might think that they are using their private software. The LLM or AI is a giant network of databases and servers. When somebody uses those systems. The system scales that interaction all over the network. So, every user in that kind of AI is participating.

 In the AI-training mission. Even if other users cannot see directly what other people ask. That data grows and refines the dataset. That. The system involves. When we ask something.  From the AI. It searches data. From its own registers. The AI searches. Is there the same question? And then the list of data sources that the LLM used. Questions. That people make. They might be from the same topics. But. They are a little. A bit differently written. 


What is the heaviest stable element? It is not the same question. As the question: “What is the heaviest non-radioactive element”?


The AI uses statistical methods to generate answers. This means that the homepages that the AI uses must be involved. A certain number of references that match the question. The AI must know the trusted data sources it can use. If the AI is used for scientific work. Things like thesis banks, universities, and national institutions offer very qualified information.  But then. We must be careful. When we use the AI. When we ask things like the heaviest known stable element. The AI might give an answer. That is wrong. It might tell.

That Oganesson is the heaviest stable element. Oganesson, element 118, is a very unstable element that decays in less than a microsecond. This means AI makes a mistake. And the reason for the mistake is in the way of thinking. The AI can “think” that the person who asks that question. Means the heaviest element that exists is confirmed. And the element is connected to the periodic table of elements. Normally, that question means the heaviest non-radioactive element. But as I just wrote. AI understands. This is the heaviest element. That has its place in the periodic table of elements. So, we should ask: what is the name of the heaviest non-radioactive element? 

This means that the AI translates and treats them as different questions. The system searches data from its internal registers. But if the questions about the topics are a little bit different. The AI also searches data from the network. This means the AI accumulates information about the source lists in its memory. And when data is accumulated. The AI will use different data sources. The other thing is that. When the LLM makes a search and opens homepages. It changes the homepage's page rank. And that also causes a situation. 

There, the AI’s answers are changing. The other thing is that the AI’s answers will be turned into the homepages. And that makes the AI recycle its output. This is the big problem for the net. The dead internet means that AI generates more and more material for it. That generates and accumulates. A data mass. The same data repeats again and again. This way of generating answers fills the servers. Another problem is that. For generating good answers. Those AIs require well-articulated questions. But another fact is that. A good-looking answer is probably not the best or right answer. 

The problem is that. The wrong answers might look funny. People share them on social media. This causes a situation. Those incorrect answers affect the page ranking. They are visible in discussion forums. If those incorrect answers are often shared. They are seen in the statistics. And that can make the AI repeat them. The problem is that. 

The AI searches data using statistical methods. Page ranking and the involvement of certain words in the list of words that are used in those pages. Make the AI select them. As the data source. Statistical methods. Don’t make a difference between right and wrong information. Or wrong information. The ability to limit. The use of data sources from trusted organizations. Makes AI more trusted.

The big problem is that. The AI causes. Those markings that people normally make become unclear. This breaks. The AI use. In medical work. If people do not make queries or make their markings as they should. That can break the AI. The AI’s purpose. It is to save the medical staff time. In the same way. It's made to make work easier. But the problem is that. Employers see the AI as a chance to fire workers. And if we think that way. The AI benefits only the owners of the companies. 


https://scitechdaily.com/chatgpt-was-asked-the-same-question-10-times-the-answers-kept-changing/

Fruit flies and artificial intelligence.

Researchers downloaded fruit fly brains to a computer. And made models of them. That means. They had a mathematical model. Or a data matrix....