Method for using human intelligence to compensate for inherent defects of machine intelligence
Through the complementarity between human intelligence and machine intelligence, the problems of uncontrollable operation quality, unreliable results and unstable performance of machine intelligence are solved, and the controllability of operations and clarity of responsibilities are achieved, and the stability of user experience and social expectations is improved.
Patent Information
- Application Number
- PCT/CN2024/099461
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2024-06-16
- Publication Date
- 2025-08-28
AI Technical Summary
In the application of machine intelligence, there are problems such as uncontrollable operation quality, unreliable operation results, and unstable operation performance. It is difficult to deal with infinite elements, noise interference, exploratory scenarios and self-referentiality, and it is impossible to clarify responsibilities independently, showing inconsistency and complexity, resulting in dissatisfaction with user experience and unstable social expectations.
Through the complementarity between human intelligence and machine intelligence, humans can use their unique capabilities to deal with infinite elements, identify noise interference, explore scenarios, and introduce self-examination and responsibility clarification mechanisms in machine intelligence to ensure controllable operational quality, credible results, and stable performance.
It realizes the stability and reliability of machine intelligent applications, improves user experience, ensures clarity of responsibilities and efficient operation, and reduces the occurrence of accidents.
Abstract
Description
How human intelligence can make up for the inherent shortcomings of machine intelligence Technical Field
[0001] The present invention relates to the fields of artificial intelligence, machine intelligence, and intelligent robots. Background Art
[0002] In recent years, the application of various intelligent robots, centered around machine intelligence, has increased significantly, including humanoid robots, non-humanoid robots, smart cars, smart aircraft, smart drones, smart ships, smart submarines, smart space vehicles, and smart satellites. Broadly speaking, smart cars can also be called road robots, smart aircraft or drones can also be called aviation robots, smart ships can also be called surface robots, smart submarines can also be called submersible robots, and smart space vehicles or satellites can also be called space robots. These various intelligent robots, with machine intelligence as their core module, bring convenience and labor substitution. They enable high-speed computing and communication, tirelessly and efficiently perform well-defined engineering tasks, and can access existing data and rapidly learn and analyze them with superhuman efficiency. Their extraordinary capabilities facilitate human labor, enhance human capabilities, and free humans from repetitive, boring, dangerous, and low-value labor to engage in more meaningful endeavors. Therefore, the freedom and well-being brought to human society by their widespread application is undoubtedly desirable to consumers and the economy.
[0003] Robots are mirror images of ourselves, imagined, designed, and manufactured by humans. By deeply caring about robots, we can deeply care about humans. Perhaps, only by caring about robots can we truly care about humans. Various intelligent robots are machine intelligence entities, but computer-based machine intelligence technology has inherent flaws. Overall, while machine intelligence is needed by humans, it is not satisfactory. Its operating quality is uncontrollable, its results are unreliable, and its performance is unstable. Sometimes human users rely too much on machine intelligence and suffer losses due to its errors. Sometimes machine intelligence makes mistakes and behaves abnormally without being aware of it. Sometimes machine intelligence goes further and further off the track without reasonable and timely human intervention to correct the errors. Sometimes, after human intervention, it is difficult for the parties involved to prove the rationality and legitimacy of their intervention. In some applications, human users perceive anomalies but are unable to effectively understand, verify, communicate, and collaborate to achieve efficient error correction. In some applications, human users try to intervene in the operation of machine intelligence but encounter unexpected failures and are unable to clarify the cause afterwards. Sometimes machine intelligence behaves abnormally but it is difficult to determine the cause and responsibility, so the same model of products behaves abnormally repeatedly. After some accidents, it is unclear which specific machine intelligence link in the operation process had a problem. After some accidents, although human users believe that the machine intelligence design is unreasonable, they can only blame themselves for their bad luck. Some machine intelligence manufacturers unreasonably increase the burden of responsibility on human users (all liability for personal and property losses is borne by human users). After some accidents, users are unable to prove the responsibility of machine intelligence and its manufacturers, forcing users and their insurance companies to take responsibility for machine intelligence errors and compensate for losses (users report to vehicle insurance companies that their accidents are caused by intelligent driving errors, but neither users nor insurance companies have evidence to reasonably claim compensation from intelligent driving manufacturers). In some incidents, economic claims cannot provide a basis for calculating the amount. The lessons learned from some typical machine intelligence error incidents or successful manual intervention incidents cannot be efficiently accumulated and disseminated as cases to help more human users establish better awareness and ability of predictive intervention and reference precedents. The entire society is far from establishing the technical concept of complementary cooperation between human intelligence and machine intelligence. Overall, the current technological, economic, and legal status and relationship of machine intelligence applications in society are not clear and complete, and cannot be carried out from beginning to end and give people peace of mind. Human users have many of the above dissatisfactions. The overall application of machine intelligence in human society is in an imperfect and unstable state, with "abnormalities may occur at any time", "it is difficult to explain after the incident", "it is frightening", "users often have to blame themselves for their bad luck", "expectations in all aspects of society are unstable", and "people do not know how to better cooperate with machine intelligence".
[0004] The reason for the instability is that "the cognition, reasoning, selection, planning, execution, performance, movement, will, storage, and rationality of machine intelligence" are often inconsistent with "the purpose, intention, conception, common sense, principle, expectation, orientation, will, memory, and sensibility of human intelligence". Machine intelligence shows its inherent defects in these inconsistencies, and human intelligence must and can only complement machine intelligence with its own inherent unique capabilities. Such human capabilities include coping with "infinity", coping with "chaos", coping with "responsibility", and coping with "human nature". Coping with "infinity": human intelligence can make flexible, large-scale, cross-domain, and leapfrog meaningful connections between infinite elements from a high and broad perspective (between the whole and the part), and then "out of nothing" discover new things, raise new questions, invent new concepts, establish new relationships, and gain new experiences, such as conjecture, hypothesis, imagination, creation, drawing inferences from one example, analogy, focusing on key points, detecting anomalies, Identify certain elements (identify noise), ignore certain elements (filter noise), and the frequency of cognitive interaction is infinite, that is, the active and continuous external interaction cognitive ability and internal self-examination ability; deal with "chaos": human intelligence can identify dilemmas and conflicts without being stuck in dilemmas and conflicts, and can reconcile dilemmas and conflicts and make (difficult) choices, such as compromise, tolerance, balance, choosing negative options, breaking rules, inventing rules, and defining and establishing oneself with its choices. It can understand that correct rules at a certain scale will be invalid at different scales and different rules must be applied at different scales. It can understand keeping pace with the times and keeping pace with the times, and can deal with ambiguity problems with a certain reasonable error approximation; deal with "responsibility": humans are social animals, living in a human society where economic and legal responsibilities are complex and intertwined but must be clear. For personal injury or property damage caused by machine intelligence, machine intelligence can neither sort out the responsibility nor bear the responsibility by itself, and ultimately it must be handled by Humans or legal persons must bear the responsibility, which requires clarifying the complex and intertwined chains of technical contracts and legal entities. This work of clarifying the unexpected responsibilities of machine intelligence in a complex society must and can only be completed by human intelligence in a complementary manner; responding to "human nature": human intelligence has common sense and real life experience, and its values are derived from the social relations of life and growth. It can spontaneously understand human life concepts such as "texture", "material", "center of gravity", and "friction" from a deep level, can deal with vague problems in life with common sense, can naturally appreciate quality, can deeply understand human motivations related to psychological emotions, can deeply and truly understand metaphors, metaphors, analogies, teasing, irony, rhetorical questions, humor, hints and other profound information transmission, can conduct profound questioning, painful exploration and lonely introspection at the historical and philosophical levels, and can naturally understand human love, beauty, hope, dreams, art, pain, imagination, curiosity, exploration, sacrifice, determination, meaning and other abstract concepts.There are no clear boundaries between these capabilities; rather, they are intertwined, forming a comprehensive complementarity between human intelligence and machine intelligence. This is the essence of the complementarity between human and machine intelligence. This essence is not simply a superficial "machines will always make mistakes, so humans are needed to compensate" (this argument is actually just as empty and well-known as "humans will always make mistakes, so machines are needed to compensate"). Instead, it emphasizes that the inherent flaws of machine intelligence are the inherent and unique strengths of human intelligence. The gaps where machine intelligence is completely incapable must and can only be filled by human intelligence. This complementary nature between human and machine intelligence is a natural law that is independent of human will and cannot be changed, created, or eliminated. Humanity must and can only adhere to this natural law to achieve perfection in the process and results of machine intelligence applications and to ensure that the social application of machine intelligence reaches a state of overall expected stability and perfection. Currently, the application of machine intelligence is booming, but human society is overly optimistic about the performance of machine intelligence and under-anticipates the problems caused by its inherent flaws. There is a lack of understanding of how to complement machine intelligence with human intelligence. The technical concept of complementarity between human and machine intelligence has not yet been established in human society. The following is a detailed introduction to the complementary nature of human intelligence and machine intelligence.
[0005] Machine intelligence interacts with humans and the world by mechanically and electronically operating electromechanical devices. In effect, it becomes an unnecessary intermediary in the interaction between humans and the world. Machines, no matter how sophisticated, are closed systems, while humans and the world are open systems rather than closed systems. The nature of humans and the world is more like a black box than a white box. Intermittent passive observation and execution are insufficient; they require continuous, active, experiential interaction for full and comprehensive cognition. Human consciousness and the physical world are both composed of infinite elements. The possibilities of human thought and behavior and the real physical world are infinitely emergent, yet machine intelligence can only achieve a one-to-one mapping with a finite set of elements. This finite set of elements is necessarily designed or input by the creators of computer machine intelligence in a limited and explicit manner. Human intelligence can inherently cope with infinite sets of elements, but computers, based on the principle of "one-to-one mapping," are inherently unable to fully cope with them. The reason why human intelligence can cope with an infinite set of elements is that it has the open interactive ability to continuously connect the finite elements in the infinite elements and give them meaning. This is something that machines cannot do. For infinite concepts such as ∞ and π that humans can accurately and completely understand and manipulate, computers can only approximate to a finite number of bits for representation, calculation, storage, and other manipulations. All machine systems can handle are "finite", finite scenarios, finite parameters, finite logic, and finite steps. Even if this "finite" is a very large number, it is still "finite", and then it can be defined, represented, quantified, calculated, analyzed, evaluated, and standardized. Correspondingly, only human intelligence can handle and cope with "infinity", which is the inherent and unique ability of human intelligence.
[0006] Broadly speaking, cognition can be divided into four quadrants: "known knowns," "unknown knowns," "known unknowns," and "unknown unknowns." Machine intelligence can tirelessly and efficiently perform repetitive work within the "known knowns" (low-level repetitive engineering). It can more efficiently transform "unknown knowns" into "known knowns" (using known methods to obtain knowledge or results that are certainly attainable but currently unknown, simply by replacing human labor and achieving them more efficiently). These two quadrants are both "finite." However, machine intelligence is powerless against the "known unknowns" and "unknown unknowns" quadrants, which are "infinite." Complementarily, only humans, with their unique ability to cope with the "infinite," can transform "unknown unknowns" into "known unknowns" (transforming perception of phenomena into conceptually grounded understanding, and experience into consciousness). The "unknown unknowns" are in a fuzzy and chaotic state. Human intelligence can make meaningful connections between these infinite elements from a broad perspective (between the whole and the parts), with flexible, large-scale, cross-domain, and leapfrogging capabilities. This allows it to "create problems out of nothing" or detect anomalies, discovering new things, raising new questions, inventing new concepts, establishing new relationships, and gaining new experiences, thereby transforming the "unknown unknowns" into "known unknowns." Next, human intelligence attempts to analyze, reason, test, and verify these "known unknowns," ultimately transforming the factual and logical parts of these unknowns into new "known knowns." Therefore, from an epistemological perspective, the inherent flaws of machine intelligence are the unique capabilities of human intelligence, and human and machine intelligence complement each other.
[0007] Machine intelligence operations are unstable and susceptible to noise interference, which can be complex and difficult to prevent. The world is constantly changing. Noise interference that humans can easily identify, distinguish, or eliminate can significantly mislead machine intelligence, leading to errors that appear foolish to humans. Furthermore, because noise interference is difficult to reproduce and constantly varies, non-expert users often struggle to specifically report, identify, and understand these technical issues. The triggering conditions for noise can either be specified only after the fact (for example, the presence of heterochromatic light spot noise that causes machine intelligence vision errors) or not at all (for example, in the case of "ghost braking" in intelligent driving, it's difficult to pinpoint the specific location of the noise in the light and shadow ahead). Low-probability random noise inevitably changes constantly, but it's impossible to anticipate or define it beforehand, making it impossible to conduct targeted system design and testing and verification when creating machine intelligence. The industry's commonly used, forward-looking product quality standardization methodology of traversing defined test scenarios and conducting thorough testing and verification is no longer applicable in the new era of computer artificial intelligence, making product quality uncontrollable. In a broad sense, "noise" refers to interference, disruptive information, and junk information. The term itself stems from human intelligence's unique ability to focus on primary cognitive objectives while filtering out distractions. Human intelligence effortlessly identifies certain elements (recognizing noise) and ignores others (filtering noise), deciding whether to focus on specific elements based on its own conscious needs. For example, in human visual cognition, distinguishing and filtering out noise interference can easily pinpoint which pixels in an image are irrelevant to the image's meaning. However, without deliberate attention, humans may not even realize they have excluded certain noise points during observation and cognition, and may not recall them later. Therefore, we cannot give a simple technical definition of "noise," nor can we write robust logic programs to enable computers to "find" and thus "eliminate" it. (Defining "noise" as "clusters of pixels with a significant color difference from the surrounding texture" is clearly incorrect.) Human intelligence's natural ability to distinguish and identify noise interference can help machine intelligence eliminate and filter it, complementing its vulnerability to noise interference.
[0008] Machine intelligence is not robust, and its emergent simulation of human intelligence and consciousness is sometimes quite fragile. The same machine intelligence can randomly emerge with different "robot personalities" during operation, lacking identity, robustness, and coherence. This can manifest as "incoherent statements," "prefaces without follow-ups," "losing watermelons while picking up sesame seeds," "split personality," "neural disorder," and other low-level errors. Normal human consciousness is inherently unified and coherent, with an infinite cognitive frequency and no time intervals. It does not miss anomalies even within even the smallest time intervals. This allows humans to complement these shortcomings of machine intelligence.
[0009] Machine intelligence is unable to handle exploratory scenarios that lack human intelligence's past experience. Machine intelligence is trained on data derived from past human experience. Therefore, it primarily imitates existing human capabilities and cannot perform purely creative work. Even if it creates a painting, composes a poem, or plays music never before created, it is essentially imitating human intelligence's past experiences and routines. This makes it incapable of handling novel situations that have never been addressed before. In the absence of previous experience and training data, it is bound to perform poorly. Exploring the unknown requires coping with "chaos," but machine intelligence is incapable of representing or solving uncertain and ambiguous issues. It lacks proactive and adaptive exploration capabilities and cannot identify problems or detect anomalies. While machine intelligence can search for connections between elements in certain areas, it cannot distinguish between useful or meaningful connections and useless or meaningless ones. This is because exploratory problems themselves are ambiguous—the computational goal is fuzzy. The ability to naturally cope with fuzzy chaos and discern which connections are useful or meaningful is a unique and inherent capability that complements human and machine intelligence.
[0010] "Self-reference" is unattainable within the realm of pure machine principles. Within the current realm of computer science, software symbol strings cannot simultaneously function as both instructions and data. Mathematical systems, through a finite number of steps, cannot produce meta-propositions that reflect on the system itself. Therefore, machine intelligence lacks a "self" and cannot spontaneously and comprehensively "examine itself." Machine intelligence can often faithfully execute instructions, but when the instructions themselves are imprecise, it may exhibit errors or anomalies. After these errors or anomalies, it often fails to recognize its own errors and anomalies and continues down its erroneous path, "ignorant and fearless." Its computational logic, operating "normally" within the digital world perceived and constructed by the machine, will continue to operate "normally" without human intervention. It will not actively self-doubt and self-correct, nor will it actively seek human intelligence to revise, update, or improve its instructions. Some machine intelligence may even persist in its errors, competing with human intelligence for control. Taking the example of humans using natural speech to send instructions to machine intelligence, machine intelligence often does not know that it has failed to fully and accurately understand the overall contextual meaning contained in the literal meaning before execution. During execution, it often does not know that its interpretation of instructions has become ambiguous and deviated further and further. If human intelligence deliberately develops a mechanism program for machine intelligence to "self-check to correct and expand the details and accuracy of instructions" in such scenarios, machine intelligence will endlessly question the underlying common sense of human life concepts. Such questions will never end and may prevent machine intelligence from ever starting execution. Humans are able to cope with "infinity," "chaos," and "human nature." They can carry out work even when the conceptual assumptions and premises of instructions are somewhat ambiguous, exploring and working, proving and verifying while testing, and working with increased vigilance and caution. Exploration and probing are essentially adaptive ways of blurring the boundaries between decisions and actions in chaos, autonomously adjusting the degree of approximation to tolerate errors. It's not about "being certain to do it" but about being able to "do it even when not quite sure." Humans are naturally able to appreciate and detect common human emotions such as anomalies, strangeness, hesitation, suspicion, interference, glaringness, fear, vigilance, tension, hesitation, and difficulty in progress. They may have similar rational and emotional impulses such as "Huh? What's going on?", "This is very unusual...", "So strange...", "Something seems wrong...", "It seems a bit dangerous...", "Be sure to be careful...", "This doesn't seem so easy, it's best to try it slowly...", but at the same time they can continue to work, which is difficult for machine intelligence to achieve.For humans, every time they feel "very strange" or "very dangerous", it will trigger human intelligence to conduct an active "self-examination", but humans cannot explicitly write all similar trigger conditions into computer programs, nor can they collect a large number of cases of all similar scenarios emerging in the outside world for machine intelligence to train and learn. This kind of self-examination exists naturally in human nature. For example, when people look down from a cliff or a tall building, they often naturally imagine "what would happen if they fell", and when people drive a car too close to other cars, they often naturally imagine "what would happen if they collided". Of course, human creators will also arrange some specific self-checking programs for machine intelligence, but these self-checking programs are still limited factors that humans tell or teach machine intelligence in advance - limited trigger conditions, limited review objects, and limited review frequencies. Machine intelligence cannot autonomously connect infinite factors for uninterrupted self-review, and cannot traverse infinite possibilities in an instant like humans. There is no time interval in the continuity of human consciousness, and the self-checking programs set for machine intelligence must have clear time intervals, such as "every 1 second" or "every 0.1 second". Such self-checking intervals are still destined to miss phenomena or signs that appear but only last for a shorter time.
[0011] The continuous use of machine intelligence in the physical world, with a methodology of testing through practice and continuous training and improvement, may be wishful thinking. It may actually cause machine intelligence programs to be drowned in a sea of code and data. Overly complex training materials are piled up in machine systems, forming increasingly complex sensor signal patterns, algorithm mappings, program logic, behavioral patterns, and consciousness simulations. This leads to increasingly frequent instances of incomprehension and conflict, a so-called "pressing down one problem only to have another pop up." Machine intelligence's skills and behaviors are bound to occasionally become inflexible, lose sight of one thing while focusing on another, or become contradictory. Human intelligence naturally has the unique ability to grasp key points, identify contradictions, draw inferences from one instance, and apply them to other cases. Human intelligence can complement machine intelligence with this inherent and unique ability.
[0012] The results of machine intelligence operations are unreliable because they are unable to independently clarify and assume responsibility. Humans are social animals, living in a society where economic and legal responsibilities are complex and yet must be clearly defined. For tort liability, personal injury, or property damage caused by machine intelligence, machine intelligence cannot independently clarify or assume responsibility. Ultimately, humans or legal entities must bear the responsibility. This requires clarifying the complex and intertwined chain of technical contracts and legal entities. On the supply side of machine intelligence, those potentially liable include the legal representative of the machine intelligence manufacturer, the manufacturer's general manager, product designers, product architects, those responsible for product safety, module developers, and component manufacturers. On the demand side of machine intelligence, those potentially seeking legal protection and compensation include users, property owners, insurance companies, and relevant government departments. Taking smart cars as an example, the entities that may bear legal responsibility on the supply side include the legal person of the vehicle manufacturer, general manager, chief engineer, designer, safety manager, safety designer, programmer, program inspector, person in charge of contracts (advertising, product manuals, sales contracts, service contracts, etc.), suppliers of smart driving solutions and their related components (AI, sensors, computing units, communication units, positioning units, maps, navigation, redundancy, electric drive, electronic control, etc.). The entities that may seek rights protection and compensation on the demand side include users (drivers, passengers) seeking rights protection and compensation after an accident, vehicle manufacturers seeking rights protection and compensation from suppliers of smart driving solutions and their related components, insurance companies that have objections or claims regarding liability sharing in single or repeated smart driving-related claims, users or owners who are damaged in accidents involving unmanned vehicles (unmanned taxis, logistics vehicles, sweepers, dock vehicles, mining vehicles, etc.), and local governments and management departments that have public safety concerns about the large-scale operation and use of smart cars. Taking smart driving liability accidents as an example, each responsible party must clearly assume responsibility according to their respective faults. This work of clarifying the responsibility for machine intelligence accidents in a complex society must and can only be completed by humans, because chain accountability and claims often appear in the technical contract chain and the legal entity chain. For example, the insurance company pays compensation for the accident first and then pursues accountability and claims against the smart car manufacturer. After the smart car manufacturer pays compensation, it pursues accountability and claims against its smart driving solution provider. After the smart driving solution provider pays compensation, it pursues accountability and claims against its sensor supplier. The determination of responsibility in such specific links involves a complete sorting out of technical logic and facts and obtaining legally effective evidence. This work must and can only be complemented by humans.
[0013] Machine intelligence is currently built on computer science and machine principles based on the von Neumann architecture, using computer language, symbols and machines as tools. Language, symbols and machines are limited, and the boundaries of language, symbols and machines are the boundaries of human thought practice within the scope of computer science. However, humans and the world itself are infinite. There are still many things that cannot be expressed in words but can only be understood through personal experience, or even destined to be unable to be expressed symbolically and can only remain silent. Many ultimate questions that are meaningful to humans have no answers. If humans cannot write answers, there is no data for machine intelligence to train and learn. The ultimate questions involve infinite elements, and computers cannot perform calculations. Computer formal systems that rely on language, symbols, and machines to simulate human conscious behavior can simulate human conscious behavior to a certain extent, no matter how much the formal system is manipulated, and no matter how complicated its logical programs and institutional systems are, they can only approximate but never reach the level of human conscious behavior. Consumers and the general public's demand and expectation for the "depth" of machine intelligence is that "high-tech robots should be as good as or better than humans." This is impossible based on computer formal systems. Machine intelligence's understanding of abstract things and their variations is a cramped, narrow, fragile, and shallow imitation, and is more based on repetition of existing human expressions. Its mining and manipulation of text or information can only be adjusted to a certain extent based on contextual background or conceptual needs, but it is difficult to conduct in-depth historical and philosophical questioning, painful quests, and lonely introspection like humans. Humans can only reason with machine intelligence to a certain extent. Beyond a certain level, machine intelligence cannot keep up with human thinking, and any in-depth interaction will be full of loopholes. Machine intelligence cannot truly possess philosophical thought. Combined with the "one-to-one mapping" and "black-or-white" principles of machine intelligence, current machine intelligence is destined to mechanically "choose one side" rather than "both sides." It is unable to define and calculate dilemmas, nor can it reconcile or decide between them. It cannot "make difficult decisions out of necessity" like humans, nor can it seek and achieve a dynamic balance between two conflicting variables like humans do. (Even if this balance point is forcibly and programmatically achieved, the method used is actually the program creator's method and is likely inconsistent with the user's own desired position. Because each person's method is unique, users may feel that their free will has been challenged and be dissatisfied with the performance of machine intelligence.) It cannot autonomously compromise, cannot autonomously choose negative options, and cannot accept moderate danger or risk, which is often unavoidable. When the actual situation requires a moderate violation of the rules, machine intelligence, due to its inherent nature, finds it difficult to break the rules. In these aspects, the inherent flaws of machine intelligence are the inherent and unique capabilities of human intelligence.
[0014] The creators of machine intelligence—humans themselves—are far from perfect. On the demand side, under the influence of ever-changing human factors, such as unintentional misleading and even malicious deception, unsatisfactory performance is inevitable. Furthermore, the business world is comprised of humans, whose needs are diverse, ever-changing, and endless. This gives rise to a wide variety of machine intelligence, provided by countless different suppliers, with varying designs, parameters, and calibrations. Combined with the inherent inexplicability of machine intelligence neural networks, this makes it difficult to predict the performance of machine intelligence applications in their interactions with humans and the world, quantify performance metrics, and standardize both processes and outcomes. On the supply side, humans in the digital age tend to over-pursue efficiency, leading to radical rather than thoughtful principles for the development of machine intelligence. Many commercial machine intelligence companies often fail to thoroughly test their products from a fundamental methodological perspective. For applications like general-purpose humanoid robots and general-purpose smart cars, real-world applications involve endless user scenario testing. Consequently, the social performance of machine intelligence applications is inherently unstable. In business, to give customers a good short-term impression, machine intelligence feedback often tends to predict user preferences. This does give customers the illusion that the machine has the ability to provide personalized feedback. However, in reality, its feedback does not necessarily point to perfection, wisdom, truth, or reality. Instead, it may be the so-called "mechanical completion of tasks" or "catering to customers." In industry, to absolve themselves of legal liability and reduce design difficulty, machine intelligence manufacturers will set machine intelligence to "fully comply with the rules." Therefore, from the outset, the meta-design concept lacks exploration capabilities. On the other hand, machine intelligence training data is often manually selected and labeled. Its resources are limited and may be localized and biased, destined to bear the mark of human imperfection. Because machine intelligence cannot generate its own values through life experiences and social relationships, it is often instilled with values by human creators, and naturally carries the one-sidedness and complacency of human creators. Its training data is often "birds that sing catch worms." ", that is, in a few unimportant fields, there will be too much training material due to the frequent and restless voices of humans, while in the majority of important fields, there is a lack of training material due to the quiet voices of humans. In essence, data only reflects part of humans and the world. The cognition, measurement and understanding of humans and the world from a mathematical perspective is itself biased and destined to bear the imperfections of computational science itself. This will cause machine intelligence to regard local judgments as holistic judgments - "seeing only the trees and not the forest", and may produce a vicious cycle and self-reinforcing resonance with local human behavior, becoming increasingly immersed in and displaying local rationality rather than overall rationality.Human nature inherently includes imperfection, uncertainty, and unfairness. The underlying assumptions and premises commonly embedded in the creation or use of machine intelligence may also be imperfect, uncertain, and unfair. For example, in some human-machine systems, human intelligence has the highest priority and authority, while in others, machine intelligence does. Humans accept human intelligence making mistakes but not machine intelligence. They also dislike the sudden, unexpected failure of machine intelligence in critical scenarios. Potential requirements for the effective lifespan or reliability of machine intelligence are extremely high but vague, making standards difficult to establish. In some real-world scenarios, humans demand maximum rationality, while in others they tolerate or even actively accept "emotional decisions trumping rational ones" or "emotional and rational decisions becoming one." However, this cannot be quantified or even characterized in advance. The performance of machine intelligence resulting from these imperfect, uncertain, and unfair underlying assumptions can only be analyzed on a case-by-case basis using human intelligence. Generalizations and set rules or standards for machine intelligence are not possible. This is how human intelligence reconciles contradictions—constantly pursuing perfection through its own imperfections—and also reflects the complementary nature of human and machine intelligence.
[0015] Machine intelligence, based on computer programs, deploys hardware and software systems, striving to achieve universal perception and action goals similar to those of human intelligence. However, due to the infinite nature of human needs and the ever-changing world, machine intelligence products are inherently unfixable. To cope with infinitely complex environments or perform infinitely complex tasks, the corresponding overall hardware and software systems will also become infinitely complex. Humans are forced to adopt a large-scale modular division of labor to address complex development. Those who master high-level programming languages may not necessarily understand the operational efficiency of low-level machine languages. Those responsible for data access modules may not necessarily coordinate with the preferences of data-using modules. Those designing hardware may not carefully and adaptively consider the software architecture. Those writing software may not understand the physical limitations of the hardware architecture. Those designing logic may not anticipate the infinite variations in real-world application scenarios. This inherently unfixable overall system is destined to remain complex to the point where the system design, material structure, program code, and comprehensive testing and verification are fundamentally impossible for a single individual to fully conceive, understand, coordinate, develop, and oversee, from the overall to the detailed. For industrial products, since system complexity stems from external engineering forces rather than inherent life-giving growth, it is destined to remain complex to the point where logic code, sensor priorities, hardware and software modules, and application scenarios hinder, contradict, and even conflict with each other. Ultimately, this complexity manifests itself primarily through system errors when responding to specific situations. Currently, the industrial community generally implements patches, with refactoring only rarely performed. However, patches only address the immediate problem, and the constant patching and patching of patches further increases the overall complexity of the system while addressing the immediate issue. Refactoring is often intended only to enable module components with performance upgrades. Because these systems are often embedded in other systems and interdependent, complete refactoring is often extremely costly, and therefore, the overall system architecture is rarely completely altered. Therefore, inherently, various errors in overly complex systems are inevitable, and potential errors are often not foreseen in advance but only identified and recorded when they occur. Human intelligence, with its inherent ability to identify, recognize, and resolve conflicts, can complement these shortcomings of machine intelligence.
[0016] The core essence of mathematics lies in precision, and it naturally resists ambiguity. Machine intelligence within the realm of computer science, which is centered around mathematics, inevitably exhibits weak generalization, versatility, and adaptability. Mathematics can only solve problems that can be solved mathematically. Mathematical models, as specific tools, originate from practical experiments and can only be applied effectively in specific scenarios to solve specific problems. Interaction with the open-ended human consciousness and the open physical world largely requires strong generalization, versatility, and adaptability. Consumers and the public demand and expect the breadth of machine intelligence to be "as good as or better than humans." Human intelligence aspires to use computers to create so-called "general intelligence" or "machine life" like itself. This is impossible in today's scientific and technological theory and practice. There has been no evidence that human intelligence and consciousness can be fully explained solely through the lens of mathematics and technology. Humans have a perceptual ability that is so profound and subtle that they cannot even explain it: people can collect subtle information from extremely subtle changes in facial expressions or extremely subtle movements of other people's bodies, and people can obtain subtle and profound information from their eyes. Human vision and hearing will automatically focus on their key points and ignore background noise. The human eye and brain will automatically cooperate to fill in the flawed and incomplete visual images. The human ear and brain will automatically cooperate to transcend the formal system of sound frequency and grasp the non-formal process to obtain meaning interpretation from a higher level (appreciating music). An experienced fur dealer can judge the quality of fur by the sound of a knife blade scraping across the fur. An experienced cotton dealer can identify the quality of cotton by rubbing it with his fingers, chewing and tasting it, and listening to the crackling sound when combing it. At the complexity level, experienced hunters can use their noses to smell impending changes in the weather with extreme sensitivity, and experienced chicken farm workers can separate large batches of newly hatched chicks by sex with extremely high accuracy without knowing the physiological details...Humans have some understanding of cognitive science, neuroscience, and brain science, but not all of them have complete analytical explanations. Since it is impossible to fully analyze, qualitate, and quantify the infinite factors involved in the profound and subtle perception of humans, and even the concept of what factors exist is not fully known, these unique human perceptual abilities are obviously mostly beyond the reach of current machine intelligence, because we don’t know what sensors to use, what variables to set for computer machine intelligence, how to quantify these variables, or what data to input into machine intelligence for learning and training.Even if robots with physical bodies were used to simulate human perception, such perception would be far more superficial and crude than that possessed by humans themselves. As creations of human intelligence, the embodiment of machine intelligence is inherently far inferior to that of humans, and it is unable to formally understand human subjectivity and intuition. Therefore, it lacks the rich common sense of the world that humans possess, whether consciously or unconsciously. Without truly understanding the real life experiences of human common sense, it is difficult to fully consider the context in which they receive instructions. It is unable to spontaneously and deeply understand human life concepts such as "texture," "material," "center of gravity," and "friction." It is unable to deeply understand human motivations related to psychological emotions. It is unable to deeply and truly understand the profound information conveyed through metaphors, allusions, analogies, teasing, irony, rhetorical questions, humor, and allusions. It is unable to understand that in many cases, human "necessity" is not "passive necessity" but "active necessity." It is unable to reconcile contradictions and independently choose from infinite options to define and construct the self. It is unable to deeply understand abstract human concepts such as love, beauty, hope, dreams, art, pain, imagination, curiosity, exploration, sacrifice, determination, and meaning. For example, a skilled poet can express with a few words what many ordinary people desire but cannot, yet with remarkable skill. This phenomenon demonstrates that profound thoughts or emotions can indeed be expressed and conveyed through written symbols. However, for a poetic machine intelligence to possess thoughts or emotions simply by probabilistically manipulating written symbols would be "confusing cause and effect" and would be impossible. It would simply be used to "pretend" to possess thoughts and emotions (it could produce neatly rhymed but bland poems). Similarly, machine intelligence often probabilistically and predictably produces content that appears correct and reasonable on the surface, but is inherently fabricated and meaningless. Human intelligence, however, naturally relies on its ability to perceive, understand, and appreciate quality, distinguishing between "serious nonsense" and "diligent and professional misconduct," complementing machine intelligence to anticipate, correct, and compensate for its inabilities and errors.
[0017] Due to the aforementioned inherent flaws in machine intelligence, the digital world created by machine intelligence and the physical world created by human intelligence often exhibit inconsistencies. Machine intelligence and human intelligence often narrate events in their respective worlds in a different way. What may appear to humans as errors or anomalies may appear normal to machine intelligence. The quality of machine intelligence's work is uncontrollable, its results unreliable, and its performance unstable. For example, after many intelligent driving accidents, users and machine intelligence often disagree, leaving victims with difficulty determining or proving that the cause of the accident was an error in their car's machine intelligence. Therefore, a well-informed approach requires proactive management of potential unsatisfactory, incompetent, or even dangerous risks associated with machine intelligence. This involves proactively assessing and maintaining vigilance beforehand, confirming and intervening during the process, documenting and verifying afterward, and pursuing accountability and compensation. Human and machine intelligence complement each other to achieve perfection, not laissez-faire.
[0018] To complement machine intelligence, humans need to efficiently organize, express, disseminate, and record vast amounts of data, information, and clues. Using only traditional media like language, text, and symbols, it's difficult to fully and comprehensively describe and express problems, reconstruct events, present clues and evidence, clarify the logic of a problem, verify the logic of suspicions, and record logical facts. Textual symbols are often too long to grasp quickly and efficiently. Presenting facts solely with words is both tedious and prone to omissions, potentially missing details that could draw attention and encourage collaborative contributions. In complex situations, describing anomalies or pointing out problems often results in repetitive, inverted, and disorganized explanations. Real-time verbal communication can't immediately, quickly, and clearly clarify both the facts and the logic. It's difficult to create fresh, living evidence that can be admitted and archived in court, reused repeatedly, stored, and disseminated, and readily understood and quickly agreed upon by others. It's also difficult to efficiently convey complex facts and logic to others and ensure consistent understanding and memory.
[0019] On the other hand, when anomalies occur when machine intelligence performs tasks, it is often necessary for multiple people (such as experts, users, and managers from different fields) to immediately collaborate from their respective professional perspectives to quickly discover and solve the problem. This requires that the preceding and following logic of the anomaly and the relationship between the factual nodes be clearly expressed in the shortest possible time, and the logic and facts be presented to all other audiences at the same time in the fastest and most efficient way, so that everyone can immediately understand and be convinced and reach a collaborative consensus, and it must also be able to be easily transmitted and archived. When reviewing and understanding the archives of abnormal events afterwards, it is often found that the audio and video archive files are complicated and lengthy, key information is submerged in the ocean of information and cannot be distinguished, and the straightforward playback lacks a sense of rhythm, which is contrary to the attention mechanism and rhythm requirements of human cognition. The audience who are confused need to use their eyes to quickly traverse all positions of the video frame for a long time, and even dare not blink for a long time to ensure that they do not miss any fleeting visual information. This makes the audience tired, difficult to understand, and misses things without knowing it. At the same time, because the audio and video archive files record and express many angles, the information is complicated, and different professions are involved, a large amount of key information and their interrelationships in the co-temporality, co-spatiality, world construction of human intelligence and machine intelligence, and event narration are not directly revealed, and the audience often does not get the point.
[0020] In addition, for many machine intelligence training and learning, there is often only positive training material but lack of negative training material (for example, there is only training material about "what is right" but lack of training material about "what is wrong"), and only behavioral training material but lack of logical training material (for example, if a user asks the machine intelligence "Why did you do something just now", the machine intelligence can only recite the perception, recognition, decision-making, planning, execution and other account information in its operation log records in chronological order or reverse order. The refresh frequency of its operation and records is extremely high, so this recitation may be extremely detailed and lengthy, and it cannot express the specific key logical links in its operation chain that lead to a certain behavior as concisely as humans).
[0021] Summary of the Invention
[0022] The technical solution of this patent is to manipulate, organize and indicate the visual images corresponding to the inconsistencies between "machine intelligence's cognition, reasoning, selection, planning, execution, performance, movement, will, storage, and rationality" and "human intelligence's purpose, intention, conception, common sense, principle, expectation, guidance, will, memory, and sensibility" in visual images reflecting the interaction between humans and machine intelligence and visual images reflecting the interaction between machine intelligence and the world. These visual images are combined to indicate the worlds perceived and constructed by machine intelligence and human intelligence respectively in the above inconsistencies, the respective event narratives of machine intelligence and human intelligence in their respective world constructions, and the time and space information in their respective event narratives. The above information is then compared and linked with each other to form a unified and integrated factual chain and logical chain reflecting the complementary nature of human intelligence and machine intelligence under the appearance of the above inconsistencies. The two chains are formed simultaneously, and the output is a visual image chain carrying instructional information that is not streamlined, condensed, refined, summarized, targeted, difficult to categorize, and difficult to store and disseminate. It can clearly and comprehensively present the facts and logic of machine intelligence application events and the appearance and nature of the human-machine inconsistency therein. If the aforementioned video images are audible, carrying and utilizing more effective information, the effect will be even better. In reality, most video images collected are also audible. When manipulating and organizing the video images and providing instructions, the video images reflecting the aforementioned inconsistencies can be selected, decomposed, organized, and arranged. The time or space reflecting the aforementioned inconsistencies can be shaped, reorganized, and connected, mimicking or guiding the human audience's attention, perceptual habits, and cognitive mechanisms, emphasizing key points and highlighting details to generate dynamic video documents. Static images can also be obtained during this process and combined with instructional information to generate static graphic documents. Such dynamic video documents or static graphic documents are streamlined, condensed, refined, summarized, targeted, easily categorized, and easily stored and disseminated outputs of this technical solution. Beyond the factual chain and logical chain, they can also be unified and integrated into a verification chain reflecting the aforementioned inconsistencies to make the verification process and results explicit. They can also be unified and integrated into an emotional chain reflecting the aforementioned inconsistencies to convey emotions to the audience and evoke group resonance. The inherent nature or potential future consequences behind the aforementioned inconsistencies can be clearly and explicitly explained. The machine intelligence's instructions can be modified or adjusted. The machine intelligence's work results can be rejected or dismissed. In the process of using this technical solution, or utilizing its outputs, multiple humans can cumulatively manipulate, organize, and instruct the visual imagery, increasing diverse perspectives and expertise and fostering collaborative collaboration.
[0023] Let's explain the term "visual" in this article. Vision, in a broad, multidisciplinary technical sense, refers to images captured by various types of imaging equipment. This includes images directly seen by the human eye and constructed and memorized by the human brain, images captured by ordinary optical imaging cameras, images captured by infrared thermal imaging cameras, images captured by low-light cameras, images captured by telescopic cameras, images captured by microscopic cameras, and so on. The vast majority of images captured by artificial electronic devices carry information about the time of capture, and some also carry information about the location (space). The captured images themselves typically indicate information about the location (space) where they were captured.
[0024] The "visual images reflecting the interaction between humans and machine intelligence" mentioned in this article are explained. On the one hand, humans input instructions, data, and information to machine intelligence through various media or methods such as hands, feet, eyes (such as binocular focus instructions), expressions, movements, language, tone, sound, symbols, and codes. This includes input operations when creating machine intelligence and input operations when using machine intelligence, including instructions, data, and information given by humans to machine intelligence, as well as the process, capabilities, methods, and nature of such giving, including those situations that should be given to machine intelligence and are successfully given to machine intelligence and those situations that should be given to machine intelligence but are not successfully given to machine intelligence, and also includes the basic assumptions and premises generally contained in human intelligence when creating or using machine intelligence; on the other hand, machine intelligence uses codes, data, signals, symbols, graphics, and other media or methods to input instructions, data, and information to machine intelligence. The cognition, reasoning, selection, planning, execution, performance, movement, will, storage, and rationality of machine intelligence are fed back to humans (humans can also ask machine intelligence to inform them) in a way that is perceptible and acceptable to human vision, hearing, touch, etc. through images, logos, animations, colors, flashes, languages, sounds, vibrations, mechanical forces, motor forces, etc., including feedback, data, information given by machine intelligence to humans and the process, ability, method and nature of such giving. This feedback is based on the digital world that machine intelligence perceives, constructs and guides its calculations and actions through its various types of sensors combined with computers. It can be information feedback that is generated and presented instantly, or it can be information feedback that is extracted and presented later after storing relevant data.For example: when humans create robots, they input instructions on the computer keyboard and after the instructions are accepted, the cursor on the computer screen automatically jumps to the beginning of the next line and flashes, meaning that it has accepted and is ready to accept the next instruction; humans give instructions to robots in natural language, and the robot repeats the corresponding text in the human language in a completely consistent way using visual display or sound, and adds a sentence "I understand" (but in fact it only accepts the instructions literally without understanding the context of the instructions); humans check the code and data of machine intelligence, and read the specific instructions executed by machine intelligence and the specific code and data of the execution process and results from the visual images of the code and data log files; when humans use smart cars, the screen inside the car displays the real vehicle in front perceived by the machine intelligence as an animated car. When driving normally, the animated car is green. When the car in front brakes suddenly and may cause the car to rear-end, the animated car turns red to warn of danger; When operating an intelligent aircraft, the operator manually pulls up the joystick, but feels the force of the motor controlled by the machine intelligence pushing down on the joystick. Once the manual pulling action stops or the force is insufficient, the actual action of the joystick switches to automatic downward push; the operator sends a return command to the drone, but the drone flies to the wrong landing airport, and keeps sending back status feedback information of "normal return". Afterwards, the navigation log data uploaded and stored by the drone via satellite are retrieved, which shows that the drone has always believed that it is landing at the correct airport; the operator asks the drone why its flight trajectory is formulated, and the drone displays its flight trajectory in three-dimensional stereoscopic form on the screen and explains the specific considerations and calculation process for formulating this trajectory; the unmanned submarine stores the sonar radar data along the way in its fuselage. After returning, technicians extract the data for analysis and reproduce a three-dimensional picture of the seabed geography and topography of the entire route it passes through. Visual images reflecting the interaction between humans and machine intelligence serve as the vehicle or medium for human observers to perceive this interaction. Visually, these images can be direct, such as a filmed audio recording of a conversation between humans and machine intelligence, or indirect, abstract forms, such as a large screen depicting the communication process as a story and displaying the text of the conversation as text symbols. When a human observer views and manipulates visual images reflecting the interaction between humans and machine intelligence due to an event, they can view and manipulate images related to the event or images not related to the event (e.g., precedents similar to the current event). Both of these images represent the interaction between humans and machine intelligence.
[0025] This article explains the "visual images reflecting the interaction between machine intelligence and the world." First, the visual images reflecting the interaction between machine intelligence and the world are essentially the physical world that human intelligence perceives and constructs through its senses and brain, guiding its thinking and behavior. These images include images of the physical world directly perceived by human intelligence, information indirectly reflecting the physical world through visual perception, images of the physical world captured by conventional optical cameras and presented to the naked eye, and images of the physical world captured by infrared thermal imaging cameras and presented to the naked eye. Second, the visual images reflecting the interaction between machine intelligence and the world encompass, on the one hand, the content, process, capabilities, methods, and nature of the autonomous expression of machine intelligence, and, on the other hand, the content, process, capabilities, methods, and nature of the corresponding changes in the physical world. Visual images reflecting the interaction between machine intelligence and the world are the carrier or medium for human observers to understand the interaction between machine intelligence and the world. Visually, they can be in the form of direct images (such as audio footage of machine intelligence performing a certain operation) or indirect abstract forms (such as a large screen displaying the operation process of machine intelligence in the form of storytelling and displaying the running status of the computer program during the operation in the form of text symbols). They can be "subjective" visual images for the machine intelligence entity or "objective" visual images for the machine intelligence entity. They can be abstract information visual images that reflect the content, process, capabilities, methods and nature of the interaction between machine intelligence and the world as seen by humans with their naked eyes. They can be images of the interaction between machine intelligence and the world as seen by human intelligence inside or outside the machine intelligence body. These images record, express or reflect how machine intelligence intervenes in the world and how the world responds to machine intelligence.For a machine agent, the "subjective" field of view refers to the field of view obtained by the machine agent's own sensors (or, when this field of view is inconvenient to obtain in actual operation, a field of view that is as close to the field of view obtained by the machine agent's own sensors as possible, obtained by engineering methods that are as close as possible to the field of view obtained by the machine agent's own sensors). This is commonly known as "what the robot sees." Examples include the field of view from the head camera of a humanoid robot, images captured from the perspectives of the front and rear cameras of a smart car, the front view captured by the nose camera of a drone performing a flight mission, the operational images captured by the camera at the tip of an extended robotic arm of a smart crane, and the three-dimensional images of the seabed topography obtained by the sonar radar of an unmanned underwater vehicle during navigation. For a machine agent, the "objective" field of view refers to the field of view captured from the outside of the machine agent's interaction with the world (not the field of view obtained from the perspective of the machine agent's own sensors). Examples include images of a humanoid robot working in the kitchen captured by a home surveillance camera, a bird's-eye view of a smart car and its entire road traffic environment captured by a roadside surveillance camera, and a satellite bird's-eye view of a smart drone, its flight environment, and the overall terrain below. The so-called abstract information visual images that reflect the content, process, capabilities, methods and nature of the interaction between machine intelligence and the world, as seen by the naked eye, include, for example, the track, time, speed, distance, coordinates, remaining power and other operating status information of a smart car on an electronic map displayed on a large screen; the track, speed, coordinates, diving depth and other operating status information of an unmanned underwater vehicle detected by an underwater sonar array displayed on a large screen; the operating performance information such as the extended distance, folding angle, and thrust generated of the robot's robotic arm displayed on a large screen; the performance indicator information such as the extended distance, folding angle, and ultimate thrust of the robot's nominal maximum capability displayed on a large screen (if there is a scenario where the robotic arm cannot push a heavy object, and the abstract information visual image clearly shows that the thrust generated by the robotic arm at this time is consistent with the nominal ultimate thrust number, then it clearly reflects that the nature of this scenario is that the robotic arm is indeed unable to push an object that is too heavy, rather than other reasons such as a hardware failure of the robotic arm). When a human observer views and manipulates visual images reflecting the interaction between machine intelligence and the world due to a certain event, he or she can view and manipulate images related to the event, or he or she can view and manipulate images related to non-events (such as precedents similar to the current event). Both of these images are visual images reflecting the interaction between machine intelligence and the world.
[0026] Explain the "instructions" described in this article. The information in "instructions" can be direct or indirect, and can be key or supporting information. Instructions can be expressed in the form of text, annotations, language, voice (screen narration), symbols, marks, graphics, animations, highlights, various colors or color changes, etc. They can be dynamic or static, and the narrative expression of the development and change of events can be linear or nonlinear. The instructions can be directed to others or oneself.
[0027] Clarify the description in this article of "selecting, decomposing, organizing, and arranging visual images reflecting the aforementioned inconsistencies, and shaping, reorganizing, and connecting the time or space reflecting the aforementioned inconsistencies, thereby mimicking or guiding the attention, perceptual habits, and cognitive mechanisms of human audiences." The so-called "human audience" refers to the audience receiving the information, including both others and oneself. Specifically, the manipulation, selection, decomposition, arrangement, organization, and interpretation of visual images, resulting in information from chaos to clarity, can be sent and explained to other audiences to achieve sudden enlightenment, or sent and explained to oneself to achieve sudden enlightenment. Methods of imitating or guiding the attention, perceptual habits and cognitive mechanisms of human audiences, such as clarifying the order of chaotic information to meet the audience's cognitive requirements for order, or finding logical information points and their logical relationships in illogical information to meet the audience's cognitive requirements for rationality, or expressing the visual picture as a whole or panoramic picture when the audience pays attention to the overall picture or needs to pay attention to the overall picture, or reducing the visual picture and integrating it into the larger overall or panoramic picture when the audience pays attention to a larger overall picture or needs to pay attention to a larger overall picture, or segmenting, enlarging, and close-up the local details of the picture when the audience pays attention to the local details or needs to pay attention to the local details, or using the switching method and rhythm of the picture to imitate the audience's unconscious cognitive process mechanism such as the change of sight, attention switching, cognitive segment switching and brain-eye burden adjustment when blinking, a human visual action, or using the perspective of the camera on the head of a machine intelligent body or the perspective of the camera on the tip of a robotic arm to simulate the human's proactive perceptual habits of using the naked eye, or playing the video in slow motion when the audience's eyes cannot keep up with the key progress of the video, etc. The shaping, reorganization, and connection of time include, for example, synchronous playback, asynchronous playback, forward playback, reverse playback, pausing the video, frequent multiple pauses, frame-by-frame playback, fast playback, slow playback, skipping clips, repeatedly playing clips, juxtaposing images at different times, adjusting the pace of the narrative, nonlinear narration, and jumpy narratives. The shaping, reorganization, and connection of space include, for example, decomposing or combining images, achieving deliberate expression through deliberate composition, juxtaposing images from different spaces, photographing the same object from different perspectives to achieve different or comprehensive interpretations of the same object, switching between multiple angles, comparing multiple angles, capturing motion trajectories within the movement of a video, transforming the dimensions of an image, magnifying or enhancing a paused image to obtain details that are invisible to the naked eye in a continuous image, and so on.
[0028] This article explains the "chain of facts and logic that reflects the complementary nature of human and machine intelligence beneath the aforementioned inconsistencies." Appearance refers to the external phenomena and surface features of things, temporary manifestations influenced by external conditions and circumstances; essence refers to the inherent properties and essential characteristics of things, their most fundamental and profound attributes and characteristics. While inconsistencies between machine and human intelligence may appear, the underlying essence is that machine intelligence's inherent flaws complement human intelligence's inherent and unique capabilities, complementing and perfecting each other. A chain refers to a number of connected nodes, with more than one node. A single node does not constitute a chain. The logical chain includes the analysis chain, reasoning chain, induction chain, deduction chain, etc., which uses the internal logical relationship of things to draw correct conclusions, find patterns from complex situations, and minimize the influence of bias, blind spots and subjective factors. The manifestation of the logical chain combined with the fact chain can be derived from correlation or causality, and the establishment of the logical chain combined with the fact chain can reflect correlation or causality. The chain can be sequential or reverse in time. The interconnection and cross-verification of the nodes in the chain can greatly improve the authenticity and reliability, because the probability of multiple nodes in a long chain being unreliable at the same time (for example, visual illusions, one-sided understandings of facts, and misled understandings of facts) is relatively low. The longer the chain, the more angles, and the more interconnected and verified the information nodes, the higher the authenticity and reliability, and the more complete and comprehensive the facts. Multiple nodes on the chain jointly constitute an event, jointly constitute the background of the event, jointly point to the same fact, and jointly reveal the essence hidden by the appearance, avoiding the shortcomings of a single node being a single one-sided fact, and a single node being easy to forge, and possibly originating from misleading or even illusion. For losses caused by errors or anomalies in machine intelligence, in legal proceedings or administrative management, under the condition that the source of the image is legal, the logical chain combined with the factual chain can form an evidence chain, provide comprehensive and complete evidence, reduce the influence of subjective factors, ensure judicial fairness, and enhance public trust. This is what is needed for the social, economic and legal governance system related to machine intelligence.
[0029] The technical effect of this patented technical solution is to enable the overall application of machine intelligence in human society to move from an imperfect and unstable state to a perfect and stable state, which is introduced in detail below.
[0030] The inconsistency between "the cognition, reasoning, selection, planning, execution, performance, movement, will, storage, and rationality of machine intelligence" and "the purpose, intention, conception, common sense, principle, expectation, guidance, will, memory, and sensibility of human intelligence" actually stems from the inconsistency between the worlds perceived and constructed by machine intelligence and human intelligence, or the inconsistency between the event narrations of machine intelligence and human intelligence in their respective world constructions. This solution manipulates, organizes, and instructs visual images reflecting the interaction between humans and machine intelligence and the interaction between machine intelligence and the world, and clarifies the time, place, object, and event process of machine intelligence operations (the worlds perceived and constructed by machine intelligence and human intelligence, and the event narrations of machine intelligence and human intelligence in their respective world constructions) by combining facts and logic. The chain of facts and logic is used to reveal the confirmed cause or the speculated possible cause of the inconsistency, to infer the exact consequences or the predicted possible consequences that the inconsistency will lead to, to improve the authenticity and reliability of the re-enactment of the event, and to reflect the complementary nature of human intelligence and machine intelligence under the appearance of the inconsistency (the inherent defects of machine intelligence are the inherent and unique strengths of human intelligence. The gap of the complete incompetence of machine intelligence must and can only be made up by human intelligence. The complementary nature of human intelligence and machine intelligence is a natural law, which is not subject to human will and cannot be changed, created or eliminated. Humans must and can only obey this natural law to make the process and results of machine intelligence application perfect, and to make the social application of machine intelligence enter an overall expected stable and perfect steady state).
[0031] Only when humans are clear about the inner nature and potential consequences behind the inconsistent appearance can they intervene in the operation of machine intelligence in a targeted manner - making timely adjustments and corrections to the instructions of machine intelligence or rejecting and refusing the work results of machine intelligence: in the operation of machine intelligence, discover problems, detect anomalies, raise doubts, express ideas, verify logic, confirm facts, clarify logic, communicate and collaborate, and then make intervention commands, adjustments and corrections efficiently and correctly. After the machine intelligence operation, reject or refute the work results of machine intelligence with reason and evidence, proving the practical effectiveness of humans interacting with the world through machine intelligence in certain usage situations. If the results are not as good as the expected effects of direct human interaction with the world, it will prove the legitimacy and rationality of the intervention in machine intelligence operations, raise reasonable doubts with data and evidence on the design and manufacture of machine intelligence, thereby promoting and helping manufacturers to improve their products, analyze and review errors or accidents that have occurred, discover problems, detect anomalies, sort out processes, find out facts, clarify logic, organize evidence, communicate and collaborate, prove responsibilities, calculate losses, and then promote the subsequent completion of chain accountability and claims in the technical contract chain and the legal entity chain, and other social and economic legal operation procedures, to fill the current imperfect social and economic legal governance system involving the responsibilities of machine intelligence entities.
[0032] This technical solution selects, decomposes, organizes and arranges inconsistent and corresponding visual images, reshapes and reorganizes time and space to form information nodes that display key points and details, and avoids the physical world and digital world that are presented continuously for a long time from becoming a flat, lengthy, boring, uninterrupted, punctuated and invalid meaningless information string that is difficult to grasp. It reorganizes the seemingly isolated and unrelated key elements of different aspects that appear in the visual image in a comprehensive and orderly manner with clear priorities, points out the key points and details of these key elements that are not easy to be understood by human intelligence or cannot be recognized by machine intelligence, and the mutual relationships between them that are not easy to be understood by human intelligence or cannot be recognized by machine intelligence, reproduces and highlights the key elements and logical relationships in the physical world and digital world of movement that are ignored by human intelligence or mistaken by machine intelligence, unifies the discontinuous audio-visual information and technical logical relationships in time or place, makes the discontinuous information fragments coherent and then produces an overall meaning that can be understood by human intelligence but not by machine intelligence, and humans The audience feels involved in the process of understanding, which enables them to break away from passive reception of boring and ineffective information and transform into active participation in high-perspective and wide-span experiential understanding. It helps the audience to penetrate the surface and understand the hidden essence, summarize and reveal the inconsistencies between human intelligence and machine intelligence that are difficult for human intelligence to understand or cannot be recognized by machines, as well as the complementary essence of human intelligence and machine intelligence behind this inconsistency, and express the abstract correlation or causality contained in the complete chain of facts and logical chains in an explicit and vivid way. At the same time and in the same process, the audience's figurative thinking and semantic thinking are mobilized at the same time, so that the audience's concrete logic and abstract logic are connected together, and then they can clearly and accurately explore, verify, recognize and judge the inconsistency between human intelligence and machine intelligence from multiple perspectives such as vision, hearing, semantics, imagination, logic, principle, significance, authenticity and reliability, and eliminate the negative impact of false appearances, illusions, misleading, ambiguity and distractions in the audience's understanding of this inconsistency. The process and output of this technical method establish verifiable exploration, logical narrative, factual evidence, emotional experience and conceptual ethical consensus for the audience, forming a verification chain, logical chain, factual chain, narrative chain, evidence chain, and emotional chain, and jointly strengthen the concept of "complementarity between human intelligence and machine intelligence". The information audience not only shares and establishes rationality in the process of logical exposition, but also shares and establishes emotions in the process of narrative substitution, and also shares and establishes ethical concepts in the process of collective resonance. Their brains simultaneously carry out efficient modeling at the rational and emotional levels, efficiently reach consensus, and enhance their understanding of the complementary nature of human intelligence and machine intelligence.Information recipients can quickly extract clear and reliable information that is difficult for each of them to extract individually from the chaotic massive information, and empathize with the machine intelligence working process that they cannot experience again in person, so that the audience can have knowledge and technical life experience that has passed or has not yet been possessed, and make the audience's understanding process conform to both rationality and sensibility, visual and auditory senses and common sense, and the operating state of human consciousness and unconsciousness. Finally, a technical document that reflects the course of events and technical concepts is formed, which can be visual, readable, understandable, disseminated, storable, evidentiary, case-based, typical, and categorized. In form, it can be either a dynamic audio video file or a static picture text file.
[0033] This approach will continue to strengthen the human capacity to constantly consider the complementary technical concepts and ethics of machine intelligence when applying it socially. The process and outputs of this technical approach can mobilize the audience's emotions, triggering immersive and spontaneous emotional empathy and psychological resonance (such as grief, fear, sympathy, self-protection, the need for security, and the pursuit of justice), thereby enhancing the collective experience of ethical concepts (such as shared social principles, norms, outlooks on life, values, and concepts of right and wrong). This approach is also based on the authenticity and credibility of facts and logic, making it particularly suitable for targeted audiences such as judges, juries, journalists, media, industry associations, and the general public, aiming to achieve factual determination in legal procedures and collective consensus on social ethical concepts. As for human perceptual habits and cognitive mechanisms, when multiple people watch images at the same time, due to a certain arrangement of the images and the consistency of the event process and narrative rhythm, the audience often shows similar or synchronous brain activities, including world modeling, event modeling, logical modeling, empathy experience, psychological resonance, mass consciousness formation, and enhanced social ethical concepts. The similarity and synchronization of such individual and collective brain activities are often most evident at the key points, climax points, and conclusion points of the narrative and logic. It also enables the audience to recall carefully organized images more easily than those that are not carefully organized, leaving a particularly deep impression and forming a particularly firm belief. Because the audience's attention is guided, memory is stimulated, imagination is aroused, logic is established, emotions are mobilized, and value beliefs are strengthened in collective ethical concepts, it is easier to eliminate ambiguity, distraction, distractions, misunderstandings, distortions or wrong imaginations, and then establish a true and correct understanding and cognition and group consensus. This is especially effective and practical when facing judges and juries in court, facing experts from different professional backgrounds in the industry, and facing the superstitious and misunderstood public in society. The common sense of immersion and substitution is more likely to trigger a consensus after cross-verification, and even to a certain extent enhance the audience's interest, favorability and solidarity with each other. This consensus not only reflects objective facts, but also takes care of subjective emotions. Rationality and sensibility are intertwined, and the audience is attracted by those they have always believed in because of superstition. The public is deeply shocked by the truth that they are unwilling to believe (the public often has unrealistic and blind beliefs in high-tech, such as the superstition that "smart cars are definitely safer than human driving" and the delusion that "computers can create so-called general artificial intelligence that is the same as human consciousness"). Breaking technological superstition or technological prejudice is an extremely important technical effect. It can continuously refresh and enhance the scientific and technological ethics of human beings to pursue truth, goodness and beauty in a rational, forceful and orderly manner in the progress of science and technology, and avoid falling into negative cognitions such as superstition, prejudice, worship, slavery, greed, delusion, fear, and credulity surrounding machine intelligence. This will be of great help to the scientific and technological concept that human intelligence should learn from each other's strengths and weaknesses, complement each other's advantages, and fight side by side with machine intelligence, and to the scientific spirit of seeking truth and being pragmatic.
[0034] The object of manipulation and review in this technical solution is audio and video, rather than the raw data from various sensors on machine intelligence. This is because the "anomalies" that occur when humans interact with the world through machine intelligence cannot be predicted in advance. Without a goal, there is no means. If it cannot be predicted in advance, it is often impossible to pre-set corresponding sensors. In addition, the types and performances of sensors are diverse. It is difficult to control the cost and unreasonable and unfeasible to continuously add various sensors to respond to new situations. Limited sensors will never be able to cope with infinite new situations. This solution chooses to use audio and video, which is the medium with the largest amount of information transmission, the richest information level, and the highest information transmission efficiency for human intelligence. It records both the information itself and the method and process of transmitting information. Audiovisual images are a comprehensive form of expression of vision, hearing, and symbols (semantics). Of the total direct information of human perception, vision accounts for more than 80% and hearing accounts for more than 10%, while the data collected by sensors are only indirect symbols. Observers and interpreters of pure data symbols not only need to be extremely familiar with the definition, background, and context of the data, but also need to translate - only through logical thinking and situational imagination from abstract to figurative can they understand trivial data. Obviously, the multiple information transmission of audiovisual images is richer, more intuitive, and more efficient than the single information transmission of pure data. In most cases, abnormal or unexpected visual cases caused by machine intelligence will be lengthy and difficult to understand when expressed in language, text, or data codes, while visual expression is clear, easy to understand, immersive, and irrefutable. The most common way for machine intelligence to perceive the outside world is vision, and the most intuitive and commonly used way for human intelligence itself to perceive the outside world is also vision. Therefore, this technical solution uses vision as the most critical medium. By manipulating, organizing, and instructing visual images, it enables the integrated comprehensive cognition of humans, machine intelligence, and the world to achieve maximum effect within reasonable cost constraints.
[0035] This technical solution promotes collaborative collaboration. Multiple human collaborators can continuously self-examine and reflect from multiple perspectives, levels, and fields, cumulatively manipulating, organizing, and instructing visual images based on their individual findings of problems or perceptions of anomalies. This creates a diverse set of identities with multiple professions, backgrounds, and perspectives for each audience member with a single expertise, background, or perspective, cross-checking and enhancing trust, leading to efficient collaboration in technical communication, exchange, and decision-making.
[0036] With this technical solution, although the application of machine intelligence in human society is still "subject to abnormalities at any time" (we don't know what we don't know), it is no longer an imperfect and unstable state of "unclear after the incident", "nervous", "users often can only blame themselves", "unstable expectations in all aspects of society", and "people don't know how to better cooperate with machine intelligence". Instead, it is a perfect and stable state that can be predicted in advance, adjusted during the process, rejected after the event, confirmed facts, clarified logic, intervened efficiently, corrected in time, understood the truth, communicated efficiently, reached group consensus, traced back, proved reasonable, divided responsibilities, determined responsibilities at each stage, cross-examined in court, calculated losses, accumulated cases, summarized experiences, created knowledge, inspired future generations, enhanced understanding of the essence of human-machine complementarity, and strengthened the scientific and technological concept of human-machine complementary cooperation, which makes human society's overall expectations for machine intelligence applications more stable. That is, this patented technical solution enables the overall application of machine intelligence in human society to move from an imperfect and unstable state to a perfect and stable state. For all mankind, this technical solution continues to obtain specific lessons, case evidence and judicial precedents on how human intelligence can complement and cooperate with machine intelligence and how to make up for the inherent defects of machine intelligence. Every application, every case, every trial, every video and every type of model increases the knowledge and experience of how the complementary cooperation between human intelligence and machine intelligence is verified, confirmed and agreed on in the fact chain and logic chain, and enhances the overall understanding of the complementary nature of human intelligence and machine intelligence by mankind. The output of this technical solution (dynamic video documents or static graphic documents) can be stored, disseminated, understood and classified conveniently and efficiently, and enhance the multi-professional and multi-angle technical and efficient cooperation among humans. By collaborating with others, we can establish, convey, and deepen the scientific and technological concept and human confidence that "humans need to use their inherent and unique capabilities of human intelligence to compensate for the inherent shortcomings of machine intelligence, and that human intelligence and machine intelligence need to work together to complement each other and work side by side." This will enable more human users to have the concept and ability to draw inferences from past experiences and more accurately predict potential errors in machine intelligence in the next human-machine collaboration, whether similar or different, so that they can decisively intervene in advance, make corrections and adjustments to machine intelligence's work instructions, reject or dismiss machine intelligence's work results, determine responsibility for accidents, provide a basis for claims, and calculate the amount of losses. This will continuously enhance the ability of human intelligence and machine intelligence to complement and cooperate with each other as a whole. Therefore, this technical solution represents the technological development trend of complementary cooperation between human intelligence and machine intelligence to achieve perfection. It will establish stable expectations in all aspects of technology, law, economy, and society. It will also ensure that consumers, society, industry, and government feel at ease in the overall life cycle of machine intelligence applications, ensuring that its application is effective, rational, and well-regulated from beginning to end. This technical solution emphasizes "end" and "rationality and coordination."Great adventures, changes, explorations, and conquests can only be initiated by humans; machine intelligence can only facilitate them. Because intellectual property rights only consider the external physical dimension, and humanity's external physical endeavors lie in exploring and harnessing natural forces to enhance human freedom and well-being, in this endeavor, human intelligence and machine intelligence should complement each other, working side by side to fulfill their respective missions and achieving near-perfect collaboration, thereby achieving and sharing achievements that neither could achieve alone. The technical solution of this invention, driven by its ideological roots and concepts, aims to provide such a complementary and collaborative approach.
[0037] Once again, it is emphasized that the "chain" mentioned in this article refers to a number of nodes that are connected together and have more than one connection. A single node does not form a chain. In terms of technical features, this patent must make the facts and logic form a chain at the same time. It is a kind of "knowing it but knowing why it is so", which can form a verification chain, a logic chain, a fact chain, an evidence chain, an emotional chain and jointly strengthen the concept of "complementarity between human intelligence and machine intelligence". Correspondingly, a single node that simply reflects the surface phenomenon of inconsistency between humans and machines is a kind of "knowing it but not knowing why it is so". It cannot sort out the logic, explain the reasons, reveal the essence, and predict the consequences. It cannot explain the specific links without forming a chain. It cannot form a chain and cannot be cross-linked and verified with each other. It may be a one-sided fact, easy to forge or mislead, and may even be based on illusions. Therefore, it is not within the scope of protection required by this patent.
[0038] In addition, for the training and learning of machine intelligence, the process and outputs of this technical solution will continue to accumulate a large amount of negative training materials and logical training materials to help machine intelligence understand "errors" and "links in the logical chain." Obviously, these technical effects also help the overall application of machine intelligence to move towards a stable state for humans. DETAILED DESCRIPTION
[0039] The method of the present invention is widely applicable throughout the entire life cycle of machine intelligence, from design and creation to use, maintenance, and retirement. Human intelligence complements and mitigates the inherent shortcomings of machine intelligence. The following are some examples.
[0040] "Humanoid robots," as used herein, refer to robots designed and manufactured with machine intelligence at its core to resemble human form, structure, size, and interaction methods. They are designed and manufactured to adapt, through training, to the full range of human-like operating mechanisms and tasks. "Intelligent driving," as used herein, broadly refers to all automated, intelligent driving functions or modes that attempt to use machine intelligence to assist humans with the observation, thinking, decision-making, and operations associated with vehicle maneuvering, freeing them from having to observe, think, make decisions, and operate all the necessary tasks themselves. These include, but are not limited to, unmanned driving, autonomous driving, assisted driving, and human-machine co-driving. "Smart cars" and "smart vehicles" as used herein refer to all vehicles with intelligent driving capabilities, and these functions or modes can be turned on or off. "Drivers," as used herein, refer to any human on or off a smart vehicle who can issue commands to control the vehicle's maneuvers. "Smart aircraft" and "drones," as used herein, refer to aircraft with machine intelligence as their core technology module.
[0041] Case 1: Machine intelligence rigidly executing instructions. A human user assigned a general-purpose humanoid robot to the driver's seat of their conventional vehicle, performing conventional driving using their eyes, hands, and feet. The user specifically instructed the robot to "brake as soon as possible when in danger." The robot's first braking action, intended to avoid a rear-end collision with the vehicle ahead on the highway, failed. The vehicle rear-ended the vehicle ahead, resulting in a chaotic collision. Upon regaining consciousness, the human user, who had been unconscious in the vehicle, recalled that the robot had not braked at all, but was unsure and could not provide evidence.
[0042] Obviously, there was a discrepancy between the "cognition, reasoning, selection, planning, execution, performance, and storage of machine intelligence" and the "purpose, intention, conception, common sense, principle, expectation, and memory of human intelligence." Afterwards, technical experts investigated the accident, retrieved various visual images reflecting the interaction between humans and machine intelligence and the interaction between machine intelligence and the world, and manipulated and organized the visual images corresponding to the discrepancy, and explained them in detail. Combined with the visual images of human users giving instructions to robots (humans gave voice instructions to robots in natural language, "step on the brakes as soon as possible when there is danger"), it was pointed out that human intelligence The world construction of the robot's machine intelligence is "If a possible danger is encountered during driving, the brake pedal should be stepped on at the maximum speed and force that is common sense for humans." Combined with the display screen of the robot's own storage of machine intelligence operation records (the robot uses voice to ask to confirm the correctness of the instruction "Step on the brake pedal as quickly as possible when encountering a dangerous situation, right?", the user answers "yes", the robot answers "Instruction received" and records the instruction [Step on the brake pedal as quickly as possible when encountering a dangerous situation] in its machine intelligence), it indicates that after the robot receives the instruction, the world construction of its machine intelligence is "If a possible danger is encountered during driving, the brake pedal should be stepped on at the maximum speed and force that is common sense for humans." The robot's own stored machine intelligence operation records indicate that the robot's event narrative is "The robot correctly and promptly perceived and identified potential traffic hazards, strictly followed user instructions to brake as quickly as possible in the face of danger, and kept the brakes pressed fully at the time of the accident." The robot's own stored footage of the accident captured by the robot's leg camera indicates that the human intelligence world construction is "The robot correctly pressed the brake pedal, but the problem was that the brake pedal was broken by the foot, separating the vehicle's brake components and effectively rendering the brakes inoperable." The robot's own stored machine intelligence operation data indicates that the human intelligence world construction is "There is no awareness of the brake mechanism's fracture and failure." The product specifications provided by the robot manufacturer indicate that the human intelligence world construction "The maximum force exerted by the robot's legs is hundreds of times greater than that of a human."Combined with the above visual images, the worlds constructed by machine intelligence and human intelligence, the respective event descriptions of machine intelligence and human intelligence in their respective world constructions, and the time and place of the above key event nodes (the time and space information in the respective event descriptions) are indicated. The above information is compared and connected with each other, and unified and integrated into a fact chain and a logic chain. The two chains are formed at the same time. This fact chain and logic chain are: at the beginning of the event, the human gave the robot a voice command in natural language, "When there is danger, step on the brakes as soon as possible." The command received by the robot is actually "When there is danger, use the fastest speed of the robot's legs and feet." and maximum power to execute the braking action"; at the middle of the incident, the robot accurately sensed and identified the danger when encountering it and strictly obeyed the command to execute the braking action at the fastest speed and maximum power of its legs and feet. However, because the robot is designed to have performance far exceeding that of humans (the electromechanical power is hundreds of times that of humans' common sense), the brake pedal was broken, resulting in brake failure. Therefore, the human user could not realize in advance that the robot had made a biased and non-common sense understanding and execution of its instructions. During the incident, it seemed that the vehicle did not slow down. In memory, it was believed that the robot did not execute the command and even suspected that it had stepped on the wrong pedal. This chain of facts and logic reflects the complementary nature of human and machine intelligence beneath this surface of inconsistency: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "humanity" because humans' interactions with the world involve infinite factors, which machine intelligence cannot account for. There are also countless possibilities for similar anomalies that machine intelligence may encounter in its autonomous actions. However, humans possess common-sense, real-life experience and can therefore use common sense to deal with ambiguous life issues. The scales that human intelligence can easily grasp with common sense are beyond the grasp of machine intelligence. Accident responsibility can only be defined, divided, and borne by humans. The victim of the accident can leverage this chain of facts and logic to file a lawsuit against the robot manufacturer, arguing that the robot's electromechanical performance design or human-machine interaction design were unreasonable and that the manufacturer should bear responsibility for the accident. The robot manufacturer can also leverage this chain of facts and logic to request design modifications, pursue accountability, or file claims against its electromechanical execution module supplier (but not the perception and recognition module supplier, the logic decision module supplier, or the electromechanical execution accuracy module supplier), thereby establishing clearly defined responsibilities at all levels, including technical, economic, and legal. The outputs of this case study provide concrete lessons and case studies on how human intelligence can complement and cooperate with machine intelligence and how to overcome inherent flaws in machine intelligence. This will help other human users consciously consider the accuracy, rationality, and enforceability of their instructions when applying machine intelligence, or refer to this case study to more efficiently handle similar incidents. It will also help manufacturers consider how to improve their robot product designs. The process and outputs of this technical solution serve as negative and logical training materials, which can be used to train robots.
[0043] This technical solution can select, decompose, organize and arrange the visual images on the above-mentioned factual chain and logical chain, shape, reorganize and connect the time or space reflecting the above-mentioned inconsistency, imitate or guide the attention, perceptual habits and cognitive mechanisms of human audiences, emphasize key points, highlight details, and clearly and explicitly indicate the inner essence or potential consequences behind the above-mentioned inconsistent appearances, and generate the following dynamic video document output (the switching of video images is indicated by "→" below): The visual display of the robot receiving the instruction process. When giving the robot an instruction, the user's original intention is simply to The robot is required to brake as quickly as possible when encountering danger to ensure safety. Therefore, a natural language voice command is given to the robot: "Apply the brakes as quickly as possible when in danger." The robot then confirms the correctness of the command by asking, "Is it correct to apply the brake pedal as quickly as possible when encountering danger?" The user answers "Yes," and the robot replies, "Command received." The time is displayed in the lower right corner of the screen, and a voice narration, combined with text below the screen, prompts, "The user's natural language was interpreted by the robot and confirmed as the literal input command [Apply the brake pedal as quickly as possible when encountering danger]. The robot's "eye" view and the machine's intelligent vision processing state are displayed side by side. Driving hazards are highlighted and highlighted in red in both screens. When the robot encounters a dangerous situation, its "eyes" (the camera on its head) immediately and correctly perceive the corresponding danger. The time is displayed in the lower right corner of the screen, and a voice narration, combined with text below the screen, prompts, "The robot accurately identified and perceived the dangerous situation immediately." → Pause the two screens and zoom in on the "brake pedal pressed" sign that indicates the machine's intelligent decision-making. The time information is displayed in the lower right corner of the screen. At the same time, the narration voice combined with the text below the screen prompts "The robot accurately identified the dangerous situation and decided to immediately press the brake pedal, indicating that its perception, recognition and logical decision-making are operating normally and appropriately." → The camera position of the robot's foot pedal position shows that the humanoid robot did immediately press the brake pedal at the fastest speed (rather than the accelerator pedal or other incorrect stepping positions). The time information is displayed in the lower right corner of the screen. At the same time, the narration voice combined with the text below the screen prompts "The robot correctly and accurately pressed the brake pedal at the first time. The robot's "eye" perspective displays a magnified image of the vehicle's exterior and speedometer, showing the vehicle continuing to move forward at a relatively high speed, without the drastic deceleration expected during emergency braking. A zoomed-in image of the speedometer indicates a sharp drop in speed, followed by a pause until the collision. The time is displayed in the lower-right corner of the screen, accompanied by a voiceover message along with text at the bottom of the screen: "The robot's perspective shows that the vehicle's speed did not continue to decrease dramatically, and the speed value displayed did not decelerate dramatically until the collision."”→The camera view at the robot's leg pedal position shows a messy collision process. The details of the broken brake pedal are partially magnified and played at 0.5 times the speed. The time information is displayed in the lower right corner of the screen. At the same time, the narration voice combined with the text at the bottom of the screen prompts "A collision occurred. The pedal was immediately broken due to the robot's stepping on it, and a large hole was stepped on in the vehicle chassis." →The camera view at the foot pedal position is magnified and quickly rewound at 8 times the speed to the footage of the robot actually stepping on the brake pedal at the fastest speed. This footage is played repeatedly (the time information is displayed in the lower right corner of the screen) at 0.5 times the speed. The maximum force value of the foot stepping on the product specifications provided by the robot manufacturer and the maximum force index value that the pedal can withstand provided by the pedal manufacturer are annotated next to the screen. Mathematical symbols are used to show that the robot's maximum force value far exceeds the force index value that the pedal can withstand provided by the pedal manufacturer. At the same time, the narration voice prompts "The maximum force value of the robot's foot provided by the robot manufacturer is far greater than the force index value that the pedal can withstand provided by the pedal manufacturer. ” → “Conclusion” is displayed in the center of the screen for 1 second → The camera position of the robot's foot pedal is played again at 0.5 times the speed, showing the collision process in a mess, the broken pedal process, the robot's machine perspective video showing the vehicle being violently collided and shattered, and the partially enlarged vehicle speed value video. The details of the broken brake pedal and the vehicle speed display are partially enlarged and played frame by frame. The audio video is played again in which the user simply asks the robot to take the fastest braking action to ensure safety when encountering danger, and the robot interprets and confirms the literal instruction "Step on the brake pedal as quickly as possible when encountering danger". The machine intelligence status display screen of the robot accurately sensing and identifying the danger and deciding to "step on the brakes immediately" is played again, and the "step on the brake pedal" sign that is lit at this time is enlarged. The narration voice combined with the text prompt at the bottom of the screen says "Preliminary judgment, the humanoid robot broke the brake pedal with one foot, which led to the failure of the brake. In order to be more advanced and superior to humans, robots are designed and manufactured with leg strength hundreds of times greater than the average human's maximum strength. However, in this case, the machine instruction the robot received from the human user, which it literally translated and accepted, was simply "press the brake pedal as quickly as possible when encountering danger," without specific instructions on how much force to use. Therefore, when correctly perceiving and identifying the danger and pressing the correct brake pedal, the robot's narrative and world-building depicted it diligently pursuing maximum speed. Consequently, its leg motors used maximum power and output maximum force, causing the pedal to break immediately due to the inappropriately large force applied, separating the vehicle's brake components and effectively rendering the brakes ineffective. On the other hand, in the narrative and world-building of human intelligence, human intelligence originally expected the robot to efficiently and effectively brake when necessary, and to do so like a human. Ordinary human users could not reasonably anticipate such inappropriate behavior from the robot.In this case, the robot's machine intelligence lacks the common sense of human intelligence, and its performance in this application was excessive. Alternatively, the robot manufacturer failed to prevent users from using robots with overly strong design specifications for inappropriate applications. Alternatively, the robot manufacturer should have more accurately and predictably provided users with a whitelist of executable tasks, and human users should have developed a scientific awareness of giving robots more rigorous and accurate instructions.
[0044] Obviously, by capturing the key images of each node of the fact chain and logic chain in the above-mentioned dynamic video document and combining them with instruction information, static graphic document outputs can also be generated.
[0045] The audiences of the outputs, including users, the public, robot suppliers, suppliers of robot suppliers, judges, juries, industry organizations, government management departments, etc., will be able to understand this incident clearly, quickly, comprehensively and efficiently. Obviously, different terms and different focuses can be used in dynamic or static outputs for different audiences to improve the efficiency of understanding of different audiences. All audiences will learn how to better complement and cooperate with robots from their respective perspectives.
[0046] If this technical solution is not adopted, the human user in the car will only think afterwards that the robot did not brake or slow down at all (did not execute the command), but will not be able to understand the details and logic of the specific process, nor can they provide specific evidence to file a claim for rights protection against the robot manufacturer. The matter will only be left unresolved, and the user can only blame himself for his bad luck. Even if the robot manufacturer admits its responsibility, it will not know which link in the chain of the machine intelligence body (electromechanical execution force module supplier, perception and recognition module supplier, logic decision module supplier or electromechanical execution accuracy module supplier) to make a reasonable and well-founded claim for accountability.
[0047] Case 2: A case where machine intelligence cannot be started due to its infinite pursuit of precision in instructions. A user instructs a humanoid robot to "catch a drop of water from the faucet with your hand as it falls." The robot asks, "How high above the ground should you catch the drop?" The user answers, and the robot asks, "Where should you catch it with your hand?" The user answers, and the robot asks, "How long should your hand wait at the point of contact before the drop makes contact?" The user answers, and the robot asks, "Is the falling time of the drop calculated based on Newton's second law?" The user answers, and the robot asks, "Does the value of the gravitational constant take into account the latitude and longitude, time, altitude, and lunar inclination at this time?" The user answers, and the robot asks, "Is the evaporation of the falling drop considered? Is the temperature drop caused by evaporation considered? Is the effect of the current temperature, humidity, and wind speed on the evaporation rate considered? Is the difference in the impact of the impact pressure increase at the bottom of the drop and the decompression of the tail flow at the top considered? Is the effect of the slight dominance of light isotopes (hydrogen) over heavy isotopes (deuterium) in the water vapor considered? Is the air considered?" Are there any considerations for the impact of dust and impurities in the water droplets on evaporation? Have you considered the impact of metals incorporated into faucets and pipes on evaporation? After the user answered, the robot asked, "At what height above the ground should the water droplets be caught? Have you considered the impact of the aerodynamics of the droplet, causing it to oscillate, deform, and rotate, affecting its falling velocity and trajectory? Have you considered the Lorentz force perturbations caused by water molecules and ions cutting through the Earth's magnetic field lines during its fall? Have you considered the changes in the droplet's shape and center of gravity caused by the behavior of microorganisms in the droplet during its fall? Have you considered the impact of the separation and combination of internal microbubbles on the droplet's shape when the droplet separates from the liquid?" The user, puzzled, asked the robot to list all the questions it would ask on the human-computer interaction screen. Seeing the robot's endless list of questions, the user realized that the robot would continue asking questions indefinitely, confirming their suspicion that the robot would fail to start executing due to the endless questions. Therefore, the user decided to modify the instructions, explicitly requiring the robot to tolerate a certain degree of error and explicitly informing the robot of this degree.
[0048] Obviously, there is a discrepancy between "the cognition, reasoning, selection, planning, execution, and performance of machine intelligence" and "the purpose, intention, conception, common sense, principles, and expectations of human intelligence." Various visual images reflecting the interaction between humans and machine intelligence and the interaction between machine intelligence and the world are retrieved, and several visual images corresponding to the inconsistency are manipulated, organized, and instructed. Combined with the external monitoring audio and video images of the entire event process, the question list images on the human-computer interaction interface screen during the event, and the images of the machine intelligence operation records during the event, it is pointed out that "the world construction of human intelligence has a certain degree of ambiguity and tolerates a certain degree of error," while the world construction of machine intelligence "pursues precision due to its mechanical design based on mathematics and physics." The event narrative of machine intelligence in its world construction is "requiring completely accurate instructions for comprehensive consideration and rigorous execution, thereby delaying the start of execution." The event narrative of human intelligence in its world construction is "expected that the robot should be able to autonomously and appropriately execute immediately with common sense in human life" and "instructions have to be modified in order for the robot to start execution immediately." The aforementioned visual images indicate the worlds constructed by machine intelligence and human intelligence, their respective event narratives in their respective world constructions, and the time and location information in their respective event narratives (delay time information caused by the robot, spatial information of the robot standing still). These pieces of information are compared and linked with each other, and unified and integrated into a chain of facts and a chain of logic. The two chains are formed simultaneously. This chain of facts and the chain of logic are: at the beginning of the event, the human issues an instruction to the robot, the robot successfully accepts the instruction and begins to attempt to execute it. In the middle of the event, the robot continuously seeks more detailed and precise instructions from the user for rigorous execution, thus delaying execution. The human user hopes that the robot can start execution quickly and cannot accept the delay caused by the robot. When it is confirmed that the robot still has a large number of questions to ask, it has to modify the instruction. This chain of facts and logic reflects the complementary nature of human intelligence and machine intelligence under the inconsistent appearance: machine intelligence cannot cope with "infinity", "chaos", "responsibility" and "human nature". Robots do not have the common sense of human life and cannot work autonomously under a certain degree of ambiguity. Every time they encounter an ambiguous problem, they need to ask human users for an approximate degree of specificity. Humans involve infinite factors in their interaction with the world, and machine intelligence cannot consider infinite factors. If they must consider them mechanically, it will infinitely require users to refine the instructions or infinitely learn human common sense. Humans have common sense life experience, and therefore can use common sense to deal with ambiguous problems in life. The scale that human intelligence can easily grasp with common sense is something that machine intelligence cannot grasp autonomously.
[0049] Users can select, decompose, organize, and arrange the visual images in the aforementioned factual and logical chains, shape, reorganize, and connect the time or space reflecting the aforementioned inconsistencies, mimicking or guiding the attention, perceptual habits, and cognitive mechanisms of human audiences, emphasizing key points and highlighting details, and providing clear and explicit instructions and explanations of the inherent nature or potential future consequences behind the aforementioned inconsistencies, thereby generating dynamic video document outputs. Users can also capture key images from each node of the factual and logical chains in the dynamic video documents and combine them with the instructional information to generate static graphic document outputs. Afterwards, users can use these deliverables to prove the legitimacy and rationality of their intervention in the robot operation process and the modification of its instructions. They can also use these deliverables to report problems to the robot manufacturer and seek assistance. If the machine intelligence operation is delayed and causes losses, users can use these deliverables to hold the robot manufacturer accountable and seek compensation. The robot manufacturer can also use these deliverables to improve the machine intelligence product. The lessons learned from these typical machine intelligence failures or successful human interventions can be effectively accumulated and disseminated through case studies to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better awareness and ability to anticipate interventions. The process and outputs of this technical solution are negative and logical training materials that can be used to train robots.
[0050] Case 3: Machine intelligence fails to understand local customs and practices. In the lobby of a public hospital in Country A, patient B and his family member, Ms. C, were sitting in the rest area, looking worried and distressed. Robot D, an information service robot deployed at the hospital's front desk, proactively "approached" and asked, "How can I help you?" Patient B complained, "How long will it take for me to recover?" Robot D, with "proactive concern," asked, "Where do you feel unwell?" The patient replied, "I've just been diagnosed with stage III cancer, and the medical consultation team is currently discussing a treatment plan..." Robot D consulted a public medical database (not the hospital's internal database) through a cloud service and responded, " Please don't worry too much. According to national statistics, the average life expectancy for patients with stage III cancer is 1.4 years, which can be extended to 3.6 years with appropriate medical care." Upon hearing this data, patient B burst into tears in despair, and his family member, Ms. C, looked annoyed. The visual machine intelligence of service robot D recognized the unhappy expressions of B and C, and proactively said, "Let me sing a song for you," and then sang and danced. The patient's family member, Ms. C, then flew into a rage and pushed and smashed robot D in public, damaging it. Afterwards, the hospital sued Ms. C for "intentional damage to property" and took her to court.
[0051] Obviously, in this incident, there was a discrepancy between "the cognition, reasoning, selection, planning, execution, performance, and rationality of machine intelligence" and "the purpose, intention, conception, common sense, principle, expectation, and sensibility of human intelligence." The defense lawyer manipulated and organized various visual images corresponding to this discrepancy: he retrieved the bird's-eye view surveillance video of the hospital lobby during the entire incident, captured the entire process of the robot proactively asking questions, providing services, singing and dancing, and being pushed over and smashed, retrieved the machine intelligence operation log data from the robot's built-in storage, recorded the process video of the robot's machine intelligence "seeing, identifying, deciding, and acting" throughout the incident, and presented the data and code of the computer program running in visual form. Combined with the visual images, the world construction and event narration of the robot's machine intelligence are as follows: according to the program requirements, "take the initiative to approach service inquiries and comfort patients", according to the program requirements, "the robot must try its best to answer human questions, cannot lie, and cannot conceal" so it searches for medical statistical data and informs the data to patient B, according to the program requirements, "continue to observe the expressions of human users", according to the program requirements, "bring joy to unhappy human users", and according to the program requirements, "play songs from Country A that express joy and auspiciousness and dance to them". Combined with the visual images, the world construction and event narration of human intelligence are as follows: Robot D was designed and produced by Country E. As a robot mainly engaged in medical services, the designers and engineers of Country E simply arbitrarily and mechanically designed the service logic of "actively approaching and identifying the customer's mood and bringing joy to the customer." The machine intelligence autonomously (rotely and arbitrarily) searched and recorded a song from Country A with the attributes of "cheerful" and "auspicious." However, this song is very inappropriate in the cultural traditions of Country A and the unique context of this hospital scene. The song is mainly used to express joy during the New Year in Country A. The lyrics contain the word "congratulations" repeated several times. Obviously, it not only fails to comfort the patient's family and bring them joy, but even carries a sense of ridicule and humiliation. The aforementioned visual images indicate the worlds perceived and constructed by machine intelligence and human intelligence, the respective event narratives within these worlds, and the time and location information within these narratives. This information is then compared and linked to one another, unified and integrated into a chain of facts and a chain of logic. The simultaneous formation of these two chains reveals the complementary nature of human and machine intelligence beneath this apparent inconsistency: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "humanity," and human intelligence must and can only compensate for this. Robot designers and engineers did not intend this scenario; the robot simply faithfully executed its program instructions. Robot designers and engineers cannot pre-program or train machine intelligence to encompass the endless elements and conditions of human life.In essence, data only ever reflects part of humanity and the world. Cognition, measurement, and understanding of humanity and the world solely from a mathematical perspective is inherently biased. This is the fundamental reason why machine intelligence does not know that the songs it chooses are inappropriate and does not understand cultural customs and common sense. In the robot's world construction and event narration, it simply mechanically and honestly serves customers and tries to bring joy to unhappy customers. In the world construction and event narration of human intelligence, humans are unacceptably offended in extremely vulnerable situations. This is a reasonable narrative about human nature.
[0052] Country A's trial utilizes a jury system. To make the technical solution deliverable more relevant for courtroom cross-examination and other technical aspects, the defense attorney zoomed in on a surveillance video from the hospital lobby showing Ms. C's expression of sadness, tears, anger, and humiliation, and explained it. The video was then played repeatedly, explaining the specific reasons for Ms. C's emotions, further demonstrating logically that these emotions led her to attack the robot in a rage, causing damage. The defense attorney also retrieved video footage of a brief, one-sided requirements-gathering interview conducted by the designer from Country E with the client at the hospital in Country A, as well as logs of the program's autonomous search, collection, and playback of the song, explaining that the song's playback was a case of "good intentions but bad consequences." When the designers and engineers from Country E created this robot for use in Country A, they should have respected Country A's cultural customs and common sense. This common sense included "not directly disclosing bad medical news to patients, and not by robots." However, such a premise was clearly not explicitly programmed into the robot, which the defense attorney pointed out as an inappropriate product design by the robot manufacturer. This carefully curated and organized video of the instructional information, featuring time, location, process, the world constructed by the robot's cognition, the program logic within the robot's narrative, and the human user's expressions and actions, logically connects and confirms the rationale behind Ms. C's actions. While the incident appears to be "intentional damage to hospital property," the simultaneous logical, factual, and emotional links demonstrate and explain the essence of the incident, which is justifiable due to "unreasonable design of the robot product." This fully explains the incident both rationally and emotionally, providing a complete chain of evidence. According to the direct legal provisions, combined with the prima facie evidence of Ms. C smashing the robot in the surveillance video, Ms. C would likely have been convicted of "intentional damage to property." However, after viewing this carefully curated video of the instructional information, the jury was moved by the rationality of the emotions expressed by Patient B and her family members during the incident. The twelve jurors resonated with righteous indignation and, collectively recognizing the factually unreasonable design of the robot, ultimately unanimously returned a not guilty verdict. Afterwards, the robot manufacturer also adjusted its machine intelligence design to refer users' medical inquiries to human physicians for answers. The lessons learned from such typical machine intelligence failures can be effectively accumulated and disseminated through case studies, helping more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better understanding and ability to reference precedents. Once the court ruling in this case is finalized, it will provide strong guidance for similar cases in the future. The process and outputs of this technical solution are negative training materials and logical training materials, which can be used to train robots.
[0053] Case 4: Machine intelligence being deceived by malicious noise. At the entrance of a large-scale social event, a security robot powered by visual AI worked alongside human security guards. Its job was to instantly capture and scan the faces of spectators entering the venue to identify suspicious individuals. The technical requirement for the scanned faces was that they were completely unobstructed. These facial scans were then compared against the public security department's database of wanted criminals. As a large number of spectators entered, the security robot scanned a large number of faces at extremely high speed. Its operational status was displayed on the security robot's human-machine interface screen. It rapidly framed, identified, and compared numerous faces captured by the robot's camera. If a wanted criminal's identity was successfully matched, an alert was issued and the comparison result displayed. If no wanted criminal's identity was matched, no special display was given. If facial obstruction, such as wearing a mask, was detected, an abnormality warning was issued. At the time of the incident, an unusual audience member entered. This audience member was wearing a headband on his forehead with a pattern of two eyes on it. The pattern was about the same size and shape as the audience member's eyes and was located directly above the audience member's eyes. If you don't pay special attention, it looks like just a decorative headband with a unique pattern. The security robot did not show any reaction to this face that was obviously obscured. At this time, there was a discrepancy between "the cognition, reasoning, selection, planning, execution, and performance of machine intelligence" and "the purpose, intention, conception, common sense, principles, and expectations of human intelligence." Machine intelligence should have issued an abnormality warning due to obscured or incomplete faces of the recognized objects, but it did not do so. The human security guard noticed this obvious inconsistency and immediately used the security robot's human-computer interaction interface to accurately review the machine intelligence operation status screen of the facial recognition scan at the time and location (space) when the audience member just entered. The screen showed that when and where the audience member's face entered the security robot's field of view, no facial recognition frame appeared on his face. The facial recognition status There was nothing displayed on the screen. These visual images, combined with the obvious inconsistency before, verified that in the world construction of machine intelligence, "the face that appeared at that time and place was not a human face at all." Its event description was "no one was seen at that time and place at all." However, in the event description and world construction of the human security guard, this was obviously a human face wearing a headband, but the pattern of this headband was particularly suspicious. The human security guard then asked the audience to take off the headband and rescan. At this time, the robot immediately issued a warning and displayed the comparison results, catching the wanted criminal who was trying to get through by disguising his appearance to create interference, deceive the machine intelligence, and sneak into the venue.
[0054] The aforementioned visuals, combined with the instructions and the time and location information indicated in each video, form a chain of facts and logic: Human users expect and require security robots to perform efficient facial recognition to identify potential wanted criminals among the audience. However, a hidden underlying expectation and requirement is the ability to "common sense-distinguish incomplete, abnormal faces, or those attempting to deceive." However, the machine's intelligence is distracted by the noise of the eyes above the eyebrows and fails to recognize the face as a human. The program logic naturally avoids comparing it to the wanted criminal database, and therefore fails to successfully identify and issue an alarm. This chain of facts and logic reflects the complementary nature of human and machine intelligence beneath this inconsistency: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "human nature," and must and can only be compensated by human intelligence. Whether a wanted criminal has dark circles under his eyes or wears heavy makeup, dyes his hair white or grows a beard, or wears a hat or turban of any unusual pattern, human users expect visual machine intelligence to efficiently and robustly recognize their faces. This is something machine intelligence cannot fully achieve and must, and can only, be compensated by human intelligence, because so-called abnormal scenarios are endless. The problem of "recognizing abnormal faces" itself is chaotic and ambiguous. This time, when programmers encounter two extra eyes on the forehead, they program in corresponding logical conditions or increase training. However, in the future, they will encounter faces with an extra nose, two extra mouths, fake eyebrows under the eyes, animal noses glued to their faces, closed eyes with heterochromatic pupils painted on their eyelids, animal ears on their heads, faces covered with Peking Opera masks, and so on. There are endless and diverse attempts to deceive machine intelligence. Machine intelligence will always encounter situations that are not pre-configured or trained in its program. "Anomalies" are endless and cannot be defined. Complementary to this, human intelligence naturally understands that a normal face is one with fully visible facial features, and can easily identify and detect anomalies and filter out noise. If the machine intelligence fails to detect the wanted criminal and the human security guard is negligent in his duties and misses this anomaly, then if criminal casualties occur during the event, the robot itself cannot be held responsible. Human intelligence must and can only be responsible for identifying the anomaly and eliminating interference.
[0055] Afterward, human security guards compiled the logical and factual chains into visual materials for communication, dissemination, reporting, and collection. This not only justified their intervention, but also served as a complete chain of evidence to accuse the wanted criminal of malicious deception. The lessons learned from these typical machine intelligence failures or successful human interventions can be effectively accumulated and disseminated as case studies to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better awareness and ability to anticipate interventions and reference precedents. The process and outputs of this technical solution are negative training materials and logical training materials that can be used to train robots.
[0056] Case 5. A case where human intelligence corrects the contradiction that machine intelligence cannot reconcile. A clothing design company uses a machine intelligence interview robot to interview clothing designers. A candidate entered the interview after scoring full marks in the clothing design skills test. The interview robot asked: "From your resume, I found that you are a recent college graduate. Your major is business administration, is that correct?" The candidate replied: "Yes, my major is business administration, and clothing design is my second major that I studied in night school after class." The interview robot then recorded "Interview conclusion 1. The candidate's clothing design skills were learned in night school, not in college." The interview robot asked: "Why didn't you take clothing design as your major in college?" The candidate replied: "Because high During the exam, my mother wanted me to inherit her trading company and wanted me to apply for business administration, but I loved fashion design, so I enrolled in night school. The interview robot then recorded, "Interview conclusion 2. The candidate may be insubordinate, not fully following instructions or orders from superiors." The interview robot asked, "From your resume, I see that your business administration grades are very poor, but your fashion design grades are very good. Why?" The candidate replied, "Because I really love fashion design, I practiced sewing late a few nights, which affected my business administration exam the next day." The interview robot then recorded, "Interview conclusion 3. The candidate's main My work performance is poor, possibly due to my extracurricular activities, and my focus on the job may be poor." The interview robot asked, "From your resume, I see that you previously worked at several trading companies, leaving each one after one to three months. Why did you change jobs so frequently?" The candidate replied, "Those were my mother's companies. She pressured me to work there after graduation, and I couldn't resist. After a few months, I just didn't like it. I still prefer fashion design." The interview robot then recorded "Interview Conclusion 4. The candidate is driven by 'whether they like it or not' and has not demonstrated a dedicated attitude of 'loving what they do.' Their frequent changes in employment relationships may indicate willfulness and indecision in their work. Lack of endurance." The interview robot asked, "Does your mother know you're applying for a job at our company?" The candidate replied, "I don't know. I kept it a secret from her. She doesn't like my career in fashion design, saying it's unrealistic and doesn't pay. Every time I talk about fashion design, she gets furious. I even hid my enrollment in night school from her." The interview robot then recorded, "Interview Conclusion 5. The candidate was dishonest with their superiors and concealed something." The interview robot asked, "How is your relationship with your mother?" The candidate, looking displeased, replied, "We haven't spoken in a long time." The interview robot then recorded, "Interview Conclusion 6. The candidate had difficulty building a harmonious relationship with their superiors."The interview robot asked, "Will you honestly tell her that you're here today seeking a fashion design position?" The candidate, looking displeased and sad, replied, "She's just been diagnosed with stage 3 cancer. I'd better not tell her. Since ancient times, it's difficult to be both loyal and filial..." The interview robot then recorded, "Interview Conclusion 7. The candidate explicitly refused to communicate candidly with her superiors for the second time." The interview robot asked, "According to statistics, the average life expectancy after a stage 3 cancer diagnosis is 0.8 years. Do you feel regret or shame for not being honest? Do you consider yourself a good daughter?" The candidate flew into a rage, leaping from her chair and slamming the table. She replied, "What does this have to do with you, a robot? What does this have to do with this job? What's the point of asking this?" After panting for several seconds, she realized her gaffe and said, "I'm sorry, I'm sorry. I really like this job. I'm sorry I couldn't help myself..." Then she sat down, looking dejected. The interview robot then recorded, "Interview Conclusion 8. The candidate failed the emotional stress test, indicating impulsiveness." "After the interview, the interview robot finally recorded "Final conclusion, in interview conclusions 1-8, all candidates performed negatively, and it is recommended to be eliminated."
[0057] The person in charge of the interview, who is experienced and has seen countless people, will check the robot's interview work afterwards, and retrieve and operate relevant videos such as the external shooting video of the entire interview process, the video of the interview process "through the eyes" of the interview robot, and the video reflecting the running status of the robot's machine intelligence program logic. Then, combined with the time and place information indicated by each video screen, the following video screen with instructions is produced: Add the instruction "The skill test has clearly shown the candidate's business skill level. Is fashion design the candidate's main major?" to the video clip screen where the robot writes the interview conclusion 1. It is only a reference factor for their professional level, and it is particularly valuable to achieve a full score on the skills test by studying only in night school. "Add instructions to the video clip where the robot writes Interview Conclusion 2:" The robot's rigid interpretation of the human mother-daughter relationship as a superior-subordinate relationship is biased. For humans, obedient children are good children, and disobedient children may be better children. Robots cannot accommodate or reconcile this logical contradiction and can only make a binary choice between 'obedient' or 'disobedient', 'obey' or 'disobey'. Robots cannot understand 'being both obedient and "Disobedient" or "both obedient and disobedient"—a machine can never 'break the rules.'" Add instructions to the video clip where the robot writes Interview Conclusion 3: "The candidate's performance in his primary job is often hindered by his amateur fashion design studies, which precisely reflects his passion for fashion design and indirectly explains his excellent performance on the skills test." Add instructions to the video clip where the robot writes Interview Conclusion 4: "This further reflects the candidate's passion for the fashion design career. His frequent changes in employment relationships are due to reasons and do not reflect his work stamina and dedication." Add instructions to the video clip where the robot writes Interview Conclusion 5: "The candidate did conceal information from his elders, but it was for a reason and does not indicate dishonesty." Add instructions to the video clip where the robot writes Interview Conclusion 6: "It cannot be concluded that the candidate cannot get along well with his superiors." Add instructions to the video clip where the robot writes Interview Conclusion 7: "The candidate concealed information that could have upset her and affected her health out of love for her mother. This does not mean he is again refusing to be candid. The robot clearly cannot understand metaphorical responses."In the video clip where the robot writes Interview Conclusion 8, the corresponding program instructions for the "stress test" deliberately fabricated by the interview robot according to the computer program and the machine intelligence program running status screen of the test generation process are added to the video screen, and the instruction is added: "The emotional stress test that the interview robot is rigidly designing and generating according to the computer program is inappropriate for this interview situation. Any normal person with flesh and blood who deeply loves their mother would tend to lose their mind under such an offense. After a few seconds, the candidate immediately realized their emotional impulse and apologized, expressing their love for the job and regret for their impulsive emotions. This is the best human performance in similar situations. The robot's observation and judgment are too extremely mechanical and rational and unrealistic." Finally, the person in charge added the instruction to the video clip where the robot writes the final conclusion: "Final conclusion. Taking all the above into consideration, the interview robot was unable to reconcile contradictions in this interview and tested and judged the candidate in an overly mechanical and rational manner. "A robot cannot understand the oftentimes difficult decisions humans must make under conflicting circumstances, nor the contradictory balance humans must embrace and bravely challenge their fate. Therefore, I have decided to reject all of their interview conclusions and admit this candidate." He signed his name.
[0058] This video, with its own time and location, corresponds to the various inconsistencies between "machine intelligence's cognition, reasoning, selection, planning, execution, performance, and rationality" and "human intelligence's purpose, intention, conception, common sense, principles, expectations, and sensibility." It illustrates the interplay of human cognition throughout the interview process, where rationality and sensibility are reconciled within the world-construction of human intelligence and the narrative of events. It also illustrates the purely rational world-construction of machine intelligence and the mechanical execution of its computer program within the narrative of events. The resulting chain of facts and logic reflects the complementary nature of human and machine intelligence beneath the surface of inconsistency: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "humanity," and must and can only be compensated by human intelligence. This video can be shared with superiors, colleagues, and others, and can be archived. The logical, factual, and emotional chains within it not only reflect the legitimacy and rationality of the person in charge's observation and judgment in identifying, understanding, and reconciling various dilemmas and conflicts, and rejecting the interview robot's conclusions, but also allows the audience to fully resonate with the legitimate human emotions within the event and the human nature that allows and encourages "breaking the rules," "pursuing ideals," and "love and struggle." The lessons learned from such typical machine intelligence errors or successful human intervention incidents can be efficiently accumulated and disseminated in the form of cases to help more people understand more deeply the complementary nature of human intelligence and machine intelligence, thereby establishing better awareness and ability of intervention judgment and reference to precedents.
[0059] A few days later, the interview robot manufacturer received this video. Robot technicians reviewed the robot's code and data, and from the log video of the code and data, they read the instructions executed by the machine intelligence, as well as the specific code and data of the execution process and results. They attached these videos to each video of the robot writing down the interview conclusion for technical and professional comparison. Based on the comments provided by the interview leader, the technicians continued to add their own technical instructions to the video, explaining that "the machine intelligence is simply executing its interview process normally based on the interview-related data. The computer program itself is correct in terms of process routine, but it clearly behaved inappropriately in this incident. In a world constructed by machine intelligence perception, candidates should possess completely rational skills and performance to maximize their competence in rational work. In a world constructed by human intelligence perception, candidates' rationality and sensibility are allowed to mix, and sensibility can even enhance the candidate's artistic expression in work such as fashion design to a certain extent. Machine intelligence cannot deeply understand the abstract concepts of human intelligence such as love, beauty, art, determination, ideals, and the meaning of life." The process and output of using this technical solution are negative training materials and logical training materials that can be used to train robots.
[0060] Case 6: Machine intelligence encounters random noise disturbances. At a railway and highway intersection, where no one was on duty and vehicles were strictly following traffic light signals, an intelligent vehicle misidentified the traffic light and ran a red light. This was particularly significant because it had correctly identified the traffic light at the same intersection dozens of times before. Even though the driver realized the error and intervened as quickly as possible, the violation had already occurred and was captured by a roadside traffic camera. The driver was left perplexed and could only blame their bad luck.
[0061] In this accident, there was a discrepancy between "the cognition, reasoning, selection, planning, execution, performance, trend, and will of machine intelligence" and "the purpose, intention, conception, common sense, principle, expectation, direction, and will of human intelligence." Afterwards, the traffic accident investigators retrieved the audio and video images reflecting the interaction between humans and machine intelligence during this incident (the driver's hand input operation process, the driver's voice input process, the driver's foot input operation process, the human-computer interaction screen image, the machine intelligence voice prompt, etc.) and the audio and video images reflecting the interaction between machine intelligence and the world (the vehicle's front view image, the overhead image of the traffic monitoring camera at the intersection where the accident occurred, the image of the smart car passing through this intersection many times before, the image of other smart cars of the same model passing through this intersection many times, etc.), and compared them with the images corresponding to the human-machine inconsistency in this accident. The images are manipulated and organized, and then combined with the time and place information indicated by each visual image, and compared and connected with each other to form a factual chain and a logical chain: at the time of running a red light, the traffic light was green in the world construction of machine intelligence, and the event narrative of machine intelligence was that the intelligent car went straight through the intersection normally, while the traffic light was red in the world construction of human intelligence, and the event narrative of human intelligence was that the intelligent car violated the law by running a red light and that even with the reasonable and fastest intervention, it was impossible to reverse the violation. At the time before running a red light, the traffic light was correctly identified as red in the world construction and event narrative of machine intelligence, and then mistakenly changed from red to green. The time when it changed to green was exactly the same as the time when a green neon light unrelated to the traffic light in the distance behind the red light in the world construction and event narrative of human intelligence turned on (noise).
[0062] In response to the needs of industry technical analysis, technical case dissemination and judicial procedures, the following key points and details are produced in the following process, sequence and manner to produce a complete chain of evidence audio-visual document with instruction information and time and place information ("→" means screen switching): Display the title "Analysis of the incident of a certain smart car running a red light at a certain time and place" with white text on a blue background for 2 seconds. → Display the path displayed on the human-computer interaction screen after the driver inputs the driving route instructions through manual operation or uses voice to input route instructions to the machine intelligence, indicating that the vehicle plans to go straight through the intersection where the incident occurred (pause for 2 seconds when the driving path is clearly displayed on the human-computer interaction screen and enlarge the human-computer interaction screen, bold the driving path lines and highlight them, and use voice narration to explain that the driving path passing the intersection where the incident occurred is planned by the machine intelligence, and write the voice narration text in the blank space in the video for the audience to listen and read, with an arrow next to the text pointing to the bold and highlighted driving path) → Connection section (play at normal speed, with voice narration and text prompts The next paragraph highlights the information about the driver's manual operation to start the automatic driving function, guiding the audience's attention to the driver's hands in advance) → The driver's hand input operation position visually shows that the driver has turned on the automatic driving function through hand operation (the specific action of the hand turning on the automatic driving function is partially magnified and displayed in close-up, played at 0.5 times slow speed, and explained with voice narration. The text of the voice narration is written in the blank space of the video for the audience to listen and read) → The transition section (played at normal speed, with voice narration and text prompts The next paragraph highlights the information about the change of the automatic driving function logo on the human-computer interaction screen, guiding the audience's attention to the human-computer interaction screen in advance The automatic driving state mark on the display of the human-computer interaction screen indicates that the vehicle has entered the automatic driving state (the automatic driving state mark on the human-computer interaction screen is partially enlarged and displayed in close-up, and the process of changing from gray to blue is repeated twice, with a pause of 1 second each in gray and blue. At the same time, a voice narration is used to explain that this color change means entering the automatic driving state, and the text of the voice narration is written in the blank space in the video) → The visual display of the driver's foot input operation position shows that the driver's foot has moved from the position of stepping on the pedal nervously to the relaxed and comfortable position next to it to start enjoying the ease and convenience brought by automatic driving (move the foot The moving action and the placement position are partially enlarged and displayed in close-up, and played at 0.5 times slow speed. The picture of the foot placed next to it is paused for 1 second, and a voice narration is used to explain it. The text of the voice narration is written in the blank space of the video for the audience to listen and read) → Connection section (the irrelevant driving process pictures in the middle are cut out and discarded, and the picture quickly switches to the eve of the red light error recognition incident, and the voice narration reminds the audience that the picture has arrived at the time and place where the incident occurred, and the text of the voice narration is written in the blank space of the video for the audience to listen and read) → A clearly visible red light begins to appear simultaneously in the vehicle's front view picture and the human-computer interaction screen picture (the two pictures are displayed side by side in a time-synchronized manner,The red lights in the physical world and the digital world are simultaneously paused and partially enlarged to indicate their correspondence. The red light display is paused for 2 seconds and an arrow flashes to guide the audience's focus. At the same time, a voice narration explains that the traffic light is red) → The vehicle's front view shows a bright green light spot in the distance that is not related to traffic (the green light spot appears and pauses for 2 seconds and an arrow flashes to guide the audience's focus. At the same time, a voice narration explains that the green light spot appears but is not related to the traffic light, and specifically points out the time when the green light spot appears. The narration explains that the driver later said that he did not notice the green light at all and could not remember it. The text of the voice narration is written in the blank space in the video). At the same time, the result of the machine intelligence's visual intelligent recognition of the traffic light on the human-computer interaction interface suddenly changes from red to green (the human The red light recognized by the machine intelligence is first displayed on the screen of the human-machine interaction interface. The red light is magnified and paused for 2 seconds. At the same time, a voice narration is used to explain that the recognized traffic light is a red light, and the text of the voice narration is written in a blank space in the video. Then the green light recognized by the machine intelligence is magnified and paused for 2 seconds. At the same time, a voice narration is used to explain the sudden recognition of the green light and the time, and it is especially eye-catchingly noted in a blank space in the video). The picture of the green light spot appearing and the picture of the machine intelligence recognizing the traffic light turning from red to green are played synchronously again, and their simultaneity is emphasized to verify the correlation (the close-up front view picture and the synchronized picture of the human-machine interaction interface video are replayed frame by frame, with a close-up of the appearance of the green light spot and the recognition of the green light displayed by the machine intelligence human-machine interaction interface, and the same time information is flashing on the screen). → The speed displayed on the human-machine interaction interface remains unchanged (the voice narration and text simultaneously indicate that the machine intelligence did not take any deceleration options but maintained forward momentum) → The vehicle's front view shows that the vehicle's front wheels have crossed the line (the crossing line image is enlarged and paused for 2 seconds, while the voice narration explains that the line has been crossed, and the text of the voice narration is written in the blank space in the image). The driver's hand operation position, foot operation position, and facial video are played synchronously to show that the driver realized that the intelligent car had no "brake" plan and execution and urgently intervened to perform manual emergency braking (the panicked movements of the hands and feet and the terrified expression on the face are framed on the screen, and the voice narration explains that the hands and feet have crossed the line). The driver then makes an emergency movement of the foot, and writes the text of the voice narration in the blank space of the video). At the same time, the change of the automatic driving status mark on the display of the human-computer interaction interface indicates that the automatic driving state has been exited due to user intervention (the process of the automatic driving exit operation mark on the human-computer interaction interface changing from blue to gray is partially enlarged and displayed in close-up and frame by frame, with voice narration and text description). Afterwards, the speed value on the display of the human-computer interaction interface drops rapidly, indicating that the vehicle has been drastically decelerated to zero due to human intervention (the speed display part on the human-computer interaction interface is partially enlarged and displayed in close-up, with voice narration and text description).At this time, a white flash appears in the vehicle's front view, indicating that the vehicle has run a red light and crossed the line and has been photographed by the traffic control camera (the time is noted in the picture, and a picture of the fine ticket sent by the traffic control department is added in the blank space of the picture. Arrows, text and voice are used to indicate that the time and place of the fine ticket are consistent with the time and place of the incident, and the information from all parties is clearly cross-verified). → The picture shows that the vehicle has decelerated sharply to a stop due to human intervention, and is dangerously stationary in the middle of the intersection and on the railway (the vehicle's front view and the overhead view of the intersection traffic monitoring camera are played at the same time, with voice narration and text to explain the danger of this state). → A few precedent videos are added for comparison (the vehicle's normal performance in responding to traffic lights at this intersection a few days ago, and other vehicles of the same model have normally responded to traffic lights at this intersection), and an instruction is given that "such green noise interference is rare, so this vehicle and other vehicles of the same model have not encountered any abnormalities before."
[0063] The above audio and video documents contain a chain of facts, logic, and emotions, forming a complete chain of evidence, which comprehensively proves that in this incident: the expectation of human intelligence is that "the smart car's machine intelligence should correctly identify any interference it encounters at the traffic light", but the performance of machine intelligence is that "the smart car identifies the red light as green and runs a red light violation"; the intention, conception, common sense and principle of human intelligence in creating, operating and using smart cars is that "safety is the highest principle of motor vehicle traffic. Smart cars speeding on the road must normally identify and respond to any red light they encounter", while machine intelligence in this case The cognition, choices, planning, and execution of a red light encounter are "severely compromised by the infinite random noise of the real world," while the average user's reasonable expectation of human intelligence is that "as long as I follow the manufacturer's instructions, maintain focused supervision, and intervene as quickly as possible upon recognizing an anomaly, I can ensure reasonable and safe driving." However, the actual performance of machine intelligence has led users to experience that "even if I follow the manufacturer's instructions, maintain focused supervision, and intervene as quickly as possible upon recognizing an anomaly, I cannot prevent the fait accompli of running a red light and the significant safety hazards it has already created." Smart driving manufacturers often create contractual clauses that force human users to bear all liability for personal injury and property damage in human-machine co-driving. This is a typical example of unreasonably increasing the burden on consumers and user responsibility while reducing or exempting the operator's own liability. This technical solution can generate comprehensive and direct evidence to support the accusation: proving that machine intelligence erred and that even the most reasonable intervention by human users was incapable of reversing the situation. The complementary nature of human and machine intelligence, beneath this surface of inconsistency, lies in the fact that machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "human nature," and that human intelligence must and can only compensate for this.
[0064] Subsequently, the visual sensor technicians of the intelligent driving manufacturer received this video data and conducted frame-by-frame analysis. They further determined that the green light spot disturbance in the picture came from random neon lights deep in the picture. Since the time when this noise disturbance appeared was highly consistent with the time when the machine intelligence recognition result of the traffic light turned from red to green on the human-computer interaction interface, there was obviously a correlation. Based on the technical principle that visual neural networks are susceptible to noise interference, the causal relationship of this anomaly is also highly suspected. Therefore, visual sensor technicians added instructions to the corresponding part of the video screen, indicating that the possible cause of the inconsistency is highly suspected to be noise disturbance caused by the green neon light. Based on this, the smart car visual neural network technicians correspondingly consulted the computer program log and conducted a more detailed manipulation and inspection of the machine intelligence working state process video, further confirming that the cause of the anomaly was indeed the visual neural network being disturbed by the green light spot. Based on the determination of responsibility for the specific technical links and the corresponding chain of evidence, the smart car manufacturer made an advance payment to the traffic fines incurred by the user, and then pursued compensation from the machine intelligence solution provider after confirming that it was responsible for the incident. The machine intelligence solution provider admitted responsibility and compensated because it knew that based on the above chain of evidence, it would definitely lose if the case was brought to court. The machine intelligence solution provider also determined internally that the accident was caused by noise disturbance in the visual neural network, and that the error did not lie in other links such as planning, decision-making, and execution. Therefore, other links do not bear responsibility or make technical modifications. If this technical solution is not used, the driver will not be able to explain the whole process of the incident, nor can he recall encountering visual noise at the time (no memory impression is left), and he will not be able to prove his innocence. He will have to accept his bad luck and pay for the loss out of his own pocket. The subsequent determination of responsibility and economic compensation cannot be determined to the specific technical links and the specific responsible legal entities.
[0065] The accumulation and dissemination of similar cases will dispel the technical superstition and prejudice that "smart cars are necessarily safer than human drivers." It will also help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better awareness and ability to proactively intervene and reference precedents. (For example, future drivers will generally be able to consciously intervene in advance to prevent machine intelligence errors when they notice the visual background near a traffic light is colorful and complex.) The process and output of this technical solution are negative training materials and logical training materials, which can be used to train the machine intelligence of smart cars.
[0066] Case 7. A case in which machine intelligence competes with human intelligence users for control. A smart car was driving normally in a tunnel. There were no obstacles in the lane ahead. However, its visual intelligence was suddenly disturbed by an unknown visual disturbance and it believed that there was an unknown dangerous obstacle ahead. The car then automatically braked. This phenomenon is commonly known in the industry as the "ghost brake" phenomenon. Since other drivers in the following cars could clearly see that there were no obstacles in front of the smart car and that it was driving at high speed, and they also saw that many vehicles in front of the smart car had passed at high speed and were moving forward, proving that there were definitely no obstacles on the road, they could not have anticipated that the smart car would suddenly brake to a stop "for no reason". Therefore, a large number of following cars kept a small distance and drove at a higher speed based on common sense and experience. The driver of the car was obviously aware of this and knew that emergency braking on the highway in such a situation could easily lead to rear-end collisions with the following cars. Therefore, when the smart car "ghost brakes", the driver urgently steps on the accelerator to try to increase the speed to avoid being rear-ended and to cancel the automatic driving state. However, at this time, the vehicle's machine intelligence uses a voice alarm "There is an unknown obstacle ahead, do not move forward". The smart car's central control screen displays a flashing red virtual obstacle animation in the lane ahead and displays "Automatic Emergency Braking AEB function activated, throttle command ignored". At this time, the driver notices that there is no response when he steps on the accelerator pedal. Due to the sharp drop in speed, the driver can only panic and repeatedly step on the accelerator and manually click "Cancel Automatic Emergency Braking AEB function" on the central control touch screen with his finger, so that his accelerator action is executed by the smart car, but it is too late. At this time, he has been rear-ended by the car behind.
[0067] In this case, there is a discrepancy between the "cognition, reasoning, selection, planning, execution, performance, movement, and will of machine intelligence" and the "purpose, intention, conception, common sense, principle, expectation, direction, and will of human intelligence." Accident investigators, users, insurance companies, and smart car technicians manipulated and organized the audio visuals reflecting the interaction between humans and machine intelligence during the incident (screens of program code and data during the creation of the smart car's machine intelligence, screens of machine intelligence operating status in the smart car's logs, screens of the driver's hand input operation process, screens of the driver's foot input operation process, screens of human-computer interaction, and voice prompts from machine intelligence) and the audio visuals reflecting the interaction between machine intelligence and the world (vehicle front-view screens, overhead screens from roadside traffic surveillance cameras). These instructions explained that in the machine intelligence's world construction, "there is a dangerous obstacle in the lane ahead of the vehicle," while in the human intelligence's world construction, "there is no obstacle ahead of the vehicle." The machine intelligence's event narrative was, "The design concept of this machine intelligence is that machines are more reliable than humans. To maximize safety, as long as the confidence coefficient for detecting obstacles ahead exceeds a certain preset value, the vehicle must be braked to ensure safety, ignoring the user's Unreliable operation." The human intelligence event narrative is, "The application concept of machine intelligence should prioritize humans over machines. Regardless of the situation, as long as the user manually operates, the machine intelligence should obey. As long as I step on the accelerator, the car must accelerate." Combined with the time and location information indicated by each video screen, the following chain of facts and logic is synthesized: The vehicle's forward-facing view and the overhead view from the roadside traffic monitoring camera clearly show that there are no obstacles in front of the vehicle (with some light and dark shadows). However, the machine intelligence operation status screen shows that the machine intelligence "very confidently" determines that there is an obstacle, triggering the automatic emergency braking (AEB) and immediately applying emergency braking. At this point, the driver, fearing a rear-end collision, immediately steps on the accelerator to try to take over and accelerate the vehicle. However, the program code screen when the smart car's machine intelligence was created shows that the machine's authority is higher than human authority. The driver's accelerator operation is ignored by the machine intelligence. The prompt message on the machine intelligence operation status screen also clearly indicates that the machine intelligence ignores human operation due to its safety policy design. The driver can only try to manually cancel the automatic driving on the central control screen, but unfortunately, he is rear-ended at this time. The complementary nature of human intelligence and machine intelligence beneath the above-mentioned inconsistent appearance is that machine intelligence cannot cope with "infinity", "chaos", "responsibility" and "human nature", and must and can only be compensated by human intelligence.
[0068] In view of the needs of technical analysis, case dissemination and judicial procedures, the above-mentioned video images and instruction information can be organized and compiled to generate targeted video documents, and the factual chain, logical chain and emotional chain of this incident can be conveyed to the judge, jury and the public. It is obvious that the smart car manufacturer should bear legal responsibility for this accident (when a vehicle on the highway brakes for no reason while driving at high speed and affects the normal driving of the following vehicle, resulting in a rear-end collision, the front vehicle is fully responsible), and the various aspects of technical and legal responsibilities will continue to be distinguished. Smart car manufacturers will clarify responsibilities with machine intelligence suppliers. The visual machine intelligence of this car obviously mistakenly identified the light and shadow of the road without obstacles as obstacles. Due to the car manufacturer and The following technical and legal clause is included in the agreements signed between machine intelligence solution providers during procurement negotiations: "If the perception confidence coefficient exceeds a pre-determined value, the machine intelligence is highly certain of the presence of an obstacle ahead. In this case, the automaker will configure the vehicle braking logic within its technical responsibility to ignore human user input." In this case, the smart car's machine intelligence operating status log shows that the obstacle perception confidence coefficient given by the machine intelligence during "ghost braking" exceeded the agreed-upon value in the signed agreement. Therefore, the technical and legal responsibility in this case lies with the machine intelligence solution provider. After the automaker assumes responsibility and compensates for the accident, it will use the above chain of evidence to pursue liability and compensation from the machine intelligence supplier. Without adopting this technical solution, it would be difficult to fully, accurately, and logically explain the events of this case and prove the unreasonable design concept of the smart car system in this case (its core design concept is that machines are more reliable than humans, and therefore machine authority prevails over human authority) and the legal fact that the user is exempt from liability even if the driver's fastest and most reasonable intervention could not have avoided the accident. It would be difficult to distinguish technical and legal responsibility in this accident, and ultimately the user's vehicle insurance company would have to make compensation without any guilt. Since "ghost braking" is a typical problem encountered from time to time by all smart car manufacturers based on artificial neural network machine intelligence, and there is no solution from the perspective of neural network technology, humans must and can only adopt this technical solution to prove the specific responsibility links of specific manufacturers, protect the legitimate rights and interests of legal entities and consumers in each link, and provide economic relief for public safety based on the principle of "whoever makes the mistake should compensate."Such cases will form judicial precedents. When smart car products from different manufacturers "ghost brake" and stop for no reason while driving at high speed on the highway, causing rear-end collisions, if the evidence is conclusive and similar to this case, the smart car manufacturers will bear full responsibility. The accumulation and dissemination of video documents in this case and repeated judicial precedents will enable smart car manufacturers to revise and adjust the authority of their machine intelligence systems so that human authority is higher than machine authority. Industry organizations and government management departments may issue corresponding industry standards. The whole society will therefore establish and share the technical concept of complementary cooperation between human intelligence and machine intelligence (and prohibit machine intelligence from excluding human intelligence). From then on, for the inevitable "ghost braking" accidents and the chaotic issues behind them (should the setting of machine intelligence system authority be higher than humans, or humans higher than machines), users, enterprises, the public, society, industry, government, and economic organizations will achieve stable and secure expectations after making a clear common choice. The lessons learned from such typical machine intelligence errors can be efficiently accumulated and disseminated in the form of cases to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby establishing a better awareness and ability to refer to precedents. This will also break the technical superstition and prejudice that "smart cars must be safer than human driving."
[0069] Case 8. Analysis of liability for conflicting machine intelligence behaviors. A smart car using "end-to-end" machine intelligence experienced a "split personality" at a traffic light and made a driving error. "End-to-end" means that the machine intelligence receives only fully digitized videos of real-world driving scenarios as training material, and directly outputs real-world driving actions. It bypasses traditional perception, prediction, planning, decision-making, and control stages, and lacks specific qualitative and variable programming for concepts like "traffic lights" and "lane lines." This smart car was waiting in line at a red light at an intersection, waiting to go straight through. The red light turned green, and the car in front crossed the line and went straight ahead. This car followed and then slowed down and stopped. When the straight traffic light turned green, it stopped within the line. After that, the straight traffic light turned from green to red, and then the left turn traffic light turned from red to green. This car, which was in the straight lane and was governed by the straight traffic light (red light), suddenly started to accelerate. Fortunately, it was stopped by the terrified driver's emergency brakes. However, the vehicle had crossed the line and was photographed by the traffic violation camera. Afterwards, the driver received a ticket and was confused. He could not remember what exactly happened before and after the accident. He only remembered that the vehicle's automatic driving was normal. He was caught off guard and suddenly realized that the car had run a red light and intervened to brake urgently. However, he did not know which link went wrong and who should be responsible for the accident.
[0070] During the incident, the investigators manipulated and organized the video images corresponding to the above-mentioned human-machine inconsistencies in the vehicle's forward-looking video, roadside traffic monitoring video, driver's hand operation position video, driver's foot operation position video, audio video of the driver's interaction with machine intelligence, central control screen video, and machine intelligence operation status record, and gave instructions. Combined with the time and place information in the video, they compared and connected each other to produce the following fact chain and logic chain: the driver gave the machine intelligence the driving destination instruction information through manual operation, the machine intelligence planned path displayed on the central control screen video indicated that it would go straight through the intersection where the incident occurred, and the driver's hand operation position video showed that the driver turned on the automatic driving function through manual operation. The automatic driving status mark on the central control screen turns from gray to blue, and the machine intelligent operation status recording video both indicate that the vehicle has entered the automatic driving state. The driver's foot operation position video shows that the driver's feet have moved from the position of nervously stepping on the pedals to a relaxed and comfortable position to start enjoying the ease and convenience brought by automatic driving. At the beginning, the vehicle's forward vision and the road monitoring vision simultaneously cross-confirm that the vehicle is in the straight lane, the vehicle and the vehicle in front are stopped together to wait for the red light, the straight light and the left turn light are both red, then the straight red light turns green, the left turn light is still red, the vehicle in front starts to move, the vehicle also starts to move, the vehicle in front goes straight across the line, but the vehicle "strangely" slows down and stops (follows the vehicle in front to pass The central control screen also cross-confirms that the speed of the vehicle briefly increased and then decreased to zero. At this time, the machine intelligence operation log video and the corresponding training material screen show that the smart car perceives that all lanes of the road ahead opposite the intersection are full of vehicles and seriously congested (the machine intelligence of this smart car has been trained to learn a large number of driving behavior videos of "very gentlemanly" car owners. Therefore, at this time, it imitates and "is also very gentlemanly" and does not follow the car in front when the light is green, going straight through the intersection and entering the opposite road, exacerbating congestion or occupying space at the intersection). The machine intelligence that makes this choice, plan and execute is specially emphasized in the screen as "Robot Cell 1". The through light turns from green to red, and a few seconds later, the left turn light turns from red to green. At this time, the vehicle's forward vision and the road monitoring vision simultaneously cross-confirm that the vehicle is in the through lane, the through light is red, the left turn light turns from red to green, the vehicle starts and crosses the line in an abnormal and illegal manner, then stops manually, and the traffic camera flashes and takes pictures. The machine intelligence operation status log video and its corresponding training material images show that the machine intelligence is now chaotically imitating the following human behavior: "The left turn light is red, the through light turns from red to green, the car in front crosses the line and goes straight, and the vehicle starts and follows and stops within the line before reaching the line. This means that the vehicle is waiting for the left turn light to turn green. That is, when encountering a green light, it does not cross the line, which means it is waiting for the other light to turn green.” (This is a common human driving behavior, and the machine intelligence is “neurologically” connected to this behavior). The screen specifically emphasizes that the machine intelligence that makes this choice, plan and executes is called “Robot Grid 2”, and then the two video clips of “Robot Grid 1” and “Robot Grid 2” are played side by side on the same screen with instructions: “In this case, although the same machine intelligence computer is performing intelligent driving, ‘Robot Grid 1’ and ‘Robot Grid 2’ are divided and do not communicate with each other. Different ‘Robot Grids’ emerge on the spot to imitate human driving behavior, but this imitation without memory communication and discontinuous logic manifests itself as a clumsy failure of mechanical confusion, resulting in a red light violation. Specifically, within the same world construction of human intelligence, the human intelligence's narrative depicts the intelligent car not crossing the line when it should, then crossing when it shouldn't. However, within the two separate world constructions of machine intelligence, the machine intelligence's narrative depicts the car 'normally and gentlemanly' passing the line in the first world construction, and 'normally starting to cross the line when the other light turned from red to green' in the second world construction. The central control screen then displays a video showing the vehicle speed increasing and then abruptly dropping to zero due to the driver's emergency intervention. The video from the driver's hands shows the driver's panicked panic as they urgently grasped the steering wheel, while the video from the driver's feet reflects the driver's quickest and most reasonable response to manually intervene and brake. Each of these video images incorporates time and location information to interconnect them.
[0071] In this incident, the purpose, intention, conception, common sense, principle, and expectation of human intelligence are "the intelligent car should correctly identify when encountering a red light or a green light and take specific correct actions according to the specific driving conditions and goals", but the cognition, reasoning, selection, planning, execution, and performance of machine intelligence are "the intelligent car correctly identifies the traffic light but logically runs a red light and violates traffic rules". The intention, conception, common sense, principle, and will of human intelligence in creating, operating, and using intelligent cars are "the robot personality of the machine intelligent body driving the vehicle must be the same personality, just like the driver in human driving behavior is the same person, and its previous and subsequent cognitions are different. , memory, logic, and behavior should be identical, coherent, stable, and consistent." However, in this incident, machine intelligence's cognition, selection, planning, execution, and will experienced "split personality," "logical disorder," or "neural network disorder." The average user's reasonable expectation of human intelligence is that "as long as I follow the manufacturer's instructions, maintain focused supervision, and intervene and take over as quickly as possible when I realize an anomaly, I can ensure reasonable and safe driving." However, the actual performance of machine intelligence has made users experience that "even if I follow the manufacturer's instructions, maintain focused supervision, and intervene and take over as quickly as possible when I realize an anomaly, I cannot prevent the fait accompli of running a red light and the safety hazards that have already resulted." Intelligent driving manufacturers often formulate contract clauses that force human users to bear all liability related to personal injury and property damage in human-machine co-driving. This is a typical example of unreasonably increasing the burden on consumers and user responsibilities while reducing or exempting the operator's own liability. The use of this technical solution can generate comprehensive and complete direct evidence for such accusations: proving that machine intelligence has erred and that human users' intervention with the most reasonable response is still powerless to reverse the situation. Beneath this apparent human-machine inconsistency, the complementary nature of human and machine intelligence lies in the fact that machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "human nature." The skills and behaviors that intelligent vehicles' machine intelligence learns to imitate after extensive training may contradict themselves when dealing with the infinite possibilities of the physical world. Human expectations of machine intelligence's performance, namely "taking specific and correct actions based on specific driving conditions and goals," are actually the unique ability of human intelligence to draw inferences from one instance to another. Machine intelligence programs, however, are drowned in a sea of code and data. Overly complex training materials lead to increasingly complex sensor signal patterns, algorithm mappings, program logic, behavioral patterns, and consciousness simulations, resulting in increasingly frequent confusion and conflicting approaches. As the saying goes, "pressing one problem only makes another one pop up." Machine intelligence's skills and behaviors are bound to occasionally lose focus or contradict each other. Furthermore, machine intelligence is incapable of self-reflection and lacks the same unified, coherent, and stable personality and will as humans. Therefore, it is unable to recognize its own logical inconsistencies. These inherent flaws must and can only be remedied by human intelligence.Playing the above-mentioned video document of the technical chain of evidence in court, along with accompanying explanations, will help the driver understand the entire incident and prove his innocence. It will also help the public, jury, and judge understand the full course of events and reach a consensus that liability should be assigned to the intelligent driving supplier or smart car manufacturer. The lessons learned from these typical machine intelligence failures can be effectively accumulated and disseminated through case studies to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better understanding and ability to reference precedents. This will also dispel the technological superstition and prejudice that intelligent cars are inherently safer than human drivers.
[0072] Case 9: Machine intelligence lacks common sense and real-life concepts. A smart car, while autonomously driving on a highway, encountered a basketball-sized, soft plastic bag, blown by the wind and floating in mid-air within its lane, at the same height as the windshield. The car's cameras, radar, and other sensors accurately detected the object within the vehicle's path. Since the plastic bag posed little threat to vehicle safety, the car could have simply passed it without reacting, and the bag would have been pushed aside by the airflow. However, the car lacked common sense and real-life concepts regarding "texture." To ensure absolute safety, its program assumed all obstacles were dangerous, prompting it to swerve to avoid the bag. However, since the car lacked common sense and real-life concepts regarding "friction," such as "slippery roads," the swerve caused the car to lose control, skidding sideways and striking a guardrail. Clearly, there was a discrepancy between the "cognition, reasoning, selection, planning, execution, and performance" of machine intelligence and the "purpose, intention, conception, common sense, principles, and expectations" of human intelligence. Investigators subsequently manipulated and organized the visual images corresponding to this discrepancy, cross-referencing temporal, spatial, visual, and instructional information to form a cross-corroborating logical chain, factual chain, and chain of evidence. This chain of facts and logic demonstrates the legal liability of the smart car manufacturer and helps to divide and pursue specific technical and legal responsibilities between the smart car manufacturer and its machine intelligence supplier. This chain of facts and logic reflects the complementary nature of human and machine intelligence behind this apparent discrepancy: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "human nature." The possible scenarios in the real physical world are infinitely variable, and the potential obstacles on the road are endless. It is impossible to equip smart cars with an infinite number of sensors and computer program logic. Machine intelligence cannot answer the chaotic question of "what is an obstacle?" Nor can it comprehend the infinite common sense of life like humans do, resulting in mechanical and rigid program responses. These shortcomings, as well as the technical, economic, and legal responsibilities for the accident, must and can only be remedied by human intelligence. The lessons learned from such typical machine intelligence errors can be efficiently accumulated and disseminated in the form of cases to help more people gain a deeper understanding of the complementary nature of human intelligence and machine intelligence, thereby establishing a better awareness and ability to refer to precedents.
[0073] Case 10. Calculation of economic losses caused by machine intelligence. During normal intelligent driving on a highway, an intelligent vehicle planned to exit at a certain ramp. Approaching the ramp, the vehicle in the middle lane automatically and correctly maneuvered to the rightmost lane (and then right to exit the ramp). However, halfway through the lane change, the vehicle inexplicably and for no apparent reason reversed back to its original lane and continued driving. The line between the middle lane and the right lane then changed from a dotted line to a solid line. The vehicle continued straight ahead in the middle lane and missed the exit ramp. Due to the suddenness of the incident, the driver had no time to react and intervene. Even if they took over driving after noticing the intelligent vehicle's erroneous return to its lane, they were unable to cross the solid line and inevitably missed the exit. During this time, when the driver realized that the smart car was abnormal, he observed the right lane with his naked eyes and did not see any vehicles or other obstacles blocking traffic. In the end, he had to detour dozens of kilometers and spend several more hours to reach his destination. The user was puzzled by this abnormal smart driving because usually this smart car was able to change lanes and exit ramps normally in other similar driving scenarios. Afterwards, the user filed an infringement lawsuit and economic claims against the smart driving manufacturer.
[0074] The investigators manipulated, organized, and explained the multi-party visual images corresponding to the human-machine inconsistency in the incident. The time information, spatial information, image information, and instruction information were compared and linked to each other to form a factual chain and a logical chain: In the detailed process of this inconsistency, the comparison between the vehicle's forward camera image and the traffic monitoring image showed that when the smart car made and initiated the lane change action normally, common and harmless road repair and filling marks appeared on the road surface in front of the rightmost lane. The vehicle human-machine interaction screen showed that in the machine intelligence's world construction, "an unknown obstacle was found in the right lane" (its visual effect looked a bit like a large area of thick branches fallen on the road in the investigators' human intelligence world construction). The machine intelligence's description of the event was "in order to avoid the unknown obstacle, the vehicle urgently returned to the lane halfway through the lane change to ensure safety." This verified that the smart car's action of returning to the original lane was justified, and the driver only noticed that there was no car in the right lane in a quick glance at the time of the incident. The driver of the vehicle, who was in the empty lane, did not notice the above-mentioned marks on the road surface on the right. This is because similar road marks (dirt, cracks, repairs, etc.) are too common for drivers and never affect their normal driving. Ordinary users would never imagine that such marks would be "identified as obstacles" by the smart car. Therefore, in the world constructed by human intelligence, "there is no obstacle in the right lane." The human intelligence narrative of the event is that "the smart car canceled the lane change for no reason, missed the ramp, and took a large detour, causing losses." In this incident, the window for changing lanes to the right and exiting the ramp was very short. Since the line between the middle lane and the right lane immediately changed from dotted to solid, according to traffic regulations, changing lanes across the solid line is prohibited. Therefore, the smart car's machine intelligence complied with the traffic regulations and did not change lanes. When the driver realized the machine intelligence error at a normal and reasonable speed, even if he immediately took over, he could not save the situation by changing lanes illegally. The corresponding in-car scene video shows that the driver immediately took over but was still unable to change lanes and was forced to miss the ramp. Next, a video clip of the driver's previous use of the smart car in a similar highway exit ramp scenario is inserted, specifically noting that no road repairs or caulking were encountered in these instances. Next, the entire process of the vehicle's significant detour to the destination is played at 32x speed. The vehicle's location and detour trajectory are displayed simultaneously on an electronic map in the lower right corner of the video. The time spent on the detour, the additional mileage traveled, and the energy consumed are simultaneously calculated, recorded, and indicated, and used as a basis for calculating economic losses.
[0075] Using this visual display, which carries instructional information, users and investigators can initiate infringement lawsuits and financial claims against smart car manufacturers, claiming reasonable and well-founded compensation based on detour time, mileage, and energy consumption, and ultimately winning their cases in court. Smart car manufacturers often create contractual clauses that force human users to bear all liability for personal injury and property damage in human-machine co-driving. This practice unreasonably increases the burden on consumers and user responsibility while mitigating or exempting the operator's own liability. This technical solution can generate comprehensive and direct evidence to support such allegations: proving that the machine intelligence erred, demonstrating that even the most reasonable intervention by human users was incapable of reversing the situation, and providing a basis for calculating the amount of compensation claimed. Obviously, in this case, there is a discrepancy between "machine intelligence's cognition, reasoning, selection, planning, execution, performance, will, and storage" and "human intelligence's purpose, intention, conception, common sense, principles, expectations, will, and memory." Unknown obstacles appear in the world construction of machine intelligence, while there are no obstacles in the world construction of human intelligence. The event narrative of machine intelligence is that everything is normal and the program is executed correctly, while the event narrative of human intelligence is that a major detour occurred that should not have happened. Under the above-mentioned inconsistent appearance, the complementary nature of human intelligence and machine intelligence is that machine intelligence cannot cope with "infinity", "chaos", "responsibility" and "human nature", and must and can only rely on human intelligence to make up for it. Legal procedures follow the principle of "he who asserts must provide proof." When making a claim, arbitrary and subjective claims will be challenged by defense attorneys, who will demand proof of the amount. Failure to provide proof will be detrimental to the claim, ultimately forcing bargaining or discretion. However, this technical solution provides a clear basis for calculating the amount, significantly increasing the probability of a successful, reasonable claim. By employing this technical solution in similar claims, we can ensure a consistent, well-reasoned, and coherent application of machine intelligence in economic and legal matters, reassuring the public and all parties involved in the industry. The lessons learned from these typical machine intelligence failures can be effectively disseminated through case studies to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby fostering a greater awareness and ability to reference precedent.
[0076] Case 11. A case where high-complexity program logic encounters complex scenarios and conflicts. A smart car planned to turn right at a certain intersection but failed to execute it. Investigators manipulated and organized the video images corresponding to the inconsistency in the vehicle's external video and the machine intelligence operation status log file during the event, and provided instructions. The time information, space information, image information, and instruction information were compared and linked to each other to form a factual chain and a logical chain of the event as follows: The urban road has three lanes, namely the leftmost lane, the middle lane, and the rightmost lane. When the event started, the smart car was in the middle lane with a high speed. It was planned to turn right at a certain intersection after a few hundred meters, so the smart car started to try to change lanes to the right early. At this time, the smart car sensed that there was a bus in the rightmost lane in front of the right and slowly started from the platform. The cyclist behind the bus entered the middle lane. The car accelerated in the middle lane and tried to overtake the slow bus from the left. At this time, the "efficiency" and "comfort" modules in the intelligent system of the smart car prevailed in the decision-making process. This was because the "efficiency" module calculated that it would be faster and more efficient to continue to maintain a high speed, overtake the bus and cyclist at high speed, and then change lanes to the right. After all, there were still several hundred meters to the right turn intersection, and there was plenty of distance and time. At the same time, the "comfort" module calculated that it would be more comfortable for the passengers to continue to maintain a high speed and not decelerate drastically because of the bus and bicycle. Therefore, the smart car chose to change lanes from the middle lane to the leftmost lane and move forward at high speed to overtake the bus and bicycle. Afterwards, when the smart car When it was almost overtaking the bus and the bicycle, it tried to change lanes to the right, but the bus had already accelerated and changed lanes to the middle lane to avoid a white vehicle that had temporarily pulled over in front of it. When the left front wheel of the bus crossed the lane line between the right lane and the middle lane and the front of the bus entered the middle lane, the "safety" module of the smart car prevailed. After perception and calculation, it believed that changing lanes to the right at this time might cause a collision with the bus, so it chose to continue to stay in the leftmost lane and drive at high speed. At this time, the "efficiency" and "comfort" modules prevailed again, planning to continue driving at high speed and speed up appropriately until it completely overtook the bus before changing lanes to the right. However, the white vehicle that had temporarily pulled over in front of the bus also started and gradually Gradually entering the middle lane, when the smart car completely left the bus behind and tried to change lanes to the right to enter the middle lane, it found that the white vehicle in the middle lane was behind it on the right, not far away from it and at a high speed. At this time, the "safety" module of the smart car prevailed again, believing that there was not enough space on the right to ensure a safe lane change. To ensure safety, the smart car "dared not" change lanes to the right. By this time, the originally spare distance and time were almost exhausted, and there were only dozens of meters left to the planned right turn intersection. At this time, in order to ensure the correct priority logic of the driving path, the "efficiency" module had to downgrade the "comfort" module, so it chose to decelerate sharply in the leftmost lane in an attempt to wait for all the vehicles behind to pass before changing lanes and turning right. However, until it almost entered the intersection,Still unable to find a safe space to change lanes right, the car's speed had already dropped to an extremely low level. The "safety" module once again took over, overwhelmed the other logic modules, believing that slowing down to an extremely low speed in the fast lane at a high-speed intersection with a green light would increase the risk of rear-end collision. Therefore, the car's speed was increased rapidly. Now that there was no room for a right turn, the "efficiency" module gave up, and the machine intelligence replanned a significantly detour. The original right turn at the intersection completely failed, forcing the driver to proceed straight through the intersection. However, the main entrance to the highway was ahead, and reversing or turning around was impossible once inside. The driver was forced to take the highway and take a significant detour, resulting in financial losses.
[0077] Similar complex traffic scenarios emerge in an endlessly changing and diverse world in real-world traffic environments. Scenario logic designers cannot anticipate all real-world application scenarios. For industrialized products, because system complexity is the product of engineering external forces rather than a life-giving internal growth, it is destined to remain complex to the point where logic code, sensor priorities, hardware and software modules, and application scenarios hinder, contradict, and even conflict with each other. Ultimately, this complexity manifests itself primarily in system errors when responding to specific situations. In this case, a human driver would instantly traverse the infinite possibilities and select the associations of particularly significant factors (no significant detours would be acceptable). They might be "timid" and slow down early (sacrificing efficiency) to ensure a safe right turn, or "bold" and force a right turn when there is almost no room left (sacrificing safety). When determined to turn right at an intersection and in the left lane, humans will resolutely and proactively choose to sacrifice some of their own safety and the smoothness of public transportation, forcibly slowing down in their lane and forcibly attempting to change lanes to the right. After all, the lane lines are all dotted at this time, and lane changes are still allowed according to the rules. Their actions and motivations reflect the driver's strong determination to turn right at this intersection. This is a common operation in human driving practice. Although it is unpleasant and risky for both drivers and traffic participants, few drivers have not experienced such difficult situations. Even if they are frightened or stressed, human drivers will take the initiative to force a right turn. Although other traffic participants may curse and criticize, they will also understand with their human empathy that "if the vehicle does not force a right turn, it will be forced to enter the highway." Human intelligence can also form a tacit understanding and understanding of this behavior based on human nature. In dilemmas like this one, human intelligence can seek and achieve a dynamic balance between two conflicting variables. This involves rapidly assessing the benefits of fully adhering to the rules and ensuring complete safety against the costs of being forced to take a detour. While this assessment may not be accurate and the choice is highly dependent on individual character, it ultimately stems from the user's free will. However, in the vast majority of cases, machine intelligence manufacturers, either to avoid legal liability or adhere to business ethics, will choose to configure machine intelligence to strictly adhere to the rules and ensure complete safety. Even if a machine intelligence program is designed to achieve a balance between these two conflicting variables, its approach is often inconsistent with the program's creator's and likely skewed from the user's own individual preferences. This can lead users to feel their free will is challenged, leading them to become dissatisfied with the machine intelligence's performance and blame it for the losses it causes. Obviously, in this case, the "cognition, reasoning, selection, planning, execution, and performance of machine intelligence" are inconsistent with the "purpose, intention, conception, common sense, principles, and expectations of human intelligence." The machine intelligence system is too complex, and the logical modules are mutually constrained. The three logical modules of "efficiency," "comfort," and "safety" compete for system priority, resulting in the failure to achieve the reasonable and satisfactory results that humans can achieve by driving themselves.In the narrative of machine intelligence, it achieved adaptable intelligent driving performance amidst competition among rational program modules focused on "safety," "efficiency," and "comfort," demonstrating normal and correct program execution. However, in the narrative of human intelligence, the intelligent vehicle's driving behavior resulted in unexpected and unreasonable planned route execution failures, which should not have occurred, resulting in significant detours and economic losses. Often, the "necessary" choices faced by human intelligence in dilemmas are not passive but active. Human consciousness defines and constructs itself through active choices. When faced with dilemmas, human intelligence will make appropriate compromises, unifying, accommodating, and reconciling them. In driving, human intelligence often adheres to rules, while occasionally transcending (disobeying) them. This is a unique, subtle yet natural, and universally manifested capability of human intelligence, one that machine intelligence cannot achieve. This apparent inconsistency reflects the complementary nature of human and machine intelligence: machine intelligence cannot cope with "infinity," "chaos," "responsibility," and "human nature," and must and can only be complemented by human intelligence. The lessons learned from such typical machine intelligence errors can be efficiently accumulated and disseminated in the form of cases to help more people understand more deeply the complementary nature of human intelligence and machine intelligence, thereby establishing a better awareness and ability to predict intervention and refer to precedents.
[0078] Case 12: An example of machine intelligence failing to meet basic human psychological and emotional needs. During the return flight of a large drone after completing a sea-skimming mission, the monitoring operator at the drone base continuously monitored the drone's real-time flight status on a large screen. The operator could see the drone's altitude, coordinates, heading, speed, autonomously planned flight path, and other machine intelligence operational status information uploaded via the satellite data link. The flight path showed the drone's roughly planned autonomous flight path passing through a cross-strait bridge. The operator could also see the forward view captured by the drone's nose camera and the corresponding planned flight path, with the planned flight path displayed virtually as a white line in the forward view. Before the bridge entered the drone's field of view, the monitoring operator only saw from the rough track screen that the drone intended to pass through the geographical location of the bridge. After the bridge clearly entered the drone's field of view, the monitoring operator saw from the screen that the white lines showed that the drone's specific intention was to continue to maintain a low-altitude sea-skimming flight through the bridge tunnel, and the aircraft was now flying towards the bridge tunnel clearly presented in the field of view. This was inconsistent with the monitoring operator's common sense that medium and large aircraft generally fly over bridges. He immediately controlled the drone's machine intelligent planning of the detailed calculation process of the long-term rough route, and displayed the planned flight trajectory on the two-dimensional screen of the electronic map. The screen showed that the latitude and longitude of the planned route were correct and reasonable, but the planned flight altitude was too low. The vertical UAV machine intelligence plans a detailed short-term detailed route through the detailed calculation process screen, and displays the planned flight trajectory in the three-dimensional screen in the front field of view. The screen clearly shows that the UAV plans to fly at low altitude through the bridge hole after measurement and calculation. The monitoring operator gives instructions based on the above pictures, "After verification by all parties, the longitude and latitude of the UAV's planned path are correct and reasonable, but the altitude is too low. After measurement and calculation, the machine intelligence believes that the bridge hole is large enough to cross safely. Its flight planning strategy takes into account the large amount of energy required to climb to a certain height, so it chooses to fly at low altitude to cross the bridge hole." At this time, the monitoring operator makes a prompt decision and issues a rejection order to the UAV's planning decision, requiring it to raise its altitude to far above the height of the bridge and to fly over the bridge instead of passing through the bridge hole.The monitoring operator manipulated and organized the various visual images corresponding to the inconsistency between the drone's performance and human expectations and gave himself sufficient instructions. The time information, spatial information, image information, and instruction information were cross-referenced and linked to form a cross-verified factual chain and logical chain that enabled him to correctly understand the event: in the world perceived and constructed by the drone's machine intelligence, the bridge hole size was measured and calculated to be suitable for flying through, and the flight trajectory was completely correct and reasonable. Therefore, the machine intelligence's event narrative was "The bridge is on the correct planned path, and the bridge hole size is safe enough to fly while ensuring a safety margin. Such a plan does not waste a lot of energy to climb altitude, and is an economical and reasonable flight planning path that tries to maintain the original flight state. The plan and strategy for crossing the bridge hole are correct." However, the monitoring operator immediately realized that such a plan and strategy were not perfect and appropriate. The potential consequence of the monitoring operator's world construction, event narrative, and inference is that "even if it is technically safe to pass through the bridge hole, this behavior will give the drone This creates a sense of insecurity for aircraft owners, commercial customers, bridge owners, and the general public. If our plane roared through the bridge arch, with so many vehicles and pedestrians on the bridge watching, it would look like a plane about to crash into the bridge, like a terrorist attack. Regardless of how safe it actually is, the public will perceive it as an "unsafe flight." This incident will likely make headlines tomorrow, and the police will be called. Such an incident clearly violates basic human psychological and emotional needs, frightening everyone and causing panic, potentially leading to far-reaching adverse consequences that far outweigh the small amount of energy consumed by climbing to higher altitudes. In this case, the "planning, execution, performance, and rationality of machine intelligence" are inconsistent with the "common sense, principles, expectations, and sensibilities of human intelligence." Machine intelligence can only understand the world through a technical and rational approach, and while it is not technically wrong, humans clearly need to consider universal human emotions when understanding the world. Universal human expectations and needs include not only "safety" but also "a sense of security." Only human intelligence with human nature can understand this, and human intelligence with human nature can universally understand this. For similar cases in real life, human intelligence "emotional decision-making surpasses rational decision-making" and "emotional decision-making and rational decision-making are combined into one", but it cannot be quantified or even qualified in advance. It is a "fuzzy problem" and is highly related to the universal instinct of human nature to empathize with each other and understand each other. It can only be decided by human intelligence through on-the-spot analysis and judgment. Therefore, the above chain of facts and logical chain reflects the complementary nature of human intelligence and machine intelligence under the appearance of inconsistency between humans and machines. Machine intelligence cannot cope with "infinity", "chaos", "responsibility" and "human nature", and must and can only be complemented by human intelligence.The monitoring operator intervened promptly during the incident. Afterward, investigators provided visual documentation of the aforementioned chain of facts and logic, combined with visual footage of the monitoring operator's intervention, to demonstrate the rationality and legitimacy of the human intelligence's impromptu rejection and dismissal of the machine intelligence's planned decision. The incident was then reported, disseminated, and preserved. The lessons learned from these typical machine intelligence failures or successful human interventions can be effectively accumulated and disseminated through case studies to help more people gain a deeper understanding of the complementary nature of human and machine intelligence, thereby developing a better understanding and ability to make informed interventions and reference precedents. The process and outputs of this technical solution are negative and logical training materials that can be used to train the machine intelligence of intelligent drones.
[0079] When reading some of the examples in this patent text, readers should be able to personally experience that understanding both facts and logic simultaneously through long paragraphs of text is time-consuming and laborious. It requires arduous reading comprehension, logical analysis, and plot imagination to fully and accurately understand. Sometimes, this can lead to misunderstandings or misconceptions. Readers with poor reading comprehension skills may need to repeatedly ask questions and communicate to accurately understand. However, the organized audio-visual files produced by this technical solution are much easier to understand. Just a few minutes of video explanation is enough to fully understand. It is immersive, understandable at a glance, with little ambiguity, and no need for translation or imagination. It can effectively understand the complementary nature of human and machine intelligence behind the inconsistency between humans and machines. Furthermore, I believe readers will also feel this way. Although some of the examples in this patent text are explained relatively briefly and the output videos are not clearly described, readers will have already fully learned how to use this patent technology to select and organize the key elements of the various videos to generate a reasonable video output. This is precisely the ability of readers to draw inferences from one example to another.
[0080] The application of this patent will continue to collect examples of complementary cooperation between human and machine intelligence. Many of these examples are inherently non-obvious—they cannot be derived through analysis, reasoning, or limited experimentation. Only through encountering them can one understand that "there are actually such cases." For ordinary machine intelligence technicians, some of the examples and their technical essences in this article are unexpected without being informed. Similar unexpected examples are endless in machine intelligence applications. Only by applying this patent solution can we address them one by one, continuously accumulate them, and enhance our understanding of the complementary nature of human and machine intelligence, making human and machine intelligence more adept at complementary cooperation, and helping the entire society establish and strengthen the technical concept of complementary cooperation between human and machine intelligence. Humans must and can only follow the natural law of complementarity between human and machine intelligence to achieve perfection in the process and results of machine intelligence applications, and to enable the social application of machine intelligence to enter an overall expected stable and perfect steady state.
[0081] In various potential cases in the future, the inconsistency between human intelligence and machine intelligence may also manifest itself as machine intelligence being superior to human intelligence, but human intelligence being unable to understand or even misunderstanding it at the time or in the short term. Humans can also use this technical solution to conduct continuous review or later review, gradually unraveling doubts and clarifying the truth. More importantly, it enables machine intelligence to learn and human intelligence to understand why the superiority of machine intelligence is manifested in various aspects such as behavior, communication, and planning, but is not effectively recognized, understood, agreed, and responded to by human intelligence at the time or in the short term, thereby enhancing humans' further understanding of human intelligence and machine intelligence. Such cases and processes will also reflect and deepen the complementarity between human intelligence and machine intelligence, and help improve various aspects of machine intelligence such as behavior, communication, and planning.
Claims
1. The method by which human intelligence compensates for the inherent shortcomings of machine intelligence is characterized by: In the visual images reflecting the interaction between humans and machine intelligence and the visual images reflecting the interaction between machine intelligence and the world, the visual images corresponding to the various aspects of the inconsistency between "machine intelligence's cognition, reasoning, selection, planning, execution, performance, movement, will, storage, and rationality" and "human intelligence's purpose, intention, conception, common sense, principle, expectation, guidance, will, memory, and sensibility" are manipulated, organized, and instructed. These visual images are combined to indicate the worlds perceived and constructed by machine intelligence and human intelligence in the above-mentioned inconsistencies, the respective event narratives of machine intelligence and human intelligence in their respective world constructions, and the time and space information in their respective event narratives. The above-mentioned information is compared and connected with each other and unified and integrated into a factual chain and a logical chain that reflect the complementary nature of human intelligence and machine intelligence under the appearance of the above-mentioned inconsistency. The two chains are formed simultaneously.
2. The method according to claim 1, wherein: The video image is an audio video with audio information.
3. The method according to claim 1, wherein: When manipulating, organizing and instructing the visual images, the visual images reflecting the above-mentioned inconsistencies are selected, decomposed, sorted and arranged, and the time or space reflecting the above-mentioned inconsistencies are shaped, reorganized and connected, imitating or guiding the attention, perceptual habits and cognitive mechanisms of the human audience, emphasizing the key points, highlighting the details, and generating dynamic video documents, or obtaining static images in the process, combining with the instruction information, to generate static graphic documents.
4. The method according to claim 1, wherein: The unified synthesis becomes a verification chain reflecting the above inconsistencies.
5. The method according to claim 1, wherein: The unified synthesis becomes an emotional chain reflecting the above inconsistencies.
6. The method according to claim 1, wherein: The instructions explain the inner nature or potential future consequences behind the above-mentioned inconsistent appearances.
7. The method according to claim 1, wherein: Make corrections or adjustments to the work instructions of artificial intelligence machines.
8. The method according to claim 1, wherein: Reject or dismiss the work results of artificial intelligence machines.
9. The method according to claim 1, wherein: Multiple humans cumulatively manipulate, organize, and instruct the visual images.
Citation Information
Patent Citations
Method for human intelligence to make up inherent defects of machine intelligence
CN117993423A
Advanced automobile accident detection, data recordation and reporting system
US20060092043A1
Legal intelligence credit business: a business operation mode of artificial intelligence + legal affairs + business affairs
US20190332983A1
Autonomous driving system and autonomous driving method capable of responding to traffic signal recognition failure
WO2023167345A1
Method for recording independent evidence and complete facts of accident in intelligent driving
WO2023174026A1
Cited By
Hospital guidance and pre-inquiry system based on artificial intelligence
CN121031601A