Personalized, stylized, specialized and localized machine intelligence training method
By using the individual experience of intelligent actors as training material for end-to-end training, the problems of inconsistent style and discontinuous logic in machine intelligence training are solved, and personalized, stylized and professional performance of machine intelligence is achieved, which reduces training costs and improves user experience.
Patent Information
- Application Number
- PCT/CN2024/099464
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-17
- Filing Date
- 2024-06-16
- Publication Date
- 2025-10-23
AI Technical Summary
The existing machine intelligence training involves a mixture of empirical data from multiple intelligent actors, resulting in inconsistent styles, discontinuous logic, and difficulty in meeting personalized, stylized, and professional needs. The quality of training materials is difficult to control, and the cost is high and the efficiency is low.
The experience created by a single individual intelligent actor is used as training material for complete end-to-end training, imitating the technical skills, experience awareness, style preferences, and rhythm and pace of the intelligent actor. The trained machine intelligence has internal consistency and pertinence.
The trained machine intelligence exhibits the individual style of a specific intelligent actor, meeting consumers' needs for personalization, stylization and specialization, reducing training costs and improving training quality and user experience.
Abstract
Description
A personalized, stylized, specialized, localized machine intelligence training method TECHNICAL FIELD
[0001] The present application relates to the field of machine intelligence and artificial intelligence. BACKGROUND
[0002] The so-called "intelligent behavior body" refers to a subject that behaves intelligently, which can be an individual human, an organization formed by a group of humans, a machine intelligence agent, a human-machine community, or the like. The essence of many behavior activities of intelligent behavior bodies is creation, rather than mechanical engineering. Instead, it is a behavior process that is not exactly the same each time under the purpose orientation, rule constraint, and environmental interaction, and eventually reaches the goal. Machine intelligence can imitate this behavior process. For example, for human driving or intelligent driving of intelligent vehicles, even if the driving task is exactly the same, the specific process of completing the same driving task each time is not the same, the action rhythm may differ, the driving route may differ, and the influence of encountering other traffic participants may also differ. Therefore, such a driving experience process is actually a kind of creation under the purpose orientation of the task, the constraint of traffic rules, and the interaction of a changing environment. Such experience data can be used as training materials for training machine intelligence.
[0003] The so-called "end-to-end" intelligence refers to an intelligent form that directly obtains a task result output from an output end without going through intermediate stages, details, logic, and concepts, after inputting raw data from an input end. For example, an infant sees an object similar to a book and says "book". Such an intelligent form can come from "end-to-end" training, that is, an overall training process of training one intelligent behavior body by another intelligent behavior body. Such training establishes and optimizes a model according to raw data. For example, a parent holds up a book and shows it to the infant "this reads book", and repeatedly pronounces "book" themselves. The input of the infant is "eyes see the parent showing the book" and "ears hear the parent pronouncing book", and the output is "mouth pronounces book". The repeated behavior training enables the infant to establish and optimize a complex system model of a long-chain giant network in its body, and from then on, the infant has the ability to see an object similar to a book and pronounce "book" in the parent's tone, intonation, rhythm, and accent. This training does not specifically teach the concept of a book or the essentials of oral muscle strength. The infant does not need to understand the sentence "this reads book" in grammar, and the sentence does not need to be consistent in the same language. Instead, it focuses on the overall "end-to-end" behavior imitation.
[0004] In the training of machine intelligence, the industry often uses "end-to-end" training based on neural network deep learning. For machine intelligence whose intelligent performance is mainly through the process of behavior to reach the goal, such as intelligent driving, intelligent painting, etc., the so-called "end-to-end" training refers to that the training material of the machine intelligence in training is the original experience data of the input and output of the purpose, state, environment, behavior, etc. of the learning object. Through training, the neural network model is established and optimized. After training is completed, in the actual application of the model, i.e. the machine operation of intelligent driving, intelligent painting, the original data of the actual application of the purpose, state, environment, etc. is input to the output end for imitative behavior data output. Taking intelligent driving as an example, the neural network machine intelligence of "end-to-end" training can use the real driving experience data created by human beings (such as collecting driving purpose data, vehicle state data, driving situation video data and corresponding acceleration, coasting, braking, steering and other maneuvering data) as training material, through training, the intelligent driving neural network model is established and optimized. After training is completed, the model can realize the output from the vehicle driving state and environment video input end to the vehicle maneuvering data output end under the guidance of the clear task purpose, i.e. "clear task purpose, state and environment input, driving behavior output", without the traditional perception, prediction, planning, decision-making, control and other sub-stage logical links in the middle of the calculation program, and without specific qualitative variable programming for the concepts of "traffic light", "lane line", "pedestrian", etc. That is, such machine intelligence and training method focuses on imitating the behavior process (purpose-oriented, rule-constrained, environment-interacted to reach the goal) of the learning object as a whole, only the input and output data are recognizable, and the neural network in the middle is a black box.
[0005] In the current machine intelligence and artificial intelligence industry, almost all manufacturers and practitioners use the behavior representation data created by multiple intelligent agents (learning objects) for complete "end-to-end" training, because, in general, the training materials available for learning created by a single intelligent agent are relatively small, which is difficult to meet the need for a large amount of training materials for complete "end-to-end" training, at the same time, the speed of a single intelligent agent to create training materials is slow, which is difficult to meet the demand of manufacturers and practitioners to train as soon as possible, more importantly, it is generally believed or expected in the industry that the use of experience data created by multiple intelligent agents (learning objects) can eliminate the individuality of experience data and move towards universality, and the long-term goal is to make machine intelligence towards so-called "general artificial intelligence", which is a technical bias gradually formed in the industry, blindly and without basis to expect that machine intelligence will one day have general intelligence like humans, such as drawing conclusions from one thing and applying them to another, making the entire industry deviate from the objective fact that humans need personalized, stylized, specialized and localized machine intelligence in some fields, especially ignoring the general demand of human users in some machine intelligence scenarios who strongly hope to have another "my intelligent clone" (rather than "general artificial intelligence") to help them work. Under such technical bias, the use of experience data created by multiple intelligent agents (learning objects) for complete "end-to-end" training in the machine intelligence industry is almost the only universal phenomenon, for example: a certain end-to-end neural network intelligent driving manufacturer has 12 test drivers around the world, covering New Zealand, Thailand, Norway and Japan, etc., collecting driving data of their test vehicles in various parts of the world as training materials, rather than using the driving data of 1 test driver or 1 test vehicle as training materials. In order to meet the large data requirement of the end-to-end training of its FSD Beta V12 autonomous driving model, 10 million selected and labeled human driving videos are input into the training process, and its nearly 2 million vehicle fleet around the world provides about 160 billion frames of video raw materials for training every day; a certain end-to-end neural network visual artificial intelligence manufacturer employs dozens of labeling teams to manually frame and label the visual targets in its training materials, rather than using only one person to frame and label; a certain end-to-end neural network content generation chatbot manufacturer uses the corpus of numerous content contributors in knowledge websites as training materials, rather than using only one person's corpus as training materials; a certain end-to-end neural network content generation poem robot manufacturer uses the poems of multiple poets in a certain dynasty as training materials, rather than using only the poems of a certain poet as training materials.
[0006] The disadvantages are: multiple learning objects are multiple different intelligent agents, each with its own style preference and rhythm pace (the unique correspondence of "the order, amplitude, weight, change of behavior" and "time"), in the process of completing the same task, some are aggressive and some are conservative, some are impatient and some are slow, some focus on speed and some focus on quality, some have fast operation rhythm and some have slow operation rhythm, some have fast operation rhythm in the front and slow operation rhythm in the back, different style preferences lead to large differences in experience data distribution, making the fitting effect of the trained model not good; the mixed experience data of multiple different intelligent agents trained machine intelligence is mediocre and has no style, the style information is mixed and mediocre, and it is difficult to express the individual style of a specific intelligent agent, it is difficult to inherit the individual experience of some excellent intelligent agents, it is difficult to meet the needs of individualization and stylization of consumers, and the machine intelligence obtained by consumers cannot have the technical skills, experience consciousness, style preference, rhythm pace of those excellent individual intelligent agents; the machine intelligence trained by the experience data of multiple different intelligent agents makes the same machine intelligence express multiple "machine personalities" with no memory communication in the work, thus showing logical discontinuity, inconsistency, and style inconsistency, users feel confused and confused, it is difficult to predict, and they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless and helpless when they feel helpless andThe collected training materials often need to be manually screened, classified and labeled, and the materials that do not meet the requirements are removed, and only the materials that meet the requirements are left to ensure that the imitated behavior produced by the machine intelligence after training also meets the requirements. The artificial screening, classification and labeling work is often not completed by the intelligent behavior body itself as the learning object, but often by employees or outsourced personnel of the machine intelligence research and development company or other subjects. The artificial screening, classification and labeling work is time-consuming and laborious, and the high cost is generally borne by the research and development subject of the machine intelligence. In addition, due to the huge workload of screening and classification and labeling and manual operation, it is difficult to check every possible human error in the training materials during quality inspection, and it is inevitable that errors in the training material delivery object are missed. The more serious human errors hidden in the massive training materials will have a significant negative impact on the training results. The responsibility subject of the training material quality is scattered and changes, the producer and the quality inspector of the training material are separated, the production structure of the training material is complex, it is difficult to trace back, and the training quality is uncontrollable. Many training materials of machine intelligence only contain results, but not detailed processes to achieve the results, which actually weakens the richness of experience and the breadth and depth of intelligent representation information. Many users of machine intelligence hope to obtain not only the results produced by machine intelligence, but also the detailed process of the results. If machine intelligence simply and roughly provides users with a result, users can only obtain shallow user experience and can only have a shallow and superficial understanding of the result, and they do not know the reason. On the other hand, if machine intelligence only gives the result without the detailed process during the operation, users cannot intervene in the process in time, errors are easy to accumulate and enlarge, and the operation quality is uncontrollable.
[0007] It is particularly important that the current users of machine intelligence are generally using machine intelligence trained with other unfamiliar intelligent agents as learning objects, which makes the users lack familiarity in predicting the performance of machine intelligence and unable to make experiential predictions from the beginning. In some machine intelligence use scenarios, users strongly hope to have a "my intelligent clone" (rather than a "general artificial intelligence") that is familiar to each other to help themselves or work together, hope that the machine intelligence "seems to be a replica of me", and hope to use an intelligent agent that is capable, familiar, and loyal "like me" rather than a general artificial intelligence agent that is unfamiliar. Such machine intelligence will particularly help itself in a way that it is familiar and preferred, and will particularly be in harmony, have an organic division of labor, and have a tacit understanding when working together, producing a unique effect of "myself working with myself". For example: an intelligent car user who is an excellent driver will hope that the machine intelligence of his intelligent car can "drive in a civilized manner like himself", "have as much driving experience as himself", and "be as familiar with local traffic conditions as himself", so that he can predict the performance of this machine intelligence according to his own experience and habits, and enjoy "myself" serving himself; an intelligent painting user hopes to have a machine intelligence that "has the same painting style as himself" and "has the same painting process style as himself" to help him establish a foundation with the same style when he is drawing a sketch, to still be able to produce a drawing draft or a guide map with the same style when he is not drawing, or even to simulate the painting process with the same style after his death to be used as a teaching example. This example is not a traditional teaching video that is the same every time, but a painting process simulation that is slightly different every time but has exactly the same style; an electronic game or online game player user hopes to be able to play the game with a "machine intelligence clone that is as good as himself and has a tacit understanding, an organic division of labor, and a tacit understanding". This machine intelligence "player clone" does not play the game according to a mechanical program, but has a flexible degree of freedom, and every time it performs a game task, the details of the process are different, but it has the same technical skills, experience awareness, style preference, and rhythm as himself.
[0008] Currently, the machine intelligence trained by humans using neural networks has a complexity that exceeds the limit of what humans can understand in a reductionist way. In this case, any effort to widen and deepen the neural network or expand the neural network model (so-called "large model") still stays at the current level of complexity that humans cannot understand. Since humans cannot understand it, they cannot continue to increase the complexity to a higher level. At the same time, such machine intelligence is based on past experience, and during the use of the fixed machine intelligence, new experiences are often not used as new training materials, making the fixed machine intelligence gradually outdated and increasingly difficult to cope with new situations.
[0009] SUMMARY
[0010] The present application is a personalized, stylized, specialized, localized machine intelligence training method, which uses only the experience created by a single individual of an intelligent agent as training material, and conducts a complete end-to-end training of machine intelligence, so that the neural network machine intelligence imitates the technical skills, experience awareness, style preference, and rhythm pace of the individual of the intelligent agent.
[0011] As training material, if the individual experience of the intelligent agent is focused on a specific professional field or a specific spatio-temporal region, a targeted training result for the specific professional field or the specific spatio-temporal region will be obtained, and the training result will be particularly suitable or applicable to the specific professional field or the specific spatio-temporal region. In a specific professional field or a specific spatio-temporal region, in addition to direct experience, indirect experience can also be added as training material. In a specific professional field or a specific spatio-temporal region, the intelligent agent can apply the machine intelligence trained with itself as the training object, and this application can be one-to-many, i.e., the intelligent agent simultaneously applies multiple machine intelligences trained with itself as the training object.
[0012] Training machine intelligence with the experience created by a single individual of an intelligent agent as training material, and then using the experience created by a single individual of the intelligent agent with the trained machine intelligence as training material to train the next machine intelligence, this process can be iteratively extrapolated.
[0013] During training, the machine intelligence can be trained with real-time experience data of the intelligent agent, or with past experience data records of the intelligent agent.
[0014] The single individual of the intelligent agent autonomously selects the experience material for training machine intelligence. Experience not only expresses results, but also emphasizes the process of achieving results, i.e., the experience material as training material is not just a final result, but also emphasizes the detailed process of gradually achieving the result step by step.
[0015] The so-called "single individual of intelligent behavior body" refers to an intelligent behavior body that can only express a unique behavior at the same time, and the behavior obviously comes from its intelligence, but the source of its intelligence does not require uniqueness. For example, a "person" can finally choose and take a unique action after complex game of multiple ideas and left and right balance of thinking before and after thinking, a "machine intelligent agent" can finally choose and take a unique action after complex game of decision emergence and priority contention and sequencing of multiple intelligent modules or robot personalities, and a "human-machine community" can finally choose and take a unique action after complex game of decision emergence and priority contention and sequencing of human decision and machine decision. That is, at the same time, "ideas" can be diverse, "intelligence" can be multiple, but "decisions" can only be unique, so "behaviors" can only be unique.
[0016] The beneficial effects brought by the patent are described as follows.
[0017] The experience created by a single individual of an intelligent agent is used as training material for the fully "end-to-end" training of machine intelligence. The neural network machine intelligence imitates the single individual of the intelligent agent, and the trained machine intelligence exhibits the technical skills, experience awareness, style preference, and rhythm pace of the single individual of the intelligent agent. Due to the obvious consistency and identity of the individual's technical skills, experience awareness, style preference, and rhythm pace, the distribution of the experience data created by a single individual of an intelligent agent is relatively small compared to the mixing of experience data from multiple different intelligent agents. The fitting effect of the training result model is better, and the machine intelligence will express the individual style of the specific intelligent agent, so that the individual style will not be lost in the mixing of experience data from multiple different intelligent agents. The individual experience of an excellent intelligent agent can be passed on, and consumers can obtain machine intelligence with the technical skills, experience awareness, style preference, and rhythm pace of the individual of an excellent intelligent agent that they want, meeting the needs of consumers for personalized and stylized machine intelligence. As training material, the experience of a single individual of an intelligent agent focuses on a specific professional field or a specific spatio-temporal region, and the trained machine intelligence is more targeted and performs better for the specific professional field or specific spatio-temporal region, meeting the needs of consumers for specialized and localized machine intelligence. Consumers can choose machine intelligence trained by a specific specialized and localized individual of an intelligent agent, and the reputation or brand of a well-tested or widely popular intelligent agent will be selected and passed on. In a specific professional field or a specific spatio-temporal region, adding indirect experience as training material in addition to direct experience will increase the amount of data expressing the technical skills, experience awareness, style preference, and rhythm pace of the single individual of the intelligent agent, and the data dimension will be more rich and the training effect will be better. The machine intelligence trained only with the experience of a single individual of an intelligent agent has natural internal consistency, and will not express multiple "robot personalities" with no memory communication and "split", but will express a single "robot personality", and will not exhibit logical discontinuity, inconsistency, and style inconsistency. The user's expectations are consistent, logical, and unified in style, making it easier to predict and more reliable for intervention or the way and degree of intervention.When the user trains the machine intelligence with himself as the training object and his own experience as the training material, he will naturally have the opportunity to use the machine intelligence later, and other users will also know that the training object is himself when using the machine intelligence. The traceability of the quality of the training material can only point to himself, so the subject of the responsibility for the quality of the training material is single and unique, the producer and the quality inspector of the training material are combined, the production structure of the training material is simplified, the user naturally has the motivation to ensure and improve the quality of the training material, actively pursues the clean, neat and pure training material produced by his experience, and actively and independently selects, chooses and cleans the experience created by himself, removes the experience fragments containing obvious major errors, and in addition, most of the time, the training material is only the original data of his normal work experience, which does not require artificial screening, labeling and other work. Therefore, for the subject of the development of machine intelligence, the artificial workload of screening, classifying, labeling and cleaning the training material is reduced or even eliminated, and the workload of quality inspection of the training material is also reduced or even eliminated, the cost of the development of machine intelligence is reduced, and the quality of the training material makes the training quality more controllable. The training material not only contains the result, but also emphasizes the detailed process of achieving the result, which will increase the richness of the experience and the breadth and depth of the intelligent representation information. In addition to the result produced by the machine intelligence, the user can also obtain, experience and understand the detailed process of the result, obtain a more complete user experience, have a deep understanding of the result, and on the other hand, since the machine intelligence not only gives the result but also gives the detailed process, it is convenient for the user to intervene in time in the process to prevent the accumulation and amplification of errors, so as to make the work quality more controllable.
[0018] Particularly important is that the user will be familiar with the performance of the machine intelligence trained by himself using himself as the learning object and his own experience as the training material, and will have a strong ability to make empirical predictions from the beginning. The machine intelligence with the same technical skills, experience awareness, style preference, and rhythm as himself will be able to establish a foundation for his work, assist him, and even try a step when he is unsure, giving him some inspiration. In some machine intelligence use scenarios, the user strongly desires a "my intelligent clone" (rather than a "general artificial intelligence") that is familiar to him to help him with his work or to collaborate with him, and hopes that the machine intelligence will be "like a re-creation of me", and that he will be able to use a capable, familiar, and loyal "machine intelligence like me" rather than a strange general artificial intelligence. Such machine intelligence will be particularly helpful to him in his own familiar and preferred way, and will be particularly in tune with him, have a natural division of labor, and be in perfect coordination when collaborating with him, producing a unique "myself working with myself" effect. Such a demand will be met by the present patent technical solution. Obviously, the user is the best collaborator. For example, an intelligent car user who is an excellent driver can use a machine intelligence that "drives like a civilized driver like himself", "has as much driving experience as himself", and "is as familiar with local traffic as himself". The user can predict the performance of the machine intelligence based on his own experience and habits, and can intervene effectively without being reckless, and can enjoy the service of "myself" to "myself" with peace of mind. An intelligent painting user can have and use a machine intelligence that "paints in the same style as himself" and "has the same style in the painting process". It can help him establish a foundation for the same style of draft when he is drawing, and can still produce the same style of draft or guide map when he is not drawing, or even after his death, the style of the painting process can be simulated to be passed down as a teaching example. This example is not a traditional teaching video that is the same every time, but a painting process simulation that is slightly different every time but completely consistent in style. An electronic game or online game player can play the game with a "machine intelligence clone like himself" that is in tune with him, has a natural division of labor, and is in perfect coordination. This machine intelligence "player clone" does not play the game according to a mechanical program, but has a flexible degree of freedom, and performs the game task in a different way every time, but has the same technical skills, experience awareness, style preference, and rhythm as himself.
[0019] The experience created by a single individual of an intelligent agent is used as training material to train a machine intelligence, and the experience created by a single individual of the machine intelligence agent of the trained machine intelligence is used as training material to train the next machine intelligence, and this process can be iteratively extrapolated. Machine intelligence can understand the neural network complexity that humans cannot understand in its inherent way, so it can continue to increase complexity to a higher and deeper level in a hierarchical leap, rather than expanding the neural network model at a low level of neural network complexity as humans do. At the same time, constantly training new machine intelligence with new experience will enable machine intelligence to keep pace with the times and respond to new situations. DETAILED DESCRIPTION
[0020] The present application is a personalized, stylized, specialized, localized machine intelligence training method, belonging to the field of machine intelligence and artificial intelligence. Only the experience created by a single individual of an intelligent agent is used as training material to train a machine intelligence completely end-to-end, so that the neural network machine intelligence imitates the individual of the intelligent agent, and the trained machine intelligence will exhibit the technical skills, experience awareness, style preference, rhythm pace of the individual of the intelligent agent, with natural internal consistency, recreating "my intelligent clone", which will be particularly intuitive, organic division of labor, and tacit coordination, and will be more targeted for specific professional fields or specific spatiotemporal regions. Consumers can choose a machine intelligence trained by a specific individual of an intelligent agent, and the reputation or brand of a well-tested or widely popular intelligent agent will be selected and passed down. Iterative extrapolation of training will enable neural network complexity to leap hierarchically to a higher and deeper level, and to keep pace with the times and respond to new situations.
[0021] One embodiment is the training of intelligent driving machine intelligence. The process of human driving a vehicle is actually a kind of human creation in the task goal-oriented, traffic rule-constrained, and changing environment interaction, even if the driving task is exactly the same, the specific process of completing the same driving task each time is different, the action rhythm may be different, the driving route may be different, and the encounter with other traffic participants may also be different. These tasks, actions, routes, states, encounters with other traffic participants, and traffic environment information can be digitized as training material for "end-to-end" training of intelligent driving machine intelligence.
[0022] The driver and the vehicle jointly constitute an intelligent agent, and a single individual of a vehicle of a driver is an intelligent agent individual, and experience data thereof are used for end-to-end training of intelligent driving machine intelligence, for example, a taxi or a network car in a city collects experience data created by a single individual thereof during driving operation to form training materials. The driving style preferences of drivers are different, some are aggressive, some are conservative, some are decisive, some are humble, some like to drive fast, and some like to drive slow, and the trained machine intelligence will have the individual style embodied by the combination of the driver and the vehicle, completely imitating the technical skills, experience awareness, style preference, rhythm pace of the intelligent agent individual, so that consumers can choose from multiple different machine intelligence styles corresponding to multiple different driving styles when selecting intelligent driving products, and have the option to choose, that is, "can eat small pot, can not eat big pot", meet the personalized and stylized user needs, rather than being forced to accept the same machine intelligence sold by the current intelligent driving manufacturers to all users.
[0023] In the current industry, the machine intelligence of intelligent vehicles often lacks pertinence when training multiple different intelligent agents (learning objects) with experience data. Vehicles and their drivers driving in different cities, different regions, and different areas face different topographies (e.g., some cities are plains, and some cities are mountainous), different traffic infrastructure forms (e.g., although there are national standards to constrain, traffic lights in different cities may have different designs and practices in terms of light heads, light posts, light covers, column decorations, interval distances, and installation angles), different traffic infrastructure styles (e.g., some cities have more viaducts, some cities have centerlines with railings, and some cities have mixed people, vehicles, and livestock on the streets), different climate environments (e.g., some cities have constant rain, some cities have constant wind and sand, and some cities have constant haze), and different lighting conditions (some cities have constant sunny days with strong light, and some cities have constant overcast days with insufficient light). Therefore, when training machine intelligence with mixed experience data collected from multiple different cities, different regions, and different areas, it is difficult to balance generalizability and pertinence. Designers of machine intelligence hope that their intelligent vehicles can automatically drive like humans in these different cities, different regions, and different areas. However, the result is that they often perform poorly and are difficult to meet the needs of consumers for specialization and localization. In this technical solution, most of the time, the operating range of taxis or online car-hailing services is generally within the same city, region, or area. Therefore, the experience data they create is often more targeted to this city, region, or area. Using this technical solution will obviously improve the pertinence of the city, region, or area, and enhance the performance of machine intelligence trained with such experience data in that city, region, or area. For example, in a city, a number of "civilized taxi" drivers and a number of "civilized online car-hailing" drivers are selected and organized, and vehicles with the same performance specifications (e.g., chassis specifications and power indicators) are used for driving operations and collecting experience data. Each "civilized taxi" driver and "civilized online car-hailing" driver corresponds to an end-to-end training of a machine intelligence. After a period of one-on-one specialized training, the machine intelligence will be able to mimic the technical skills, experience awareness, style preferences, and rhythm pace of the driver. These intelligent driving machine intelligences can then be sold, and they can be labeled with tags such as "a certain person trained by a certain person who won the title of 'civilized taxi' in a certain city in a certain year," "a certain person trained by a certain person who has operated a taxi safely for a certain number of days without accidents," "a certain person trained by a certain person who is a 'March 8 Red Flag' driver," "a certain person trained by a certain person who is the best online car-hailing driver evaluated by consumers in a certain city in a certain year," and others to meet the needs of consumers for personalized and stylized selection of intelligent driving machine intelligences.For consumers, their experience in intelligent driving is "just like the best local taxi or online car driving person X X X is driving for me", for example, users will experience the specific civilized driving habits of the civilized taxi driver X X X: 1 meter away from the zebra crossing before the zebra crossing, slightly early stop to avoid scaring the pedestrians crossing the street.
[0024] The driver of the taxi or online car can also use the existing intelligent driving based vehicle to drive in the way of man-machine co-driving when creating experience data in the driving process, that is, the vehicle behavior is driven by the data output by the intelligent car in the usual automatic driving of the intelligent car, and the vehicle behavior is driven by the data output by the driver when the driver intervenes and takes over due to his own needs and style, at this time the data output by the driver covers the data output by the intelligent car in the instruction chain. For man-machine co-driving, whether the immediate decision is human intelligence or machine intelligence, the final driving behavior is unique, and the training material of the collected experience data is also unique, including the above-mentioned collection of behavior data output by the intelligent car when not intervening and the behavior data output by the driver when intervening, this data collection still represents the technical skill, experience consciousness, style preference and rhythm of the driver with assistance, that is, the single intelligent car of man-machine co-driving also belongs to the "single individual of intelligent behavior body" described in the technical solution.
[0025] When the training module and the trained model are on the car, the vehicle is in the effective driving process (working state, such as normal driving process, stopping at red light for tens of seconds), and the model is trained in real time with the real-time environment and action data of the vehicle. The training module uses the stored data record to train in the non-effective driving process (stop state, such as the driver's lunch break for 1 hour, the driver's return home for 10 hours).
[0026] When the training module and the trained model are on the cloud, the vehicle is in the effective driving process (working state, such as normal driving process, stopping at red light for tens of seconds), and the real-time experience data is uploaded to the cloud, and the training module on the cloud trains the model in real time. The training module on the cloud uses the data record stored on the cloud to train in the non-effective driving process (non-working state, such as the driver's lunch break for 1 hour, the driver's return home for 10 hours).
[0027] In a specific professional field or specific space-time region, in addition to direct experience, indirect experience can also be added as training material. For intelligent driving machine intelligence, direct experience data in end-to-end training material obviously contains task information, state information, environment information and driving action information, for example, the specific environment cognition of the driver during driving directly leads to the task purpose orientation, traffic rule constraint and specific driving action under the interaction of changing environment. In addition to these direct experiences, indirect experience data can also be collected, such as collecting GPS geographic positioning latitude and longitude data of the vehicle during driving as training material. The reason why it is called "indirect" is that the driver never consciously directly perceives and considers the latitude and longitude data of the location during driving, but because the vehicle operates and collects experience data in a specific space-time region (for example, limited to a city), its driving behavior and latitude and longitude data can often be approximately matched in fact, with indirect relevance, for example, only straight or right turn actions occur at a certain latitude and longitude position, and only forward driving occurs at a certain latitude and longitude to a certain latitude and longitude. Adding indirect experience as training material will increase the amount of data expressing the technical skills, experience awareness, style preference and rhythm pace of the individual of the intelligent behavior body, the data dimension is more rich, and the training effect is better.
[0028] In the current real case, a brand of intelligent vehicle intelligent driving system FSD v12 that uses full end-to-end training of machine intelligence has the following real performance: the intelligent vehicle is queuing in front of a traffic light at a crossroads to straighten through the intersection, the straight traffic light turns green, the front car straightens through the line, the car slows down and stops first, and then stops in the line when the straight traffic light is green, and then the straight traffic light turns red and the left turn traffic light turns green. The car in the straight lane and under the jurisdiction of the straight traffic light (red light) starts to accelerate immediately, and fortunately the driver stops in panic, but the vehicle has crossed the line and violated the law. After technical investigation, it is found that the intelligent vehicle was originally prepared to follow the front car straight through the line, but its machine intelligence perceived that all lanes in the road in front of the crossroads opposite were full of vehicles and severely congested. This intelligent vehicle machine intelligence has been trained and learned many drivers' driving behaviors, so it imitates the "gentlemanly" behavior of not following the front car to straighten through the intersection into the opposite road to intensify congestion or occupy the intersection space when the light is green. We call this "machine personality 1" for convenience. Then, the straight light turns red and the left turn light turns green a few seconds later. At this time, the intelligent vehicle imitates the following human behavior according to the driving behavior of many "ordinary normal" drivers it has learned: "the left turn light is red, the straight light turns green, the front car straightens through the line, the car starts to follow and stops in the line, which means that the car is waiting for the left turn light to turn green, that is, it is waiting for another light to turn green when the green light is encountered." This is a common and normal human driving behavior. The machine intelligence "neurological disorder" and "personality split" connect this imitated behavior. We call this "machine personality 2" for convenience. Obviously, in this real case, although the same machine intelligence is driving, "machine personality 1" and "machine personality 2" are split, learning and imitating from different drivers' vehicles, and the two "machine personalities" do not communicate with each other. Multiple different "machine personalities" emerge on the spot and imitate human driving experience, but this kind of imitation without memory communication, logical discontinuity, inconsistency, and unified style, shows disordered and poor intelligent driving behavior, leading to red light violation, at the same time, because the machine intelligence performs better in "machine personality 1", the user cannot predict the low-level violation action of "machine personality 2" that follows, so it is unable to intervene predictively.Obviously, in this case, if we only use the experience created by a single "vehicle driven by an extremely gentlemanly driver" as training material (the same straight-moving vehicle must be consistent and logically continuous after "gentlemanly" stopping and waiting and continue to move straight after the road ahead is no longer congested), and conduct complete end-to-end training on the intelligent driving machine intelligence, the neural network machine intelligence will have the internal consistency to imitate the single individual of this intelligent actor. The trained intelligent driving machine intelligence will exhibit the consistent and continuous technical skills, experience awareness, style preferences, and rhythm and pace of this "vehicle driven by an extremely gentlemanly driver", and emerge as a single "robot personality". There will no longer be memoryless communication, logical discontinuity, inconsistency, and inconsistent style. Users will no longer experience the "unexpected" and "inexplicable" performance of the vehicle in this case, but will be more likely to predict its internally consistent performance, and will be more confident in whether to intervene or the method and degree of intervention.
[0029] When the driver trains the intelligent driving machine intelligence with himself as the training object and his experience as the training material, the training material quality responsibility subject is single and unique, the producer and quality inspector of the training material are combined, the production structure of the training material is simplified, considering that he will naturally have the opportunity to use this machine intelligence in the future, and considering that other users will know that he is the training object when using the machine intelligence, the training material quality can only be traced back to himself, the driver will naturally have the motivation to ensure and improve the quality of the training material, and will actively pursue the clean, neat and pure training material generated from his experience, and will actively and autonomously filter, classify, select and clean the experience fragments he created, remove experience fragments containing obvious major errors, for example: if the driver accidentally runs a red light during the operation and data collection process, it is a major mistake, and the driver will obviously not want to mislead the machine intelligence training with it as training material, so the driver actively and autonomously deletes the experience fragments collected within several minutes before the violation through relevant operations; in the event of transporting a critically ill patient to the hospital for emergency treatment, the driver intentionally runs a red light, which is a special behavior that is not universal in extreme small probability events, and the driver will obviously not want to mislead the machine intelligence training with it as training material, so the driver actively and autonomously deletes the experience fragments related to the event through relevant operations. In addition, most of the time, the training material is only the raw data of his normal operation experience, and does not require human filtering, classification, labeling and cleaning, so for the research and development subject of intelligent driving machine intelligence, it not only reduces or even eliminates the artificial workload of filtering, classifying, labeling and cleaning the training material, but also reduces or even eliminates the workload of quality inspection of the training material, the research and development subject of intelligent driving machine intelligence reduces the cost, and the high-quality training material also makes the training quality more controllable. In such a technical and economic structure, the individual driver trains the machine intelligence generated by the individual driver, and the individual driver's technical skills, experience awareness, style preference and rhythm form a reputation or brand that has been tested for a long time or is widely welcomed, and is specified and selected by consumers and passed down, consumers can pay for and purchase from the driver (the accumulators and providers of training material), and intelligent driving manufacturers and drivers can also be responsible for or compensate for problems, and consumers, intelligent driving manufacturers and individual drivers (accumulators and providers of training material) can share and share the benefits and risks of intelligent driving.
[0030] The training of machine intelligence can keep pace with the times. For example, driver Li buys the "civilized taxi title holder Zhang's exclusive full-time training" intelligent driving vehicle, conducts taxi operation, and trains the next (next generation) new intelligent driving machine intelligence with the experience data created in his operation. After a period of time, Li's taxi also wins the "civilized taxi" title, so the new generation of intelligent driving machine intelligence can be sold, with a label indicating "from a certain month to a certain month of a certain year to a certain month of a certain year by the exclusive full-time training of the civilized taxi title holder Zhang, from a certain month to a certain month of a certain year to a certain month of a certain year by the continued full-time training of the civilized taxi title holder Li", meeting the needs of consumers' individualization and stylization. For consumers, their experience of intelligent driving is "the best taxi drivers Zhang and Li in the local continue to serve me, their training time is very long so their experience is very rich, and they have been training recently so their experience is very fresh and can handle new situations". In this way, the reputation or brand of individuals who have been tested for a long time or are widely popular will be selected and passed down. Machine intelligence can understand the neural network complexity that humans cannot understand in its own way. The new machine intelligence can continue to increase complexity to a higher and deeper level in a hierarchical leap, rather than expanding the neural network model at a low level of the current level of neural network complexity as humans do. New training of new machine intelligence with new experience will enable machine intelligence to keep pace with the times and cope with new situations. For example, during Li's operation and training period, some new traffic sign infrastructure appears in the urban area. These new situations did not occur during Zhang's operation and training period, so the "old" intelligent driving vehicle of "civilized taxi title holder Zhang's exclusive full-time training" is unable to cope with it. The new generation of intelligent driving vehicle of "from a certain month to a certain month of a certain year to a certain month of a certain year by the exclusive full-time training of the civilized taxi title holder Zhang, from a certain month to a certain month of a certain year to a certain month of a certain year by the continued full-time training of the civilized taxi title holder Li" will have the ability to cope with it.
[0031] For a certain city, a certain civilization taxi title winner, the driver can use the machine intelligent vehicle trained by himself as the training object and his own experience as the training material to carry out intelligent driving operation. The excellent driver will experience the machine intelligence of "civilized driving habits are as good as his own", "driving experience is as rich as his own", and "familiarity with local traffic conditions is as good as his own" when using such intelligent driving, and enjoy the service provided by "himself" to himself. Since the machine intelligent vehicle is a "my intelligent driving clone", the driver is familiar with his own technical skills, experience awareness, style preference, and rhythm, so the driver will have a high degree of confidence and operation execution ability in predicting when to intervene and take over and when not to intervene. The more familiar the driver is with the place, scene, and action, the less he needs to intervene, and the more he needs to pay special attention and be vigilant in the less familiar place, scene, and action. When using his own machine intelligent "intelligent driving clone" intelligent car, the driver can supervise and control from inside the car or remotely from outside the car. Because he is particularly familiar, proficient, and confident, he can apply, monitor, and manage multiple "intelligent driving clone" intelligent cars at the same time. When the driver is relatively old, he can use his own machine intelligent "intelligent driving clone" to operate multiple "intelligent driving clone" intelligent cars remotely in a comfortable room, rather than sitting in the car and driving one taxi. He can successfully predict with high confidence and a high probability which situations the intelligent car can handle and which situations the intelligent car may have difficulty handling based on his life experience, because he is familiar with the local street environment and knows what experience he used to train the intelligent car. His "intelligent driving clone" will help him in a way he is familiar with and prefers, and their cooperation will be in perfect harmony, with clear division of labor and tacit understanding. For example, one day, a street is affected by a flower car parade, and the normal traffic environment has changed, but the road structure of the street itself has not changed. At this time, the driver confidently knows that his experience-trained intelligent car can turn reliably, but he is worried about the intelligent car's response to the extremely complex traffic participants (flower cars, pedestrians, dancers, waving colorful flags, and flying paper streamers balloons) in the sudden flower car parade (worried that the intelligent car will brake too frequently). Then the driver can instruct the machine intelligence of the intelligent car to be responsible for turning, while he temporarily takes care of the accelerator and brake of the vehicle, avoiding collisions with non-collidable traffic participants and not being forced to stop by essentially harmless streamers balloons and other pseudo obstacles. This tacit understanding and organic division of labor is obviously the best cooperation between the driver and his own machine intelligence.
[0032] One embodiment is the training of intelligent painting machine intelligence. The process of human painting is actually a kind of human creation under the guidance of expression purpose, the constraint of tool rule, and the limitation of changing environment. Even if the painting expression task is completely the same, the specific process of completing the same task each time is different, the rhythm of the pen movement may be different, the specific painting process may be different, and the influence of encountering painting limitation environment may also be different. These task, action, process, and environment information can be digitized as training materials for end-to-end training of intelligent painting machine intelligence.
[0033] Only the experience data of the personal creation process of a certain painting master is used as training material to train machine intelligence in a completely end-to-end manner. For example, the action sequence and picture area video pictures reflecting the entire painting process of a certain painting master are collected. These materials will reflect the technical skills, experience awareness, style preference, and rhythm pace of the painting master, including the composition of the entire picture, the selection of pictures and blank spaces, the order and darkness selection of each stroke, the selection and matching of colors, the lightness and change of pen force, etc. The trained intelligent painting machine intelligence will have superb technical skills, experience awareness, style preference, and rhythm pace to imitate this master. Since the training material comes from one person, the distribution difference of the experience data is relatively small, and the fitting effect of the training result model is better. Consumers can choose the painting machine intelligence trained with this painting master as the learning object to meet the consumer's demand for individualization and stylization of machine intelligence. The reputation or brand that has been tested for a long time or widely welcomed will be selected and passed down.
[0034] A certain painting master in a certain professional field can be taken as a learning object, such as a certain painting master who specializes in painting flower and bird Chinese paintings. Then the trained machine intelligence will also have superb technical skills, experience awareness, style preference, and rhythm pace to imitate this master in the professional field of flower and bird Chinese paintings, and will perform better in this professional field of flower and bird Chinese paintings. Consumers can choose the painting machine intelligence trained with this flower and bird Chinese painting master as the learning object to meet the consumer's demand for specialization and localization of machine intelligence. The reputation or brand that has been tested for a long time or widely welcomed will be selected and passed down. For painting, in addition to the direct experience of the composition of the entire picture, the selection of pictures and blank spaces, the order and darkness selection of each stroke, the selection and matching of colors, and the lightness and change of pen force, there is also the indirect experience of the painting master's creation, such as the thickness of the paper, the viscosity of the pigment, and the size of the gold powder particles used for the edge of the flower petals and bird tails. These indirect experiences can be quantified into data and entered into the training materials. The amount of training data in the professional field is increased, the data dimension is more rich, and the training effect is better.
[0035] When the painting master trains the machine intelligence by taking himself as the training object and using his own experience as the training material, the training material quality responsibility subject is single and unique, the producer and quality inspector of the training material are combined, the production structure of the training material is simplified, considering that he himself will naturally have the opportunity to use the machine intelligence in the future, and considering that other users will know that the training object of the machine intelligence is himself when using the machine intelligence, the reverse trace of the quality of the training material can only point to himself, and he will naturally have the motivation to guarantee and improve the quality of the training material, actively pursue the clean, neat and pure training material generated from his experience, and actively and autonomously filter, classify, select and clean the experience he created, remove experience fragments containing obvious major errors, such as major mistakes that may occur occasionally during the painting process, which may have a major adverse impact on the training results, and the painting master will actively and autonomously delete this experience data. In addition, most of the time, the training material is only the raw data of his normal painting experience, and does not require artificial screening, labeling and other work, so for the machine intelligence research and development subject, it not only reduces or even eliminates the artificial workload of screening, classifying, labeling and cleaning the training material, but also reduces or even eliminates the workload of quality inspection of the training material, and the machine intelligence research and development subject reduces the cost, and the high-quality training material also makes the training quality more controllable.
[0036] The painting master can use the machine intelligence trained by taking himself as the training object, and can have and use the machine intelligence with the same style as his own painting and the same style as his own painting process, which can help him establish a similar style draft foundation when drawing a draft, and can still produce a similar style draft or guide drawing when he does not paint, or even after his death, the painting style process can be simulated to be passed down as a teaching example. This example is not a traditional teaching video that is the same every time, but a painting process simulation that is slightly different every time but has exactly the same style. When hesitating about the next stroke in the painting process, the machine intelligence like himself can try to draw a stroke first and give him some inspiration. If the painting master needs to mass-produce paintings, he can monitor and manage multiple machine intelligences trained by taking himself as the training object to draw paintings at the same time, and use his experience to monitor and guide the painting process and screen and perfect the painting results.
[0037] The experience created by a single individual of the intelligent agent in which the machine intelligence trained on a certain master painter resides can be used as training material to train the next (next generation) machine intelligence. Machine intelligence can understand the neural network complexity that humans cannot understand in its own way, so it can continue to increase complexity to a higher and deeper level in a hierarchical leap, rather than simply expanding the neural network model at the same level of neural network complexity as humans do. At the same time, constantly training new machine intelligence with new experience will enable machine intelligence to keep pace with new situations, such as adapting to and training on a new type of paint that has emerged.
[0038] Machine intelligence can be trained in real time with the artist's creative process experience data during painting, or it can be trained using past recorded creative process experience data.
[0039] The training material formed by the painting experience of the master painter not only contains the result, but also emphasizes the detailed process of achieving the result, that is, it focuses on collecting the entire painting process, which will increase the richness of the painting experience and the breadth and depth of the painting intelligence representation information. In addition to the painting results produced by the machine intelligence, the intelligent painting user can also obtain, experience and understand the detailed process of the output results, such as every stroke and every stroke of the painting process, which color covers which color, and how the specific outline is obtained by dyeing, to obtain a more complete user experience and a deep level of understanding of the result. On the other hand, since such a painting machine intelligence not only gives the painting result but also gives the detailed process, it is convenient for the user to intervene in the machine intelligence painting process in time, such as stopping and updating the instructions immediately when a certain stroke deviates from the user's expectations, so as to prevent errors from accumulating and amplifying, thereby making the painting quality more controllable.
[0040] One embodiment is the training of "player avatar" machine intelligence in electronic games or online games. The process of human playing games is actually a kind of human creation under the guidance of task purpose, the constraint of game rules, and the limitation of changing environment. Even if the game task is exactly the same, the specific process of completing the same task each time is different, the action rhythm may be different, the skill selection may be different, the route may be different, and the encounter with other game players and environmental obstacles may also be different. These tasks, actions, skills, routes, encounters with other players, and environmental obstacles can be digitized as training materials to train machine intelligence game players end to end.
[0041] Only the game experience created by a certain game player is used as training material to train machine intelligence in a completely end-to-end manner. For example, the information data of a certain famous game player playing the game, such as tasks, actions, skills, routes, encounters with other players, environmental obstacles, etc. during the whole process of playing the game, the game player is a game star or a game player who is widely liked by game players. The game machine intelligence trained will have superb technical skills, experience awareness, style preference, rhythm pace of the famous player. Since the training material comes from one person, the distribution difference of experience data is relatively small, the fitting effect of the training result model is better, and the consumer can choose the machine intelligence player avatar trained with the famous player as the learning object. When playing the game, get the assistance and guidance from the game star or the master, meet the consumer's demand for machine intelligence individualization and stylization, and the reputation or brand that has been tested for a long time or widely welcomed will be selected and passed down.
[0042] A famous player can focus on a certain fixed game, or a certain fixed role in a certain game, or a certain fixed task in a certain fixed map version in a certain game. The machine intelligence player avatar produced by training with such experience will be particularly focused on these local areas, and will not only have good game performance, but also show the individual preferences of the famous player, meeting the consumer's demand for machine intelligence specialization and localization.
[0043] When a famous player trains machine intelligence with himself as the training object and his experience as the training material, the training material quality responsibility subject is single and unique, the producer and quality inspector of the training material are combined, and the output structure of the training material is simplified. Considering that he will naturally have the opportunity to use this machine intelligence in the future, and considering that other users will know that the training object of the machine intelligence is himself when using it, the training material quality can only be traced back to himself, and he will naturally have the motivation to ensure and improve the quality of the training material. He will actively pursue the cleanliness, tidiness and purity of the training material produced by his experience, and will actively and autonomously filter, classify, select and clean the experience he creates, remove experience fragments containing obvious major errors, such as accidental friendly fire in the game process, which may have a major adverse impact on the training result. This famous player will actively and autonomously delete this experience data. In addition, most of the time, his training material is only the raw data of his normal game experience, and does not require artificial filtering, classification, labeling, etc. Therefore, for the machine intelligence research and development subject, it not only reduces or even eliminates the artificial workload of filtering, classifying, labeling and cleaning the training material, but also reduces or even eliminates the workload of quality inspection of the training material. The machine intelligence research and development subject reduces the cost, and the high-quality training material also makes the training quality more controllable.
[0044] Game players can use the machine intelligence trained with themselves as training objects to play games cooperatively with "machine intelligence avatars like themselves, who play games as well as they do, have the same understanding, organic division of labor, and tacit coordination". This machine intelligence player avatar does not play games according to mechanical programming, but has flexible freedom, and each time it performs game tasks, the process details are different, but it has the same technical skills, experience awareness, style preference, and rhythm pace as the player. Under the allowed conditions, a player can use one or more "machine intelligence avatars like themselves" to play games without personally operating the game. Many network games do not allow players to use ordinary program robots to play games without personally operating the game. Such ordinary program robots completely follow fixed programs to operate games. Network game operators can screen out players who output repetitive, programmed, and patterned game operations and then automatically remove them from the game server. However, if the "machine intelligence avatars like themselves" trained by the present technical solution are used, even if the game tasks are exactly the same, the specific process of each "game avatar" completing the same task is different. The action rhythm may differ, the skill selection may differ, the route may differ, and the corresponding response to changes in other game players and environmental obstacles may also differ. It has the game technical skills, experience awareness, style preference, and rhythm pace of the human player, so it will not be identified as a program robot by the game operator server and will not be removed from the game server. Users can also "play games with themselves", that is, they and one or more "machine intelligence avatars like themselves" form a team to play multiplayer games, or they can use "machine intelligence avatars like themselves" to operate mobile positioning in single-player games while concentrating on manually operating attack actions or skill casting. Such organic division of labor can make individual players balance ideal displacement (such as "body movement") and action (such as "attack and defense") in many complex scenarios, that is, without explicitly instructing the game avatar how to position, the player knows how it will run (just like the player knows how he will run), and concentrates all attention on manually operating game attack actions or game skill casting, without being affected by the need to consider positioning. This is a tacit coordination of organic division of labor, and obviously, the player and himself are the best coordination.
[0045] The game experience created by a game player can be used to train a machine intelligent player, the experience created by the machine intelligent player can be used to train the next (next generation) machine intelligent player, and this process can be iterated and extrapolated, machine intelligence can understand the neural network complexity that humans cannot understand in its own way, so it can continue to increase complexity to higher and deeper levels in a hierarchical leap, rather than expanding the neural network model at the same level as humans, while constantly training new game machine intelligences with new emerging experiences, which will enable game machine intelligences to keep pace with the times and adapt to new situations, such as adapting and training to new tasks, new characters, and new ways of playing in new versions of the game.
[0046] The game experience data created by the game player can be used for immediate training or stored for training.
Claims
1. A method for training machine intelligence for individualized stylized specialization localization, characterized by: Fully end-to-end training of machine intelligence with experience created by a single individual of the intelligent agent as training material.
2. The method of claim 1, wherein: Individual experience of the intelligent agent as training material focuses on a specific professional field or a specific spatio-temporal region.
3. The method of claim 2, wherein: In addition to direct experience, indirect experience is added as training material in a specific professional field or a specific spatio-temporal region.
4. The method of claim 2, wherein: The intelligent agent applies the machine intelligence trained with itself as training material.
5. The method of claim 4, wherein: The intelligent agent applies the machine intelligence trained with itself as training material to multiple individuals.
6. The method of claim 1, wherein: Iterative extrapolation of the process of training machine intelligence with experience created by a single individual of the intelligent agent as training material, and then training the next machine intelligence with experience created by a single individual of the intelligent agent of the machine intelligence trained before.
7. The method of claim 1, wherein: Training of machine intelligence with real-time experience data of the intelligent agent.
8. The method of claim 1, wherein: Training of machine intelligence with past experience data records of the intelligent agent.
9. The method of claim 1, wherein: Autonomous selection of experience material for training machine intelligence by a single individual of the intelligent agent.
10. The method of claim 1, wherein: Experience not only expresses results, but also emphasizes the process of achieving results.
Citation Information
Patent Citations
End-to-end automatic driving model pre-training method based on asynchronous supervised learning
CN112508164A
Multi-machine collaborative air combat planning method and system based on deep reinforcement learning
CN112861442A
Deep reinforcement learning intelligent vehicle behavior decision-making method based on path planning
CN114153213A
Intelligent machine training method for sleeving black box with black box
CN118014038A
Reinforcement machine learning for personalized intelligent alerting
US20170083929A1