Large-model and value-driven three-dimensional scene anthropomorphic agent behavior planning method
Through large models and value-driven methods, anthropomorphic agent behavior planning is solved in three-dimensional scenarios, and the problem of difficulty in dealing with the three-dimensional environment and complex human behavior intentions in the existing technology is solved, and efficient and automated behavior planning data generation is achieved.
Patent Information
- Application Number
- CN202510073086.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology is difficult to implement anthropomorphic agent behavior planning in three-dimensional scenarios, and cannot effectively handle dynamic and static environmental information and complex human behavior intentions.
Using a large model and a value-driven method, the original goal-behavior planning text is acquired and processed, and aligned with three-dimensional action clips and scenes is combined to construct behavior planning data, and on this basis, the execution module is trained to generate a three-dimensional human action sequence.
Achieve high-level, long-term and abstract goal-driven human behavior planning, solves the shortcomings of existing methods in plain text planning or low-level action generation, and generates behavior planning data completely automated, reducing the cost of high-quality data acquisition.
Smart Images

Figure CN119991951A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and three-dimensional computer vision, and in particular relates to a large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method. Background Art
[0002] The latest artificial intelligence technology represented by large models is becoming the core driving force leading the world's new generation of industrial transformation, deeply affecting all areas of society, and thus greatly promoting the social application of intelligent machines. As an important carrier of artificial intelligence, intelligent machines are facing major changes and challenges in the relationship between man and machine in the process of deep participation in social division of labor: machines have developed from functional tool attributes to social role attributes. The interaction between humans and intelligent machines can no longer stay at the level of simple tool use, but needs to be closer and closer to the interaction between people in the real world; intelligent machines will no longer be limited to structured, single stable, and interactive limited task environments, but will face a wider, more open, and human-centered social scenario. This requires intelligent machines to be able to fully understand, foresee and visualize human behavioral intentions through interaction with people, and plan and adjust their own behaviors accordingly, so as to achieve human-centered intelligent decision-making capabilities. For example, when a person makes a request of "needing to rest", the intelligent robot anticipates a series of behavioral steps for people to go from the living room to the bedroom to rest, and makes decisions such as turning off the TV in the living room and adjusting the brightness of the bedroom lights based on this. The new generation of three-dimensional scene anthropomorphic intelligent agent system needs to complete the paradigm shift from perceptual intelligence to interactive intelligence and decision-making intelligence, realize anthropomorphic behavior planning in three-dimensional scenes, and provide key technical guarantees for the social application of intelligent machines.
[0003] At present, there are relatively few studies on the problem of behavior planning for anthropomorphic agent systems in three-dimensional scenes, and there are many deficiencies: (1) Existing human behavior planning technologies mainly study simple mappings from high-level task descriptions to low-level action sequences based on text. For example, "getting ready for bed" is decomposed into a sequence of three low-level actions, namely "walk to the bedside" -> "turn off the bedside lamp" -> "lie down on the bed", which completely ignores the three-dimensional scene and the current state of humans. It is also difficult to visualize and concretize the actions that humans are going to perform based on "turning off the bedside lamp"; (2) Existing three-dimensional human action generation technologies focus more on how to generate three-dimensional human forms corresponding to simple actions such as "sit", "stand", and "walk", while the expression of human behavior intentions during the interaction process is often a high-level, abstract task description, which is far more complex than "sit", "stand", and "walk" (such as "planning to go back to the bedroom to rest"). Therefore, existing behavior planning technologies and three-dimensional human action generation technologies are difficult to solve the problem of behavior planning for anthropomorphic agents in three-dimensional environments. The reason is that current research work faces the following two challenges: (1) It is difficult to obtain 3D data. The collection of real data is time-consuming and laborious, and there is a lack of data soil for algorithm development; (2) Human behavior is changeable and the expression is abstract. The planning space in the 3D scene is huge and the modeling is difficult. How to solve the shortcomings of the above-mentioned existing technologies and realize the behavior planning of the 3D scene anthropomorphic intelligent system is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The purpose of the present invention is to solve the following problems in the prior art: 1) existing human behavior planning technologies are all based on text-based steps and are unable to process dynamic and static environmental information in three-dimensional scenes; 2) existing three-dimensional human behavior generation technologies can only model simple, independent human movements and are unable to meet the planning requirements of complex real-world scenes; 3) there is currently a lack of data for three-dimensional anthropomorphic intelligent agent behavior planning, and manual annotation of data sets is time-consuming and labor-intensive, making it difficult to achieve these three problems, and to provide a large-model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method.
[0005] In order to achieve the above-mentioned invention object, the present invention specifically adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method, which comprises the following steps:
[0007] S1, obtaining an original target-behavior planning text and processing the original target-behavior planning text, obtaining a three-dimensional action segment and a three-dimensional scene and aligning the two, and the processed target-behavior planning text, the aligned three-dimensional action segment and the three-dimensional scene constitute behavior planning data;
[0008] S2. Train the execution module on the behavior planning data. The perception module, planning module and the trained execution module together constitute a 3D scene anthropomorphic intelligent agent system. The perception module is used to receive the 3D point cloud and the first The planning module is used to receive the execution target, the 3D point cloud of the 3D scene, the 3D scene perception description and the previous action planning description sequence in the behavior planning and output the first The action planning description sequence of time steps is used to receive the trained execution module. The action plan description sequence of time steps, The 3D human body state of the time step is taken as the last frame of the 3D human body action sequence. 3D human body state in time steps When the 3D scene anthropomorphic intelligent agent system completes the execution goal in the behavior planning, the 3D human body state finally output by the trained execution module is used as the human body state prediction result; among them, the preceding action planning description sequence is composed of the perception module The action plan description sequence output by all time steps before the time step.
[0009] Based on the above scheme, each step can be implemented in the following preferred specific manner.
[0010] As a preferred embodiment of the above-mentioned first aspect, in step S1, the original goal-behavior planning text is a tree structure, which includes a root node, a middle-layer planning node and a leaf node. The root node represents the execution goal in the behavior plan, each planning node represents each execution step in the behavior plan, and each leaf node is composed of each execution action in the behavior plan and the object that interacts with the execution action.
[0011] As a preferred embodiment of the first aspect, the specific process in step S1 is as follows:
[0012] S11, inputting a first prompt word containing multiple target-behavior planning text examples into a pre-trained large-scale language model, and the large-scale language model generates an original target-behavior planning text based on the first prompt word using a bootstrapping method;
[0013] S12, inputting the second prompt word into the large-scale language model, marking the relative order of the planning nodes in the original goal-behavior planning text, obtaining the first goal-behavior planning text, and completing the missing planning nodes in the first goal-behavior planning text based on the common sense inside the large-scale language model, enhancing the abstractness of the planning node description through the prompt engineering, and obtaining the second goal-behavior planning text;
[0014] S13, calculating the similarity between the second goal-behavior planning text and the goal-behavior planning text example, retaining the second goal-behavior planning text whose similarity is less than a preset similarity threshold, deleting the second goal-behavior planning text whose similarity is greater than or equal to the similarity threshold, inputting the third prompt word into the large-scale language model, filtering the retained second goal-behavior planning text by the large-scale language model, deleting the second goal-behavior planning text that cannot complete the behavior planning, obtaining the final goal-behavior planning text and adding it to the text data set;
[0015] S14, repeating steps S11-S13 until the number of final goal-behavior planning texts in the text data set reaches a preset sample quantity threshold;
[0016] S15, obtaining a human action resource dataset, an indoor scene dataset, an action category label of the executed action, and an object category label of the object, sampling actions with the same action category label from the human action resource dataset and forming a three-dimensional action segment, taking objects with the same object category label in the indoor scene dataset as reference objects, and sampling a three-dimensional scene containing the reference object from the indoor scene dataset;
[0017] S16, taking the reference object as an anchor point, taking the cumulative SDF volume of all contact vertices of the 3D human body mesh and the 3D scene in each frame of the 3D action clip as the cumulative SDF volume corresponding to each frame, taking the cumulative SDF volume of all frames of the 3D action clip as the global cumulative SDF volume, adjusting the translation parameter matrix and the rotation parameter matrix of the 3D action clip to minimize the global cumulative SDF volume, taking the translation parameter matrix with the minimum global cumulative SDF volume as the optimized translation parameter matrix, and taking the rotation parameter matrix with the minimum global cumulative SDF volume as the optimized rotation parameter matrix;
[0018] S17. Use the optimized translation parameter matrix and the optimized rotation parameter matrix to transform the three-dimensional action fragment so that the three-dimensional action fragment is aligned with the three-dimensional scene. The behavior planning data is composed of the final goal-behavior planning text, the aligned three-dimensional action fragment and the three-dimensional scene.
[0019] As a preferred embodiment of the above-mentioned first aspect, in the training process of the execution module in step S2, the three-dimensional scene, the leaf nodes in the final goal-behavior planning text, and the three-dimensional human body state are first input into the pre-trained first diffusion model to predict the human body behavior trajectory, and then the predicted human body behavior trajectory, the three-dimensional scene, and the three-dimensional human body state are input into the pre-trained second diffusion model to predict the three-dimensional human body action sequence. With the goal of minimizing the distance between the predicted three-dimensional human body action sequence and the aligned three-dimensional action fragments, the first diffusion model and the second diffusion model are fine-tuned. When the distance between the three-dimensional human body action sequence and the aligned three-dimensional action fragments is minimized, a trained execution module is obtained.
[0020] As a preferred embodiment of the first aspect, the specific processing process of the perception module is as follows:
[0021] AS21, input the 3D point cloud of the 3D scene into the first 3D point cloud understanding model, output the global text description, The 3D human body coordinates are obtained from the 3D human body state of the time step, the 3D human body coordinates are used as the coordinates of the center of the perception ball, and the perception ball is generated with a preset radius. The 3D point cloud overlapping with the perception ball in the 3D point cloud of the 3D scene is used as a local 3D point cloud, and the local 3D point cloud is input into the second 3D point cloud understanding model to output a local text description;
[0022] AS22, the global text description and the local text description and The three-dimensional human body states at each time step are taken as the input of the heuristic module. The description of the three-dimensional scene layout is extracted from the global text description and used as the scene layout description. The description of the object category and object state of the objects in the perception sphere is extracted from the local text description and used as the object global description. The description of the three-dimensional human body state is extracted from the three-dimensional human body state and used as the human body global description. The three-dimensional scene perception description is composed of the scene layout description, the object global description and the human body global description.
[0023] As a preferred embodiment of the first aspect, the specific processing process of the perception module is as follows:
[0024] BS21, input the execution target in the behavior planning, the 3D point cloud of the 3D scene and the 3D scene perception description into the large-scale language model, and the large-scale language model uses a preset window size Take samples and get candidate behaviors and the confidence level corresponding to each candidate behavior;
[0025] BS22, each candidate behavior, the 3D point cloud of the 3D scene and the The 3D human body state of each time step is input into the trained execution module to obtain the execution result of each candidate behavior, and the first The action plan describes the sequence of time steps, and the numerical value evaluation module calculates the The spatial distance between the action plan description sequence of time steps and the execution result of each candidate behavior is calculated, and the normalized spatial distance is used as the output result of the numerical value evaluation module. In the description value evaluation module, the probability value output by the large-scale language model is used as the description evaluation probability. The description evaluation probability output by the description value evaluation module is multiplied by the normalized spatial distance and the confidence corresponding to the candidate behavior to obtain the value probability corresponding to each candidate behavior. The candidate behavior with the largest value probability is taken as the first The action plan describes a sequence of time steps.
[0026] As a preferred embodiment of the first aspect, in the description value evaluation module of step BS22, the specific processing process is as follows: The character corresponding to the three-dimensional human body state at time steps is used as the reference character, and the text describing the reference character's preferences is obtained and used as the character preference description. The character preference description, the execution goal in the behavior planning, the three-dimensional point cloud of the three-dimensional scene, the three-dimensional scene perception description, and each candidate behavior are input into the large-scale language model. The large-scale language model determines the probability of each candidate behavior occurring under the given character preference description and outputs the probability value of each candidate behavior occurring.
[0027] In a second aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any of the solutions in the first aspect above.
[0028] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any of the schemes in the first aspect above is implemented.
[0029] In a fourth aspect, the present invention provides a computer electronic device comprising a memory and a processor;
[0030] The memory is used to store computer programs;
[0031] The processor is used to implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any of the schemes of the first aspect above when executing the computer program.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. The method of the present invention is based on a large-scale language model and realizes high-level, long-term and abstract goal-driven human behavior planning, which solves the problem that existing methods only focus on pure text planning or low-level action generation.
[0034] 2. The present invention fully automatically generates behavior planning data, fully considers the complexity of the structured human behavior planning, and avoids the high cost of acquiring and expanding high-quality behavior planning data. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flow chart of the steps of the method of the present invention;
[0036] Figure 2 A schematic diagram of a process for generating a first prompt word according to the method of the present invention;
[0037] Figure 3 A schematic diagram of a process for generating behavior planning data for the method of the present invention;
[0038] Figure 4 A schematic diagram of a process for realizing behavior planning in the present invention;
[0039] Figure 5 It is a composition diagram of the computer electronic equipment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.
[0041] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.
[0042] like Figure 1 As shown, in a preferred implementation of the present invention, the above large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method includes the following steps S1 to S2. The specific implementation process is described in detail below.
[0043] S1. Obtain the original target-behavior planning text and process the original target-behavior planning text, obtain the three-dimensional action fragment and the three-dimensional scene and align the two, and the processed target-behavior planning text, the aligned three-dimensional action fragment and the three-dimensional scene constitute the behavior planning data.
[0044] It should be noted that in step S1 of the present invention, the above-mentioned original goal-behavior planning text is a tree structure, which includes a root node, a middle-layer planning node and a leaf node. The root node represents the execution goal in the behavior planning, each planning node represents each execution step in the behavior planning, and each leaf node is composed of each execution action in the behavior planning and the object that interacts with the execution action.
[0045] It should be noted that in step S1 of the present invention, the original goal-behavior planning text can be obtained by a large-scale language model and processed. Here, a large graphic model and a large three-dimensional multimodal model can also be used to replace the large-scale language model. In an embodiment of the present invention, a small amount of manual annotation prompts are used based on a large-scale language model to complete the generation of goal-behavior planning text, quality assessment, and three-dimensional scene and three-dimensional action fragment generation, and then the existing human action resource data set and indoor scene data set are used to align the goal-behavior planning text, three-dimensional scene and three-dimensional action fragments, and finally a large amount of high-quality, complex task-oriented three-dimensional scene anthropomorphic intelligent system behavior planning data is generated at a relatively low cost. Taking a large-scale language model as an example, the process of step S1 is further described below. Figure 3 As shown, the details are as follows:
[0046] S11. Input a first prompt word containing multiple target-behavior planning text examples into a pre-trained large-scale language model, and use a bootstrapping method based on the first prompt word to generate an original target-behavior planning text by the large-scale language model.
[0047] In step S11 of the present invention, an example of generating the first prompt word is as follows: Figure 2 shown.
[0048] S12. Input the second prompt word into the large-scale language model, mark the relative order of the planning nodes in the original goal-behavior planning text, obtain the first goal-behavior planning text, and complete the missing planning nodes in the first goal-behavior planning text based on the common sense within the large-scale language model, enhance the abstractness of the planning node description through the prompt engineering, and obtain the second goal-behavior planning text.
[0049] S13. Calculate the similarity between the second goal-behavior planning text and the goal-behavior planning text example, retain the second goal-behavior planning text whose similarity is less than a preset similarity threshold, delete the second goal-behavior planning text whose similarity is greater than or equal to the similarity threshold, input the third prompt word into the large-scale language model, filter the retained second goal-behavior planning text by the large-scale language model, delete the second goal-behavior planning text that cannot complete the behavior planning, obtain the final goal-behavior planning text and add it to the text data set.
[0050] In step S13 of the present invention, the similarity between the second goal-behavior planning text and the goal-behavior planning text example is calculated by the BERT Score method. Other similarity calculation methods can also be selected here. In addition, the similarity threshold can be selected by those skilled in the art according to actual needs. In step S13 of the present embodiment, the similarity threshold is set to 0.5. After the similarity comparison is completed, a large-scale language model is used in the form of a question-and-answer prompt, that is, the large-scale language model is asked "Is this a valid goal-behavior planning text?" to filter out the second goal-behavior planning text marked as "invalid" to promote the generation of high-quality data.
[0051] S14. Repeat steps S11-S13 until the number of final goal-behavior planning texts in the text data set reaches a preset sample quantity threshold.
[0052] In step S14 of the present invention, the sample quantity threshold can be set by those skilled in the art according to actual needs, and is therefore not limited.
[0053] S15. Obtain a human action resource dataset, an indoor scene dataset, the action category label of the executed action, and the object category label of the object, sample the action with the same action category label from the human action resource dataset and form a three-dimensional action segment, use the object with the same object category label in the indoor scene dataset as a reference object, and sample the three-dimensional scene containing the reference object from the indoor scene dataset.
[0054] In step S15 of the present invention, both the human action resource dataset and the indoor scene dataset adopt existing datasets, such as AMASS, BABEL, GRAB, etc., and the indoor scene datasets such as ScanNet, Replica, HM3D, etc., and the two datasets are sampled using corresponding category labels to obtain available three-dimensional scenes and three-dimensional action clips.
[0055] S16. Taking the reference object as an anchor point, taking the cumulative SDF volume of all contact vertices of the three-dimensional human body mesh and the three-dimensional scene in each frame of the three-dimensional action clip as the cumulative SDF volume corresponding to each frame, taking the cumulative SDF volume of all frames of the three-dimensional action clip as the global cumulative SDF volume, adjusting the translation parameter matrix and the rotation parameter matrix of the three-dimensional action clip to minimize the global cumulative SDF volume, taking the translation parameter matrix with the minimum global cumulative SDF volume as the optimized translation parameter matrix, and taking the rotation parameter matrix with the minimum global cumulative SDF volume as the optimized rotation parameter matrix.
[0056] S17. Use the optimized translation parameter matrix and the optimized rotation parameter matrix to transform the three-dimensional action fragment so that the three-dimensional action fragment is aligned with the three-dimensional scene. The behavior planning data is composed of the final goal-behavior planning text, the aligned three-dimensional action fragment and the three-dimensional scene.
[0057] S2. Training the execution module on the behavior planning data By the perception module Planning Module Together with the trained execution module, it constitutes a 3D scene anthropomorphic intelligent system. The perception module is used to receive the 3D point cloud and the The planning module is used to receive the execution target, the 3D point cloud of the 3D scene, the 3D scene perception description and the previous action planning description sequence in the behavior planning and output the first The action planning description sequence of time steps is used to receive the trained execution module. The action plan description sequence of time steps, The 3D human body state of the time step is taken as the last frame of the 3D human body action sequence. 3D human body state in time steps When the 3D scene anthropomorphic intelligent agent system completes the execution goal in the behavior planning, the 3D human body state finally output by the trained execution module is used as the human body state prediction result; among them, the previous action planning describes the sequence By the perception module The action plan description sequence output by all time steps before the time step.
[0058] It should be noted that in the training process of the execution module in step S2 of the present invention, the three-dimensional scene, the leaf nodes in the final goal-behavior planning text, and the three-dimensional human body state are first The human body behavior trajectory is input into the pre-trained first diffusion model to predict the human body behavior trajectory, and the predicted human body behavior trajectory, three-dimensional scene and three-dimensional human body state are then input into the pre-trained second diffusion model to predict the three-dimensional human body action sequence. With the goal of minimizing the distance between the predicted three-dimensional human body action sequence and the aligned three-dimensional action fragments, the first diffusion model and the second diffusion model are fine-tuned. When the distance between the three-dimensional human body action sequence and the aligned three-dimensional action fragments is minimized, a trained execution module is obtained.
[0059] It should be noted that in the three-dimensional scene anthropomorphic agent system in step S2 of the present invention, Figure 4 As shown, firstly, a perception module based on a heuristic module is used to perceive and understand the three-dimensional scene, and the three-dimensional point cloud and three-dimensional human body state of the three-dimensional scene are input, and a text description of the key objects in the indoor environment, the location information of these key objects and the corresponding state (such as refrigerator: open) is output, and for the three-dimensional human body state, the information describing the three-dimensional human body posture, location and surrounding environment is supplemented (such as, if a sitting person wants to go to the next place, he should stand up first).
[0060] Specifically, the specific processing process of the above perception module is as follows:
[0061] AS21, input the 3D point cloud ε of the 3D scene into the first 3D point cloud understanding model, output a global text description, and 3D human body state in time steps The 3D human coordinates are obtained from the 3D human coordinates, and the 3D human coordinates are used as the coordinates of the center of the perception ball. The perception ball is generated with a preset radius. The 3D point cloud that overlaps with the perception ball in the 3D point cloud of the 3D scene is used as a local 3D point cloud. The local 3D point cloud is input into the second 3D point cloud understanding model, and a local text description is output. is the step identifier.
[0062] It should be noted that in step AS21 of the present invention, the first three-dimensional point cloud understanding model adopts 3DCaptioning, and the second three-dimensional point cloud understanding model adopts 3D Dense Captioning. Both implementation methods belong to the existing technology and will not be repeated here.
[0063] AS22, the global text description and the local text description and The 3D human body states at each time step are taken as the input of the heuristic module. The description of the 3D scene layout is extracted from the global text description and used as the scene layout description. The description of the object category and object state of the objects in the perception sphere is extracted from the local text description and used as the object global description. The description of the 3D human body state is extracted from the 3D human body state and used as the human body global description. The 3D scene perception description e is composed of the scene layout description, the object global description and the human body global description.
[0064] It should be noted that, in step AS22 of the present invention, the global text description is a text description of all objects in the three-dimensional scene and their position information and corresponding states.
[0065] It should be noted that in the three-dimensional scene anthropomorphic intelligent agent system in step S2 of the present invention, the planning module is centered on a large-scale language model, which receives the execution target l in the behavior planning, the three-dimensional point cloud and three-dimensional scene perception description of the three-dimensional scene, and the previous action planning description sequence as input, and finally outputs the next action planning description sequence based on the large-scale language model and value-driven action planning to realize behavior planning.
[0066] Specifically, if Figure 4 As shown, the specific processing process of the above planning module is as follows:
[0067] BS21, input the execution target in the behavior planning, the 3D point cloud of the 3D scene and the 3D scene perception description into the large-scale language model, and the large-scale language model uses a preset window size Take samples and get candidate behaviors and the confidence level corresponding to each candidate behavior.
[0068] BS22, each candidate behavior, the 3D point cloud of the 3D scene and the The 3D human body state of each time step is input into the trained execution module to obtain the execution result of each candidate behavior, and the first The action plan describes a sequence of time steps Calculate the value of the module in the numerical value The action plan describes a sequence of time steps The spatial distance between the execution results of each candidate behavior and the normalized spatial distance is used as the output result of the numerical value evaluation module. In the description value evaluation module, the probability value output by the large-scale language model is used as the description evaluation probability. The description evaluation probability output by the description value evaluation module is multiplied by the normalized spatial distance and the confidence corresponding to the candidate behavior to obtain the value probability corresponding to each candidate behavior. The candidate behavior with the largest value probability is taken as the first The action plan describes a sequence of time steps
[0069] In this embodiment, the numerical value assessment module is used to calculate numerical values, such as indicators such as the shortest path, and the indicator calculation results are normalized into values between 0 and 1 as the output of the numerical value assessment module; the descriptive value assessment module is used to calculate descriptive values, and the human habits described in language, such as "love cleanliness", are evaluated for relative likelihood through a large-scale language model and converted into probability values accordingly, such as 1.0 / 0.7 / 0.3 / 0.01.
[0070] It should be noted that in step BS22 of the present invention, the specific processing process in the description value assessment module is as follows:
[0071] Will be with The character corresponding to the three-dimensional human body state at time steps is used as the reference character, and the text describing the reference character's preferences is obtained and used as the character preference description. The character preference description, the execution goal in the behavior planning, the three-dimensional point cloud of the three-dimensional scene, the three-dimensional scene perception description, and each candidate behavior are input into the large-scale language model. The large-scale language model determines the probability of each candidate behavior occurring under the given character preference description and outputs the probability value of each candidate behavior occurring.
[0072] The present invention will now use a specific example to demonstrate the application effect of the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method described in S1 to S2 in the above embodiments on a specific data set, so as to facilitate understanding of the essence of the present invention.
[0073] Example
[0074] The specific implementation process of the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method used in this embodiment is as described above and will not be repeated here.
[0075] This embodiment uses a large-scale anthropomorphic agent behavior planning dataset. The dataset contains 10K human behavior samples in 1.5K three-dimensional scenes, covering 0.1K actions and 2K unique activities on 1K objects, which is an order of magnitude larger than the manually collected Activity Programs. In this dataset, on average, each high-level goal has 15.7 steps, resulting in a total of 10.1 corresponding 83.3 frame motion sequences, spanning a total of 8.6 million motion frames. Each activity corresponds to 4.1 different motion sequences in 1.7 rooms.
[0076] In order to comprehensively evaluate the performance of the method of the present invention, this embodiment tests the algorithm performance from two key aspects: (1) The simulation should generate linguistically reasonable behavior plans: that is, use Sentence-BLEU and BERT Score to measure the semantic similarity between the real action plan description sequence and the predicted action plan description sequence. (2) Make the three-dimensional human body state at each time step consistent with the action sequence in the three-dimensional environment: Consider the step success rate (SSR) to record the percentage of successful completion of the step target by the three-dimensional scene anthropomorphic intelligent system, defined by the contact distance threshold. For example, if the human body's hips and head are within 30 cm of the target position, lying down is considered successful. Measuring the goal success rate (GSR) is to determine whether all steps in the entire behavior plan have been successfully executed. The experimental results obtained are shown in Table 1, and the results show that the method of the present invention has better planning effects.
[0077] Table 1. Quantitative evaluation results of planning performance of the proposed method on the anthropomorphic agent behavior planning dataset
[0078]
[0079]
[0080] It is understandable that the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method described in S1-S2 above can be substantially implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method provided in the above embodiment, which includes a computer program / instruction. When the computer program / instruction is executed by a processor, the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method described in the above embodiment can be implemented.
[0081] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method provided in the above embodiment, such as Figure 5 As shown, it includes a memory and a processor;
[0082] The memory is used to store computer programs;
[0083] The processor is used to implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method in the above-mentioned embodiment when executing the computer program.
[0084] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0085] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method provided in the above-mentioned embodiment, and a computer program is stored on the storage medium. When the computer program is executed by the processor, it can implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method in the above-mentioned embodiment.
[0086] It is understandable that the above storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. The storage medium may also be a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc., which can store program codes.
[0087] It is understandable that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0088] It should also be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or steps can be combined or integrated together, and a module or step can also be split.
[0089] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.
Claims
1. A large model and value-driven three-dimensional scene anthropomorphic agent behavior planning method, characterized in that: The following steps are involved: S1, obtaining an original target-behavior planning text and processing the original target-behavior planning text, obtaining a three-dimensional action segment and a three-dimensional scene and aligning the two, and the processed target-behavior planning text, the aligned three-dimensional action segment and the three-dimensional scene constitute behavior planning data; S2. Train the execution module on the behavior planning data. The perception module, the planning module and the trained execution module together constitute a three-dimensional scene anthropomorphic intelligent agent system. The perception module is used to receive the three-dimensional point cloud of the three-dimensional scene and the three-dimensional human body state at the i-th time step and output the corresponding three-dimensional scene perception description. The planning module is used to receive the execution target in the behavior planning, the three-dimensional point cloud of the three-dimensional scene, the three-dimensional scene perception description and the previous action planning description sequence and output the action planning description sequence of the i+1-th time step. The trained execution module is used to receive the action planning description sequence of the i+1-th time step, the three-dimensional human body state at the i-th time step and use the last frame of the three-dimensional human body action sequence as the three-dimensional human body state a at the i+1-th time step. i+1 When the three-dimensional scene anthropomorphic intelligent agent system completes the execution goal in the behavior planning, the three-dimensional human body state finally output by the trained execution module is used as the human body state prediction result; among them, the preceding action plan description sequence is composed of the action plan description sequence output by the perception module in all time steps before the i+1th time step.
2. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 1, characterized in that: In step S1, the original goal-behavior planning text is a tree structure, which includes a root node, a middle-layer planning node and a leaf node. The root node represents the execution goal in the behavior planning, each planning node represents each execution step in the behavior planning, and each leaf node is composed of each execution action in the behavior planning and the object that interacts with the execution action.
3. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 2, characterized in that: The specific process in step S1 is as follows: S11, inputting a first prompt word containing multiple target-behavior planning text examples into a pre-trained large-scale language model, and the large-scale language model generates an original target-behavior planning text based on the first prompt word using a bootstrapping method; S12, inputting the second prompt word into the large-scale language model, marking the relative order of the planning nodes in the original goal-behavior planning text, obtaining the first goal-behavior planning text, and completing the missing planning nodes in the first goal-behavior planning text based on the common sense inside the large-scale language model, enhancing the abstractness of the planning node description through the prompt engineering, and obtaining the second goal-behavior planning text; S13, calculating the similarity between the second goal-behavior planning text and the goal-behavior planning text example, retaining the second goal-behavior planning text whose similarity is less than a preset similarity threshold, deleting the second goal-behavior planning text whose similarity is greater than or equal to the similarity threshold, inputting the third prompt word into the large-scale language model, filtering the retained second goal-behavior planning text by the large-scale language model, deleting the second goal-behavior planning text that cannot complete the behavior planning, obtaining the final goal-behavior planning text and adding it to the text data set; S14, repeating steps S11-S13 until the number of final goal-behavior planning texts in the text data set reaches a preset sample quantity threshold; S15, obtaining a human action resource dataset, an indoor scene dataset, an action category label of the executed action, and an object category label of the object, sampling actions with the same action category label from the human action resource dataset and forming a three-dimensional action segment, taking objects with the same object category label in the indoor scene dataset as reference objects, and sampling a three-dimensional scene containing the reference object from the indoor scene dataset; S16, taking the reference object as an anchor point, taking the cumulative SDF volume of all contact vertices of the 3D human body mesh and the 3D scene in each frame of the 3D action clip as the cumulative SDF volume corresponding to each frame, taking the cumulative SDF volume of all frames of the 3D action clip as the global cumulative SDF volume, adjusting the translation parameter matrix and the rotation parameter matrix of the 3D action clip to minimize the global cumulative SDF volume, taking the translation parameter matrix with the minimum global cumulative SDF volume as the optimized translation parameter matrix, and taking the rotation parameter matrix with the minimum global cumulative SDF volume as the optimized rotation parameter matrix; S17. Use the optimized translation parameter matrix and the optimized rotation parameter matrix to transform the three-dimensional action fragment so that the three-dimensional action fragment is aligned with the three-dimensional scene. The behavior planning data is composed of the final goal-behavior planning text, the aligned three-dimensional action fragment and the three-dimensional scene.
4. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 3, characterized in that: In the training process of the execution module of step S2, the three-dimensional scene, the leaf nodes in the final goal-behavior planning text, and the three-dimensional human body state are first input into the pre-trained first diffusion model to predict the human body behavior trajectory, and then the predicted human body behavior trajectory, the three-dimensional scene, and the three-dimensional human body state are input into the pre-trained second diffusion model to predict the three-dimensional human body action sequence. With the goal of minimizing the distance between the predicted three-dimensional human body action sequence and the aligned three-dimensional action fragments, the first diffusion model and the second diffusion model are fine-tuned. When the distance between the three-dimensional human body action sequence and the aligned three-dimensional action fragments is minimized, a trained execution module is obtained.
5. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 4, characterized in that: The specific processing process of the perception module is as follows: AS21, input the 3D point cloud of the 3D scene into the first 3D point cloud understanding model, output a global text description, obtain the 3D human body coordinates from the 3D human body state at the i-th time step, use the 3D human body coordinates as the spherical center coordinates of the perception sphere, and generate a perception sphere with a preset radius size, use the 3D point cloud overlapping with the perception sphere in the 3D point cloud of the 3D scene as a local 3D point cloud, input the local 3D point cloud into the second 3D point cloud understanding model, and output a local text description; AS22. The global text description, the local text description and the three-dimensional human body state at the i-th time step are taken as the input of the heuristic module. The description of the three-dimensional scene layout is extracted from the global text description and used as the scene layout description. The description of the object category and the object state of the object in the perception sphere is extracted from the local text description and used as the object global description. The description of the three-dimensional human body state is extracted from the three-dimensional human body state and used as the human body global description. The three-dimensional scene perception description is composed of the scene layout description, the object global description and the human body global description.
6. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 5, characterized in that: The specific processing process of the perception module is as follows: BS21, input the execution target in the behavior planning, the 3D point cloud of the 3D scene and the 3D scene perception description into the large-scale language model, and the large-scale language model samples with a preset window size w to obtain w candidate behaviors and the confidence corresponding to each candidate behavior; BS22. Input each candidate behavior, the three-dimensional point cloud of the three-dimensional scene, and the three-dimensional human body state at the i-th time step into the trained execution module to obtain the execution result of each candidate behavior, obtain the action planning description sequence of the i-th time step in the previous action planning description sequence, calculate the spatial distance between the action planning description sequence of the i-th time step and the execution result of each candidate behavior in the numerical value evaluation module, and use the normalized spatial distance as the output result of the numerical value evaluation module, use the probability value output by the large-scale language model as the description evaluation probability in the description value evaluation module, multiply the description evaluation probability output by the description value evaluation module with the normalized spatial distance and the confidence corresponding to the candidate behavior, and obtain the value probability corresponding to each candidate behavior, and use the candidate behavior with the largest value probability as the action planning description sequence of the i+1-th time step.
7. A large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as claimed in claim 6, characterized in that: In the description value assessment module of step BS22, the specific processing process is as follows: the person corresponding to the three-dimensional human body state at the i-th time step is taken as the reference person, the text describing the reference person's preferences is obtained and used as the person's preference description, the person's preference description, the execution target in the behavior planning, the three-dimensional point cloud of the three-dimensional scene, the three-dimensional scene perception description, and each candidate behavior are input into the large-scale language model, and the large-scale language model determines the probability of each candidate behavior occurring under the given person's preference description and outputs the probability value of each candidate behavior occurring.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, it can implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any one of claims 1 to 7.
10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the large model and value-driven three-dimensional scene anthropomorphic intelligent agent behavior planning method as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Human body action generation method capable of interacting with three-dimensional scene target and user
CN117953113A
Body-equipped agent training system and method
CN118194966A
Exhibition hall robot visual language navigation method based on large model
CN119309580A
System and Method for Controlling Behavior of a Robotic Character
US20140288704A1
Multi-robot control method, apparatus and system, and storage medium, electronic device and program product
WO2022134732A1