AI large model agent construction method and system

By calculating multimodal perception coefficients and weight coefficients, the training data set of AI large model agents is optimized, and the problem of environmental perception ability decline caused by multimodal information is solved, and the response speed and accuracy of the agents are improved.

CN120277512AActive Publication Date: 2025-07-08SHANDONG HAILIANXUN INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510771564.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Too much multimodal information will lead to a decrease in the environmental perception ability of AI large model agents, and the accuracy of response cannot be guaranteed.

Method used

By constructing the initial agent, obtaining multiple environmental multimodal data corresponding to multiple interaction behaviors, calculating multimodal perception coefficients and weight coefficients, optimizing the training data set to adjust the order of call of the agent to the target API interface, and obtaining the final agent that meets the user's needs.

Benefits of technology

It improves the environmental perception ability of the agent, enhances the response speed and accuracy to user needs, and reduces unnecessary information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277512A_ABST
    Figure CN120277512A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to an AI large model agent construction method and system.The AI large model agent construction method comprises the steps that a target API interface is called through a constructed initial agent, and multiple pieces of environment multi-modal data corresponding to multiple interactive behaviors of a user are obtained, the multiple interactive behaviors occur in the same session established between the user and the initial agent; according to a plurality of environmental multi-modal data corresponding to multiple interactive behaviors, calculating a multi-modal sensing coefficient when the initial agent performs environment sensing for different modal data types under each interactive behavior; calculating weight coefficients of different modal data types based on the multi-modal sensing coefficient; and constructing a pre-loading training data set based on the weight coefficient, and correcting and training a calling sequence of the initial agent for the target API interface during environment perception by using the pre-loading training data set to obtain a final agent meeting user requirements. According to the invention, the response accuracy of the intelligent agent can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method and system for constructing an AI large model intelligent agent. Background Art

[0002] An artificial intelligence (AI) intelligent agent is actually a software tool driven by AI. With minimal supervision, it can perform multi-step tasks. In addition to natural language processing, AI intelligent agents can also make decisions, solve problems, and interact with the environment when performing tasks. The intelligent agent itself is a small program that can accept natural language commands, interact with scenarios, and has a preliminary thought chain. It can split tasks and has the ability to remember, plan, call tools, and execute actions. Moreover, through closed-loop thinking during actions, the intelligent agent can transform the knowledge of the large model into long-term memory or even perception, and can perform specific tasks independently of the large model.

[0003] In the actual use of an AI large model intelligent agent, it depends on the environmental perception ability of multi-modal data. For example, by simultaneously analyzing the images and audio in a video during a conversation, the intelligent agent can more accurately understand the behavioral intentions of humans, realize scenario recognition and reasoning. The environmental perception information is usually real-time or near real-time, and the intelligent agent needs to be able to quickly respond to changes in the environment. However, too much multi-modal information will cause a decline in the environmental perception ability and cannot guarantee the accuracy of the intelligent agent's response. Summary of the Invention

[0004] In order to solve the technical problem that too much multi-modal information will cause a decline in the environmental perception ability and cannot guarantee the accuracy of the intelligent agent's response, the purpose of the present invention is to provide a method and system for constructing an AI large model intelligent agent. The specific technical solutions adopted are as follows:

[0005] In a first aspect, the present invention provides a method for constructing an AI large model intelligent agent, including:

[0006] Using the constructed initial intelligent agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors with the user, where the multiple interaction behaviors occur in the same session created between the user and the initial intelligent agent;

[0007] According to the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the multi-modal perception coefficient when the initial intelligent agent performs environmental perception for different modal data types under each interaction behavior;

[0008] Based on the multi-modal perception coefficient, calculate the weight coefficient of the different modal data types, and the weight coefficient is used to represent the contribution degree of the initial intelligent agent to the environmental perception of the different modal data types;

[0009] Construct a pre-loaded training dataset based on the weight coefficients, and use the pre-loaded training dataset to correct and train the calling order of the initial agent for the target API interface during environmental perception, so as to obtain a final agent that meets the user's needs.

[0010] Optionally, use the constructed initial agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors with the user, including:

[0011] Use the constructed initial agent to call the neural network to analyze whether each interaction behavior input by the user is effective;

[0012] If it is determined that the interaction behavior is effective, use the initial agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors with the user.

[0013] Optionally, according to the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the multi-modal perception coefficients of the initial agent for different modal data types during environmental perception for each interaction behavior, including:

[0014] Based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the user demand bias coefficient for each interaction behavior;

[0015] Based on the environmental multi-modal data and the user demand bias coefficient, calculate the multi-modal perception coefficients of the initial agent for different modal data types during environmental perception for each interaction behavior.

[0016] Optionally, the calculating the user demand bias coefficient for each interaction behavior based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors includes:

[0017] Successively determine each interaction behavior in the multiple interaction behaviors as the current interaction behavior, and determine at least one historical interaction behavior corresponding to the current interaction behavior in the multiple interaction behaviors;

[0018] Based on the environmental multi-modal data corresponding to the at least one historical interaction behavior, repeatedly execute the calculation process of the user demand bias coefficient for the current interaction behavior until the user demand bias coefficients for each interaction behavior in the multiple interaction behaviors are obtained.

[0019] Optionally, the calculation process of the user demand bias coefficient for the current interaction behavior includes:

[0020] Conduct input demand analysis for each historical interaction behavior to determine the demand quantity of the modal data type corresponding to each historical interaction behavior;

[0021] Based on the environmental multimodal data corresponding to the at least one historical interaction behavior, determine the number of perceptions of the initial agent for the modal data type in each of the historical interaction behaviors;

[0022] Based on the required quantity of the modal data type corresponding to each of the historical interaction behaviors and the number of perceptions of the initial agent for the modal data type in each of the historical interaction behaviors, calculate the first user demand bias coefficient under the current interaction behavior.

[0023] Optionally, based on the environmental multimodal data and the user demand bias coefficient, calculate the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior, including:

[0024] According to the acquisition order of different modal data types in the environmental multimodal data corresponding to each interaction behavior, quantify the environmental multimodal data into an environmental perception vector of the initial agent for different modal data types;

[0025] Based on the environmental perception vector and the user demand bias coefficient, calculate the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior.

[0026] Optionally, based on the environmental perception vector and the user demand bias coefficient, calculate the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior, including:

[0027] Successively determine each interaction behavior in the multiple interaction behaviors as the current interaction behavior, and determine the previous interaction behavior corresponding to the current interaction behavior, and repeatedly execute the calculation process of the multimodal perception coefficient for the current interaction behavior until the multimodal perception coefficients for each interaction behavior in the multiple interaction behaviors are obtained;

[0028] Among them, the calculation process of the multimodal perception coefficient includes:

[0029] Determine the first environmental perception vector and the first user demand bias coefficient corresponding to the current interaction behavior, and the second environmental perception vector and the second user demand bias coefficient corresponding to the previous interaction behavior;

[0030] Based on the first environmental perception vector, the first user demand bias coefficient, the second environmental perception vector, and the second user demand bias coefficient, calculate the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in the current interaction behavior.

[0031] Optionally, calculating the weight coefficients of the different modal data types based on the multi-modal perception coefficients includes:

[0032] Obtaining multiple multi-modal perception coefficients when the initial agent performs environmental perception for each modal data type under the multiple interaction behaviors;

[0033] Calculating the weight coefficients of the corresponding modal data types based on the multiple multi-modal perception coefficients.

[0034] Optionally, the method further includes:

[0035] Integrating API interfaces from different content providers and service providers, where the API interfaces are used to support the invocation of external service resources to obtain environmental multi-modal data under the different modal data types;

[0036] Creating an interface invocation relationship between the initial agent and each of the API interfaces.

[0037] In a second aspect, an embodiment of the present invention further provides an AI large model agent construction system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method described in any one of the above are implemented.

[0038] The present invention has the following beneficial effects: Through the technical solution provided by the present invention, first, the initial agent is used to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors with the user; then, according to the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the multi-modal perception coefficients when the initial agent performs environmental perception for different modal data types under each interaction behavior; further calculate the weight coefficients of the different modal data types based on the multi-modal perception coefficients; finally, construct a pre-loaded training data set based on the weight coefficients, and use the pre-loaded training data set to correct the call order of the initial agent for the target API interface during environmental perception, so as to obtain a final agent that meets the user's needs. Through the calculation of the multi-modal perception coefficients, the present invention can accurately quantify the perception ability of the agent for different environmental multi-modal data, and clarify the importance and sensitivity of different environmental multi-modal data in the process of meeting the user's needs. Further calculating the weight coefficients of different modal data types based on the multi-modal perception coefficients, and using the weight coefficients to optimize the call order of the training agent for the target API interface during environmental perception, can enable the finally obtained agent to improve the environmental perception ability, more efficiently obtain multi-modal information valuable for the user's needs, reduce unnecessary information acquisition, and thus improve the response speed and accuracy of the agent.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present invention. Other features and advantages of the present invention will be described in detail in the following specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 FIG. is a schematic flow chart of a method for constructing an AI large model intelligent agent provided by an embodiment of the present invention;

[0042] Figure 2 FIG. is a schematic flow chart of a method for constructing an AI large model intelligent agent provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail a method and system for constructing an AI large model intelligent agent according to the present invention, including its specific implementation, structure, features and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0045] The following will specifically describe the specific solutions of a method and system for constructing an AI large model intelligent agent provided by the present invention with reference to the accompanying drawings.

[0046] Please refer to Figure 1 , which shows a flowchart of a method for constructing an AI large model intelligent agent provided by an embodiment of the present invention. The method includes the following steps:

[0047] Step 110: Use the constructed initial intelligent agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors with the user, and the multiple interaction behaviors occur in the same session created between the user and the initial intelligent agent.

[0048] Among them, the environmental multi-modal data includes various modal data in the interactive environment, such as text data generated when the user inputs a question in text form, voice data when asking questions by voice, image data when sending pictures, etc.

[0049] In the process of constructing an AI large model agent, an initial agent is first built. This initial agent has the ability to call relevant Application Programming Interfaces (APIs). The API is like a channel connecting the agent to various external information resources. A session is established between the user and this initial agent. A session can be understood as an interaction process, in which the user and the initial agent continuously communicate and interact. During the same session, the user will have multiple interaction behaviors. For example, at the beginning, the user may ask a question by text input, then send a voice command, and may also upload a picture, etc. Each interaction will generate different types of data, which contain various modal information in the environment, such as semantic information of text, audio information of voice, visual information of pictures, etc. The initial agent can collect all kinds of environmental multi-modal data corresponding to multiple interaction behaviors generated in the same session by calling the API related to the interaction behavior (i.e., the target API). The API supported by the initial agent can be the API related to the Content Provider (CP) or Service Provider (SP). These APIs are key external resources for constructing an AI large model agent, which enable the agent to access and utilize various online services and information, including hardware devices in the actual environment to obtain video, audio and other information, so as to be able to understand and respond to user input.

[0050] In a specific application scenario, before using the initial agent to call the target API, it is also necessary to create an interface call relationship between the initial agent and each API. Correspondingly, the steps of the embodiment may further include: integrating APIs from different content providers and service providers, where the APIs are used to support the call of external service resources to obtain environmental multi-modal data under different modal data types; creating an interface call relationship between the initial agent and each API.

[0051] Step 120: Calculate the multi-modal perception coefficient of the initial agent for environmental perception for different modal data types under each interaction behavior according to the multiple environmental multi-modal data corresponding to multiple interaction behaviors.

[0052] During the operation of the initial agent, the user interacts with the initial agent multiple times. Each interaction generates environmental data containing multiple modalities, which cover various types such as text, speech, and images, that is, environmental multimodal data. Based on these rich environmental multimodal data, it is necessary to quantitatively evaluate the environmental perception ability of the initial agent for different modality data types (such as text modality, speech modality, image modality, etc.) during each interaction behavior. The result of this quantitative evaluation is the multimodal perception coefficient. Environmental perception refers to the process by which the initial agent understands and analyzes different environmental multimodal data to obtain information about the user's needs and the surrounding environment. For example, by analyzing the text and speech input by the user, understanding the user's question intention; identifying relevant objects or scenes from images to assist in judging the environmental situation, etc.

[0053] When calculating the multimodal perception coefficient, multiple factors are comprehensively considered. For example, the richness of the modality data types included in the user input behavior, and the number of modality data types perceived by the initial agent after each interaction. Through a specific calculation method, these complex information are converted into a specific value to measure the comprehensive ability of the initial agent to perform environmental perception for different modality data during each interaction behavior. This coefficient is very important for optimizing the environmental perception ability of the initial agent. Subsequently, it can be used to adjust the processing strategies of the initial agent for different modality data, improving the performance of the agent and the response accuracy to the user's needs.

[0054] Step 130: Calculate the weight coefficients for different modality data types based on the multimodal perception coefficient.

[0055] In a specific application scenario, the multimodal perception coefficient reflects the environmental perception situation of the initial agent for different modality data types during each interaction behavior. Through specific calculation and processing of the multimodal perception coefficient, the weight coefficients corresponding to different modality data types can be obtained. This calculation process comprehensively considers various factors, such as the change trend of the perception coefficient of a certain modality data in multiple interactions, the importance difference of different modality data in meeting the user's needs, etc.

[0056] Among them, the weight coefficient is used to characterize the contribution degree of the initial agent to the environmental perception of different modal data types. In practical applications, different modal data have different contribution degrees to the initial agent's understanding of user needs and making accurate responses. A modal data type with a higher weight coefficient means that when the initial agent performs environmental perception, it will give priority to processing and paying attention to this type of data and place it in a more important perception order; while for a modal data type with a lower weight coefficient, the initial agent has a relatively lower priority when processing it. In this way, the initial agent can reasonably allocate resources and attention according to the weight coefficient, more efficiently process environmental multimodal data, improve the accuracy and efficiency of environmental perception, and better meet the needs of users.

[0057] Step 140: Construct a pre-loaded training data set based on the weight coefficient, and use the pre-loaded training data set to correct the calling order of the initial agent for the target API interface during environmental perception, so as to obtain a final agent that meets the user's needs.

[0058] For the embodiments of the present disclosure, various types of multimodal data can be sorted and combined according to corresponding proportions and rules based on the weight coefficient to construct a pre-loaded training data set. For example, if the weight coefficient of text modal data is relatively high, then in the data set, the proportion of training data related to text may increase accordingly. The data set constructed in this way contains various data samples that match the weights of different modal data, providing targeted materials for subsequent training of the initial agent. The initial agent's call to the target API interface may not fully consider the importance differences of multimodal data under different user needs. When training the initial agent using the pre-loaded training data set, the initial agent can learn the relationship between different modal data and user needs. Through continuous training, the initial agent adjusts the calling order of the target API interface during environmental perception according to the weight information reflected in the data set. For example, if the data set shows that voice modal data is more critical for a certain type of user needs, the initial agent will give priority to calling the API interface that can obtain voice-related information after training, instead of randomly or indifferently calling the interface as before.

[0059] After using the pre-loaded training data set to correct and train the initial agent, the initial agent will become more accurate and efficient in environmental perception and response to user needs. Since the initial agent has learned to reasonably call the target API interface according to the weight coefficient and obtain the multimodal information most valuable to user needs, when processing user input, it can more accurately understand the user's intention and give a response that better meets the user's expectations. In this way, the initial agent optimized through training becomes the final agent that meets the user's needs, which can improve the user experience and the practicality of the agent.

[0060] In summary, according to an AI large model agent construction method provided by the present invention, first, an initial agent is used to obtain multiple environmental multimodal data corresponding to multiple interaction behaviors with a user; then, according to the multiple environmental multimodal data corresponding to the multiple interaction behaviors, the multimodal perception coefficients of the initial agent for environmental perception for different modal data types are calculated for each interaction behavior; further, the weight coefficients of different modal data types are calculated based on the multimodal perception coefficients; finally, a pre-loaded training data set is constructed based on the weight coefficients, and the pre-loaded training data set is used to correct the call order of the initial agent for the target API interface during environmental perception, so as to obtain a final agent that meets the user's needs. Through the calculation of the multimodal perception coefficients, the present invention can accurately quantify the perception ability of the initial agent for different modal data, and clarify the importance and sensitivity of different modal data in the process of meeting the user's needs. Further, the weight coefficients of different modal data types are calculated based on the multimodal perception coefficients, and the call order of the initial agent for the target API interface during environmental perception is optimized using the weight coefficients, so that the finally obtained agent can improve the environmental perception ability, more efficiently obtain multimodal information valuable to the user's needs, reduce unnecessary information acquisition, and thus improve the speed and accuracy of the agent's response.

[0061] Based on Figure 1 the embodiment shown, as a refinement and extension of the above embodiment, in order to fully illustrate the specific implementation process of the method of this embodiment, this embodiment provides the following Figure 2 shown specific method. Figure 2 Based on Figure 1 the embodiment shown. As Figure 2 shown, the method includes the following steps:

[0062] Step 210: Use the constructed initial agent to call the target API interface to obtain multiple environmental multimodal data corresponding to multiple interaction behaviors with the user, and the multiple interaction behaviors occur in the same session created between the user and the initial agent.

[0063] For the embodiments of the present disclosure, the step of using the constructed initial agent to call the target API interface to obtain multiple environmental multimodal data corresponding to multiple interaction behaviors with the user in step 210 may include the following steps:

[0064] Step 210-1: Use the constructed initial agent to call a neural network to analyze whether each interaction behavior input by the user is valid.

[0065] For the embodiments of the present disclosure, when the initial agent invokes the neural network, each interaction behavior input by the user can be used as input data and input into the neural network. The neural network analyzes these input data to determine whether the interaction behavior is effective. For example, if the user input is meaningless garbled characters, repetitive and worthless content, or information that is completely irrelevant to the interaction scenario set by the initial agent, the neural network may determine that the interaction behavior is invalid; while if the user input is content that conforms to the function positioning of the agent and can guide the initial agent to perform effective information processing, such as clear questions, reasonable instructions, etc., it will be determined as an effective interaction behavior.

[0066] Step 210-2: If it is determined that the interaction behavior is effective, use the initial agent to invoke the target API interface to obtain multiple environmental multi-modal data corresponding to the multiple interaction behaviors of the user.

[0067] After the neural network determines that the user's interaction behavior is effective, the initial agent will perform the next operation, that is, invoke the target API interface. The API interface serves as a bridge for the agent to interact with the external environment and can be connected to various different data sources and services. By invoking these API interfaces, the initial agent can obtain multiple environmental multi-modal data corresponding to the multiple interaction behaviors of the user. The environmental multi-modal data covers various forms, such as text data generated when the user inputs a question in text form, voice data when asking a question by voice, image data when sending a picture, etc. These data contain rich environmental information.

[0068] Step 220: Based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the user demand bias coefficient for each interaction behavior.

[0069] For the embodiments of the present disclosure, calculating the user demand bias coefficient for each interaction behavior based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors in step 220 may include the following steps:

[0070] Step 220-1: Sequentially determine each interaction behavior in the multiple interaction behaviors as the current interaction behavior, and determine at least one historical interaction behavior corresponding to the current interaction behavior in the multiple interaction behaviors.

[0071] Step 220-2: Based on the environmental multi-modal data corresponding to at least one historical interaction behavior, repeatedly execute the calculation process of the user demand bias coefficient for the current interaction behavior until the user demand bias coefficient for each interaction behavior in the multiple interaction behaviors is obtained.

[0072] Among them, the calculation process of the user demand bias coefficient for the current interaction behavior includes the following steps:

[0073] Step 220-2-1: Conduct input requirement analysis for each historical interaction behavior to determine the demand quantity of the modal data type corresponding to each historical interaction behavior.

[0074] Every time the user interacts with the initial intelligent agent, such as inputting text, issuing voice commands, uploading pictures, etc., the initial intelligent agent needs to conduct in-depth analysis on these inputs. This analysis process is to clarify what the user wants to achieve through this input, for example, asking for information, requesting to execute a certain task, or having an emotional communication, etc. Based on the analysis of the user's input requirements, the initial intelligent agent needs to determine the demand quantity of the modal data type involved in each historical interaction behavior. Multimodal data includes various types such as text, voice, image, video, etc. For example, when the user asks "What good movies are there in the nearby cinema today", this historical interaction behavior may involve not only text data (to understand the content of the question), but also may require image data (to show movie posters) and voice data (to play sound clips of movie trailers) to more comprehensively meet the user's needs. Here, text, image, and voice are three different modal data types, and the demand quantity is 3.

[0075] Step 220-2-2: Based on the environmental multimodal data corresponding to at least one historical interaction behavior, determine the perception quantity of the initial intelligent agent for the modal data type in each historical interaction behavior.

[0076] In a specific application scenario, the initial intelligent agent needs to analyze and process the environmental multimodal data received in each historical interaction behavior. The perception quantity refers to the number of types of modal data that the intelligent agent can recognize and understand. For example, in a historical interaction behavior, the user sends both a text message and a picture, and also shares location information (belonging to the sensor data modality). If the initial intelligent agent can recognize and process these three different types of data, then in this historical interaction behavior, the perception quantity of the initial intelligent agent for the modal data type is 3. This quantity reflects the ability of the initial intelligent agent to capture and process different types of environmental information in a specific interaction.

[0077] Step 220-2-3: Based on the demand quantity of the modal data type corresponding to each historical interaction behavior, and the perception quantity of the initial intelligent agent for the modal data type in each historical interaction behavior, calculate the first user demand preference coefficient under the current interaction behavior.

[0078] Among them, the user demand bias coefficient is used to measure the user's demand bias situation during the interaction process. The first user demand bias coefficient is the user demand bias coefficient, which is used to distinguish the user demand bias coefficients of the current interaction behavior and the previous interaction behavior. It comprehensively considers the demand quantity and the perceived quantity, and can reflect the relationship between the user's input demand and the actual perception ability of the initial intelligent agent. In this embodiment, the situation where the demand quantity and the perceived quantity are zero is not analyzed, so generally there is no possibility that the demand quantity and the perceived quantity are zero. When the demand quantity and the perceived quantity are more matched, and the correspondence relationship between the overall feedback and the environmental perception information types is more orderly, it indicates that the initial intelligent agent has a higher degree of understanding and satisfaction of the user's needs, and the user demand bias coefficient is larger; on the contrary, if the difference between the two is large and the correspondence relationship is chaotic, the coefficient is smaller.

[0079] For the embodiments of the present disclosure, the demand quantity of the modal data type corresponding to each historical interaction behavior, and the perceived quantity of the modal data type by the initial intelligent agent in each historical interaction behavior can be substituted into the user demand bias coefficient calculation formula to obtain the user demand bias coefficient for each historical interaction behavior. Through this formula, the demand quantity and the perceived quantity in multiple historical interaction behaviors are quantitatively calculated, and the obtained user demand bias coefficient can accurately reflect the user demand bias characteristics, providing a basis for subsequent adjustment of the intelligent agent's perception method to achieve the efficient and accurate construction of the initial intelligent agent. Among them, the formula characteristics of the user demand bias coefficient calculation formula are described as:

[0080]

[0081] In the formula, represents the user demand bias coefficient under the th interaction behavior of the user (i.e., the current interaction behavior); represents the th interaction behavior in the same conversation; j represents the jth historical interaction behavior before the th interaction behavior; represents the number of historical interaction behaviors included corresponding to the th interaction behavior; [ ] is the exponential function with the natural constant as the base; represents the demand quantity of the modal data type corresponding to the jth historical interaction behavior before the th interaction behavior of the user; represents the perceived quantity of the modal data type by the initial intelligent agent in the jth historical interaction behavior before the th interaction behavior; represents the consistency relationship of the input and output category quantities in the jth historical interaction behavior before the th interaction behavior; Indicates the maximum value between the required quantity and the perceived quantity of the modal data type in the j-th historical interaction behavior before the i-th interaction behavior.

[0082] Step 230: Based on the environmental multimodal data and the user demand bias coefficient, calculate the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior.

[0083] Among them, the multimodal perception coefficient is used to measure the comprehensive ability of the initial agent to perform environmental perception for different modal data types in each interaction behavior. By calculating this coefficient, the initial agent can understand its perception effect on different modal data, judge which modal data is more critical to meeting user needs in the current interaction, and which modal data needs to improve its perception ability. This helps the initial agent optimize its perception strategy and improve the response accuracy and efficiency to user needs.

[0084] For the embodiments of the present disclosure, calculating the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior based on the environmental multimodal data and the user demand bias coefficient in step 230 may include the following steps:

[0085] Step 230-1: Quantize the environmental multimodal data into an environmental perception vector of the initial agent for different modal data types according to the acquisition order of different modal data types in the environmental multimodal data corresponding to each interaction behavior.

[0086] Among them, the environmental multimodal data includes various modal data types such as text, voice, image, sensor data (such as temperature, humidity, etc.). When the user interacts with the initial agent, different modal data will be acquired by the initial agent in a certain order. For example, if the user first sends a voice command and then uploads a picture, the acquisition order of the voice data is before, and the acquisition order of the image data is after. This acquisition order contains rich information, which may reflect the focus of user needs or the logical relationship of information transmission. Quantization is to convert non-numerical information into a numerical form for computer processing and analysis. For environmental multimodal data, quantization is to convert it into a digital representation according to the acquisition order of different modal data types, which helps the agent understand and process these data more precisely.

[0087] The environmental perception vector is formed based on the quantified environmental multi-modal data. It describes the perception of the initial agent of the environment under different modal data types in vector form. Similarly, the following first environmental perception vector and second environmental perception vector are environmental perception vectors, which are also used to distinguish the environmental perception vectors of the current interaction behavior and the previous interaction behavior. After quantization according to the acquisition order of different modal data types, each modal data type corresponds to an environmental perception vector, which means that for each different type of modal data, the initial agent will create a dedicated vector to describe its perception of this type of data. Through the environmental perception vector, the initial agent can more intuitively analyze and compare environmental information, understand the differences and characteristics of environmental information under different interaction behaviors, and thus provide strong support for subsequent decision-making and responses.

[0088] Step 230-2: Based on the environmental perception vector and the user demand bias coefficient, calculate the multi-modal perception coefficient when the initial agent performs environmental perception for different modal data types in each interaction behavior.

[0089] For the embodiments of the present disclosure, the environmental perception vector and the user demand bias coefficient can be substituted into the multi-modal perception coefficient calculation formula to calculate the multi-modal perception coefficient when the initial agent performs environmental perception for different modal data types. The larger this coefficient, the higher the sensitivity of the user demand bias in the process of the agent's environmental perception implementation in the current interaction, and the higher the importance of the corresponding modal data in meeting the user's needs; otherwise, it is lower. Through such a calculation method, the environmental perception vector and the user demand bias coefficient are organically combined, providing a quantitative basis for the initial agent to accurately evaluate its perception ability of different modal data, and thus laying a foundation for optimizing the environmental perception method of the initial agent and improving its performance.

[0090] Among them, the formula feature description of the multi-modal perception coefficient calculation formula is:

[0091]

[0092] In the formula, represents the multi-modal perception coefficient when the initial agent performs environmental perception for the modal data type in the th interaction behavior of the user; represents the maximum-minimum normalization function; represents the th interaction behavior in the same conversation; represents the first user demand bias coefficient of the user in the th interaction behavior; represents the The second user demand bias coefficient under the previous interaction behavior corresponding to the current interaction behavior; Indicates the First environmental perception vector of the initial agent for the modal data type a under the Indicates the Second environmental perception vector of the initial agent for the modal data type a under the previous interaction behavior corresponding to the Indicates the influence of the initial agent on the order of multimodal types in the environmental perception situation. The greater the influence, the higher the sensitivity of the user demand bias to the environmental perception implementation process of the initial agent; for , if an extreme situation makes it zero, then add a non-zero constant, such as 0.01, after its position, that is .

[0093] Correspondingly, for the embodiments of the present disclosure, the embodiment steps may include: sequentially determining each interaction behavior in multiple interaction behaviors as the current interaction behavior, and determining the previous interaction behavior corresponding to the current interaction behavior, and repeatedly executing the following calculation process of the multimodal perception coefficient for the current interaction behavior until the multimodal perception coefficients under each interaction behavior in multiple interaction behaviors are obtained; wherein, the calculation process of the multimodal perception coefficient includes: determining the first environmental perception vector and the first user demand bias coefficient corresponding to the current interaction behavior, and the second environmental perception vector and the second user demand bias coefficient corresponding to the previous interaction behavior; based on the first environmental perception vector, the first user demand bias coefficient, the second environmental perception vector, and the second user demand bias coefficient, calculating the multimodal perception coefficient when the initial agent performs environmental perception for different modal data types under the current interaction behavior.

[0094] Step 240, calculating the weight coefficients of different modal data types based on the multimodal perception coefficients.

[0095] For the embodiments of the present disclosure, calculating the weight coefficients of different modal data types based on the multimodal perception coefficients in step 240 may include the following steps:

[0096] Step 240-1, obtaining multiple multimodal perception coefficients when the initial agent performs environmental perception for each modal data type under multiple interaction behaviors.

[0097] In multiple interaction behaviors, for each modal data type, the initial agent calculates the corresponding multi-modal perception coefficients. For example, in an interaction, if the user first enters text and then sends voice, the initial agent will calculate the multi-modal perception coefficients for the text modality and the voice modality in this interaction respectively; in the next interaction, if image data is involved, the initial agent will also calculate the multi-modal perception coefficient for the image modality. By obtaining multiple multi-modal perception coefficients for each modal data type under multiple interaction behaviors, the initial agent can more comprehensively understand the impact of different modal data on its own environmental perception ability in different scenarios, and then optimize its processing strategy for multi-modal data, improving the quality and efficiency of interaction with users.

[0098] Step 240-2: Calculate the weight coefficient for the corresponding modal data type based on multiple multi-modal perception coefficients.

[0099] Among them, the weight coefficient is used to represent the contribution degree of the initial agent's environmental perception for different modal data types, and it can be used to determine the priority order of the initial agent's environmental perception for different modal data types. Different modal data play different roles in meeting user needs and realizing the functions of the agent. By calculating the weight coefficient, the relative importance of each modal data in the agent's decision-making and information processing process can be clarified. For example, if the weight coefficient of a certain modal data type is high, it means that the initial agent should give higher priority to this modal data when processing information and responding to user needs; conversely, the modal data with a low weight coefficient has a relatively lower priority during processing. This can help the initial agent reasonably allocate computing resources and attention, improving the overall operation efficiency and the response accuracy to user needs.

[0100] For the embodiments of the present disclosure, multiple multi-modal perception coefficients of the initial agent for environmental perception of each modal data type can be substituted into the weight coefficient calculation formula to calculate the weight coefficient of this modal data type. Among them, the formula feature description of the weight coefficient calculation formula is:

[0101]

[0102] In the formula, represents the weight coefficient of the modal data type ; represents the number of times the user sends interaction behaviors to the initial agent in the same conversation; represents the th interaction behavior in the same conversation; represents the linear normalization function; represents the th interaction behavior of the user, and under this interaction behavior, the initial agent's environmental perception for the modal data type The multi-modal perception coefficient during environmental perception.

[0103] Step 250: Construct a pre-loaded training data set based on the weight coefficients, and use the pre-loaded training data set to correct the calling order of the training initial agent for the target API interface during environmental perception, so as to obtain the final agent that meets the user's needs.

[0104] For the embodiments of the present disclosure, the specific implementation process can refer to the relevant descriptions in step 140 of the embodiments, and will not be elaborated here.

[0105] In summary, in the technical solution of the present application, according to multiple environmental multi-modal data corresponding to multiple interaction behaviors obtained, the multi-modal perception coefficient of the initial agent for different modal data types during environmental perception is calculated for each interaction behavior. This coefficient comprehensively considers the user demand bias coefficient and the environmental perception vector. The user demand bias coefficient is obtained by analyzing factors such as the demand quantity of the modal data types included in the user's multiple input demands and the perception quantity of the modal data types by the agent, and reflects the bias situation of the user's needs. The environmental perception vector quantifies the environmental multi-modal perception results in the order of data type acquisition, and the magnitude of its difference reflects the strength of environmental correlation. Through the calculation of the multi-modal perception coefficient, the perception ability of the initial agent for different modal data can be accurately quantified, and the importance and sensitivity of different modal data types in the process of meeting user needs can be clarified. Through the calculation of the multi-modal perception coefficient, the perception ability of the initial agent for different environmental multi-modal data can be accurately quantified, and the importance and sensitivity of different environmental multi-modal data in the process of meeting user needs can be clarified. Further, based on the multi-modal perception coefficient, the weight coefficients of different modal data types are calculated, and the calling order of the training agent for the target API interface during environmental perception is optimized by using the weight coefficients, so that the finally obtained agent can improve the environmental perception ability, more efficiently obtain multi-modal information valuable to the user's needs, reduce unnecessary information acquisition, and thus improve the speed and accuracy of the agent's response.

[0106] Based on the same inventive concept as the above method, an AI large model agent construction system is also provided in an embodiment of the present invention, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above methods for constructing an AI large model agent are implemented.

[0107] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. And the above specific embodiments of the present specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.

[0108] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0109] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing an AI large model intelligent agent, characterized in that, The method includes: Using the constructed initial agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors of the user, where the multiple interaction behaviors occur in the same session created between the user and the initial agent; According to the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the multi-modal perception coefficients when the initial agent performs environmental perception for different modal data types for each interaction behavior; Calculate the weight coefficients of the different modal data types based on the multi-modal perception coefficients, where the weight coefficients are used to represent the contribution degree of the initial agent's environmental perception for the different modal data types; Construct a pre-loaded training dataset based on the weight coefficients, and use the pre-loaded training dataset to correct and train the calling order of the initial agent for the target API interface during environmental perception to obtain a final agent that meets the user's needs.

2. The method for constructing an AI large model intelligent agent according to claim 1, wherein Using the constructed initial agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors of the user, including: Using the constructed initial agent to call a neural network to analyze whether each interaction behavior input by the user is valid; If it is determined that the interaction behavior is valid, use the initial agent to call the target API interface to obtain multiple environmental multi-modal data corresponding to multiple interaction behaviors of the user.

3. The method for constructing an AI large model intelligent agent according to claim 1, wherein According to the multiple environmental multi-modal data corresponding to the multiple interaction behaviors, calculate the multi-modal perception coefficients when the initial agent performs environmental perception for different modal data types for each interaction behavior, including: Calculate the user demand bias coefficient for each interaction behavior based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors; Calculate the multi-modal perception coefficients when the initial agent performs environmental perception for different modal data types for each interaction behavior based on the environmental multi-modal data and the user demand bias coefficient.

4. The method for constructing an AI large model agent according to claim 3, wherein The calculating the user demand bias coefficient for each interaction behavior based on the multiple environmental multi-modal data corresponding to the multiple interaction behaviors includes: Sequentially determine each interaction behavior in the multiple interaction behaviors as the current interaction behavior, and determine at least one historical interaction behavior corresponding to the current interaction behavior in the multiple interaction behaviors; Based on the environmental multi-modal data corresponding to the at least one historical interaction behavior, repeatedly execute the calculation process of the user demand bias coefficient for the current interaction behavior until the user demand bias coefficients for each interaction behavior in the multiple interaction behaviors are obtained.

5. The method for constructing an AI large model intelligent agent according to claim 4, wherein, The calculation process of the user demand bias coefficient for the current interaction behavior includes: Perform input demand analysis for each of the historical interaction behaviors to determine the demand quantity of the modal data type corresponding to each of the historical interaction behaviors; Based on the environmental multi-modal data corresponding to the at least one historical interaction behavior, determine the perception quantity of the initial agent for the modal data type in each of the historical interaction behaviors; Calculate a first user demand bias coefficient for the current interaction behavior based on the required quantity corresponding to the modal data type of each of the historical interaction behaviors and the perceived quantity of the modal data type by the initial agent in each of the historical interaction behaviors.

6. The method for constructing an AI large model intelligent agent according to claim 3, wherein Based on the environmental multimodal data and the user demand bias coefficient, calculate a multimodal perception coefficient for the initial agent to perform environmental perception for different modal data types in each interaction behavior, including: Quantify the environmental multimodal data into an environmental perception vector of the initial agent for different modal data types according to the acquisition order of different modal data types in the environmental multimodal data corresponding to each interaction behavior; Based on the environmental perception vector and the user demand bias coefficient, calculate a multimodal perception coefficient for the initial agent to perform environmental perception for different modal data types in each interaction behavior.

7. The method for constructing an AI large model intelligent agent according to claim 6, wherein Based on the environmental perception vector and the user demand bias coefficient, calculate a multimodal perception coefficient for the initial agent to perform environmental perception for different modal data types in each interaction behavior, including: Successively determine each interaction behavior in the multiple interaction behaviors as the current interaction behavior, and determine the previous interaction behavior corresponding to the current interaction behavior, and repeatedly execute the following calculation process of the multimodal perception coefficient for the current interaction behavior until the multimodal perception coefficients for each interaction behavior in the multiple interaction behaviors are obtained; Wherein, the calculation process of the multimodal perception coefficient includes: Determine a first environmental perception vector and a first user demand bias coefficient corresponding to the current interaction behavior, and a second environmental perception vector and a second user demand bias coefficient corresponding to the previous interaction behavior; Based on the first environmental perception vector, the first user demand bias coefficient, the second environmental perception vector, and the second user demand bias coefficient, calculate a multimodal perception coefficient for the initial agent to perform environmental perception for different modal data types in the current interaction behavior.

8. The method for constructing an AI large model intelligent agent according to claim 1, characterized in that Calculate a weight coefficient for different modal data types based on the multimodal perception coefficient, including: Obtain multiple multimodal perception coefficients for the initial agent to perform environmental perception for each modal data type under the multiple interaction behaviors; Calculate a weight coefficient for the corresponding modal data type based on the multiple multimodal perception coefficients.

9. The method for constructing an AI large model intelligent agent according to claim 1, wherein The method further includes: Integrate API interfaces from different content providers and service providers, and the API interfaces are used to support the invocation of external service resources to obtain environmental multimodal data under different modal data types; Create an interface call relationship between the initial agent and each of the API interfaces.

10. An AI large model intelligent agent construction system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of an AI large model intelligent agent construction method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Man-machine interaction method and system based on AI large model

    CN119292453A

  • Man-machine interaction method and system based on multiple modes

    CN119806335A

  • Multi-modal large model-based intent understanding interaction method

    CN119807811A

  • Large language model dialogue intention recognition method and device, equipment and storage medium

    CN119903850A

  • System

    JP2025053720A