Multi-model cooperation method and device, storage medium and program product
Multi-model collaboration is achieved through unified memory database and protocol conversion middleware, which solves the problem of memory data silos between models and improves resource utilization efficiency and model performance.
Patent Information
- Application Number
- CN202510548980.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
In a multi-model collaborative architecture, each model cannot share or interoperate memory resources, resulting in data islands and affecting resource utilization efficiency.
Through the unified memory database and protocol conversion middleware, memory data matching the model input is obtained, prompt words are created and inputted to the target model for processing, new memory data is generated and stored in the unified memory database, and memory data sharing between different models is realized.
Reduced resource waste of multi-model collaborative methods, improved interoperability and utilization efficiency of memory data between models, and optimized model performance and user experience.
Smart Images

Figure CN120471168A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a multi-model collaboration method, device, storage medium, and program product. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, especially the popularization of large models, more and more applications have begun to integrate multiple models to achieve collaborative work on complex tasks.
[0003] Currently, in a multi-model collaborative architecture, each model generates its own specific memory resources during runtime, but the models cannot share or interoperate memory resources, resulting in data silos. Summary of the Invention
[0004] The present application provides a multi-model collaboration method, device, storage medium and program product for sharing memory data between different models.
[0005] In a first aspect, the present application provides a multi-model collaboration method, which is applied to an application server, where multiple models are deployed. The multi-model collaboration method includes:
[0006] Obtaining model input for a target model among multiple models;
[0007] According to the model input, a prompt word corresponding to the model input is obtained. The prompt word is created based on the target memory data matching the model input obtained from the unified memory database. The unified memory database is used to store the memory data generated by each model in multiple models;
[0008] The prompt word and model input are input into the target model for processing to obtain the model output corresponding to the model input.
[0009] In a possible implementation, obtaining a prompt word corresponding to the model input according to the model input includes:
[0010] According to the model input, a memory data acquisition instruction is sent to the unified protocol conversion middleware to obtain the prompt word corresponding to the model input, wherein the memory data acquisition instruction carries the model input, and the unified protocol conversion middleware is used to obtain the target memory data matching the model input from the unified memory database and create a prompt word based on the target memory data.
[0011] In a possible implementation, the method further includes:
[0012] After obtaining the model output corresponding to the model input, new memory data is generated based on the model input and model output;
[0013] Store new memory data into the unified memory database.
[0014] In one possible implementation, generating new memory data based on model input and model output includes:
[0015] Generate a memory to be stored containing model input and model output;
[0016] Set the importance level for the memory to be stored;
[0017] The importance level is added to the memory to be stored to obtain new memory data.
[0018] In one possible implementation, storing new memory data into a unified memory database includes:
[0019] Based on the new memory data, a memory data storage instruction is sent to the unified protocol conversion middleware, the memory data storage instruction carries the new memory data, and the unified protocol conversion middleware is used to store the new memory data in the unified memory database.
[0020] In a second aspect, the present application provides a multi-model collaboration method, which is applied to a unified protocol conversion middleware. The multi-model collaboration method includes:
[0021] Receive a memory data acquisition instruction sent by the application server, where the memory data acquisition instruction carries a model input, and the model input is for a target model among multiple models deployed in the application server;
[0022] Obtain target memory data that matches the model input from a unified memory database, where the unified memory database is used to store memory data generated by each of the multiple models;
[0023] Create prompt words corresponding to model input based on target memory data;
[0024] The prompt word is sent to the application server, and the application server is used to input the prompt word and the model input into the target model for processing to obtain the model output corresponding to the model input.
[0025] In one possible implementation, obtaining target memory data matching the model input from the unified memory database includes:
[0026] According to the model input, a plurality of model memory data matching the target scene corresponding to the model input is queried in the unified memory database, and the model memory data and the target model have the same operation object;
[0027] For any model memory data among the multiple model memory data, determining the priority corresponding to the model memory data;
[0028] Based on the priority, the target memory data that matches the model input is selected from multiple model memory data according to the longest context window supported by the target model.
[0029] In one possible implementation, each memory data stored in the unified memory database includes a storage time, a model name, a model input, and an importance level. Determining the priority corresponding to the model memory data includes:
[0030] Determine the model name similarity between the target model name included in the model memory data and the model name of the target model;
[0031] Determine the target model input included in the model memory data and the model input similarity of the model input;
[0032] The priority of the model memory data is determined based on the target retention time and target importance level, model name similarity, and model input similarity contained in the model memory data.
[0033] In one possible implementation, the memory data stored in the unified memory database is table item data. Based on priority and the longest context window supported by the target model, target memory data that matches the model input is selected from multiple model memory data, including:
[0034] Based on the priority, the target model memory data is selected from multiple model memory data according to the longest context window supported by the target model;
[0035] Convert the target model memory data into text data as the target memory data that matches the model input.
[0036] In a possible implementation, the method further includes:
[0037] Receive a memory data storage instruction sent by the application server, the memory data storage instruction carries new memory data, and the new memory data is generated by the application server based on the model input and model output;
[0038] Store new memory data into the unified memory database.
[0039] In a possible implementation, the new memory data is text data, and storing the new memory data in a unified memory database includes:
[0040] Add the saving time and the model name of the target model to the new memory data to obtain the text data to be stored;
[0041] The text data to be stored is processed by text conversion and the processing results are stored in a unified memory database.
[0042] In a possible implementation, the method further includes:
[0043] If the target memory data matching the model input is not obtained from the unified memory database, and / or the new memory data is not stored in the unified memory database, the corresponding steps are re-executed after the first random sleep period until success or the number of executions reaches the execution number threshold.
[0044] In a third aspect, the present application provides a multi-model collaboration method, which is applied to a unified memory database. The multi-model collaboration method includes:
[0045] receiving multiple requests sent by the application server, the multiple requests including a memory data acquisition request and / or a memory data storage request, the memory data acquisition request being used to request acquisition of the first aspect and / or various possible implementations of the first aspect, and the memory data storage request being used to request storage of new memory data in the second aspect and / or various possible implementations of the second aspect;
[0046] Lock the memory data through distributed locks to respond to a target request among multiple requests;
[0047] When the target request is satisfied, the lock is released to process other requests in the multiple requests.
[0048] In a possible implementation, the method further includes:
[0049] If it is detected that the lock has timed out and is not released, the lock is directly reclaimed;
[0050] And / or, the memory data is sorted according to the storage time and importance level, and the memory data at the end of the sort is deleted according to its own maximum capacity.
[0051] In a fourth aspect, the present application provides a multi-model collaborative system, comprising:
[0052] An application server deployed with multiple models, the application server being configured to execute the first aspect and / or various possible implementations of the first aspect;
[0053] Unified protocol conversion middleware, used to implement the second aspect and / or various possible implementation methods of the second aspect;
[0054] A unified memory database is used to store memory data generated by each model in multiple models, and to execute the third aspect and / or various possible implementation methods of the third aspect.
[0055] In a fifth aspect, the present application provides an electronic device, comprising: a memory, a processor;
[0056] Memory stores computer-executable instructions;
[0057] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect and / or the second aspect and / or the third aspect and / or various possible implementations of the first aspect and / or various possible implementations of the second aspect and / or various possible implementations of the third aspect as above.
[0058] In a sixth aspect, the present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed, they are used to implement the first aspect and / or the second aspect and / or the third aspect and / or various possible implementations of the first aspect and / or various possible implementations of the second aspect and / or various possible implementations of the third aspect as described above.
[0059] In the seventh aspect, the present application provides a computer program product, including a computer program, which, when executed, implements the first aspect and / or the second aspect and / or the third aspect and / or various possible implementations of the first aspect and / or various possible implementations of the second aspect and / or various possible implementations of the third aspect.
[0060] The multi-model collaboration method, device, storage medium and program product provided by the present application relate to the field of artificial intelligence technology. The multi-model collaboration method is applied to an application server, in which multiple models are deployed. The method includes: obtaining a model input for a target model among multiple models; obtaining a prompt word corresponding to the model input based on the model input, wherein the prompt word is created based on target memory data matching the model input obtained from a unified memory database, and the unified memory database is used to store the memory data generated by each model among multiple models; inputting the prompt word and the model input into the target model for processing to obtain a model output corresponding to the model input. The present application obtains a model input for a target model among multiple models, and based on the obtained model input, determines the prompt word corresponding to the model input, wherein the prompt word is created based on target memory data matching the model input obtained from a unified memory database storing the memory data generated by each model among multiple models, thereby sharing memory data between different models; finally, inputting the obtained prompt word and model input into the target model for processing to obtain a model output corresponding to the model input, thereby reducing the waste of resources in the multi-model collaboration method. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0062] Figure 1 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 1 ;
[0063] Figure 2 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 2 ;
[0064] Figure 3 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 3 ;
[0065] Figure 4 A schematic diagram of the structure of a multi-model collaborative system provided in an embodiment of the present application;
[0066] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application.
[0067] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0068] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0069] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0070] In recent years, artificial intelligence (AI) technology has made significant progress in areas such as natural language processing, intelligent interaction, and knowledge analysis. As application scenarios become increasingly complex, the limitations of single models in terms of knowledge coverage, analytical capabilities, and task adaptability are becoming increasingly apparent. This has led to the emergence of multi-model collaboration as a key trend in solving complex tasks. In this context, how to efficiently integrate the memory resources of different models and enable cross-model knowledge sharing and collaborative analysis has become a key technical challenge in the field of AI.
[0071] Currently, in multi-model collaborative architectures, each model typically operates independently and generates and maintains private memory resources during the analysis process, such as context information, operating status, and knowledge base updates. However, these memory resources often use model-specific storage formats and access mechanisms, resulting in the isolation of memory data between different models, making it impossible to directly share or interoperate, leading to the emergence of data silos.
[0072] In response to the above problems, this application proposes a multi-model collaboration method, which determines the prompt word corresponding to the model input based on the obtained model input, wherein the prompt word is created based on the target memory data matching the model input obtained from a unified memory database that stores the memory data generated by each model in multiple models, so as to share the memory data between different models; finally, the obtained prompt word and model input are input into the target model for processing to obtain the model output corresponding to the model input.
[0073] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0074] An embodiment of the present application provides a multi-model collaboration method, which is applied to an application server in which multiple models are deployed. Figure 1 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 1 ,like Figure 1 As shown, the method includes:
[0075] S101. Obtain model input for a target model in the deployed multiple models.
[0076] Step S101 is intended to obtain the model input of the target model. The model input is the basis for the target model to perform subsequent processing and decision-making. The form and source of the model input can be selected according to actual needs.
[0077] In one implementation, the source of model input can be relevant information directly input by the user. Specifically, the source of model input is text, images, audio, video or other forms of data directly provided to the model by the user through various input devices, where the input device can be a keyboard, mouse, touch screen, etc.
[0078] In another implementation, the source of model input may be information items recommended to the user by other models when the user clicks on them.
[0079] S102. Obtain prompt words corresponding to the model input according to the model input. The prompt words are created based on target memory data matching the model input obtained from a unified memory database. The unified memory database is used to store memory data generated by each model in the multi-model.
[0080] In some examples, based on the model input, a prompt word corresponding to the model input is obtained, including: based on the model input, a memory data acquisition instruction is sent to the unified protocol conversion middleware to obtain the prompt word corresponding to the model input, wherein the memory data acquisition instruction carries the model input, and the unified protocol conversion middleware is used to obtain target memory data matching the model input from the unified memory database, and create a prompt word based on the target memory data.
[0081] In this example, it's clear that to obtain the prompt word corresponding to the model input, a memory data retrieval instruction must be sent to the unified protocol conversion middleware based on the model input. This is the only way to obtain the prompt word corresponding to the model input. It's important to note that the memory data retrieval instruction carries information related to the model input.
[0082] Among them, the unified protocol conversion middleware is used to obtain target memory data that matches the model input from the unified memory database, and create prompt words based on the obtained target memory data.
[0083] The embodiment of the present application obtains prompt words corresponding to model input based on model input based on unified protocol conversion middleware, thereby enabling the sharing of memory data between different models.
[0084] S103: Input the prompt word and the model input into the target model for processing to obtain a model output corresponding to the model input.
[0085] In this step, it can be understood that the model input obtained in S101 and the prompt word obtained in S102 are input into the target model together for processing, so as to obtain the model output corresponding to the model input.
[0086] The form of the model output is not fixed, but depends on the received model input and the design of the target model itself. Different model inputs may trigger the target model to perform different operations, thereby producing different forms of model output.
[0087] In the first implementation, the target model uses its internal knowledge and algorithms to perform logical analysis based on the model inputs it receives, thereby drawing conclusions or predictions. For example, in a risk control prediction model, the model inputs include information such as a user's transaction history and credit score. The target model analyzes this information and outputs the user's credit risk level.
[0088] In the second implementation, the target model generates natural speech text as a response or answer to the user input. For example, the target model generates corresponding chat content based on the user input.
[0089] In the third implementation, the target model decides to call external tools or services to complete a specific task based on the model input received.
[0090] In summary, it can be understood that model input is a key factor in determining the form of model output. For example, if the model input is a question raised by the user, then the model output is more likely to be an answer to the question; if the model input is a user's control instruction, then the model output is more likely to be a call to a third-party tool.
[0091] An embodiment of the present application obtains model input for a target model among multiple models, and determines a prompt word corresponding to the model input based on the obtained model input, wherein the prompt word is created based on target memory data that matches the model input obtained from a unified memory database that stores memory data generated by each model among multiple models, thereby sharing memory data between different models; finally, the obtained prompt word and model input are input into the target model for processing to obtain a model output corresponding to the model input, thereby reducing resource waste in the multi-model collaborative method.
[0092] Based on the above embodiment, the multi-model collaboration method also includes: after obtaining the model output corresponding to the model input, generating new memory data according to the model input and model output; and storing the new memory data in a unified memory database.
[0093] In this embodiment, it can be understood that after obtaining the model output corresponding to the model input, it is necessary to store the generated new memory data into the unified memory database, wherein the new memory data is generated by the model input and the model output.
[0094] Specifically, new memory data is generated based on the model input and model output, including: generating a memory to be stored that contains the model input and model output; setting an importance level for the memory to be stored; and adding the importance level to the memory to be stored to obtain new memory data. This means that the new memory data contains the importance level of the memory to be stored. Among them, the importance level refers to a graded evaluation of the memory to be stored based on the degree of impact of the memory to be stored on the performance of the target model. The importance level can be further divided into short-term importance and long-term importance. This distinction can help the model better adapt to the memory needs of different time scales, thereby optimizing model performance and user experience.
[0095] Short-term importance refers to the short-term impact of a to-be-stored memory on model performance or user experience. Memories with high short-term importance are often closely related to the current task or session and can immediately improve model performance.
[0096] Long-term importance refers to the extent to which a to-be-stored memory impacts model performance or user experience over the long term. Highly important to-be-stored memories typically represent stable user preferences, general model knowledge, or core system rules, and can continuously improve model performance.
[0097] Short-term importance is time-sensitive and its importance may decline rapidly over time; long-term importance is stable and will not decline rapidly over time.
[0098] For example, in a conversation model, contextual information in the current conversation turn has high short-term importance because it directly affects the model's understanding of user intent and response generation; in a conversation model, user personal information has high long-term importance because it can help the model better understand the user's background and needs.
[0099] Furthermore, the importance level of a memory to be stored is not static. Instead, the short-term and long-term importance of the memory to be stored can be dynamically updated over time and based on changes in user behavior. For example, the importance level of a memory to be stored can be adjusted based on information such as the frequency of use of the memory to be stored and user feedback.
[0100] In summary, by setting the importance level for the memory to be stored and adding the importance level to the memory to be stored, new memory data is obtained, which can better manage and utilize the memory data, thereby bringing significant benefits in improving the performance of the target model, optimizing the user experience, and reducing computing costs.
[0101] Furthermore, after obtaining the new memory data carrying the importance level, the new memory data needs to be stored in the unified memory database. Storing the new memory data in the unified memory database includes: based on the new memory data, sending a memory data storage instruction to the unified protocol conversion middleware, the memory data storage instruction carrying the new memory data, and the unified protocol conversion middleware being used to store the new memory data in the unified memory database.
[0102] It is understood that if new memory data is to be stored in the unified memory database, a memory data storage instruction must be sent to the unified protocol conversion middleware based on the new memory data, so that the new memory data can be stored in the unified memory database. It should be noted that the memory data storage instruction carries the new memory data.
[0103] Among them, the unified protocol conversion middleware is used to store new memory data into the unified memory database.
[0104] One thing that needs to be explained is that before storing new memory data in a unified memory database, it is necessary to first determine the storage quantity of the memory data contained in the unified memory database; if the storage quantity is less than the set maximum storage quantity, the new memory data is saved in the unified memory database based on the unified protocol conversion middleware; if the storage quantity is equal to the set maximum storage quantity, the memory data to be deleted is determined in the unified memory database; the memory data to be deleted is deleted from the unified memory database, and the new memory data is saved in the unified memory database based on the unified protocol conversion middleware.
[0105] The embodiment of the present application saves new memory data into a unified memory database based on a unified protocol conversion middleware, thereby providing a data basis for sharing memory data between different models.
[0106] Based on the above embodiments, an embodiment of the present application provides a multi-model collaboration method applied to unified protocol conversion middleware. Figure 2 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 2 ,like Figure 2 As shown, the method includes:
[0107] S201: Receive a memory data acquisition instruction sent by an application server. The memory data acquisition instruction carries a model input. The model input is for a target model among multiple models deployed in the application server.
[0108] In S201, it can be understood that the application server sends a memory data acquisition instruction to the unified memory conversion middleware according to the model input, and the unified protocol conversion middleware receives the memory data acquisition instruction sent by the application server, wherein the memory data acquisition instruction carries the model input.
[0109] S202. Obtain target memory data that matches the model input from a unified memory database, where the unified memory database is used to store memory data generated by each of the multiple models.
[0110] After receiving the memory data acquisition instruction through S201, the unified protocol conversion middleware acquires the target memory data that matches the model input from the unified memory database. Acquiring the target memory data that matches the model input from the unified memory database includes: according to the model input, querying and obtaining multiple model memory data that match the target scene corresponding to the model input in the unified memory database, and the model memory data and the target model have the same operation object; for any model memory data in the multiple model memory data, determining the priority corresponding to the model memory data; based on the priority, according to the longest context window supported by the target model, selecting the target memory data that matches the model input from the multiple model memory data. This means that the multiple model memory data that match the target scene corresponding to the model input obtained by querying and have the same operation object as the target model. For example, assuming that user A inputs relevant information to the target model, it is necessary to query the memory data saved by user A in the unified memory database, and query the multiple model memory data that match the target scene corresponding to the model input from the memory data saved by user A.
[0111] It should be noted that each memory data stored in the unified memory database contains a storage time, a model name, a model input and an importance level. For any model memory data, the priority corresponding to the model memory data is determined, including: determining the target model name contained in the model memory data, and the model name similarity with the model name of the target model; determining the target model input contained in the model memory data, and the model input similarity with the model input; determining the priority corresponding to the model memory data based on the target storage time and target importance level, model name similarity and model input similarity contained in the model memory data. It can be understood that the priority corresponding to the model memory data is obtained by constructing a priority weighting function, and the priority weighting function is composed of the dimensions of model name similarity, target storage time, model input similarity, and target importance level, wherein the target importance level includes short-term importance level and long-term importance level. Specifically, the priority weighting function can be expressed by the following formula:
[0112] ;
[0113] Among them, f represents the priority weighting function; time represents the target storage time; name - simi indicates the model name similarity; input - simi represents the model input similarity; short - Import indicates the short-term importance level; long - Import indicates the long-term importance level.
[0114] The embodiment of the present application can provide a data basis for subsequent acquisition of target memory data that matches the model input by determining the priority corresponding to the model memory data.
[0115] After determining the priorities corresponding to the model memory data, the longest context window supported by the target model is determined based on the priorities, and target memory data that matches the model input is selected from the multiple model memory data. It will be appreciated that after determining the priorities, the multiple model memory data are sorted based on the priorities. Then, the target memory data that matches the model input is selected from the sorted multiple model memory data. This facilitates the rapid and accurate determination of target memory data that matches the model input.
[0116] S203: Create prompt words corresponding to the model input according to the target memory data.
[0117] In this step, it is understood that after determining the target memory data in S202, prompt words corresponding to the model input need to be created. The prompt words are used to guide the target model in generating the desired output. By incorporating the target memory data into the prompt words, the target model's contextual understanding ability can be enhanced, improving the target model's performance.
[0118] For example, constructing the prompt words corresponding to the model input can be performed by the following steps:
[0119] Convert the target memory data into text form. If the target memory data is already in text form, use the target memory data directly. If the target memory data is structured data or vector representation, convert the target memory data into text form.
[0120] The target memory data in text form is inserted into the predefined prompt word template to construct the complete prompt word. The prompt word template can be customized according to different tasks and models.
[0121] S204: Send the prompt word to the application server, and the application server is used to input the prompt word and the model input into the target model for processing to obtain a model output corresponding to the model input.
[0122] After completing the construction of the prompt word in S203, the unified protocol conversion middleware sends the prompt word to the application server. After receiving the prompt word, the application server inputs the prompt word and the model input into the target model for processing, thereby obtaining the model output corresponding to the model input.
[0123] The embodiment of the present application utilizes a unified protocol conversion middleware to obtain target memory data that matches the model input from a unified memory database, thereby enabling the sharing of memory data between different models, avoiding the existence of data islands, and reducing resource waste.
[0124] In some embodiments, the memory data stored in the unified memory database is table item data, and based on the priority and the longest context window supported by the target model, target memory data that matches the model input is selected from multiple model memory data, including: based on the priority and the longest context window supported by the target model, target model memory data is selected from multiple model memory data; and the target model memory data is converted into text data as the target memory data that matches the model input. In these embodiments, it can be understood that when the memory data stored in the unified memory database is table item data, determining the target memory data requires performing a vector-to-text operation on the determined target model memory data to obtain text data, and using the obtained text data as the target memory data that matches the model input.
[0125] The embodiment of the present application can ensure that different types of models can utilize shared memory data by converting table item data into text data, thereby avoiding the situation where the model cannot utilize memory data due to incompatible memory data formats, and improving the interoperability of memory data between different models.
[0126] Based on the above embodiment, the multi-model collaboration method also includes: receiving a memory data storage instruction sent by an application server, the memory data storage instruction carries new memory data, and the new memory data is generated by the application server based on the model input and model output; storing the new memory data in a unified memory database.
[0127] In this embodiment, it can be understood that the unified protocol conversion middleware also needs to receive a memory data storage instruction sent by the application server, wherein the memory data storage instruction carries new memory data, and the new memory data is generated by the application server based on the model input and model output. After receiving the memory data storage instruction, the unified protocol conversion middleware stores the new memory data carried by the memory data storage instruction in a unified memory database. It should be noted that the new memory data is text data, and storing the new memory data in the unified memory database includes: adding the storage time and the model name of the target model to the new memory data to obtain the text data to be stored; performing text conversion processing on the text data to be stored, and storing the processing results in the unified memory database.
[0128] When storing new memory data, it is necessary to add the storage time and the model name of the target model to the new memory data to obtain the text data to be stored. Since the new memory data is text data, it is necessary to perform text conversion processing on the text data to be stored, and then store the processing results in the unified memory database. The storage principle of the unified memory database has been explained in detail in the previous embodiment, so it will not be repeated here in this embodiment of the present application.
[0129] The embodiment of the present application can better retain the semantic information of the memory data and avoid information loss by performing text conversion operations on the new memory data and storing the processing results in a unified memory database.
[0130] Furthermore, the multi-model collaboration method also includes: if the target memory data that matches the model input is not obtained from the unified memory database, and / or the new memory data is not stored in the unified memory database, then after the first random sleep period, the corresponding steps are re-executed until success or the number of executions reaches the execution number threshold. It can be understood that when the unified protocol conversion middleware fails to obtain the target memory data that matches the model input from the unified memory database, and / or the new memory data is not stored in the unified memory database, it is necessary to randomly sleep the unified protocol conversion middleware for the first period, then re-execute the corresponding steps until the corresponding steps are successfully executed or the number of executions reaches the execution number threshold, then stop all steps. In other words, when the unified protocol conversion middleware fails to call due to a conflict, the unified protocol conversion middleware will randomly sleep for a period of time according to the demand and then re-apply until the application is successful or the number of applications exceeds the maximum number of applications, wherein the random sleep period does not exceed 1 second and the maximum number of applications is 10.
[0131] The embodiment of the present application can improve the storage success rate and reduce resource usage by randomly dormant the unified protocol conversion middleware for a first period of time when the target memory data that matches the model input is not obtained from the unified memory database and / or when the new memory data is not stored in the unified memory database, thereby re-executing the corresponding steps. This mechanism can enable the multi-model collaborative method to better cope with various abnormal situations.
[0132] Based on the above embodiments, an embodiment of the present application provides a multi-model collaboration method applied to a unified memory database. Figure 3 Schematic diagram of the process of the multi-model collaborative method provided in this application Figure 3 ,like Figure 3 As shown, the method includes:
[0133] S301. Receive multiple requests sent by the application server, where the multiple requests include a memory data acquisition request and / or a memory data storage request. The memory data acquisition request is used to request acquisition of target memory data in the method described in the above embodiment, and the memory data storage request is used to request storage of new memory data in the method described in the above embodiment.
[0134] In this step, it is understood that the unified memory database receives multiple requests from the application server, including memory data acquisition requests and / or memory data storage requests. The memory data acquisition request is used to request the target memory data described in the previous embodiment, and the memory data storage request is used to request the new memory data described in the previous embodiment.
[0135] S302: Lock the memory data by means of a distributed lock to respond to a target request among the multiple requests.
[0136] In this step, it can be understood that when the unified memory database receives multiple requests, it locks the memory data using a distributed lock, ensuring that only one request, i.e., a target request, will receive the memory data. Distributed locks are a mechanism used in distributed systems to control concurrent access to shared resources by multiple nodes. Distributed locks ensure that only one node can access shared memory data at a time, thus avoiding data contention and inconsistency.
[0137] S303: When the target request is satisfied, release the lock to process other requests in the multiple requests.
[0138] In this step, it can be understood that when the target request is satisfied, the lock is released, and then other requests among the multiple requests can be processed.
[0139] The embodiment of the present application uses a distributed lock to ensure that only one node can access shared memory data at the same time, thereby avoiding data competition and data inconsistency problems.
[0140] Based on the above embodiment, it also includes: if it is detected that the lock has timed out and is not released, the lock is directly reclaimed; and / or the memory data is sorted according to the storage time and importance level, and the memory data at the end of the sort is deleted according to its own maximum capacity.
[0141] In this embodiment, it is understood that when it is detected that the lock added in S302 has timed out and is not released, the lock is directly reclaimed. For example, when the lock is not released due to the target model, the unified memory database will directly reclaim the lock. In other words, the embodiment of the present application introduces an automatic lock release mechanism upon timeout. When a node acquires a distributed lock, if it fails to actively release the lock within the preset timeout period, the unified memory database will automatically reclaim the lock so that other nodes can access the shared memory data.
[0142] Furthermore, the unified memory database will sort the memory data according to the storage time and importance level, and when the unified memory database is in an idle state, it will delete the memory data at the end of the sorting according to its maximum capacity.
[0143] The embodiments of the present application effectively avoid deadlock by directly reclaiming the lock upon detecting that the lock has timed out and has not been released. Even if the node acquiring the lock fails or experiences a program anomaly, the lock will automatically be released after the timeout, thus preventing other nodes from waiting and causing a deadlock. Furthermore, the embodiments of the present application can reduce the storage space occupied by the memory data and lower storage costs by deleting the memory data at the end of the sort according to its maximum capacity.
[0144] It should be noted that the unified protocol conversion middleware and multiple models can be deployed in different application servers or in the same application server.
[0145] Figure 4 This is a schematic diagram of the structure of the multi-model collaborative system provided in the embodiment of the present application, such as Figure 4 As shown, the multi-model collaborative system 400 provided in this embodiment includes:
[0146] An application server 401 deployed with multiple models, the application server being used to execute the method described in the previous embodiment;
[0147] Unified protocol conversion middleware 402, used to execute the method described in the previous embodiment;
[0148] The unified memory database 403 is used to store the memory data generated by each model in the multiple models and to execute the method described in the previous embodiment.
[0149] The multi-model collaborative system provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principles and technical effects are similar and will not be described in detail in this embodiment.
[0150] Next, we will give an example of how to use the multi-model collaborative system provided by the embodiment of this application. The strategies for each model are as follows: every time the user inputs, the memory acquisition method is executed, the capabilities provided by the unified protocol conversion middleware are called, and unified, context-appropriate, and consistent with the current scene memory data is obtained; the local or cloud-based large model analyzes, responds, and calls third-party tools based on the acquired memory data and user input; after a single round of question and answer, the memory preservation method is executed to set the short-term importance and long-term importance of the current memory data, and the capabilities provided by the unified protocol conversion middleware are called to save the new memory data generated by the current question and answer in the unified memory database, that is, save it locally.
[0151] The strategy for the unified protocol conversion middleware is as follows: provide memory data acquisition and memory data storage methods for large models; convert data formats, convert table data into text data, and convert text data into table data.
[0152] The steps for obtaining memory data are as follows:
[0153] Here are the steps:
[0154] 1. Obtain the memory data saved by the current user in the unified memory database;
[0155] 2. Construct a priority weighting function based on five dimensions: large model storage time, large model name similarity, user input similarity, short-term importance, and long-term importance, and perform priority sorting;
[0156] 3. Get the first n table item data based on the longest context window supported by the local or cloud large model
[0157] 4. Construct the n table item data into text data, and further construct it into prompt words, and provide it to the large model.
[0158] The steps for saving memory data are as follows:
[0159] 1. Generate new memory data based on user input and the large model’s answer text;
[0160] 2. Add two items, question and answer time and large model name, to the new memory data to obtain the memory data to be stored;
[0161] 3. Perform text transformation operations on the stored memory data, i.e. embedding operations;
[0162] 4. The memory data to be stored after the text-to-quantity operation is completed is saved in a unified memory database.
[0163] It's important to note that when the unified protocol conversion middleware receives requests to save or retrieve memory data from multiple large models, it executes all requests and calls the unified memory database. Data consistency is guaranteed by the database itself. If a call fails due to a conflict, the unified protocol conversion middleware will randomly hibernate the request for a period of time and then reapply until the request succeeds or the maximum number of requests is exceeded.
[0164] For the unified memory database, the strategy is as follows:
[0165] 1. When receiving multiple requests, the memory data is locked through distributed locks so that only one request will get the resource. When this request is satisfied, the lock is released and other requests can be processed.
[0166] 2. When the lock is not released for a long time, for example, when a large model crashes and the lock is not released, the unified memory database will pre-set a countdown, for example, set the countdown to 3 seconds, and directly reclaim the lock when the timeout expires;
[0167] 3. When idle, sort the memory data according to its retention time and long-term importance, and delete the n table items at the end of the sort based on its maximum capacity.
[0168] In summary, the embodiments of the present application have designed a unified and standardized memory resource management protocol that can support memory resource sharing and interoperability between models of different manufacturers and different architectures. Furthermore, the embodiments of the present application also define unified interface standards and data formats to ensure that different models can understand and use the same memory resources, thereby solving the problem of resource competition and data consistency caused by multiple models accessing or modifying the same memory resources at the same time in the scenario of multi-model collaboration. At the same time, a protocol conversion mechanism is provided to be compatible with the private protocols of different manufacturers to achieve seamless docking.
[0169] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 500 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 500 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.
[0170] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that the at least one processor 501 performs the above method.
[0171] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0172] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0173] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0174] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0175] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0176] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0177] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0178] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0179] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.
[0180] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0181] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0182] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0183] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0184] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.
Claims
1. A multi-model collaborative method, characterized in that: Applied to an application server, where multiple models are deployed, the multi-model collaboration method includes: Obtaining a model input for a target model among the multiple models; Obtaining, based on the model input, a prompt word corresponding to the model input, the prompt word being created based on target memory data matching the model input obtained from a unified memory database, the unified memory database being used to store memory data generated by each of the multiple models; The prompt word and the model input are input into the target model for processing to obtain a model output corresponding to the model input.
2. The method according to claim 1, characterized in that The step of obtaining a prompt word corresponding to the model input according to the model input includes: According to the model input, a memory data acquisition instruction is sent to the unified protocol conversion middleware to obtain a prompt word corresponding to the model input, wherein the memory data acquisition instruction carries the model input, and the unified protocol conversion middleware is used to obtain target memory data matching the model input from the unified memory database, and create the prompt word according to the target memory data.
3. The method according to claim 1 or 2, characterized in that Also includes: After obtaining the model output corresponding to the model input, generating new memory data according to the model input and the model output; The new memory data is stored in the unified memory database.
4. The method according to claim 3, characterized in that The generating new memory data according to the model input and the model output includes: generating a memory to be stored comprising the model input and the model output; Setting an importance level for the memory to be stored; The importance level is added to the memory to be stored to obtain the new memory data.
5. The method according to claim 3, characterized in that Storing the new memory data in the unified memory database includes: Based on the new memory data, a memory data storage instruction is sent to the unified protocol conversion middleware, the memory data storage instruction carries the new memory data, and the unified protocol conversion middleware is used to store the new memory data in the unified memory database.
6. A multi-model collaborative method, characterized in that: Applied to unified protocol conversion middleware, the multi-model collaboration method includes: Receive a memory data acquisition instruction sent by an application server, wherein the memory data acquisition instruction carries a model input, and the model input is for a target model among multiple models deployed in the application server; Acquire target memory data that matches the model input from a unified memory database, wherein the unified memory database is used to store memory data generated by each model in the multiple models; Creating a prompt word corresponding to the model input according to the target memory data; The prompt word is sent to the application server, and the application server is used to input the prompt word and the model input into the target model for processing to obtain a model output corresponding to the model input.
7. The method according to claim 6, characterized in that The acquiring target memory data matching the model input from the unified memory database includes: According to the model input, querying in the unified memory database to obtain a plurality of model memory data that matches the target scene corresponding to the model input, wherein the model memory data and the target model have the same operation object; For any model memory data among the plurality of model memory data, determining a priority corresponding to the model memory data; Based on the priority, target memory data that matches the model input is selected from the multiple model memory data according to the longest context window supported by the target model.
8. The method according to claim 7, characterized in that Each memory data stored in the unified memory database includes a storage time, a model name, a model input, and an importance level. The step of determining the priority corresponding to the model memory data includes: Determine the model name similarity between the target model name included in the model memory data and the model name of the target model; determining a model input similarity between a target model input included in the model memory data and the model input; The priority corresponding to the model memory data is determined according to the target preservation time and target importance level contained in the model memory data, the model name similarity, and the model input similarity.
9. The method according to claim 7, characterized in that The memory data stored in the unified memory database is table item data, and the selecting, based on the priority and according to the longest context window supported by the target model, target memory data that matches the model input from the plurality of model memory data includes: Based on the priority, and according to the longest context window supported by the target model, target model memory data is selected from the plurality of model memory data; The target model memory data is converted into text data as target memory data matching the model input.
10. The method according to any one of claims 6 to 9, characterized in that Also includes: receiving a memory data storage instruction sent by the application server, wherein the memory data storage instruction carries new memory data, and the new memory data is generated by the application server according to the model input and the model output; The new memory data is stored in the unified memory database.
11. The method according to claim 10, characterized in that The new memory data is text data, and storing the new memory data in the unified memory database includes: Adding a storage time and a model name of the target model to the new memory data to obtain text data to be stored; The text data to be stored is subjected to text conversion processing, and the processing result is stored in the unified memory database.
12. The method according to claim 10, characterized in that Also includes: If the target memory data matching the model input is not obtained from the unified memory database, and / or the new memory data is not stored in the unified memory database, the corresponding steps are re-executed after the first random sleep period until success or the number of executions reaches the execution number threshold.
13. A multi-model collaborative method, characterized in that: Applied to a unified memory database, the multi-model collaboration method includes: receiving a plurality of requests sent by an application server, the plurality of requests comprising a memory data acquisition request and / or a memory data storage request, the memory data acquisition request being used to request acquisition of target memory data in the method according to any one of claims 6 to 12, and the memory data storage request being used to request storage of new memory data in the method according to any one of claims 10 to 12; Locking the memory data by means of a distributed lock to respond to a target request among the multiple requests; When the target request is satisfied, the lock is released to process other requests in the plurality of requests.
14. The method according to claim 13, characterized in that Also includes: If it is detected that the lock has timed out and is not released, the lock is directly reclaimed; And / or, the memory data is sorted according to the storage time and importance level, and the memory data at the end of the sort is deleted according to its own maximum capacity.
15. A multi-model collaborative system, characterized in that: include: An application server deployed with multiple models, the application server being configured to execute the method according to any one of claims 1 to 5; Unified protocol conversion middleware, configured to execute the method according to any one of claims 6 to 12; A unified memory database is used to store the memory data generated by each model in the multiple models and to execute the method according to claim 13 or 14.
16. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1-5, 6-12, 13 or 14.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 or 6 to 12 or 13 or 14 when executed by a processor.
18. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 5 or 6 to 12 or 13 or 14 when executed by a processor.
Citation Information
Cited By
Memory-based man-machine interaction method, electronic equipment, medium and program product
CN121560166A