Information processing method applied to large model system and related device

By introducing module-level information storage and task disassembly mechanisms in the large model system and cache processing results, the delay and resource consumption problems of the large model system are solved, and more efficient information processing and diversified output are achieved.

CN120297416APending Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396877.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The processing delay and computing resource consumption of large-model systems are high, making it difficult to effectively reduce and maintain the diversity of output results.

Method used

The module-level information storage mechanism is adopted to cache the processing results of the target module, and the repeated calculation is reduced for multiplexing historical information, especially the processing results of small models without output randomness, and combined with task disassembly and multimodal information processing, the utilization rate of information sets is optimized.

Benefits of technology

Effectively reduce the processing delay of large-model systems, save computing resources, maintain the diversity and accuracy of output results, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297416A_ABST
    Figure CN120297416A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method applied to a large model system and a related device. The invention relates to the technical field of computers, in particular to the technical fields of artificial intelligence, deep learning, large models, computer vision, voice technologies, intelligent search and the like. According to the specific implementation scheme, for a target module in a plurality of processing modules in the large model system, under the condition that input information of the target module is obtained, the input information is searched for in a stored information set; under the condition that the input information is found, obtaining a processing result of the target module pre-stored in the information set on the input information; and using the processing result as input information of a next processing module of the target module in the large model system, and sending the input information to the next processing module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to technologies such as artificial intelligence, deep learning, large models, computer vision, speech technology, and intelligent search. Background Art

[0002] A large model refers to a large-scale deep learning model in the field of artificial intelligence that usually has a large number of parameters. These large models can learn complex patterns and features in the data through training on a large scale of data, and thus demonstrate powerful performance and generalization ability in various tasks.

[0003] A large model system, that is, a processing system built based on a large model, can respond to user input, generate corresponding content and feedback it to the user. Summary of the Invention

[0004] The present disclosure provides an information processing method and related device applied to a large model system.

[0005] According to one aspect of the present disclosure, there is provided an information processing method applied to a large model system, including:

[0006] When obtaining the input information of a target module among multiple processing modules in the large model system, searching for the input information in the stored information set;

[0007] When the input information is found, obtaining the processing result of the target module on the input information pre-stored in the information set;

[0008] Taking the processing result as the input information of the next processing module of the target module in the large model system and sending it to the next processing module.

[0009] According to another aspect of the present disclosure, there is provided an information processing device applied to a large model system, including:

[0010] A search module, configured to search for input information in a stored information set when obtaining the input information of a target module among multiple processing modules in the large model system;

[0011] An obtaining module, configured to obtain the processing result of the target module on the input information pre-stored in the information set when the input information is found;

[0012] A processing module, configured to take the processing result as the input information of the next processing module of the target module in the large model system and send it to the next processing module.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 is a schematic flowchart of an information processing method applied to a large model system according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic diagram of determining input information of a target module based on the conversation content of the current round of conversation according to an embodiment of the present disclosure;

[0023] Figure 3 is a module structure diagram included in a target model according to an embodiment of the present disclosure;

[0024] Figure 4 is another schematic flowchart of an information processing method applied to a large model system according to an embodiment of the present disclosure;

[0025] Figure 5 is another schematic diagram of the information processing flow of a large model system according to an embodiment of the present disclosure;

[0026] Figure 6 is another schematic diagram of the information processing flow of a large model system according to an embodiment of the present disclosure;

[0027] Figure 7It is a schematic structural diagram of an information processing device applied to a large model system according to an embodiment of the present disclosure;

[0028] Figure 8 It is another schematic structural diagram of an information processing device applied to a large model system according to an embodiment of the present disclosure;

[0029] Figure 9 It is a block diagram of an electronic device for implementing an information processing method applied to a large model system according to an embodiment of the present disclosure. Detailed implementation manners

[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.

[0032] Terms such as "first" and "second" in the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included. A method, system, product, or device is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0033] In order to handle complex tasks and improve the quality of the results generated by large models, large model systems are often built based on multiple processing modules. It can be understood that these processing modules can include models built based on neural networks, or models built in other ways. Among them, models built based on neural networks usually have a much smaller number of parameters than large models, and their main role is to assist large models in performing task reasoning and analysis, improving the accuracy and standardization of the results generated by large models.

[0034] It can be seen that a large model system is a huge and complex system, which occupies a lot of computing resources to complete tasks. Due to the extremely high complexity of large model systems, their processing latency and resource consumption costs are often highly concerned.

[0035] In view of this, embodiments of the present disclosure provide an information processing method applied to a large model system, which can effectively reduce the processing latency of the large model system and save computing resources.

[0036] As Figure 1 shown, it is a schematic flowchart of the method, including the following content:

[0037] S101, for a target module among multiple processing modules in the large model system, when obtaining the input information of the target module, search for the input information in the stored information set.

[0038] Among them, after processing the historical input, the target module can store the historical input and its corresponding processing result in the information set. So that in subsequent tasks, when receiving the same input, the processing result corresponding to the input can be reused.

[0039] To improve the processing efficiency, the information set can be stored using a cache.

[0040] It can be understood that multiple processing modules construct a large model system according to a certain topological structure. The target module among them can be part of the processing modules in the large model system, and this part of the processing modules can meet the preset conditions. For example, the preset conditions include that the target module is a model in the preprocessing link before the large model in the large model system. Each target module corresponds to its own information set. Embodiments of the present disclosure can be described by taking one target model as an example, and it can be understood that the operating principles of other target models are the same.

[0041] S102, when the input information is found, obtain the processing result of the target module for the input information pre-stored in the information set.

[0042] S103, use the processing result as the input information of the next processing module of the target module in the large model system and send it to the next processing module.

[0043] It can be understood that if the next processing module is also a model that meets the preset conditions, then the next processing module serves as a new target module and executes the operations of S101 - S103.

[0044] In summary, in the embodiments of the present disclosure, for a large model system with extremely high complexity, some of its target modules can adopt a module-level information storage mechanism so that the processing results of the information that has been processed in the previous task can be reused. Based on this, the operation of reprocessing the input information by the target module can be omitted, especially for some small models that are relatively small compared to the large model based on neural networks. The computational amount of this type of small model should not be ignored. Especially when the large model system has multiple such small models, this type of small model can reuse the processing results of the previous task, saving the entire reasoning process of the small model, and can generally reduce the processing delay of the large model system. Since the output results of the template module are stored in the information set, the data volume of the output results is relatively small, and the storage space requirements are not high. It can also save the consumption of computing resources brought by the processing process and reduce the cost of computing resources. In addition, when the solution provided by the embodiment of the present disclosure is used in the preprocessing link in the large model system, it is possible to reduce the delay while retaining the diversity requirements of the large model generated content, ensure the quality of the processing results of the large model system, and improve the user experience.

[0045] In some embodiments, while reducing processing latency and saving computing resources, it is also necessary to retain the diversity of the output results of the large model system as much as possible, that is, for the same input, the large model system can provide different output results. In this regard, in the embodiments of the present disclosure, the preset conditions satisfied by the target module may include the characteristic that the target module has no output randomness. This type of target module satisfies at least one of the following conditions:

[0046] Condition 1) The output diversity control parameters of the target module are set so that the output results of the target module for the same input data are consistent; the output diversity control parameters are used to control the output results of the model to meet the diversity requirements.

[0047] Wherein, when the target module has an output diversity control parameter, the output diversity control parameter is used to control the output result of the target module to have randomness or diversity. The input diversity control parameter may be, for example, a sampling parameter, a temperature parameter, and the like.

[0048] However, these parameters can be set to reduce the randomness of the output results of the target module, or even make them stable and deterministic. For example, the closer the temperature parameter is to 0, the lower the randomness. In the sampling parameter top-k, if k is 1, the randomness determined only by the sampling parameters is 0, and the output result is deterministic.

[0049] In some large model systems, by setting diversity control parameters, it is possible to ensure that the large model has better output and avoid some erroneous outputs.

[0050] Condition 2), the target module does not have output diversity control parameters.

[0051] That is, for the same input, the target module itself will output the same processing result. Such a target module naturally satisfies the characteristic of no output randomness.

[0052] In the embodiments of the present disclosure, a processing module that satisfies no output randomness is used as the target module. In particular, a processing module that originally has output randomness also has a considerable complexity. When the output diversity control parameter is set so that such a module can also reuse historical processing results, the processing delay of the large model system can be effectively reduced, and the processing efficiency of the large model system can be improved. Moreover, by reducing the processing process of the processing module with no output randomness, not only can computing resources be saved, but also the diversity of the output results of the large model system can be ensured as much as possible.

[0053] In the embodiments of the present disclosure, in order to improve the utilization rate of its historical processing results, that is, the information set, for the target model, the processing tasks of the large model system are disassembled, and then the input information of the target module is constructed by using the disassembled tasks.

[0054] In some embodiments, obtaining the input information of the target module by means of task disassembly can be implemented as: when the original information before inputting into the target module includes multimodal information, the information of the target modality applicable to the target module is screened out from the multimodal information to obtain the input information of the target module.

[0055] Wherein, the original information is the information before being processed by the target module. If the target module is the first processing module in the large model system, the original information is the information to be processed provided by the target object.

[0056] If the target module is an intermediate processing module other than the first processing module in the large model system, the original information is the processing result of the previous processing module of the target module.

[0057] For example, if the original multimodal information includes a text file and an image, the text file and the image need to be preprocessed separately. At this time, the text file can be extracted from the original information as the input information of the text processing module (i.e., the target module). The image can be extracted from the original information as the input information of the image processing module (i.e., another target module).

[0058] It can be seen from this that when the original information is used as the complete input information, if there is a slight change in the original information, the matching historical input cannot be found in the information set, and thus the historical processing result cannot be obtained and reused.

[0059] Compared with using the original information as the complete input information, in the embodiments of the present disclosure, through task decomposition, when the information of some modalities in the original information is the same, the information of this part of the modality can still reuse the historical processing results of the target module, thereby improving the utilization rate of the information set and the processing efficiency of the large model system.

[0060] In the embodiments of the present disclosure, the multi-modal information includes at least one of the following information: text information, audio information, image information, etc.

[0061] It can be understood that if the input of the target module includes information of multiple modalities, the target modality can also include information of multiple modalities.

[0062] In some embodiments, obtaining the input information of the target module by means of task decomposition can also be implemented as: in the case of multi-round conversations, determining the input information of the target module based on the conversation content of the current round of conversation.

[0063] That is, in the case of multi-round conversations, tasks are decomposed according to the conversation rounds. Each round of conversation corresponds to a task, and each task corresponds to a single input information for retrieval in the information set.

[0064] In the embodiments of the present disclosure, in the scenario of multi-round conversations, each round of conversation is separately decomposed into a task. Compared with jointly constructing the input information for multi-round conversations, it can improve the utilization rate of the information set, improve the processing efficiency of the large model system, and save computing resources.

[0065] In some embodiments, determining the input information of the target module based on the conversation content of the current round of conversation can be implemented as: when the input information of the target module needs to include the original conversation content, obtaining the conversation content of the current round of conversation as the input information of the target module.

[0066] For example, the content detection module needs to detect whether the conversation content provided by the target object meets the corresponding requirements. Therefore, it is necessary to detect the original conversation content. During implementation, each round of conversation content can be detected separately, which can improve the retrieval hit rate of the information set and the utilization rate of the information set.

[0067] For another example, some tasks need to determine whether to perform information retrieval. For the same conversation content, the judgment results are almost the same in a short period of time. Therefore, the historical result utilization rate of the retrieval classification module can be reused to improve the processing efficiency of the large model system.

[0068] In the embodiments of the present disclosure, for the target module that needs to process the original conversation content, using the corresponding single-round conversation content to construct the input information can improve the hit rate of the information set, reduce the processing delay of the large model system, and save computing resources compared with using multi-round conversation content.

[0069] In some embodiments, determining the input information of the target module based on the conversation content of the current round of conversation can also be implemented as follows:

[0070] Step A1, when the target module is a processing module for intermediate processing results, determine the input source of the target module.

[0071] Step A2, construct the input information of the target module based on the output results of at least one processing module corresponding to the input source for the conversation content of the current round of conversation.

[0072] For example Figure 2 As shown, the large model system includes Module 1, Module 2, Module 3, Module 4, Module 5, and a large language model. Among them, the information to be processed input by the target object is processed by Module 1 and Module 2 respectively. The output result of Module 1 is input to Module 3 for processing, the output result of Module 2 is input to Module 4 for processing, the output results of Module 3 and Module 4 are input to Module 5 for processing, and the output result of Module 5 is handed over to the large language model for inference and analysis to obtain the final processing result.

[0073] Among them, in the process from the information to be processed input by the target object to obtaining the processing result of the large language model, it passes through the processing of multiple modules in the middle, and the output result of each module can be called an intermediate processing result.

[0074] Taking Module 5 as an example, its corresponding input sources include the output results of Module 3 and Module 4. Then, the input information of Module 5 is the information constructed based on the output results of Module 3 and Module 4.

[0075] In some embodiments, the output results of Module 3 and Module 4 can be concatenated to obtain the input information of Module 5. Of course, in some business scenarios, it is necessary to optimize the output results of Module 3 and / or Module 4, such as adding special characters, to improve the accuracy of the output results of the large model. In this case, special characters can be added to the output results of Module 3 and the output results of Module 4 to obtain the input information of Module 5.

[0076] In the embodiments of the present disclosure, for the intermediate processing module in the large model system (i.e., the target module for intermediate processing results), the input information for query can be accurately constructed according to the topological structure of the large model system, so as to improve the utilization rate of the information set of the target model.

[0077] Under normal circumstances, for a large model system, the information to be processed input by the target object can be task-split so that the processing module serving as the entry in the large model system can construct the input information according to the split tasks, thereby improving the utilization rate of the information set, reducing the processing latency of the large model, and saving computing resources. Among them, the information to be processed input by the target object can include a short query (query) preset by the large model system, or the information obtained by processing the user's original input using a preset prompt. In this way, for each round of conversation, the preset prompt can be used as the input information of the target model alone to utilize the historical processing results of the preset prompt and improve the processing efficiency of the large model system.

[0078] If the intermediate processing module of the large model system can also split tasks in the above manner, it can be further split, and this disclosure's embodiments do not limit this.

[0079] In some embodiments, task splitting can also be jointly performed according to the conversation rounds and multimodal information. For example, when the current round of conversation includes multimodal information, the current round of conversation is first split into an initial task, and then the initial task is further split into subtasks according to the modalities required by the target model, so as to facilitate constructing the input information of the target model based on the subtasks. For example, if the current round of conversation in multiple rounds of conversation includes a text file and an image, the text file and the image in the current round of conversation can be split into a subtask respectively, and each subtask corresponds to the input information of its respective target model.

[0080] In some embodiments, when the input information is not found in the stored information set, the input information is sent to the target model to obtain the processing result of the target model for this input information; then, the processing result is used as the input information of the next processing module of the target module in the large model system and sent to the next processing module.

[0081] Among them, in order to reduce the processing latency, the input information and its processing result can be stored in the information set for use in processing subsequent tasks.

[0082] In the embodiments of this disclosure, some processing modules consume a large amount of machine resources of the GPU (Graphics Processing Unit) or CPU (Central Processing Unit), and due to being in the link before the large model inference, it will cause a certain latency overhead. Among them, there are some small models that originally have output randomness. By setting output diversity control parameters, these small models are made to have no output randomness, thereby ensuring the stability of the output results of the large model system and reducing unexpected effects (bad cases). Often, when these small models are used as the target model, such asFigure 3 As shown, it includes a pre-filling module 301 and a decoding module 302. Among them, the pre-filling module 301 is used to process the input information of the target model to obtain feature information, and this feature information needs to be input to the decoding module 302 for processing to obtain the processing result of the target model.

[0083] During implementation, in order to reduce the processing delay, the feature information output by the pre-filling module can also be stored to facilitate the reuse of this feature information, and it can also be implemented as shown in Figure 4 as follows:

[0084] S401. For the target module among multiple processing modules in the large model system, when the input information of the target module is obtained, search for the input information in the stored information set.

[0085] S402. When the input information is found, obtain the processing result of the target module for the input information pre-stored in the information set.

[0086] S403. When the input information is not found in the stored information set, search for the feature information corresponding to the input information in the pre-stored feature set; the feature information is generated based on the pre-filling module.

[0087] S404. When the feature information corresponding to the input information is found, input the feature information into the encoding module to obtain the processing result of the target module.

[0088] S405. Store the corresponding relationship between the processing result and the input information in the information set to facilitate reuse in subsequent tasks.

[0089] In the embodiments of the present disclosure, a caching mechanism that can provide diverse information can be provided. For the target module with a pre-filling module and a decoding module, since the decoding module consumes more computing resources and has a relatively longer delay compared to the pre-filling module, the processing links of the pre-filling module and the decoding module can be omitted through the information set of the target module, thereby reducing the delay and saving processing resources. By further adding the storage of the feature information of the pre-filling module, it is possible to further reduce the delay appropriately when encountering new tasks.

[0090] It should be noted that the feature information of the pre-filling module usually has a much larger amount of information than the entire processing result output by the target module. This is because the feature information is the intermediate processing result of the model and retains a large amount of information. Therefore, during implementation, if the input information of the target model and its corresponding processing result are used, compared with storing the feature information, the storage resource amount can be effectively reduced, thereby saving storage resources.

[0091] During implementation, the module-level caching mechanism of the target module and the caching mechanism of the pre-filling module inside the module can be determined according to actual needs.

[0092] Large model systems require a large amount of storage resources, but storage resources are limited. If all historical processing results are stored, a large amount of information may have low utilization. To further save storage resources, in the embodiments of the present disclosure, each piece of information in the information set has a corresponding life cycle, so as to facilitate cleaning up information with low utilization according to the life cycle. The expiration time of the input information in the information set can be set to clean up the input information according to the expiration time.

[0093] In some embodiments, when the input information is associated with a time parameter, the expiration time of the input information in the information set is determined based on the time parameter.

[0094] Among them, the time parameter can be a time parameter used to reduce the hallucination problem of the large model system.

[0095] For example, the target object asks: "What's the weather like today and how to plan a one-day tour around the area". The large model may generate hallucinations based on the training data and output the weather in the training data. Therefore, the time of today needs to be forced into the input information so that the large model system can pay attention to the specific time and output accurate results.

[0096] This time parameter is generally set within the natural day of the current day. Therefore, the early morning of the current day can be set as the expiration time of the input information to facilitate timely release of useless information.

[0097] Of course, the time parameter can also include the time parameter explicitly required by the target object.

[0098] In the embodiments of the present disclosure, based on the time parameter in the input information, the expiration time of the input information can be reasonably set, while reducing the processing delay of the large model system, cleaning up the input information and its corresponding processing results in a timely manner, and releasing the storage space.

[0099] In addition, for the processing module after the target module in the large model system, when the input of the processing module includes at least one of the input information and the processing result obtained by the target module based on the input information, the expiration time of the information corresponding to the input information in the information set of the processing module is also determined based on the time parameter.

[0100] In the foregoing example, it includes not only asking about the weather but also planning a travel route. Thus, during implementation, the task can be split into different subtasks, and a reasonable expiration time can be determined for each subtask separately. In this way, the cleaning time of the planned one-day tour route can be different from the cleaning time of the weather query task, improving the reuse rate of the one-day tour route.

[0101] In some embodiments, when the input information is not associated with a time parameter, the relationship between the utilization rate of the information set of the target module and the duration can be determined, and the duration when the utilization rate reaches the peak value can be selected to determine the expiration time.

[0102] In addition, when updating the target module, it is necessary to clean the data in the information set.

[0103] In order to be able to reasonably select the target module, in the embodiments of the present disclosure, reference information can also be given based on the following methods, including:

[0104] Step C1, based on multiple input samples, determine the utilization rate of the information set of the target module.

[0105] When the target module adopts module-level caching, the utilization rate of the information set of the target module can be statistically calculated, and the hit rate of the information set can be statistically calculated as the utilization rate. For example, for m pieces of input information, where n pieces can be found in the information set, the hit rate can be the percentage of (n / m).

[0106] Step C2, based on the utilization rate, determine the first resource consumption when the target module adopts the information set;

[0107] Among them, the first resource consumption mainly focuses on computing resources, and is represented by the average computing resource cost saved by a single piece of data in the information set. For example, if the computing resource cost consumed by a single piece of data processed by the target model is y, the first resource consumption C1 can be expressed as (n / m)*100%*y.

[0108] Step C3, compare the first resource consumption with the second resource consumption corresponding to the information set to obtain a comparison result.

[0109] Among them, the second resource consumption C2 mainly focuses on the cache cost required to cache a piece of data. Assuming that the cache cost of a single piece of data is C2, when C1 is greater than C2, it means that computing resources can be saved through the information set.

[0110] Step C4, based on the comparison result and the latency information when the target module adopts the information set, determine whether the target module adopts the information set.

[0111] When making a decision, comprehensively consider cost optimization and latency information to determine whether the target module adopts the information set for optimization. When it is determined based on the comparison result that resource cost can be saved and the latency information of adopting the information set indicates that the latency can be effectively reduced, it can be determined that the target module needs to adopt the information set.

[0112] If special attention is paid to latency information, it can also be determined to adopt the information set within the acceptable range of cost consumption.

[0113] Of course, the present disclosure is particularly applicable to processing modules with a large amount of computation and a high utilization rate of information sets. The larger the amount of computation, the more the latency can be reduced, and the higher the superposed utilization rate, the more cost can be saved.

[0114] In the embodiments of the present disclosure, a standard reference method is provided for whether to enable information set acceleration for the target module, so as to guide the screening of which processing modules to be used as the target module during implementation.

[0115] In some embodiments, for multi-turn conversations, module-level caching can also be used to further optimize the processing efficiency of subsequent turns. In the embodiments of the present disclosure, when the large model system executes a content generation task, it may be necessary to generate multiple independent elements according to the intention of the target object, and then fuse and process the multiple independent elements to obtain a response result fed back to the target object. For example, in a scenario of graphic and text creation, a piece of content including graphics and text can be generated according to the requirements of the target object. Then the independent elements include images and text. Even the text can be further divided into smaller independent elements according to chapters, paragraphs, etc. Each independent element can be edited and processed separately.

[0116] Therefore, during implementation, in the content generation scenario of multi-turn conversations, the target module can also be an element generation module. Based on this, it can also be implemented as:

[0117] Step D1, when the large model system executes a content generation task, for the target turn conversation in the multi-turn conversation, store the multiple independent elements generated by the large model system for the target turn conversation into the element set;

[0118] That is, store the independent elements generated by the element generation module into their respective information sets.

[0119] In multi-turn conversations, when the initial turn requires the large model system to generate content, the initial turn can be used as the target turn conversation to store each independently generated independent element for subsequent content editing.

[0120] The independent elements generated by subsequent turns also need to be stored for subsequent use.

[0121] Step D2, when the current turn conversation is a subsequent turn and the intention of the current turn conversation is to update the target element corresponding to the target turn conversation, find the independent elements other than the target element from the element set to obtain the fixed elements.

[0122] For example, the generated independent elements include element 1, element 2, and element 3. When the target object needs to perform content editing and update element 1, then element 2 and element 3 are the fixed elements.

[0123] Step D3: Control the processing modules related to the target elements in the large model system to process the current round of conversation and obtain the updated elements of the target elements.

[0124] Step D4: Based on the element fusion model of the large model system, fuse the updated elements and the fixed elements to obtain the response result of the current round of conversation.

[0125] Continuing with the previous example, generate a new Element 1 as the updated element. The updated element and the fixed elements are fused together and fed back to the target object as the response result of the new round of conversation.

[0126] In the large model system, for example, in the scenario of image content editing, assume that the initial turn requires the large model system to generate butterflies flying among mountains and flowers. Then the generated independent elements may include mountains, flowers, and butterflies. In the subsequent turn, the target object hopes to replace the generated butterflies with elephants. Then it is necessary to change the butterflies in the picture to elephants while keeping other elements. The element generation module in the large model system can be used to generate the image of an elephant according to the description and related features of the elephant. During the generation process, specific conditions (such as posture, size, etc.) can be set to better adapt to the scene of the original picture. When generating the elephant image, information such as the style and color of the original picture can be referred to so that the generated elephant is visually consistent with the original picture.

[0127] Finally, through the fusion module of the large model system, the generated elephant image is fused with other elements of the original picture. According to the structure and semantic information of the original picture, the element generation module can be guided to generate in specific areas while maintaining consistency with other areas.

[0128] During the fusion process, perform a masking operation on the elephant image, only retaining the elephant part, while keeping the other parts of the original picture unchanged. At the same time, the edge of the elephant can be feathered to make it blend more naturally with the background.

[0129] The embodiments of the present disclosure can store the independent elements of the element generation module, so that in the scenario of multi-round conversation, only part of the element content needs to be generated, which can reduce latency and improve the processing efficiency of the large model system.

[0130] In summary, in the embodiments of the present disclosure, the information processing flow of the large model system that caches and stores the information set can be as Figure 5 shown:

[0131] Based on the input of the target object, generate the information to be processed, disassemble the information to be processed to obtain files of each modality, preset prompt words, the conversation content of the current round of conversation, etc. The disassembled tasks construct the input information of the corresponding target modules. As Figure 5As shown in the figure, the input information of preprocessing module 1, preprocessing module 2, and preprocessing module 3 is constructed. For each preprocessing module, the corresponding input information can be looked up based on the cache. In the case where the input information is hit in the cache, the processing operation of the corresponding preprocessing module can be directly omitted. Finally, after being processed by the subsequent preprocessing link, it is handed over to the multi-modal understanding and generation large model for processing.

[0132] The above process can improve the processing efficiency of the large model system, save the computing resources consumed by the preprocessing module, and ensure the diversity of the output results of the multi-modal understanding and generation large model.

[0133] Taking the large model system as an example again, the entire processing flow can be represented as Figure 6 As shown in the figure, assume that before the preprocessing link of the large model inference, there are preprocessing module 1, preprocessing module 2, and preprocessing module 3. Among them, preprocessing module 1 and preprocessing module 2 are modules without output randomness and can be used as target modules respectively. Preprocessing module 1 and preprocessing module 2 each have their own cache to store historical processing results and obtain their respective information sets. When the large model system receives a user query (Query), if the Query has been processed and the processing result is stored in the corresponding cache, then preprocessing module 1 and preprocessing module 2 can be directly skipped, and the processing result is obtained from the cache of preprocessing module 2 and handed over to preprocessing module 3 with output randomness. After being processed by preprocessing module 3, it is input to the large model inference to generate the result feedback to the user.

[0134] It can be seen from this that the entire process is lossless, and the module with output randomness still retains the ability of output diversity. Moreover, multiple modules in the preprocessing link do not need to perform processing operations, which can especially save computing resources and reduce processing latency for some small models with relatively small number of parameters.

[0135] In summary, in the complex large model system of the present disclosure embodiment, based on the characteristics of some GPU models without output randomness and the situation that the preprocessing before the large model system requires a large amount of complex CPU / GPU, the computing cost and latency are relatively high. By introducing the module-level cache technology in the present disclosure embodiment, the response speed of the large model system can be improved and the cost can be reduced. In addition, on the basis of the module-level cache, the present disclosure embodiment adds a task decomposition function, which can improve the cache hit rate, thereby further improving the response speed and reducing the cost.

[0136] Based on the same technical concept, the present disclosure embodiment also provides an information processing device 700 applied to a large model system, as Figure 7 shown, including:

[0137] The search module 701 is configured to search for the input information in the stored information set when obtaining the input information of the target module among multiple processing modules in the large model system;

[0138] The acquisition module 702 is configured to acquire the processing result of the target module on the input information pre-stored in the information set when the input information is found;

[0139] The processing module 703 is configured to send the processing result as the input information of the next processing module of the target module in the large model system to the next processing module.

[0140] In some embodiments, the target module satisfies at least one of the following conditions:

[0141] The target module does not have an output diversity control parameter; the output diversity control parameter is used to control the output result of the model to meet the diversity requirement;

[0142] The setting of the output diversity control parameter of the target module is such that the output results of the target module for the same input data are consistent.

[0143] In some embodiments, as Figure 8 shown, it further includes a determination module 704, configured to screen out the information of the target modality applicable to the target module from the multimodal information to obtain the input information of the target module when the original information before inputting the target module includes multimodal information.

[0144] In some embodiments, the determination module 704 is further configured to determine the input information of the target module based on the conversation content of the current round of conversation in the case of a multi-round conversation.

[0145] In some embodiments, as Figure 8 shown, the determination module 704 includes:

[0146] The first determination unit 7041 is configured to obtain the conversation content of the current round of conversation as the input information of the target module when the input information of the target module needs to include the original conversation content.

[0147] In some embodiments, as Figure 8 shown, the determination module 704 includes:

[0148] The second determination unit 7042 is configured to determine the input source of the target module when the target module is a processing module for intermediate processing results;

[0149] The construction unit 7043 is configured to construct the input information of the target module based on the output results of at least one processing module corresponding to the input source for the conversation content of the current round of conversation.

[0150] In some embodiments, as Figure 8 shown, it further includes:

[0151] A storage module 705, configured to store multiple independent elements generated by the large model system for a target round of conversation in a multi-round conversation into an element set when the large model system executes a content generation task;

[0152] An extraction module 706, configured to, when the current round of conversation is a subsequent turn and the intention of the current round of conversation is to update the target element corresponding to the target round of conversation, find independent elements other than the target element from the element set to obtain fixed elements;

[0153] An update module 707, configured to control a processing module related to the target element in the large model system to process the current round of conversation to obtain an updated element of the target element;

[0154] A fusion module 708, configured to fuse the updated element and the fixed element based on the element fusion model of the large model system to obtain a response result for the current round of conversation.

[0155] In some embodiments, the target module includes a pre-fill module and a decoding module. As Figure 8 shown, the device further includes:

[0156] A query module 709, configured to, when the input information is not found in the stored information set, find the feature information corresponding to the input information in a pre-stored feature set; the feature information is generated based on the pre-fill module;

[0157] An input module 710, configured to, when the feature information corresponding to the input information is found, input the feature information into an encoding module to obtain a processing result of the target module;

[0158] A maintenance module 711, configured to store the corresponding relationship between the processing result and the input information into the information set.

[0159] In some embodiments, as Figure 8 shown, the device further includes:

[0160] A statistics module 712, configured to determine the utilization rate of the information set of the target module based on multiple input samples;

[0161] A first consumption determination module 713, configured to determine a first resource consumption amount when the target module uses the information set based on the utilization rate;

[0162] A comparison module 714, configured to compare the first resource consumption amount with a second resource consumption amount corresponding to the information set to obtain a comparison result;

[0163] A reference module 715 is configured to determine whether the target module adopts the information set based on the comparison result and the latency information when the target module adopts the information set.

[0164] In some embodiments, when the input information is associated with a time parameter, the expiration time of the input information in the information set is determined based on the time parameter.

[0165] For the specific functions and examples of the modules and sub-modules of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.

[0166] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0167] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0168] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0169] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0170] Multiple components in device 900 are connected to I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0171] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the information processing method applied to the large model system. For example, in some embodiments, the information processing method applied to the large model system can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the information processing method applied to the large model system described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the information processing method applied to the large model system in any other suitable manner (e.g., by means of firmware).

[0172] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0173] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0174] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0175] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0176] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0177] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0178] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0179] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An information processing method applied to a large model system, comprising: When obtaining the input information of a target module among multiple processing modules in the large model system, searching for the input information in a stored information set; When the input information is found, obtaining the processing result of the target module on the input information pre-stored in the information set; Taking the processing result as the input information of the next processing module of the target module in the large model system and sending it to the next processing module.

2. The method according to claim 1, wherein the target module satisfies at least one of the following conditions: The target module does not have an output diversity control parameter; the output diversity control parameter is used to control the output result of the model to meet the diversity requirement; The output diversity control parameter of the target module is set such that the output results of the target module for the same input data are consistent.

3. The method according to claim 1 or 2, wherein Obtaining the input information of the target module includes: When the original information before inputting into the target module includes multimodal information, screening out the information of the target modality applicable to the target module from the multimodal information to obtain the input information of the target module.

4. The method according to claim 1 or 2, wherein Obtaining the input information of the target module includes: In the case of a multi-round conversation, determining the input information of the target module based on the conversation content of the current round of conversation.

5. The method according to claim 4, wherein The determining the input information of the target module based on the conversation content of the current round of conversation includes: When the input information of the target module needs to include the original conversation content, obtaining the conversation content of the current round of conversation as the input information of the target module.

6. The method according to claim 4, wherein The determining the input information of the target module based on the conversation content of the current round of conversation includes: When the target module is a processing module for intermediate processing results, determining the input source of the target module; Based on the output results of at least one processing module corresponding to the input source for the conversation content of the current round of conversation, constructing the input information of the target module.

7. The method according to any one of claims 4-6, further comprising: When the large model system executes a content generation task, for a target round of conversation in the multi-round conversation, storing multiple independent elements generated by the large model system for the target round of conversation into an element set; When the current round of conversation is a subsequent conversation turn and the intention of the current round of conversation is to update the target element corresponding to the target round of conversation, searching for independent elements other than the target element from the element set to obtain fixed elements; Controlling the processing module related to the target element in the large model system to process the current round of conversation to obtain an updated element of the target element; Based on the element fusion model of the large model system, fusing the updated element and the fixed element to obtain the response result of the current round of conversation.

8. The method according to any one of claims 1-7, wherein the target module includes a prefill module and a decoding module, and further comprises: In the case where the input information is not found in the stored information set, in the pre-stored feature set, search for the feature information corresponding to the input information; The feature information is generated based on the pre-population module; In the case where the feature information corresponding to the input information is found, input the feature information into the encoding module to obtain the processing result of the target module; Store the correspondence between the processing result and the input information in the information set.

9. The method according to any one of claims 1-8, further comprising: Based on multiple input samples, determine the utilization rate of the information set of the target module; Based on the utilization rate, determine the first resource consumption when the target module uses the information set; Compare the first resource consumption with the second resource consumption corresponding to the information set to obtain a comparison result; Based on the comparison result and the delay information when the target module uses the information set, determine whether the target module uses the information set.

10. The method according to any one of claims 1-9, wherein, In the case where the input information is associated with a time parameter, the expiration time of the input information in the information set is determined based on the time parameter.

11. An information processing apparatus applied to a large model system, comprising: A search module, configured to search for the input information in the stored information set when obtaining the input information of the target module among multiple processing modules in the large model system; An acquisition module, configured to, when the input information is found, acquire the processing result of the target module on the input information pre-stored in the information set; A processing module, configured to send the processing result as the input information of the next processing module of the target module in the large model system to the next processing module.

12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

14. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-10.