Streaming output method of model and electronic equipment

By allocating independent asynchronous threads and event loops to each model and utilizing message queues for priority management and merging, the stability and efficiency issues of streaming output in multi-model collaborative work are resolved, improving system performance and user experience.

CN121050892AActive Publication Date: 2025-12-02INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511586865.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-12-02
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

In scenarios where multiple models work together, existing technologies struggle to achieve stable, real-time, and efficient streaming output, resulting in poor user experience and decreased system performance. This is mainly due to coroutine execution conflicts, resource management issues, and a lack of robust flow control logic.

Method used

Each model is assigned an independent asynchronous thread and event loop. Priority management and merging of asynchronous streaming output tasks are implemented through message queues to ensure the orderly transmission and synchronous output of streaming data.

Benefits of technology

It improves the stability and system performance of multi-model concurrent execution of streaming output tasks, enhances the user's receiving experience, and solves the problem of overall streaming output instability caused by differences in model processing speed and output frequency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121050892A_ABST
    Figure CN121050892A_ABST
Patent Text Reader

Abstract

The invention provides a streaming output method of a model and electronic equipment, and is applied to the technical field of artificial intelligence, the method comprises the following steps: in response to a plurality of models, receiving a target problem sent by a user side, and respectively creating a corresponding event loop for each model in an asynchronous thread; executing an asynchronous streaming output task of the corresponding model based on an event loop in the asynchronous thread, and generating streaming data for the target problem; writing the streaming data output by each asynchronous thread into a message queue, wherein each streaming data has different first priorities for executing the asynchronous streaming output task based on the corresponding model; in response to the situation that the streaming data in the message queue reaches a preset capacity threshold value, determining target streaming data of which the first priority is lower than a preset priority threshold value from the message queue, and merging at least two lexical elements in the target streaming data; and obtaining lexical elements from the merged message queue, and sequentially returning the lexical elements to the user side in a synchronous stream output mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method for streaming output of a model and an electronic device. Background Technology

[0002] With the widespread deployment of large-scale language models in practical applications, more and more scenarios require the simultaneous invocation of multiple models to handle complex tasks in order to improve response quality and processing efficiency. However, when multiple models execute streaming output tasks concurrently, the differences in processing speed and output frequency among the models can easily lead to instability in the overall streaming output, affecting the user's receiving experience and the overall performance of the system. Summary of the Invention

[0003] In view of the above problems, this application provides a method for streaming output of a model and an electronic device.

[0004] According to a first aspect of this application, a model streaming output method is provided, comprising: in response to multiple models receiving a target question sent by a user terminal, creating a corresponding event loop for each model in an asynchronous thread; executing the asynchronous streaming output task of the corresponding model based on the event loop in the asynchronous thread, generating streaming data for the target question respectively; writing the streaming data output by each asynchronous thread into a message queue, wherein each streaming data has a different first priority based on the importance of executing the asynchronous streaming output task of the corresponding model; in response to the streaming data in the message queue reaching a preset capacity threshold, determining target streaming data with a first priority lower than the preset priority threshold from the message queue, and merging at least two tokens in the target streaming data; sequentially retrieving tokens from the message queue after token merging, and sequentially returning them to the user terminal in a synchronous streaming output manner.

[0005] According to a second aspect of this application, an electronic device is provided, comprising: one or more processors; a memory for storing one or more computer programs; the one or more processors executing the one or more computer programs to implement the above-described model streaming output method. Attached Figure Description

[0006] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0007] Figure 1 This application provides a method for streaming output of a model and an application scenario diagram of an electronic device.

[0008] Figure 2 A flowchart of a streaming output method for a model provided in an embodiment of this application;

[0009] Figure 3A This is one of the scenario illustrations of a multi-model question answering method provided in the embodiments of this application;

[0010] Figure 3B This is a second schematic diagram of a multi-model question answering scenario provided in an embodiment of this application;

[0011] Figure 4 A block diagram of an electronic device for a streaming output method suitable for a model according to an embodiment of this application is illustrated schematically. Detailed Implementation

[0012] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0013] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0014] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0015] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0016] With the rapid development and widespread application of large-scale language model technology, a single model can no longer meet the diverse needs of complex business scenarios. In intelligent dialogue platforms, it is common not to run just one model instance, but to run multiple models of different types and complementary functions to form a complete intelligent service ecosystem.

[0017] For example, these models typically include: a main model for core question answering and text generation, responsible for handling the user's main dialogue needs; an auxiliary model for sentiment analysis, style correction, and tone adjustment, ensuring that the output content conforms to a specific language style and sentiment tendency; a knowledge retrieval model for external knowledge retrieval and fact-checking, enhancing the accuracy and timeliness of the answers; and a service model for generating structured responses or performing specific actions.

[0018] In the aforementioned multi-model collaborative application scenarios, to provide a smooth user experience and efficient system performance, the system must meet the following technical requirements: achieve synchronous output to ensure that users can receive model responses in a timely manner; support asynchronous generation of multiple models, fully consider the different inference rates of different models, and avoid affecting the overall system efficiency due to the slow processing of a certain model; have a sound flow control method to prevent excessive system load caused by high-frequency output, and at the same time prevent service interruption caused by network congestion; ensure that the output content is orderly, smooth, and semantically coherent, and maintain a good user interaction experience.

[0019] However, related technologies often suffer from serious coroutine execution conflicts and resource management issues when handling multi-model streaming output. Existing server-side implementations of streaming output functionality require either synchronous or asynchronous coroutine functionality. When the system attempts to use both synchronous and asynchronous coroutine mechanisms simultaneously, it throws coroutine exception errors due to execution context conflicts, preventing multi-model collaborative work from proceeding normally.

[0020] Meanwhile, all coroutine tasks share the same event loop for scheduling and execution by default. When the internal processing logic of a coroutine in a certain model is unusually complex, it will occupy a lot of processing time in the event loop, causing other model tasks to be blocked and seriously affecting the overall system's concurrent processing capabilities.

[0021] More importantly, the relevant technologies generally lack sound traffic control logic, which can lead to system resource exhaustion and message backlog at the producer level, data processing problems and frequent network connection establishment and disconnection at the consumer level, network channel congestion at the network transmission level, and frequent data parsing can also consume a lot of computing resources, further exacerbating the deterioration of system performance.

[0022] Due to the aforementioned technical problems, related technologies struggle to simultaneously meet the stability, real-time performance, and efficiency requirements of multi-model streaming output. To address these issues, this application proposes a model streaming output method and electronic device.

[0023] Figure 1 The diagram illustrates an application scenario of a streaming output method for a model and an electronic device provided in an embodiment of this application.

[0024] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. In this embodiment, it can be understood as a network infrastructure supporting multi-model streaming output data transmission. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, to ensure that streaming data can be transmitted in real-time and stably between the user end and the server.

[0025] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. In this embodiment, the first terminal device 101, the second terminal device 102, and the third terminal device 103 serve as user terminals, used to send target questions to the server 105 and receive synchronous streaming output results from the collaborative processing of multiple models. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (only examples), which can trigger the concurrent processing requirements of multiple models and present streaming output effects.

[0026] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays that support progressive rendering of streaming content, including but not limited to smartphones, tablets, laptops, and desktop computers. For example, these terminal devices have the ability to receive and render streaming data in real time, and can progressively display integrated output content from multiple models in a typewriter effect.

[0027] Server 105 can be a server providing multi-model streaming output services, such as an AI service platform configured with multiple large language model instances and supporting asynchronous streaming to synchronous streaming output (for example only). As a carrier for multi-model collaborative work, Server 105 can respond to the received target problem by creating corresponding event loops for each model in asynchronous threads, and use message queues to aggregate and control the streaming data output from each asynchronous thread. The backend management server can analyze and process received user requests and other data, identify the types and number of models to be called, dynamically allocate computing resources, and synchronously feed back the merged streaming output results to the terminal devices.

[0028] It should be noted that the streaming output method of the model provided in this application embodiment can generally be executed by server 105, where server 105 internally runs multiple model instances, each model instance corresponding to an independent asynchronous thread and event loop, and the technical conversion from asynchronous streaming output to synchronous streaming output is achieved through message queues. The streaming output method of the model provided in this application embodiment can also be executed by a distributed server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or server 105, where each server node can respectively host different types of model instances, and cross-node streaming data aggregation and collaborative processing are achieved through network communication.

[0029] Optionally, in a distributed deployment scenario, the message queue can be implemented using a distributed message middleware, which supports streaming data transmission and merging processing across server nodes, further improving the system's scalability and fault tolerance.

[0030] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0031] The following will be based on Figure 1 The following describes the streaming output method of the model in the embodiments of this application in detail, based on the described scenario.

[0032] Figure 2 This is a flowchart illustrating a streaming output method for a model provided in an embodiment of this application. Figure 2 As shown, the method may include the following operations.

[0033] Operation S210 responds to multiple models receiving the target question sent by the user terminal by creating a corresponding event loop in an asynchronous thread for each model.

[0034] Operation S220 executes the asynchronous streaming output task of the corresponding model based on the event loop in the asynchronous thread, generating streaming data for the target problem respectively.

[0035] Operation S230 writes the streaming data output by each asynchronous thread into the message queue. Each streaming data has a different first priority based on the importance of executing the asynchronous streaming output task according to the corresponding model.

[0036] Operation S240, in response to the streaming data in the message queue reaching a preset capacity threshold, determines the target streaming data with a first priority lower than the preset priority threshold from the message queue, and merges at least two tokens in the target streaming data.

[0037] Operation S250 retrieves tokens sequentially from the message queue after token merging and returns them to the user terminal in a synchronous streaming manner.

[0038] In operation S210, multiple models refer to language model instances of different functional types running in parallel on the server. In this application embodiment, they can be understood as a set of models with different professional capabilities and processing characteristics, used to collaboratively process the target question sent by the user.

[0039] The user terminal refers to a client device or application that initiates a query request to the server. In this embodiment, it can be understood as a terminal that interacts with the server and receives streaming output results. For example, the user terminal includes, but is not limited to, mobile application clients, web browsers, desktop applications, smart hardware devices, etc.

[0040] The target problem refers to the specific query content or task instruction submitted by the user to the server. In this embodiment, it can be understood as the input information that triggers the collaborative work of multiple models.

[0041] For example, the target questions include, but are not limited to, natural language question answering, text creation tasks, knowledge query requests, sentiment analysis needs, and multi-turn dialogue instructions.

[0042] After receiving the target question, the server creates a corresponding event loop for each model in an asynchronous thread. An asynchronous thread is an execution unit that runs independently of the main thread and can process tasks concurrently without blocking the execution of other threads. An event loop is responsible for scheduling and managing asynchronous task execution units in asynchronous programming. It can be understood as an independent running environment and task scheduler for each model instance, ensuring that multiple models can execute streaming output tasks in parallel without interfering with each other.

[0043] For example, asynchronous threads provide independent runtime spaces for each model, preventing multiple model tasks from blocking each other in a single thread; the event loop is responsible for the creation, scheduling, and destruction of coroutine tasks in each asynchronous thread, ensuring that the streaming output tasks of the model can be executed efficiently and orderly. Asynchronous threads and event loops work together to form concurrent execution of multiple models, where each asynchronous thread carries an event loop, and each event loop manages all asynchronous tasks of the corresponding model.

[0044] It should be noted that in related technologies, multiple models typically share the same event loop, which can easily cause the entire system to block when one model is handling a complex task. This application's embodiments achieve operational isolation between models by allocating independent asynchronous threads and event loops to each model, effectively avoiding single-point-of-blockage issues and improving the system's concurrent processing capabilities and stability.

[0045] In one feasible implementation, dedicated asynchronous threads can be created for each model sequentially based on preset model configuration information, and an independent event loop instance can be initialized within each asynchronous thread. The model configuration information includes key attributes such as model identifier, resource configuration parameters, and priority settings.

[0046] In another feasible implementation, multiple asynchronous threads can be pre-created. When a target problem is received, available threads are allocated from the thread pool to each model, and an event loop is dynamically created in the allocated thread to improve the efficiency of thread creation.

[0047] In operation S220, the asynchronous streaming output task refers to the computational task that generates and outputs data content step by step in an asynchronous execution environment in a non-blocking manner. In the embodiments of this application, it can be understood as the asynchronous execution process in which each model performs inference calculations based on the target problem and generates output results in real time, which is used to realize the continuous generation and instant transmission of model output content.

[0048] Asynchronous streaming output tasks continuously generate output content during execution, which is presented in the form of a data stream, hence the term streaming data. Streaming data is a sequence of data units with temporal characteristics continuously generated by the model during inference. It can be understood as the real-time output content generated by the model in response to the target question, which is gradually generated in the form of words, phrases, or sentences to provide continuous and progressive answers to the user.

[0049] In one feasible implementation, the inference task of the corresponding model can be registered in the event loop of each asynchronous thread, and the inference calculation process of the model can be started by coroutine scheduling. During the inference process, the model encapsulates the generated data units into streaming data objects in real time, and passes the streaming data to the subsequent processing module through a callback function.

[0050] In operation S230, the message queue refers to an intermediate buffer structure used for storing and transmitting data. In this embodiment, it can be understood as a data transmission channel connecting the asynchronous streaming output module and the synchronous streaming output module, used to temporarily store the streaming data generated by each model and manage it in a specific order.

[0051] The first priority refers to the level of streaming data allocation based on the importance of the asynchronous streaming output task performed by the model. It is a numerical identifier used to determine the processing order and merging strategy of streaming data in the message queue, and to realize the flow control of multi-model output data.

[0052] Optionally, the numerical range of the first priority can be set to an integer from 1 to 10, where a larger value indicates a higher priority and a smaller value indicates a lower priority.

[0053] Similarly, importance refers to the evaluation index of the task value and urgency of each model in dealing with the target problem. It can be understood as the basis for determining the first priority allocation of streaming data and is used to distinguish the processing priority of the output content of different models.

[0054] After each asynchronous thread generates streaming data, the streaming data needs to be written to a message queue according to a preset data format. During the writing process, the system assigns a corresponding first priority identifier to each piece of streaming data based on the importance of the asynchronous streaming output task performed by each model.

[0055] In one feasible implementation, a basic importance weight can be pre-set for each model. When a model generates streaming data, the first priority of the streaming data is calculated based on the basic weight and the urgency of the current task, and the streaming data object containing the priority information is added to the tail of the message queue.

[0056] In another feasible implementation, the execution status and output quality of each model can be dynamically monitored, and the importance evaluation results of the model can be adjusted in real time. This allows for a more accurate first priority allocation to the subsequently generated streaming data, adapting to changes in processing requirements under different application scenarios.

[0057] In operation S240, the preset capacity threshold refers to the upper limit of the amount of streaming data that the message queue is allowed to store. In this embodiment, it can be understood as the critical condition for triggering the flow control mechanism to prevent the system resources from being exhausted due to excessive data backlog in the message queue.

[0058] For example, the preset capacity threshold can be set to 100 or 500 streams of data depending on the server's memory configuration and processing capacity.

[0059] When the streaming data in the message queue reaches the aforementioned capacity threshold, the system needs to determine the processing strategy by using a preset priority threshold. The preset priority threshold refers to the first priority threshold used to filter the streaming data to be processed, which can be understood as the dividing standard for distinguishing between high-priority and low-priority data.

[0060] Based on a preset priority threshold, the system can identify target streaming data from the message queue. Target streaming data refers to the set of streaming data selected from the message queue whose first priority is lower than the preset priority threshold; it can be understood as data objects that need to be merged.

[0061] In this context, a lexical unit refers to the smallest semantic unit in natural language processing. In this embodiment, it can be understood as the basic data unit output by the model, including text components such as words, characters, and punctuation marks.

[0062] For the tokens in the target streaming data, the system performs a merging operation. Merging refers to the process of combining multiple independent tokens into a single data unit, which is used to reduce the storage pressure on the message queue while maintaining the integrity of the output content.

[0063] In one feasible implementation, when the amount of streaming data in the message queue reaches a preset capacity threshold, flow control is initiated. By comparing the first priority of each streaming data with a preset priority threshold, target streaming data with relatively lower priority is filtered out. Subsequently, the system performs word merging processing on the target streaming data, integrating at least two adjacent or related words.

[0064] In one feasible implementation, the current data volume of the message queue can be continuously monitored. When the data volume reaches 80% of the preset capacity threshold, an alert is triggered. When it reaches 100%, a merging operation is immediately performed by grouping and merging continuous words in the target streaming data according to semantic relevance.

[0065] In operation S250, the process of obtaining tokens can adopt the first-in-first-out principle of the queue, or it can be sorted according to the first priority of the streaming data and then extracted sequentially.

[0066] To achieve a consistent user experience, the system can adopt a synchronous streaming output method. Synchronous streaming output refers to the method of transmitting data to the user terminal according to a uniform time tick and output format. It can be understood as integrating streaming data from different asynchronous threads into a single, continuous output stream to eliminate the output timing differences caused by multi-model parallel processing.

[0067] In one feasible implementation, a fixed output time interval can be set, for example, retrieving one or more tokens from the message queue every 50 milliseconds and immediately transmitting them to the user's display interface via a network connection, thereby achieving a typewriter-like progressive content presentation.

[0068] In another feasible implementation, the output rate can be dynamically adjusted according to the network conditions and processing capabilities of the user terminal. When the network latency is high, the number of tokens transmitted in a single transmission can be appropriately increased, and when the network conditions are good, a faster output frequency can be maintained to optimize the user experience under different network environments.

[0069] Please refer to Figure 3A , Figure 3A This is one of the schematic diagrams illustrating a multi-model question-answering scenario provided in an embodiment of this application. For example... Figure 3A As shown, user Tom sent target question 301 to the education platform through the learning application on tablet 300: "Explain the wave-particle duality of quantum mechanics and provide related exercises and learning suggestions."

[0070] Figure 3B This is a second schematic diagram of a multi-model question-answering scenario provided in an embodiment of this application. For example... Figure 3B As shown, after receiving the target question 301, the server 310 of the education platform identifies that three professional models need to be invoked: the physics concept explanation model 311a, the exercise generation model 311b, and the learning path planning model 311c. The server 310 then creates asynchronous threads 312a, 312b, and 312c for each model respectively, and creates a corresponding event loop in each asynchronous thread to realize the independent running environment of each model.

[0071] Based on the event loops in each asynchronous thread, the three models launch asynchronous streaming output tasks 320a, 320b, and 320c in parallel. The physics concept explanation model 311a generates the corresponding streaming data 321a in the event loop: "Wave-particle duality is a fundamental concept in quantum mechanics...". The exercise generation model 311b generates streaming data 321b in the event loop: "Exercise 1: The double-slit experiment of light demonstrates...". The learning path planning model 311c generates streaming data 321c in the event loop: "Suggested learning order: 1. Review of classical physics...".

[0072] Each asynchronous thread writes the generated streaming data to message queue 330. During the writing process, the system assigns a first priority based on the importance of the asynchronous streaming output task performed by the model: the first priority of streaming data 321a is 9, the first priority of streaming data 321b is 7, and the first priority of streaming data 321c is 5. The storage format in message queue 330 is model ID, streaming content, and first priority.

[0073] When the streaming data in message queue 330 reaches a preset capacity threshold, the system sets a preset priority threshold of 6 and identifies target streaming data with a first priority lower than 6, mainly from the learning path planning model 311c. Therefore, at least two terms in the target streaming data are merged, for example, merging the terms "suggestion," "learning," and "order" into "suggest learning order."

[0074] The system retrieves words from the merged message queue 330 at fixed time intervals and returns them to the tablet 300 in logical order using synchronous streaming output. The tablet 300's display interface presents the complete answer in a progressive manner, first showing the explanation of the physical concept, then the related exercises, and finally the learning suggestions.

[0075] By adopting the embodiments of this application, when multiple models receive the target question sent by the user terminal, corresponding event loops are created in asynchronous threads for each model, so that each model can execute asynchronous streaming output tasks based on independent event loops without blocking each other, effectively solving the problem of mutual interference caused by the difference in processing speed among multiple models.

[0076] The streaming data output by each asynchronous thread targeting the target problem is written to a message queue, and different first priorities are set based on the importance of the asynchronous streaming output task for the corresponding model. This establishes a streaming data aggregation and scheduling method, with the message queue acting as a buffer layer to effectively coordinate the output frequency differences between different models. When the streaming data in the message queue reaches a preset capacity threshold, target streaming data with a first priority lower than the preset priority threshold is identified from the message queue, and at least two terms in the target streaming data are merged. This avoids queue overflow and achieves priority-based dynamic load balancing.

[0077] Finally, the tokens are obtained from the merged message queue and returned to the user terminal in a synchronous streaming output manner. This transforms the asynchronous concurrent processing of multiple models in the backend into a stable synchronous streaming output in the frontend, thereby improving the overall stability when multiple models execute streaming output tasks concurrently. It effectively solves the technical problem of unstable overall streaming output caused by differences in the processing speed and output frequency of each model in related technologies, improves the user terminal's receiving experience, and enhances the overall performance of the system.

[0078] Based on the above embodiments, as an optional implementation method, the first priority of the model executing asynchronous streaming output tasks can be determined in the following way.

[0079] The first parameter representing the importance of the model in the system and the second parameter representing the dynamic performance of the model during streaming output are weighted and summed to determine the first priority of each model for the streaming data. The first parameter includes at least one of the following: the task type of the asynchronous streaming output task, the user level of the user terminal, and the performance level of the model. The second parameter includes at least one of the following: the output rate of the streaming data, the waiting time corresponding to the model, the amount of streaming data, and the system load.

[0080] The first parameter refers to the set of static attributes that characterize the importance of the model in the system. In this embodiment, it can be understood as a quantitative indicator that reflects the relative position and role of the model in the entire service system, and can be used to assign basic weight coefficients to the model.

[0081] For example, there is a positive correlation between task type and user level; higher-level users typically correspond to more important task types, and both together determine the basic weight of business priority. There is a matching relationship between model performance level and task type; complex task types require models with high-performance levels to ensure output quality and processing efficiency. A service quality correspondence is formed between user level and model performance level; higher-level users are given priority in allocating high-performance model resources.

[0082] Similarly, the second parameter refers to the set of real-time attributes that characterize the dynamic performance of the model during streaming output. In this embodiment, it can be understood as a dynamic indicator that reflects the current running state of the model and the system load, and the priority weight of the model can be adjusted in real time.

[0083] For example, there is an inverse relationship between output rate and waiting time; a higher output rate usually corresponds to a shorter waiting time, and both reflect the model's response performance. Data volume and output rate are related; excessive data volume may lead to a decrease in output rate, requiring a balance to be struck between the two. System load is a constraint on all other dynamic parameters; under high load, the system will simultaneously affect output rate, waiting time, and data volume performance.

[0084] In one feasible implementation, a weighted summation method can be used to establish the relationship between the first parameter and the second parameter. Specifically, the indicators in the first parameter form a static weight base through weighted averaging, and the indicators in the second parameter form a dynamic weight adjustment through weighted averaging. The two are linearly combined according to a preset ratio to obtain the final first priority, where the static weight provides a stable basic assessment, and the dynamic weight provides real-time performance adjustment.

[0085] In another feasible implementation, when the system state reflected by the second parameter is good, the influence weight of the first parameter is increased to highlight the inherent importance of the model and the task; when the system state reflected by the second parameter is tense, the influence weight of the second parameter is increased to give priority to system load balancing and response efficiency.

[0086] By adopting the embodiments of this application, a multi-level parameter association relationship is established, which realizes the organic combination of static importance and dynamic performance, provides an adjustable priority allocation method, and effectively improves the system adaptability in multi-model concurrent scenarios.

[0087] Based on the above embodiments, as an optional embodiment, in order to further optimize the merging and processing strategy of streaming data and ensure the semantic coherence and temporal correctness of the output content, a second priority can be determined for the terms in the streaming data, which may specifically include the following operations.

[0088] The second priority of each word is determined based on its position in the corresponding model output sequence; the second priority is used to determine the merging order of at least two words in the target streaming data; wherein, the second priority of the word located at the beginning of the output sequence is greater than the second priority of the word located in the middle of the output sequence.

[0089] The second priority refers to the weight level assigned to a single word based on its positional features in the model's output sequence. It can be understood as a numerical identifier used to determine the retention priority of words in the merging process, thus ensuring the semantic integrity of the output content after merging.

[0090] Similarly, the position number of a word in the output sequence refers to the order number of the word in the complete output sequence of the corresponding model. In this embodiment, it can be understood as a numerical index that identifies the generation time and semantic position of the word, and is used to reflect the importance of the word in the overall output content.

[0091] For example, the numerical range of the second priority can be set to an integer from 1 to 100, where a larger value indicates a higher positional importance of the word, and a smaller value indicates a lower positional importance of the word.

[0092] The second priority of a word at the beginning of the output sequence is higher than that of a word in the middle of the output sequence. Words at the beginning typically carry key semantic information and contextualization functions, playing a crucial role in maintaining the logical coherence of the output content. In contrast, words in the middle are mostly connective or descriptive content, and their transmission frequency can be reduced through merging if necessary.

[0093] In one feasible implementation, the second priority can be calculated using a decreasing function based on the position number of the word. Specifically, the total length of the output sequence is used as the baseline value. Words with smaller position numbers are assigned higher second priority values, while words with larger position numbers are assigned lower second priority values. At the same time, a higher second priority is set for the end marker word at the end of the sequence to ensure output integrity.

[0094] In another feasible implementation, semantic importance analysis in natural language processing can be combined to add a second priority weight to words in special positions such as sentence-initial words, keywords, and semantic boundary words, adding semantic weight on the basis of standard position weight.

[0095] By adopting the embodiments of this application, a second priority system based on position number is established for the lexical units in the streaming data. When performing the merging process, the lexical units at the semantic key positions can be retained first. This effectively avoids the semantic breakage and logical confusion of the output content caused by random merging, improves the streaming data merging process in the scenario of concurrent output of multiple models, and ensures that the content transmitted to the user always maintains good readability and semantic coherence.

[0096] Based on the above embodiments, as an optional embodiment, the operation of merging at least two tokens in the target streaming data in operation S240 may further include the following operations.

[0097] Operation S310 identifies words in the target streaming data whose second priority is lower than the preset word priority threshold, and selects at least two adjacent words from the identified words as objects to be merged.

[0098] Operation S320 merges the words in the object to be merged in order based on their positional order in the original output sequence.

[0099] In operation S310, the object to be merged refers to the set of words selected from the target streaming data that need to be merged. In this embodiment, it can be understood as a combination of adjacent words that meet the merging conditions.

[0100] After identifying the target streaming data, the system can further analyze the word composition within each target streaming data. For each target streaming data, the system compares the second priority of each word it contains with a preset word priority threshold, identifying words with a second priority lower than the threshold. Among the identified low-priority words, the system selects at least two words that are adjacent in position in the output sequence as the objects to be merged.

[0101] In one feasible implementation, a minimum number of merging units parameter can be set, such as requiring each merge to contain at least 2 tokens and at most 5 tokens, in order to avoid semantic ambiguity caused by excessive merging or limited effectiveness caused by insufficient merging.

[0102] In operation S320, the original output sequence refers to the complete word sequence generated by the corresponding model when performing the asynchronous streaming output task. In this embodiment, it can be understood as a baseline sequence that maintains the word generation time order and semantic logic.

[0103] The system connects the terms in the original output sequence according to their position numbers, from beginning to end, to form a new merged term. The merging process strictly follows the positional order of the original output sequence to ensure that the merged content maintains the same semantic logic and temporal relationship as the original output.

[0104] By adopting the embodiments of this application and establishing a priority-based merging strategy, the system can selectively perform merging operations on low-priority content while maintaining the integrity of high-priority content. This effectively alleviates the storage pressure on the message queue, ensures the semantic coherence and logical integrity of the output content, and improves the stability of streaming output and the quality of user experience in multi-model concurrent scenarios.

[0105] Based on the above embodiments, as an optional embodiment, the streaming output method of the above model may further include the following operations.

[0106] Operation S410: In response to the ratio of the current capacity of the message queue to the preset maximum capacity exceeding a preset threshold, the number of tokens contained in the object to be merged is dynamically adjusted based on the capacity ratio, wherein the number of tokens is positively correlated with the capacity ratio.

[0107] Operation S420 responds to the fact that the overall output rate of each asynchronous thread is greater than the output rate of the synchronous streaming mode, and adjusts the number of tokens contained in the object to be merged based on the rate difference between the overall output rate and the output rate, wherein the number of tokens is positively correlated with the rate difference.

[0108] Operation S430 responds to the fact that the proportion of target streaming data with a first priority lower than a preset priority threshold in the message queue is greater than a preset proportion threshold. Based on the difference between the proportion and the preset proportion threshold, the number of tokens contained in the object to be merged is adjusted proportionally.

[0109] In operation S410, the capacity ratio refers to the ratio between the amount of streaming data currently stored in the message queue and the preset maximum capacity. In this embodiment, it can be understood as a quantitative indicator reflecting the current load level of the message queue, used to assess the storage pressure of the system.

[0110] When the capacity ratio exceeds a preset threshold, the system initiates a dynamic adjustment strategy, adjusting the number of tokens contained in the objects to be merged based on the specific value of the capacity ratio. There is a positive correlation between the number of tokens and the capacity ratio; that is, the higher the capacity ratio, the more tokens are included in a single merge operation, thus more effectively alleviating the storage pressure on the message queue.

[0111] In one feasible implementation, a linear adjustment function can be established to map the capacity ratio to a word quantity adjustment coefficient, and different capacity ratio ranges can be set to correspond to different degrees of word quantity growth.

[0112] In operation S420, the overall output rate refers to the total amount of streaming data generated by each asynchronous thread per unit time. In this embodiment, it can be understood as an evaluation index reflecting the concurrent output capability of multiple models, used to measure the overall performance of the data production end.

[0113] The rate difference refers to the numerical difference between the overall output rate and the synchronous streaming output rate. In this embodiment, it can be understood as a quantitative indicator reflecting the degree of imbalance between data production and consumption, used to determine whether it is necessary to balance the input and output speeds through merging processing.

[0114] When the overall output rate is greater than the output rate of the synchronous streaming mode, it indicates that the data production rate exceeds the consumption rate, which can easily lead to message queue backlog. At this time, the system adjusts the number of tokens in the objects to be merged based on the rate difference. There is a positive correlation between the rate difference and the number of tokens, that is, the larger the rate difference, the more tokens are merged, in order to reduce the output frequency and balance the production and consumption rates.

[0115] In one feasible implementation, the output rate changes of each asynchronous thread can be monitored in real time, the average overall output rate within the sliding window can be calculated, and it can be compared with a fixed synchronous output rate. When the rate difference continues to exceed a preset range, the number of tokens can be adjusted.

[0116] In operation S430, target streaming data refers to streaming data in the message queue whose first priority is lower than the preset priority threshold. It can be understood as a set of low-priority data that needs to be focused on and optimized.

[0117] When the percentage exceeds a preset threshold, it indicates that too much target streaming data has accumulated in the message queue, requiring increased merging intensity to alleviate storage pressure. The system calculates the difference between the current percentage and the preset threshold, and adjusts the number of terms contained in the objects to be merged proportionally, maintaining a proportional relationship between the adjustment magnitude and the difference.

[0118] In one feasible implementation, a proportional adjustment strategy based on the percentage difference can be established, and the corresponding word element quantity growth ratio can be determined according to the specific value of the difference.

[0119] By adopting the embodiments of this application, the system can flexibly adjust the intensity and scope of merging processing according to the real-time running status by establishing a multi-dimensional dynamic adjustment based on message queue capacity, output rate differences and target streaming data distribution. This ensures the stable operation of the message queue, maintains the continuity and real-time performance of streaming output, and improves the system's adaptability and overall performance in multi-model concurrent scenarios.

[0120] Based on the above embodiments, as an optional embodiment, operation S230 may further include the following operations:

[0121] Operation S510 assigns instance indexes with one-to-one correspondence with the corresponding models to the streaming data output by each asynchronous thread. The instance indexes are used to identify and distinguish the data sources of different models in scenarios where multiple models are output concurrently.

[0122] The S520 operates by assembling the content information, instance index, and first priority of streaming data into standardized message format data based on a preset data structure format.

[0123] Operate S530 to write the assembled message format data into the tail position of the message queue in the order of the generation time of the streaming data.

[0124] In operation S510, the instance index refers to the unique identifier assigned to each model instance. In this embodiment, it can be understood as a numerical or string marker used to distinguish data sources from different asynchronous threads, in order to ensure the traceability of data when multiple models are output concurrently.

[0125] For example, when initializing the asynchronous threads of each model, the system assigns a unique instance index to each asynchronous thread. This instance index remains unchanged throughout the entire lifecycle of the asynchronous thread, ensuring that all streaming data generated by that thread can be correctly identified and categorized.

[0126] In one feasible implementation, instance index allocation rules can be predefined when the system starts up, and ordered index values ​​can be set according to the type, function or importance of the model. When a new model instance is added, the next available instance index can be automatically allocated according to the preset rules.

[0127] In another feasible implementation, an index allocation method based on model configuration files can be adopted to extract information such as model name and version number from the model's configuration parameters to generate a unique and readable instance index.

[0128] In operation S520, the data structure format refers to the standard template that defines the organization method of message content and the order of field arrangement. In this embodiment, it can be understood as a format standard that ensures the consistency of data storage in the message queue and is used to unify the representation method of data output by different asynchronous threads.

[0129] Among them, message format data refers to standardized data objects assembled according to a preset data structure format. In this application embodiment, it can be understood as a message queue storage unit containing complete information elements and with a unified format, used for orderly storage and retrieval in the message queue.

[0130] For example, the system can assemble the content information of streaming data, the corresponding instance index, and the calculated first priority in a specified order according to a preset data structure format. The assembly process needs to ensure the integrity of each field and the uniformity of the format, generating a standardized data object that meets the storage requirements of the message queue.

[0131] For example, the structure of the message format data can be designed as a dictionary format of {instance index, streaming data content, first priority, generation timestamp} or as a tuple format.

[0132] In one feasible implementation, a custom data class object can be defined as the message format, and various information can be stored through the class's attribute fields.

[0133] In operating the S530, the system continuously monitors the streaming data generation status of each asynchronous thread. When new streaming data is detected, the assembled message-formatted data is immediately written to the message queue. The write operation is strictly executed in the actual generation time order of the streaming data, ensuring that data generated earlier is written to the message queue first, and data generated later is written to the message queue later.

[0134] In one feasible implementation, a timestamp can be added to each message format data. When multiple asynchronous threads generate streaming data simultaneously within a very short time, the accurate writing order can be determined based on the difference in timestamps.

[0135] By employing the embodiments of this application, a data traceability system is established by assigning a unique instance index to the streaming data of each asynchronous thread, effectively solving the problem of unclear data ownership when multiple models output concurrently. A standardized assembly method based on a preset data structure format ensures the consistency of data format in the message queue. An ordered writing method according to the generation time sequence guarantees the temporal correctness of data in the message queue, avoiding output logic chaos caused by disordered writing order.

[0136] Based on the above embodiments, as an optional embodiment, the streaming output method of the above model may further include the following operations.

[0137] Operation S610 creates an iterable class object, which includes an initialization method and a method for retrieving the next element.

[0138] Operate the S620 to start asynchronous threads corresponding to each model to execute streaming output tasks by calling the initialization method.

[0139] Operation S630 retrieves words sequentially from the head of the message queue by repeatedly calling the next element retrieval method in a blocking and waiting manner.

[0140] Operation S640 repeats the process of calling the next element retrieval method until it is detected that all streaming data output from each model in the message queue has been completed.

[0141] In S610 operation, an iterable class object refers to a class instance that implements the iterator protocol. It can be understood as a program object that supports accessing data elements one by one. It is used to encapsulate the control logic and data access interface for converting asynchronous streaming output to synchronous streaming output.

[0142] The initialization method refers to the constructor that is automatically executed when an iterable class object is created. It can be understood as the program entry point responsible for setting the initial state of the object and starting the necessary resources, and is used to establish the connection between the asynchronous thread and the message queue.

[0143] The next element retrieval method refers to the interface function provided by the iterable class object for retrieving the next data unit. It can be understood as a program method that implements synchronous blocking access to message queue data and is used to return ordered streaming data content to the caller.

[0144] The system creates iterable class object instances based on preset class definition templates. During the creation process, the necessary attribute variables and method interfaces are automatically initialized. The iterable class object acts as a bridge between asynchronous and synchronous streaming output, encapsulating complex thread coordination and data synchronization operations.

[0145] For example, iterable class objects can be designed as objects containing core attributes such as queue references, thread state identifiers, and timeout configurations, or the corresponding class structure can be implemented using other programming languages ​​that support iterator protocols.

[0146] In one feasible implementation, the connection parameters of the message queue, the thread configuration information of each model, and the control parameters of the streaming output can be pre-configured in the constructor of the iterable class object to ensure that the object has complete working capabilities after creation.

[0147] In operating the S620, the system triggers the startup process of the asynchronous threads corresponding to each model by calling the initialization method of the iterable class object. During the execution of the initialization method, independent asynchronous threads are created for each model in turn, and a corresponding event loop is established in each asynchronous thread to ensure that each model can concurrently execute streaming output tasks.

[0148] The initialization method is also responsible for establishing a data transmission channel between the asynchronous thread and the message queue, configuring the necessary inter-thread communication parameters, and preparing for subsequent streaming data transmission.

[0149] In one feasible implementation, thread pool technology can be used in the initialization method to manage the asynchronous threads corresponding to each model, and the creation, execution and destruction of threads can be controlled through a unified thread scheduling strategy.

[0150] In another feasible implementation, thread monitoring and exception handling logic can be set in the initialization method. When an asynchronous thread encounters an exception or fails to execute, the thread can be automatically restarted or the error can be logged.

[0151] In S630 operation, blocking wait refers to a synchronization control method that suspends the current thread during program execution until a specific condition is met before resuming execution. It can be understood as a program control means to ensure that data acquisition operations and data production operations are synchronized and coordinated, and is used to avoid empty queue access and data loss problems.

[0152] The system achieves continuous access to streaming data in the message queue by repeatedly calling the next element retrieval method. Each time the method is called, it checks if there is available data at the head of the message queue. If so, it returns the data immediately; otherwise, it enters a blocking wait state until new data arrives.

[0153] In one feasible implementation, polling logic can be set in the next element retrieval method to check the message queue status at fixed time intervals. When new data is detected at the head of the queue, the process returns immediately, and when the queue is detected to be empty, the process continues to wait for the next check cycle.

[0154] In another feasible implementation, an event-driven approach can be used to achieve blocking and waiting. The execution of the next element retrieval method is triggered by the arrival of data in the message queue, thus avoiding unnecessary polling overhead.

[0155] In operation of S640, the system continuously repeats the call process of the next element retrieval method until it receives a streaming data output completion signal from each model. The detection of output completion can be achieved through special end-of-line marker data or asynchronous thread state changes.

[0156] When it is detected that the streaming data output of all models has been completed, the next element retrieval method stops returning new data content, the iteration process of the iterable class object ends, and the synchronous streaming output operation is completed.

[0157] In one feasible implementation, a unified output completion flag can be set for each model. After all models have sent the completion flag, the next element acquisition method returns an iteration end signal to notify the caller to stop the data acquisition operation.

[0158] By employing the embodiments of this application, a unified access method between asynchronous and synchronous streaming output is established through the creation of iterable class objects and the implementation of a standardized iterator interface, effectively resolving compatibility issues between different programming models. The calling method of the initialization method ensures the orderly startup of asynchronous threads and the correct allocation of resources, avoiding race conditions and resource conflicts during thread creation. The blocking and waiting method for retrieving the next element enables secure access and orderly retrieval of message queue data, guaranteeing the integrity and continuity of streaming data during synchronous output, and improving the data transmission reliability and user experience consistency of multi-model concurrent streaming output.

[0159] Based on the above embodiments, as an optional embodiment, operation S250 may further include the following operations.

[0160] If the message queue is not empty and the time elapsed since the last word was read from the message queue is greater than a preset time interval, a word is read from the head of the merged message queue.

[0161] The preset time interval refers to the minimum time difference between two consecutive word reading operations set by the system. It can be understood as a time base for controlling the synchronous streaming output rate, which is used to ensure the stability of streaming data transmission to the user end.

[0162] For example, the system records the timestamp of the last time a word was read from the message queue, and calculates the duration difference between the current time and the last read time each time a new word is attempted to be read. The system only performs the actual word reading operation when the duration difference reaches or exceeds a preset time interval and there is a word to be read in the message queue.

[0163] The above reading method can be implemented based on the flow control principle of the "barrel algorithm". The barrel algorithm refers to an algorithmic strategy that controls the data flow rate through token generation and consumption. In this embodiment, it can be understood as a flow management method that generates reading permissions at fixed time intervals to ensure that token reading operations are executed at a constant rate, thereby smoothing the irregularity of asynchronous output from multiple models.

[0164] Specifically, the system maintains a virtual token bucket, adding tokens to the bucket at a rate equal to the reciprocal of a preset time interval. Each token read operation consumes one token. When there are no available tokens in the bucket, the read operation is paused until a new token is generated. The token bucket's capacity can be set to 1 to ensure that read operations are strictly performed according to the preset time interval, avoiding sudden high-frequency reads.

[0165] In one feasible implementation, the current timestamp can be obtained using the system clock. By comparing the timestamp difference with the value of a preset time interval, it can be determined whether to allow the current word reading operation to be performed, while updating the last reading time record.

[0166] In another feasible implementation, a timer can be set to trigger the word reading task at a preset time interval. When the timer is triggered, the message queue is checked for non-empty status. If the condition is met, a word is immediately read from the head of the queue and returned to the user.

[0167] When the message queue is empty, the system pauses the word reading operation but continues to calculate the time interval to ensure that the reading process can be resumed immediately according to the correct time tick when the message queue receives data again.

[0168] In one feasible implementation, a polling check logic can be set when the message queue is empty, with half of the preset time interval as the check frequency, to continuously monitor changes in the message queue status. When new data is detected to have arrived and the time interval condition is met, a read operation is immediately performed.

[0169] In another feasible implementation, an event notification method can be combined. When the message queue changes from an empty state to a non-empty state, a notification event is triggered. During the event processing, the time interval condition is checked, and if the condition is met, a word reading operation is performed.

[0170] When the duration difference is less than the preset time interval, the system enters a waiting state until the duration difference reaches the preset time interval requirement. During the waiting process, the system does not perform any word reading operations, maintaining the rhythm stability of synchronous streaming output.

[0171] By employing the embodiments of this application, a stable and controllable synchronous streaming output rhythm is established through reading control conditions based on preset time intervals, effectively solving the problem of fluctuations in user reception experience caused by inconsistent asynchronous output rates of multiple models. The application of the "barrel algorithm" ensures that the token reading operation is executed at a constant time interval, avoiding the processing burden on the user end caused by sudden high-frequency data transmissions, while also preventing response delays due to excessively slow reading rates. The dual conditional judgment of message queue non-empty checks and time interval control ensures that the system maintains a stable output rate under various operating states, improving the reliability of concurrent streaming output of multiple models and the consistency of user experience.

[0172] Based on the above embodiments, as an optional embodiment, in order to provide a more refined timeout control process and exception handling steps, and to ensure the controllability and maintainability of the streaming output operation, operation S250 may further include the following operations.

[0173] Operate S710 to set the waiting time threshold as a timeout control parameter.

[0174] Operation S720 returns a timeout exception and terminates the corresponding streaming output process when the message queue is empty and the current waiting time exceeds the waiting time threshold.

[0175] Operation S730: When the message queue is empty and the current waiting time is less than or equal to the waiting time threshold, keep the message queue in a blocked waiting state until each asynchronous thread generates new streaming data and writes it to the message queue.

[0176] When operating the S710, the system can either use a static configuration method to pre-set a fixed waiting time threshold, or use a dynamic adjustment method to automatically adjust the threshold value based on the real-time system load and model performance.

[0177] In one feasible implementation, the average data generation interval and standard deviation can be calculated based on the statistical analysis of the historical output data of each model. The waiting time threshold can be set to the average interval plus a number of times the standard deviation to ensure that false alarms and timeout anomalies are not triggered under normal circumstances.

[0178] In another feasible implementation, a hierarchical waiting time threshold system can be set, including a warning threshold and a termination threshold. When the waiting time reaches the warning threshold, a warning message is recorded but the waiting continues. When the termination threshold is reached, a timeout termination operation is performed.

[0179] In S720 operation, the system continuously monitors the duration of the message queue's empty state and compares the current waiting time with a set waiting time threshold. When the waiting time is detected to exceed the threshold limit, a timeout handling process is initiated.

[0180] The timeout handling process includes generating detailed timeout exception information, recording a snapshot of the current system state, cleaning up related resources, and safely terminating the streaming output process. The system ensures that timeout handling does not affect other concurrently running streaming output tasks.

[0181] In one feasible implementation, detailed status information of each asynchronous thread, historical write records of the message queue, current system resource usage, and the precise time of timeout can be recorded in the timeout exception information to provide information for subsequent fault analysis.

[0182] In another feasible implementation, a gradual resource cleanup operation can be performed before terminating the streaming output process, including stopping related asynchronous threads, releasing occupied memory resources, closing network connections, and cleaning up temporary files to ensure complete reclamation of system resources.

[0183] In S730 operation, when the message queue is empty but the waiting time has not exceeded the threshold, the system enters a controlled blocking wait state. In this state, the system suspends lexical reading operations but maintains continuous monitoring of asynchronous thread states and message queue changes.

[0184] The system employs an efficient event-driven approach in the blocked waiting state, triggering state transitions through asynchronous thread data write events or periodic check events. When new streaming data is detected in the message queue, the system immediately exits the blocked waiting state and resumes normal word reading and output operations.

[0185] In one feasible implementation, a multi-level waiting strategy can be set in the blocked waiting state. Initially, short-interval active waiting is adopted, and the inspection interval is gradually extended as the waiting time increases, so as to reduce system resource consumption while ensuring responsiveness.

[0186] In addition, the system also needs to handle external interrupt requests and cancellation operations while in a blocked waiting state to ensure timely response and safe exit from the waiting state when the user actively cancels the streaming output request.

[0187] By adopting the embodiments of this application, a timeout handling workflow is established, improving the traceability and maintainability of system anomaly handling. Precise setting of the waiting time threshold ensures the accuracy and rationality of timeout judgment, avoiding both premature termination and excessive waiting. Step-by-step timeout checks and blocking wait control combine timeout protection functions with normal data processing flows, maximizing data output integrity while ensuring system stability.

[0188] Based on the above embodiments, as an optional embodiment, operation S210 may further include the following operations:

[0189] Operate the S810 to allocate independent asynchronous threads to each model.

[0190] Operate the S820 to create event loops in each asynchronous thread.

[0191] The S830 operates by initializing streaming output tasks in the corresponding asynchronous threads based on the event loops of each asynchronous thread, using the configuration parameters of each model. The configuration parameters include at least one of the model interface address, authentication information, and streaming output parameters.

[0192] In S810 operation, the system can create a dedicated asynchronous thread for each model based on a preset multi-model list, ensuring that the streaming output operations of each model can run independently in an isolated execution environment. An independent asynchronous thread refers to an independent execution unit specifically allocated to a single model. It can be understood as a thread entity that implements concurrency isolation and resource exclusivity between models, avoiding mutual interference and resource contention between different models.

[0193] When creating an asynchronous thread, the system assigns a unique identifier to each thread and establishes a mapping relationship between the thread and the model to ensure that subsequent task scheduling and resource management can accurately locate the corresponding thread entity.

[0194] In one feasible implementation, different priorities and resource quotas can be set for asynchronous threads based on the computational complexity and expected output frequency of each model, ensuring that high-priority models can obtain more system resources and execution time.

[0195] In operation S820, the system creates a dedicated event loop structure within each asynchronous thread. The event loop refers to a loop scheduling structure used to handle asynchronous events and callback functions. In this embodiment, it can be understood as a scheduler that manages the execution order and timing of asynchronous tasks, used to implement non-blocking asynchronous operation execution and event response processing.

[0196] The creation process of an event loop includes initializing the event queue, setting up the event dispatcher, and configuring the loop's execution parameters. Each event loop runs independently, handling only events and tasks within its own asynchronous thread, thus avoiding interference from events across threads.

[0197] In one feasible implementation, different event loop parameters can be configured for different types of models, including event queue capacity, loop check interval and timeout handling strategy, to optimize event processing performance according to model characteristics.

[0198] In operating the S830, the system utilizes the event loops created in each asynchronous thread, combined with the specific configuration parameters of the model, to complete the initialization of the streaming output task. Configuration parameters refer to a set of parameters describing the model access method and output behavior; they can be understood as configuration information defining the model's call interface, authentication, and output format, used to ensure that the streaming output task can correctly connect to and call the corresponding model service.

[0199] The model interface address refers to the network address information used to access the model service. It can be understood as a complete address containing the protocol type, server address, port number, and interface path, used to establish a network connection with the model service.

[0200] The authentication information refers to the credential data used to verify the identity and permissions of the caller. In this embodiment, it can be understood as including authentication credentials in the form of keys, access tokens, or digital certificates, which are used to ensure the security and legitimacy of model service calls.

[0201] Among them, the streaming output parameters refer to the parameter settings that control the output behavior and format of the model. In the embodiments of this application, they can be understood as a combination of parameters including configuration items such as output block size, encoding format, compression options and transmission protocol, which are used to customize the streaming data output method of the model.

[0202] When initializing the streaming output task, the system verifies the completeness and validity of the configuration parameters, establishes a connection with the model service, and configures the corresponding data receiving and processing logic. The initialization process is executed asynchronously under the scheduling of the event loop to avoid blocking the concurrent execution of other tasks.

[0203] In one feasible implementation, connection tests and parameter verifications can be performed during the initialization process to ensure the accessibility of the model service and the correctness of the configuration parameters. When problems are found, exceptions can be reported in a timely manner and appropriate actions can be taken.

[0204] After the streaming output task is initialized, each asynchronous thread enters the ready state, waiting to receive specific user query requests and begin executing the corresponding model calls and data processing operations.

[0205] By employing the embodiments of this application, a completely isolated concurrent execution environment is established by allocating independent asynchronous threads to each model and creating a dedicated event loop, effectively avoiding resource contention and mutual interference between different models. The event loop-based streaming output task initialization method achieves efficient scheduling and management of asynchronous operations, ensuring that each model can execute streaming output operations in the optimal operating environment.

[0206] Figure 4 A block diagram of an electronic device for a streaming output method suitable for a model according to an embodiment of this application is illustrated schematically.

[0207] like Figure 4As shown, an electronic device according to an embodiment of this application includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage portion 408 into a random access memory (RAM) 403. The processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 may also include onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0208] RAM 403 stores various programs and data required for the operation of the electronic device. Processor 401, ROM 402, and RAM 403 are interconnected via bus 404. Processor 401 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 402 and / or RAM 403. It should be noted that the programs may also be stored in one or more memories other than ROM 402 and RAM 403. Processor 401 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0209] According to embodiments of this application, the electronic device may further include an input / output (I / O) interface 405, which is also connected to a bus 404. The electronic device may also include one or more of the following components connected to the input / output (I / O) interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output (I / O) interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 410 as needed so that computer programs read from it can be installed into the storage section 408 as needed.

[0210] This application also provides a computer-readable storage medium, which may be included in the device / system described in the above embodiments; or it may exist independently and not assembled into the device / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0211] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 402 and / or RAM 403 and / or one or more memories other than ROM 402 and RAM 403 described above.

[0212] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the streaming output method of the model provided in the embodiments of this application.

[0213] When the computer program is executed by the processor 401, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0214] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via communication section 409, and / or installed from removable medium 411. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0215] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by processor 401, it performs the functions defined in the system of this application embodiment. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0216] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0217] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0218] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A method for streaming output of a model, characterized in that, include: In response to multiple models receiving the target question sent by the user, a corresponding event loop is created for each of the models in an asynchronous thread; Based on the event loop in the asynchronous thread, the corresponding asynchronous streaming output task of the model is executed to generate streaming data for the target problem respectively; The streaming data output by each of the asynchronous threads is written to a message queue, wherein each of the streaming data has a different first priority based on the importance of executing the asynchronous streaming output task according to the corresponding model; In response to the streaming data in the message queue reaching a preset capacity threshold, target streaming data with a first priority lower than the preset priority threshold is determined from the message queue, and at least two words in the target streaming data are merged. The tokens are retrieved sequentially from the message queue after token merging and returned to the user terminal in a synchronous streaming manner.

2. The method according to claim 1, characterized in that, The method further includes: A second priority is determined for each word based on its position number in the corresponding model output sequence; the second priority is used to indicate the merging processing order of at least two words in the target streaming data; In this context, the second priority of a word located at the beginning of the output sequence is greater than the second priority of a word located in the middle of the output sequence.

3. The method according to claim 2, characterized in that, The process of merging at least two terms in the target streaming data includes: Identify words in the target streaming data whose second priority is lower than a preset word priority threshold, and select at least two adjacent words from the identified words as objects to be merged; Based on the positional order of each word in the object to be merged in the original output sequence, the words in the object to be merged are merged sequentially.

4. The method according to claim 1, characterized in that, The method further includes: The first parameter representing the importance of the model in the system and the second parameter representing the dynamic performance of the model in the streaming output process are weighted and summed to determine the first priority of each model for the corresponding streaming data. The first parameter includes at least one of the following: the task type of the asynchronous streaming output task, the user level of the user terminal, and the performance level of the model; the second parameter includes at least one of the following: the output rate of the streaming data, the waiting time corresponding to the model, the data volume of the streaming data, and the system load.

5. The method according to claim 1, characterized in that, The step of writing the streaming data output by each of the asynchronous threads into the message queue includes: Each asynchronous thread outputs streaming data and assigns an instance index that has a one-to-one correspondence with the corresponding model. The instance index is used to identify and distinguish the data source of different models in a scenario where multiple models output concurrently. Based on a preset data structure format, the content information of the streaming data, the instance index, and the first priority are assembled into standardized message format data. The assembled message format data is written sequentially to the tail of the message queue according to the generation time order of the streaming data.

6. The method according to claim 1, characterized in that, The method further includes: Create an iterable class object, wherein the iterable class object includes an initialization method and a method for retrieving the next element; The initialization method is invoked to start asynchronous threads corresponding to each model to execute streaming output tasks. The word is retrieved sequentially from the head of the message queue by repeatedly calling the next element retrieval method in a blocking and waiting manner; Repeat the call process of the next element retrieval method until it is detected that all streaming data output from each of the models in the message queue has been completed.

7. The method according to claim 1, characterized in that, The step of retrieving tokens from the merged message queue includes: If the message queue is not empty and the time elapsed since the last word was read from the message queue is greater than a preset time interval, a word is read from the head of the merged message queue.

8. The method according to claim 7, characterized in that, The method further includes: Set a waiting time threshold as a timeout control parameter; If the message queue is empty and the current waiting time is greater than the waiting time threshold, return a timeout exception and terminate the corresponding streaming output process. If the message queue is empty and the current waiting time is less than or equal to the waiting time threshold, the message queue remains in a blocked waiting state until each of the asynchronous threads generates new streaming data and writes it into the message queue.

9. The method according to claim 1, characterized in that, The step of creating corresponding event loops in asynchronous threads for each of the aforementioned models includes: Allocate independent asynchronous threads for each of the aforementioned models; Create an event loop in each of the aforementioned asynchronous threads; Based on the event loop in each of the asynchronous threads, the streaming output task is initialized in the corresponding asynchronous thread using the configuration parameters of each of the models, wherein the configuration parameters include at least one of the model interface address, authentication information, and streaming output parameters.

10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs. The one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cloud-edge collaborative big language model intelligent customer service deployment optimization method

    CN117808481A

  • Dialogue backflow data processing method and training method based on large model and related device

    CN119166388A

  • Multi-agent cooperative task planning method, system and device and storage medium

    CN119917319A

  • Human-computer interaction method, computing device, server, storage medium and program product

    CN120123501A

  • Collaborative response method, device and equipment for large and small models, medium and program

    CN120430386A

Cited By

  • Data transmission method and system based on decorator drive and electronic equipment

    CN121486092A