Large model streaming output termination method

By introducing a two-threading mechanism in large-scale applications, listening to user termination requests in real time and interrupting processing threads, the "pseudo-termination" problem in the existing technology is solved, and effective savings of computing resources and improvement of user experience is achieved.

CN119938247APending Publication Date: 2025-05-06SHENZHEN WORKEC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411775013.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

Smart Images

  • Figure CN119938247A_ABST
    Figure CN119938247A_ABST
Patent Text Reader

Abstract

The invention discloses a termination method for large model streaming output, and aims to solve the problem of'false termination 'in the prior art, namely, after a system stops displaying content at a front end, a background still continues to execute processing logic, so that resource waste and poor user experience are caused. The method comprises the steps that a user initiates a questioning request, and a system responds and creates a first thread to execute related processing of the questioning request; and meanwhile, creating a second thread to monitor a termination request of the user. And when the second thread detects the termination request, the first thread is immediately interrupted, subsequent processing of the first thread is stopped, and it is ensured that the system stops outputting new content to the user. By monitoring the termination request in real time and quickly responding, real output termination is realized, computing resources are effectively saved, user experience is improved, and a user can flexibly control the interaction process with the system at any time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method for terminating streaming output of a large model. Background Art

[0002] With the development of artificial intelligence technology, natural language processing applications based on large models are becoming more and more widespread. In scenarios such as intelligent customer service and virtual assistants, users interact with the system more and more frequently. However, in most existing large model application scenarios, after the user initiates a question request, the system usually continues to generate and output answers until the entire answer process is completed. Although this design meets the needs in general, it exposes obvious deficiencies in certain specific scenarios. For example, when the user realizes that the question is incorrectly stated or no longer needs an answer, he or she hopes to stop outputting immediately, but most existing systems do not support this operation.

[0003] Most of the large model application products on the market currently provide a function that seems to stop output, but in fact it only stops displaying new content on the user interface, while the large model behind it continues to execute the original processing logic until the entire task is completed. This method not only wastes computing resources, but may also cause delays or freezes on the user interface, affecting the user experience.

[0004] The main problem with the above-mentioned existing technologies is that the termination answer function they provide is actually a "pseudo termination", that is, only the display of new content on the front-end interface stops, while the actual operation of the large model does not stop, which leads to ineffective occupation of resources and a decline in user experience. Summary of the invention

[0005] The purpose of the present invention is to provide a termination method for large model streaming output, which can not only effectively save computing resources, but also significantly improve user experience and ensure that users can flexibly control the interaction process with the system at any time to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention adopts the following technical solution: a method for terminating streaming output of a large model, comprising the following steps:

[0007] The user initiates a question request, and the system responds and creates a first thread to execute the processing related to the question request; while creating the first thread, a second thread is created, and the second thread is used to listen to the user's termination request; when the second thread listens to the user's termination request, the second thread immediately interrupts the first thread to prevent the first thread from continuing to execute the subsequent processing; the processing process of the first thread includes but is not limited to splitting the sequence input by the user into several vectors, performing attention calculation on all vectors, and generating content output according to the calculation results; after the termination request is listened and processed, the system stops outputting new content to the user.

[0008] Preferably, the first thread executes the processing related to the question request including the following steps:

[0009] Split the text sequence input by the user into multiple tokens, and convert each token into a vector form to form a vector set V = {v1, v2, ..., v n}, where n represents the number of tokens;

[0010] Position encoding is applied to the vector set V, and the updated vector set is recorded as V′={v1′,v1′,...,v1′}, where each vector v i ′=v i +PE (i) , PE (i) is a position encoding function used to maintain the contextual meaning of the vector;

[0011] Use the updated vector set V′ to calculate attention, the calculation formula is Q = V′W Q , K=V′W K , V=V′W V , where W Q , W K , W V The weight matrices of query, key, and value are obtained respectively, and the attention score matrix is ​​obtained d k is the dimension of the key vector;

[0012] By multiplying the attention score matrix S with the value vector matrix V, the weighted vector set C = SV is calculated, and this vector set C is used as the input for the next step of processing.

[0013] Preferably, the second thread is used to monitor the user's termination request and includes the following steps:

[0014] When the first thread is started, the second thread is initialized and started, and the second thread is set to the daemon state to ensure that it automatically exits when the main process ends;

[0015] The second thread enters the monitoring loop and continuously checks whether there is a termination request signal from the user. If no termination request is detected, it continues to execute the waiting operation in the loop body. Set the waiting time to t wait Second;

[0016] If the second thread detects the termination request signal, it immediately sets the termination flag flag stop Set to true, flag stop =1, and record the current timestamp T stop ;

[0017] The second thread uses flag stop Status Check and T stop Timestamp information, sending an interrupt instruction to the first thread, ensuring that the first thread stops executing at a safe point after receiving the interrupt instruction, to prevent data corruption or abnormal state.

[0018] Preferably, the second thread immediately interrupts the first thread to prevent the first thread from continuing to execute the subsequent processing process, including the following steps:

[0019] When the second thread detects a termination request from the user, an interrupt flag I is set. flag =1, marking the need to interrupt the first thread;

[0020] The second thread sends an interrupt signal S to the first thread int , the signal format is S int =(I flag ,T int ), where T int The timestamp for sending the interrupt signal;

[0021] The first thread receives S int After that, check I flag If the value of I flag =1, the first thread starts looking for the nearest safe point P safe ;

[0022] The first thread reaches P safe After that, perform cleanup operations, including releasing resources and saving necessary states. After completion, set the execution state E status =0, indicating that the first thread has stopped safely.

[0023] Preferably, the processing of the first thread includes the following steps:

[0024] Receive the text sequence input by the user and decompose it into a series of basic units U = {u1u2,...,u m}, each unit u i corresponds to a word or character in the text;

[0025] Convert these basic units into corresponding vector representations to form a vector set VU = {vu1, vu2, ..., u m}, where each vector vu i Represents the corresponding basic unit u i ;

[0026] Apply position encoding to vector set VU to generate a new vector set V containing position information PE = {v pe1 ,v pe2 ,...,v pem}, where v pei =v ui +PE(i), PE(i) is a position-based encoding function used to preserve the relative position relationship between vectors;

[0027] Use the new vector set V PE Perform attention mechanism calculations, using the calculation formula Where Q, K, V are respectively PE The query, key, and value matrices obtained from k is the dimension of the key vector, and the attention-weighted vector set V is obtained. A = {v a1 ,v a2 ,...,v am};

[0028] Based on the attention weighted vector set V A , using feedforward neural network processing to generate the final output vector O = FFN (VA), where FFN represents the feedforward neural network, and finally generating the corresponding text output based on the output vector O.

[0029] Preferably, after the termination request is monitored and processed, the system stops outputting the new content to the user, including the following steps:

[0030] When the second thread receives the user's termination request, it immediately sends the termination signal S term Sent to the first thread, where term = (T req ,F term ), T req Indicates the request time, F term It is the termination flag, set to 1;

[0031] The first thread receives the termination signal S term After that, check F term If the value of F term =1, then prepare to interrupt the current task;

[0032] The first thread is at the nearest safe point P safeThe task is interrupted at the safe point to ensure that it does not stop in the middle state and cause data inconsistency. This safe point is pre-defined by the system.

[0033] After the task is interrupted, the first thread performs cleanup operations, including but not limited to releasing memory resources, closing file handles, and setting the cleanup completion flag C. done =1;

[0034] Set the output termination flag O stop =1, prohibiting the first thread from continuing to generate new output content, ensuring that no new content is generated after the request is terminated;

[0035] Check output termination flag O stop If the value of O stop =1, stop sending new content to the user, ensuring that the user interface is no longer updated until the user issues a new instruction or request.

[0036] Preferably, the method further comprises the following steps:

[0037] Define output buffer B out , used to store the content to be output, the initial state is empty;

[0038] When the first thread generates output content, add the content fragments one by one to the output buffer B out , that is, B out =B out +C i , where C i represents the content fragment generated for the i-th time;

[0039] When the first thread receives the termination signal and completes the cleanup operation, it checks the output buffer B out Is it empty? If not, the remaining contents in the buffer are output to the user at once, and then the buffer is cleared;

[0040] Set the output completion flag O end =1, indicating that the output process has completely stopped. After the output stops completely, the system enters standby mode and waits for the next request or instruction from the user.

[0041] Preferably, the output buffer B out Management includes:

[0042] In each out Before adding content, check the output termination flag O stop , if O stop =1, skip the current content adding operation and directly enter the next loop;

[0043] If O stop=0, then continue to perform the content addition operation and update the status of the output buffer;

[0044] In output buffer B out Before clearing, lock the buffer to prevent data competition in a multi-threaded environment and ensure data consistency and integrity. After clearing the buffer, unlock the buffer to allow other threads to access or modify it.

[0045] Preferably, an error handling mechanism is also included:

[0046] If an unrecoverable error occurs during the execution of any step, the system shall record an error log and set the error code E to code and error description E desc Send it to the monitoring system. After receiving the error information, the monitoring system takes corresponding measures according to the error type;

[0047] If an error is encountered during the cleanup operation, the system executes the backup cleanup plan to ensure that resource release and status recovery can be completed even in abnormal situations;

[0048] After the error processing is completed, set the error processing completion flag H done =1, indicating that the error has been properly handled

[0049] Preferably, the response of the monitoring system includes:

[0050] After receiving the error information, the monitoring system first determines the severity of the error. For minor errors, only logs are recorded. For serious errors, the emergency response process is immediately executed.

[0051] The emergency response process includes but is not limited to automatically restarting the service, sending alarm emails or text messages to the administrator, and recording detailed error site information;

[0052] After the service is successfully restarted, the monitoring system automatically checks the service status to ensure that the service resumes normal operation;

[0053] If the service restart fails, the monitoring system will continue to try to restart or switch to the backup system until the service is restored; after the service is restored to normal, the monitoring system clears the error status flag E status =0, and notify all relevant parties that the service has returned to normal.

[0054] Technical effects and advantages of the present invention: Compared with the prior art, the termination method of a large model streaming output proposed by the present invention has the following advantages:

[0055] The present invention introduces a dedicated listening thread to monitor the user's termination request in real time. Once a request is detected, the ongoing task thread can be quickly interrupted, thereby achieving true output termination. This method can not only effectively save computing resources, but also significantly improve the user experience, ensuring that users can flexibly control the interaction process with the system at any time. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A flowchart of a method for terminating a large model streaming output according to the present invention;

[0057] Figure 2 A block diagram of a termination system for large model streaming output according to the present invention;

[0058] Figure 3 It is the work flow chart of the termination system of the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0060] The present invention provides a termination method for streaming output of a large model, aiming to solve the "pseudo-termination" problem in the prior art, that is, after the system stops displaying content on the front end, the background continues to execute processing logic, resulting in resource waste and poor user experience. The method includes: the user initiates a question request, the system responds and creates a first thread to execute related processing of the question request; at the same time, a second thread is created to listen to the user's termination request. When the second thread detects the termination request, it immediately interrupts the first thread, prevents its subsequent processing, and ensures that the system stops outputting new content to the user. The processing process of the first thread involves splitting the user input into vectors, performing attention calculations, and generating content output. The present invention realizes true output termination by real-time monitoring of termination requests and responding quickly, effectively saving computing resources, improving user experience, and allowing users to flexibly control the interaction process with the system at any time. The details are as follows:

[0061] like Figure 1-3 As shown, a method for terminating a large model streaming output in this embodiment includes the following steps:

[0062] S101: A user initiates a question request, and the system responds and creates a first thread to execute a processing process related to the question request; further comprising:

[0063] Split the text sequence input by the user into multiple tokens. Each token can be a word, punctuation mark, or other meaningful language unit. And convert each token into a vector form to form a vector set V = {v1,v2,...,v n}, where n represents the number of tokens; the vector set here contains the vector representations of all tokens, where n represents the number of tokens. This process converts natural language into a machine-understandable form to facilitate subsequent computational processing.

[0064] In order to maintain the contextual meaning between vectors, position encoding is applied to the vector set V, and the updated vector set is recorded as V′={v1′,v1′,...,v1′}, where each vector v i ′=v i +PE (i) , PE (i) is a position encoding function used to maintain the contextual meaning of the vector; the position encoding is a special vector that is added to each original vector to reflect the position information of the vector in the sequence. The updated vector set not only retains the semantic information of the original vector, but also adds position information, which is crucial for understanding sentence structure and context.

[0065] The updated vector set will then be used for attention calculation, using the updated vector set V′ for attention calculation, the calculation formula is Q = V′W Q , K=V′W K , V=V′W V , where W Q , W K , W V The weight matrices of query, key, and value are obtained respectively, and the attention score matrix is ​​obtained d k is the dimension of the key vector; the attention mechanism allows the model to focus on different parts of the input sequence when generating output, thereby improving the relevance and accuracy of the output. In this process, the system transforms the vector set using the weight matrices of the query, key, and value to generate an attention score matrix. These scores reflect the relevance of each input vector to other vectors, helping the model determine which parts of the information are more important.

[0066] By multiplying the attention score matrix S with the value vector matrix V, we calculate the weighted vector set C = SV, which is used as the input for the next step. This weighted vector set combines all the information of the input sequence and is weighted by the attention mechanism, so that the model can generate output content more accurately. This vector set will be used as the input for the next step to generate the final answer or output.

[0067] Through the above-mentioned specific implementation methods, the present invention can effectively process the user's question request, and introduces position encoding and attention mechanism in the processing process, thereby improving the model's understanding ability and generation quality. In particular, when a user initiates a question request, the system can respond quickly and create a dedicated processing thread to ensure that the processing process is efficient and orderly. In addition, by splitting the text sequence input by the user into tokens and vectorizing them, the system can better capture the semantic information of the text, and the application of position encoding further enhances the model's ability to understand the context.

[0068] The introduction of the attention mechanism enables the model to more flexibly focus on different parts of the input when generating output, thereby improving the relevance and accuracy of the output. Overall, these technical means work together to significantly improve the performance of the system and user experience.

[0069] S102, while creating the first thread, creating a second thread, the second thread being used to monitor the user's termination request, thereby ensuring that the user can flexibly control the interaction process with the system at any time; further comprising:

[0070] When the first thread is started, the second thread is initialized and started, and the second thread is set to the daemon state to ensure that it automatically exits when the main process ends; setting the second thread to the daemon state means that when the main process ends, the second thread will automatically exit, avoiding the problem of resource leakage.

[0071] The second thread enters the monitoring loop and continuously checks whether there is a termination request signal from the user. If no termination request is detected, it continues to execute the waiting operation in the loop body. Set the waiting time to t wait seconds; if no termination request is detected, the second thread will execute the waiting operation in the loop body and set a reasonable waiting time (for example, 1 second) to reduce unnecessary resource consumption.

[0072] If the second thread detects the termination request signal, it immediately sets the termination flag flag stop Set to true, flag srop =1, and record the current timestamp T stop ; This step ensures that the system can respond quickly to the user's termination request and record the specific time of the request to provide a basis for subsequent processing.

[0073] The second thread uses flag stop Status Check and T stopThe timestamp information is used to send an interrupt instruction to the first thread, ensuring that the first thread stops executing at a safe point after receiving the interrupt instruction to prevent data corruption or abnormal state. After receiving the interrupt instruction, the first thread stops executing at the nearest safe point to ensure that it will not stop in an intermediate state and cause data corruption or abnormal state.

[0074] Safety points refer to some key points predefined by the system. Stopping execution at these points can ensure the consistency and integrity of data.

[0075] S103, when the second thread monitors the user's termination request, the second thread immediately interrupts the first thread to prevent the first thread from continuing to execute the subsequent processing process; further comprising the following steps:

[0076] When the second thread detects a termination request from the user, an interrupt flag I is set. flag =1, marking the need to interrupt the first thread; this step ensures that the system can quickly recognize and respond to the user's termination request.

[0077] The second thread sends an interrupt signal S to the first thread int , the signal format is S int =(I flag ,T int ), where T int The timestamp for sending the interrupt signal; the recording of the timestamp helps with subsequent logging and debugging, ensuring that the system can trace the specific time of the interrupt when necessary.

[0078] The first thread receives S int After that, check I flag If the value of I flag =1, the first thread starts looking for the nearest safe point P safe ;Safe points refer to some key points predefined by the system. Stopping execution at these points can ensure the consistency and integrity of the data.

[0079] The first thread reaches P safe After that, perform cleanup operations, including releasing resources and saving necessary states. After completion, set the execution state E status = 0, indicating that the first thread has been safely stopped. This step ensures that the system can clearly know that the first thread has been stopped and can perform subsequent processing or restart.

[0080] In some embodiments, the processing process of the first thread includes but is not limited to splitting the sequence of user input into a plurality of vectors, performing attention calculation on all vectors, and generating content output according to the calculation results; and further includes:

[0081] Receive the text sequence input by the user and decompose it into a series of basic units U = {u1, u2,..., u m}, where each unit u i corresponds to a word or character in the text; decompose this text sequence into a series of basic units, and each unit corresponds to a word or character in the text. For example, the input text "Hello, World!" will be decomposed into basic units such as "Hello", "World", "!", etc.

[0082] Convert these basic units into corresponding vector representations to form a vector set VU = {vu1, vu2,..., u m}, where each vector vu i represents the corresponding basic unit u i ; these vectors are usually obtained from a pre-trained word embedding model and can capture the semantic information of each word or character.

[0083] To preserve the relative position relationship between vectors, apply positional encoding to the vector set VU to generate a new vector set V PE = {v pe1 , v pe2 ,..., v pem}, where v pei = v ui + PE(i), and PE(i) is a position-based encoding function used to preserve the relative position relationship between vectors;

[0084] Use the new vector set V PE to perform attention mechanism calculations. The attention mechanism allows the model to focus on different parts of the input sequence when generating the output, thereby improving the relevance and accuracy of the output. Through the calculation formula where Q, K, and V are the query, key, and value matrices obtained from V PE respectively, and d k is the dimension of the key vector, to obtain the attention-weighted vector set V A = {v a1 , v a2 ,..., v am};

[0085] Based on the attention-weighted vector set V A , use a feed-forward neural network to process and generate the final output vector O = FFN(VA), where FFN represents the feed-forward neural network, and finally generate the corresponding text output according to the output vector O. These text outputs can be answers, explanations, or other relevant information.

[0086] S104: After the termination request is monitored and processed, the system stops outputting the new content to the user; further, the following steps are included:

[0087] When the second thread receives the user's termination request, it immediately sends the termination signal S term Sent to the first thread, where term = (T req ,F term ), T req Indicates the request time, F term The termination flag is set to 1; this step ensures that the first thread can quickly receive the termination request.

[0088] The first thread receives the termination signal S term After that, check F term If the value of F term =1, then prepare to interrupt the current task; this step ensures that the first thread can respond to the termination request in time and start the interrupt processing process.

[0089] The first thread is at the nearest safe point P safe The task is interrupted at the point where the system is running to ensure that it will not stop in the middle and cause data inconsistency. The safety point is pre-defined by the system. Safety points refer to some key points pre-defined by the system. Stopping execution at these points can ensure data consistency and integrity. This step ensures data integrity and system stability.

[0090] After the task is interrupted, the first thread performs cleanup operations, including but not limited to releasing memory resources, closing file handles, and setting the cleanup completion flag C. done =1; These cleanup operations ensure that system resources are effectively released to avoid resource leakage and data corruption.

[0091] Set the output termination flag O stop =1, prohibiting the first thread from continuing to generate new output content, ensuring that no new content is generated after the termination request; this step ensures that the system will not generate new output after receiving the termination request, avoiding unnecessary resource consumption and user confusion.

[0092] Check output termination flag O stop If the value of O stop = 1, stop sending new content to the user, ensuring that the user interface is no longer updated until the user issues a new instruction or request. This step ensures that the user interface remains stable after the request is terminated, preventing the user from seeing incomplete or inconsistent output.

[0093] In summary, the present invention realizes true output termination by introducing termination signals, safe point mechanisms and cleanup operations, which not only improves system performance and resource utilization, but also significantly improves user experience and system stability.

[0094] In another embodiment, the above-mentioned method for terminating the streaming output of a large model further includes the following steps:

[0095] Define output buffer B out , used to store the content to be output, and its initial state is empty; the output buffer is used to temporarily store content fragments during the process of generating content to ensure that the content can be output to the user in an orderly manner.

[0096] When the first thread generates output content, add the content fragments one by one to the output buffer B out , that is, B out =B out +C i , where C i represents the content fragment generated for the i-th time; each time a new content fragment is generated, it is added to the output buffer. This step ensures that the generated content can be stored in order to facilitate subsequent output processing.

[0097] When the first thread receives the termination signal and completes the cleanup operation, it checks the output buffer B out Is it empty? If not, the remaining contents in the buffer are output to the user at once, and then the buffer is cleared. This step ensures that after the request is terminated, the user can see all the contents that have been generated but not yet output, thus avoiding the loss of content.

[0098] Set the output completion flag O end =1, indicating that the output process has completely stopped. This step ensures that the system can clearly know that the output process has ended and can proceed with subsequent processing or restart. After the output is completely stopped, the system enters standby mode and waits for the next request or instruction from the user. This step ensures that the system can enter a low-power state after processing the current task, saving resources and being ready to respond to the next request from the user.

[0099] Furthermore, the output buffer B out Management includes:

[0100] In each out Before adding content, check the output termination flag O stop , if O stop =1, skip the current content addition operation and directly enter the next loop; this step ensures that after the request is terminated, no new content fragments will be added to the output buffer, avoiding unnecessary resource consumption and data inconsistency.

[0101] If O stop = 0, the content adding operation continues and the output buffer status is updated; specifically, each time a new content fragment is generated, it is added to the output buffer. This step ensures that the generated content can be stored in an orderly manner, which is convenient for subsequent output processing.

[0102] In output buffer B oit Before clearing, lock the buffer to prevent data competition in a multi-threaded environment and ensure data consistency and integrity. After clearing the buffer, unlock the buffer to allow other threads to access or modify it. The operation of locking the buffer can be achieved through a mutex to ensure that only one thread can access and modify the buffer at the same time. This step avoids data competition and inconsistency problems that may occur in a multi-threaded environment.

[0103] Furthermore, it also includes error handling mechanisms:

[0104] If an unrecoverable error occurs during the execution of any step, the system shall record an error log and set the error code E to code and error description E desc The error message is sent to the monitoring system. After receiving the error message, the monitoring system takes corresponding measures according to the error type. This step helps with subsequent troubleshooting and system maintenance.

[0105] If an error is encountered during the cleanup operation, the system executes the backup cleanup plan to ensure that resources can be released and status restored even in abnormal situations; for example:

[0106] For minor errors, the monitoring system only records logs for subsequent analysis.

[0107] For serious errors, the monitoring system will immediately execute emergency response processes, such as restarting services, sending alarm notifications to administrators, etc.

[0108] Alternative cleanup options may include:

[0109] Try the cleanup operation multiple times until it succeeds.

[0110] Use a different cleanup strategy, such as releasing some resources instead of all.

[0111] Records detailed information about cleanup failures for subsequent analysis and repair.

[0112] After the error processing is completed, set the error processing completion flag H done =1, indicating that the error has been properly handled. This step ensures that the system can clearly know that the error handling process has ended and can continue to perform subsequent normal operations.

[0113] Specifically, the response of the monitoring system includes:

[0114] After receiving the error information, the monitoring system first determines the severity of the error. For minor errors, it only records the log. For serious errors, it immediately executes the emergency response process.

[0115] The emergency response process includes but is not limited to the following steps:

[0116] Automatically restart services: The monitoring system automatically triggers service restart operations to try to restore normal operation of the service.

[0117] Send an alarm email or text message to the administrator: The monitoring system sends an alarm email or text message to the administrator to inform the administrator that a serious error has occurred in the system and needs to be handled in a timely manner.

[0118] Record detailed error site information: The monitoring system records detailed error site information, including the time when the error occurred, error code, error description, system status, etc., for subsequent analysis and troubleshooting.

[0119] After the service is successfully restarted, the monitoring system automatically checks the service status to ensure that the service resumes normal operation; the inspection content includes but is not limited to: whether the service can respond to user requests normally, whether the system resource usage is normal, and whether there is new error information in the log.

[0120] If the service restart fails, the monitoring system will continue to try to restart or switch to the backup system until the service is restored; after the service is restored to normal, the monitoring system clears the error status flag E status =0, and notify all relevant parties that the service has returned to normal.

[0121] To summarize, this embodiment achieves more efficient error handling and system management by introducing a detailed monitoring system response mechanism, including determining the severity of errors, executing emergency response procedures, checking service status, handling service restart failures, and clearing error status flags. This not only improves the stability and reliability of the system, but also significantly improves user satisfaction and system maintenance efficiency.

[0122] As 2 and Figure 3 As shown, in this embodiment, a termination system for large model streaming output is proposed, including:

[0123] Business layer: various businesses completed by users through large models, such as chat robots, etc.

[0124] Embedding layer: splits the sequence tokens input by the user and converts them into several vectors in sequence.

[0125] Position encoding: ensures that the transformed vector still has contextual semantics.

[0126] Guardian layer: Start a daemon thread to terminate the main thread that is outputting.

[0127] Attention layer: Perform attention calculation on all vectors so that the position of the vector changes according to the semantics and obtains several corresponding vectors.

[0128] Feedforward layer: Add several vectors together to get the final vector.

[0129] SoftMax: Get the probability distribution of the next most likely output vocabulary based on the final vector.

[0130] Output: Output content based on the obtained vocabulary probability distribution table.

[0131] By adding a thread guard layer in front of the attention layer to listen to the termination request issued by the user, the output thread is stopped, so that the large model can truly terminate the output.

[0132] The detailed implementation is as follows:

[0133] When a user initiates a question request, a thread (thread A) will normally be started to execute the logic related to the large model, such as content embedding, attention calculation, content output, etc. When thread A is created, an additional thread (thread B) is also created to listen for the stop signal. When the user sends a request to stop output, thread B can listen to the signal.

[0134] At the same time, the subsequent logic of thread A is terminated, thereby terminating the output. In this way, the output termination operation of the large model is realized.

[0135] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for terminating streaming output of a large model, characterized in that: The following steps are involved: When receiving a question request initiated by a user, creating a first thread to execute a processing process related to the question request; While creating the first thread, creating a second thread, the second thread being used to monitor a termination request from a user; When the second thread monitors the termination request initiated by the user, the second thread interrupts the first thread to prevent the first thread from continuing to execute the subsequent processing of the question request; After the termination request is monitored and processed, outputting the new content to the user stops.

2. A method for terminating streaming output of a large model according to claim 1, characterized in that: The first thread performs the processing related to the question request, including the following steps: Split the text sequence input by the user into multiple tokens, and convert each token into a vector form to form a vector set V = {v1, v2, ..., v n }, where n represents the number of tokens; Position encoding is applied to the vector set V, and the updated vector set is recorded as V′={v1′,v1′,...,v1′}, where each vector v i ′=v i +PE (i) , PE (i) is a position encoding function used to maintain the contextual meaning of the vector; Use the updated vector set V′ to calculate attention, the calculation formula is Q = V′W Q , K=V′W K , V=V′W V , where W Q , W K , W V The weight matrices of query, key, and value are obtained respectively, and the attention score matrix is ​​obtained d k is the dimension of the key vector; By multiplying the attention score matrix S with the value vector matrix V, the weighted vector set C = SV is calculated, and this vector set C is used as the input for the next step of processing.

3. A method for terminating streaming output of a large model according to claim 2, characterized in that: The second thread is used to monitor the user's termination request and includes the following steps: When the first thread is started, the second thread is initialized and started, and the second thread is set to a daemon state so that it automatically exits when the main process ends; The second thread enters the monitoring loop and continuously checks whether there is a termination request signal from the user. If no termination request is detected, it continues to execute the waiting operation in the loop body. Set the waiting time to t wait Second; If the second thread detects the termination request signal, the termination flag flag is set stop Set to true, flag stop =1, and record the current timestamp T stop ; The second thread uses flag stop Status Check and T stop The timestamp information sends an interrupt instruction to the first thread, so that the first thread stops executing at a safe point after receiving the interrupt instruction.

4. A method for terminating large model streaming output according to claim 3, characterized in that: The second thread immediately interrupts the first thread to prevent the first thread from continuing to execute a subsequent processing process, including the following steps: When the second thread detects a termination request from the user, an interrupt flag I is set. flag =1, marking the need to interrupt the first thread; The second thread sends an interrupt signal S to the first thread int , the signal format is S int =(I flag ,T int ), where T int The timestamp for sending the interrupt signal; The first thread receives S int After that, check I flag If the value of I flag =1, the first thread starts looking for the nearest safe point P safe ; The first thread reaches P safe After that, perform cleanup operations, including releasing resources and saving necessary states. After completion, set the execution state E status =0, indicating that the first thread has stopped safely.

5. A method for terminating streaming output of a large model according to claim 4, characterized in that: The processing process of the first thread includes the following steps: Receive the text sequence input by the user and decompose it into a series of basic units U = {u1u2,...,u m }, each unit u i corresponds to a word or character in the text; Convert these basic units into corresponding vector representations to form a vector set VU = {vu1, vu2, ..., u m }, where each vector vu i Represents the corresponding basic unit u i ; Apply position encoding to the vector set VU to generate a new vector set V containing position information PE = {v pe1 ,v pe2 ,...,v pem }, where v pei =v ui +PE(i), PE(i) is a position-based encoding function used to preserve the relative position relationship between vectors; Use the new vector set V PE Perform attention mechanism calculations, using the calculation formula Where Q, K, V are respectively PE The query, key, and value matrices obtained from k is the dimension of the key vector, and the attention-weighted vector set V is obtained. A = {v a1 ,v a2 ,...,v am }; Based on the attention weighted vector set V A , using feedforward neural network processing to generate the final output vector O = FFN (VA), where FFN represents the feedforward neural network, and finally generating the corresponding text output based on the output vector O.

6. A method for terminating large model streaming output according to claim 5, characterized in that: After the termination request is monitored and processed, the output of new content to the user is stopped, including: When the second thread receives the user's termination request, it sends the termination signal S term Sent to the first thread, where term = (T req ,F term ), T req Indicates the request time, F term It is the termination flag, set to 1; The first thread receives the termination signal S term After that, check F term If the value of F term =1, then prepare to interrupt the current task; The first thread is at the nearest safe point P safe Interrupt tasks; After the task is interrupted, the first thread performs a cleanup operation, which includes releasing memory resources, closing file handles, and setting a cleanup completion flag C. done =1; Set the output termination flag O stop =1, prohibit the first thread from continuing to generate new output content; Check output termination flag O stop If the value of O stop =1, stop sending new content to the user until the user issues a new instruction or request.

7. A method for terminating large model streaming output according to claim 6, characterized in that: The following steps are also included: Define output buffer B out , used to store the content to be output, the initial state is empty; When the first thread generates output content, add the content fragments one by one to the output buffer B out , that is, B out =B out +C i , where C i represents the content fragment generated for the i-th time; When the first thread receives the termination signal and completes the cleanup operation, it checks the output buffer B out Is it empty? If not, the remaining contents in the buffer are output to the user at once, and then the buffer is cleared; Set the output completion flag O end =1, indicating that the output process has completely stopped. After the output stops completely, the system enters standby mode and waits for the next request or instruction from the user.

8. A method for terminating large model streaming output according to claim 7, characterized in that: The output buffer B out Management includes: In each out Before adding content, check the output termination flag O stop , if O stop =1, skip the current content adding operation and directly enter the next loop; If O stop =0, then continue to perform the content addition operation and update the status of the output buffer; In output buffer B out Before clearing, lock the buffer to prevent data competition in a multi-threaded environment and ensure data consistency and integrity. After clearing the buffer, unlock the buffer to allow other threads to access or modify it.

9. A method for terminating streaming output of a large model according to claim 8, characterized in that: It also includes error handling mechanisms: If an unrecoverable error occurs during the execution of any step, the system shall record an error log and set the error code E to code and error description E desc Send it to the monitoring system. After receiving the error information, the monitoring system takes corresponding measures according to the error type; If an error is encountered during the cleanup operation, the system executes the backup cleanup plan to ensure that resource release and status recovery can be completed even in abnormal situations; After the error processing is completed, set the error processing completion flag H done =1, indicating that the error has been properly handled.

10. A method for terminating streaming output of a large model according to claim 9, characterized in that: The responses of the monitoring system include: After receiving the error information, the monitoring system first determines the severity of the error. For minor errors, only logs are recorded. For serious errors, the emergency response process is immediately executed. The emergency response process includes but is not limited to automatically restarting the service, sending alarm emails or text messages to the administrator, and recording detailed error site information; After the service is successfully restarted, the monitoring system automatically checks the service status to ensure that the service resumes normal operation; If the service restart fails, the monitoring system will continue to try to restart or switch to the backup system until the service is restored; after the service is restored to normal, the monitoring system clears the error status flag E status =0, and notify all relevant parties that the service has returned to normal.

Citation Information

Cited By

  • Text-to-voice conversion method and device

    CN121640987A

  • AI dialogue flow control method, system and device supporting real-time interruption of user, medium and program product

    CN121641022A