Intelligent session task processing method and system based on multi-model collaboration

Through the intelligent session task processing method of multi-model collaboration, the problem of difficult user input in the prior art is solved, and more accurate and efficient session task processing is achieved, and user experience and business efficiency are improved.

CN120373469AInactive Publication Date: 2025-07-25HUNAN XINGONG BOTE INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510828248.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing smart session task processing methods are difficult to fully and accurately understand user input and give appropriate responses, especially when facing diverse and complex user information, resulting in reduced business processing efficiency and user experience.

Method used

An intelligent session task processing method based on multi-model collaboration is adopted. By obtaining user's session information, determining the session task type, and selecting a matching model from the preset model library for processing, integrating multiple model results to generate the final session task results.

Benefits of technology

It improves the accuracy and efficiency of user session task processing, can meet the needs of different users in different scenarios, provides more comprehensive and personalized solutions, and improves business processing efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373469A_ABST
    Figure CN120373469A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent session task processing method and system based on multi-model collaboration, and relates to the field of data processing, and the method comprises the steps: obtaining session information of a user in response to a received session task initiated by the user; based on the session information, determining a session task type of the user; selecting at least one matching model in a preset model library according to the session task type; inputting the session information into at least one matching model, and outputting at least one session task result; and under the condition that the session task results are multiple, integrating all the session task results to obtain a final session task result. The service processing efficiency and the user experience can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of data processing, and in particular, to an intelligent conversation task processing method and system based on multi-model collaboration. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent conversation task processing has been widely applied in many fields such as intelligent customer service, smart home, and autonomous driving. The core of intelligent conversation task processing lies in understanding the user's intention and giving an accurate and efficient response. Therefore, improving the understanding ability and response quality of intelligent conversation systems has become the focus of research and practice.

[0003] Existing intelligent conversation task processing methods still face many challenges. The information input by users is often diverse and complex, including different language styles, expression methods, and potential intentions. This makes it difficult for a single model to comprehensively and accurately understand the user input and give an appropriate response. For example, the retrieval model may be limited by the coverage of the predefined corpus and cannot handle novel or special user queries. Although language models and dialogue generation models can generate natural language responses, they still have certain limitations in dealing with complex contexts and implicit intentions, resulting in users being unable to accurately obtain the required response or solution, thereby reducing the business processing efficiency and user experience. Summary of the Invention

[0004] Embodiments of the present application provide an intelligent conversation task processing method and system based on multi-model collaboration, which are used to improve the business processing efficiency and user experience.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions: In a first aspect, there is provided an intelligent conversation task processing method based on multi-model collaboration, which is applied to a conversation task processing system. The method includes: Responding to receiving a conversation task initiated by a user, obtaining the conversation information of the user; Based on the conversation information, determining the type of the conversation task of the user; In a preset model library, selecting at least one matching model according to the type of the conversation task; Inputting the conversation information into at least one of the matching models, and outputting at least one conversation task result; In the case where there are multiple conversation task results, integrating all the conversation task results to obtain a final conversation task result.

[0006] In another possible implementation manner of the first aspect, the type of the conversation task includes picture task processing, and the picture task processing includes image recognition and image generation, including: In response to receiving the picture uploaded by the user, obtain the URL and format type of the picture; Store the URL and the format type of the picture in a fixed entity class; In the case where the picture task is processed into the image recognition or the image generation, call the image AI model in the preset model library to process the picture to obtain a picture processing result, where the image AI model is used to process the instructions of the image recognition or the image generation; Save the picture processing result in the fixed entity class and transmit it to the user side through a first instruction.

[0007] In another possible implementation manner of the first aspect, the session task type includes document task processing, including: In response to receiving the document uploaded by the user, obtain the URL and document type of the document; Store the URL and the document type of the document in a fixed entity class; Call a file download instruction and a file reading instruction to download and read the document content of the document; In the case where the document content is blank, generate a session title and record the session title; In the preset model library, call the document AI model corresponding to the document task processing; Input the document content into the document AI model to obtain a document processing result; Save the document processing result in the fixed entity class and transmit it to the user side through a second instruction.

[0008] In another possible implementation manner of the first aspect, the session task type includes long text task processing, including: In response to receiving the long text uploaded by the user, obtain the text length of the long text; In the case where the text length is less than the preset text length threshold, perform segmentation processing on the long text through the session task processing system to obtain segmented text; In the preset model library, call the long text AI model corresponding to the long text task processing; Input the segmented text into the long text AI model to obtain a long text processing result; Save the long text processing result in the fixed entity class and transmit it to the user side through a third instruction.

[0009] In another possible implementation manner of the first aspect, the method further includes: When the text length is greater than a preset text length threshold, in response to receiving abnormal text, obtain the abnormal information of the abnormal text; Through the session task processing system, issue a preset abnormal class; Package the abnormal information into a preset abnormal class and return the abnormal class to the client.

[0010] In another possible implementation manner of the first aspect, determining the session task type of the user based on the session information includes: Extract keywords in the session information and count the frequency of occurrence of the keywords in the session information to obtain task type indicator words in the session information; Based on the task type indicator words, use a preset natural language processing technology to identify the session task type of the user.

[0011] In another possible implementation manner of the first aspect, extracting keywords in the session information and counting the frequency of occurrence of the keywords in the session information to obtain task type indicator words in the session information includes: Through a preset word segmentation algorithm, split the session information into multiple phrases to obtain keywords in the session information; Traverse each keyword and count the number of times the keyword appears; Use the keywords whose number of occurrences exceeds a preset threshold as candidate task type indicator words and obtain the types of the candidate task type indicator words; Perform context analysis on the session information to obtain the context of the session information; Combine the context and the types of the candidate task type indicator words to obtain task type indicator words in the session information.

[0012] In another possible implementation manner of the first aspect, selecting at least one matching model according to the session task type in a preset model library includes: When the session task type is single, select a corresponding matching model from a preset model library according to the session task type; When the session task type is multiple, use each session task type as multiple sub-task types; For any one of the sub-task types, select a corresponding matching model in a preset model library; Arrange each of the matching models in a preset order to obtain a matching model combination, and the matching model combination is used to represent the matching model.

[0013] In a second aspect, the present application provides a machine-readable storage medium having instructions stored thereon for causing a machine to execute the above-mentioned intelligent conversation task processing method based on multi-model collaboration.

[0014] In a third aspect, the present application provides an intelligent conversation task processing system based on multi-model collaboration, including: a memory configured to store instructions; and a processor configured to call the instructions from the memory and capable of implementing the above-mentioned intelligent conversation task processing method when executing the instructions.

[0015] Through the above technical solutions, obtaining the user's conversation information and determining the conversation task type can more accurately understand the user's needs and intentions. And selecting a matching model from the preset model library can ensure that the most suitable model is used to process specific conversation tasks, thereby improving the accuracy and efficiency of user conversation task processing. In addition, by processing various types of conversation tasks, such as long text, document, image recognition and other tasks, the needs of different users in different scenarios can be met. At the same time, by integrating the results of multiple models, tasks involving multiple fields or multiple skills can be comprehensively processed, which can not only provide users with more comprehensive and personalized solutions, but also improve the business processing efficiency and user experience.

[0016] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flowchart of an intelligent conversation task processing method based on multi-model collaboration provided by an embodiment of the present application; Figure 2 is a schematic flowchart of a picture task processing provided by an embodiment of the present application; Figure 3 is a schematic flowchart of a document task processing provided by an embodiment of the present application; Figure 4 is a schematic flowchart of a long text task processing provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the specific implementation manners described herein are only for explaining and interpreting the embodiments of this application, and are not used to limit the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0019] It should be noted that if there are directional indications (such as up, down, left, right, front, back,...) involved in the embodiments of this application, then such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If this specific posture changes, then the directional indications will also change accordingly.

[0020] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0021] Figure 1 A flowchart schematically shows a method for processing intelligent conversation tasks based on multi-model collaboration according to an embodiment of this application. As Figure 1 shown, an embodiment of this application provides a method for processing intelligent conversation tasks based on multi-model collaboration, which is applied to a conversation task processing system. The method may include the following steps.

[0022] S110. In response to receiving a conversation task initiated by a user, obtain the user's conversation information; S120. Based on the conversation information, determine the type of the user's conversation task; S120. In a preset model library, select at least one matching model according to the type of the conversation task; S130. Input the conversation information into at least one matching model and output at least one conversation task result; S140. In the case where there are multiple conversation task results, integrate all the conversation task results to obtain a final conversation task result.

[0023] First, in response to receiving a session task initiated by a user, the system obtains the user's session information. That is to say, when the system detects a user's request, it needs to obtain the specific information provided by the user. In this embodiment, the session information may include various forms such as the user's text input, voice input, and image input. After obtaining the user's session information, the system will parse and process this information so that it can correctly understand and respond to the user's needs later.

[0024] Subsequently, based on the session information, determine the type of the user's session task, which can be achieved through natural language processing technology. Natural language processing technology is an important branch in the fields of artificial intelligence and computer science. It can use a computer to process, analyze, understand, and generate natural language text and speech, so as to achieve interaction with humans in various applications. Specifically, allocate the session information input by the user to predefined categories, and each category corresponds to a type of session task. The session information may include text, images, etc. Text input refers to the text information input by the user through the keyboard, and image input refers to the images provided by the user through the camera or other image input devices, which may be used for tasks such as image recognition. Next, identify the intention in the text input by the user through natural language processing technology, that is, what operation the user wants to perform or what information the user wants to obtain.

[0025] In the preset model library, select at least one matching model according to the type of the session task. In this embodiment, the preset model library can be determined according to the actual situation and may store a collection of various NLP models. The type of the session task refers to the task category to which the session initiated by the user belongs. For example, the user may be having various types of conversations such as chatting, asking questions, requesting services, expressing emotions, etc. The system needs to determine the type of the session task according to the input content of the user. Specifically, first, the system needs to identify the type of the user's session task, which can be achieved through text classification or intention recognition methods in natural language processing technology. The system will match the input text of the user with the predefined task types to determine the type of the session task. Next, the system will search for a model matching this task type in the preset model library. After finding a matching model, it is necessary to select the most suitable model to process the current session task. The selection criteria may include multiple aspects such as the accuracy, efficiency, and resource consumption of the model. Subsequently, load this model and prepare it to process the session task.

[0026] After determining the matching model, the conversation information is input into at least one matching model, and at least one conversation task result is output. Specifically, based on the task type of the conversation information, at least one matching model is selected from a preset model library. The matching model may be constructed based on methods such as deep learning, statistical learning, or traditional rules, and is specifically used to process specific types of natural language tasks. Next, the conversation information is input into the matching model for processing. The matching model will parse, understand, and generate the input according to its internal structure and training data. After processing the conversation information, the matching model will generate at least one conversation task result, which may be a generated text (such as answering a user's question), a structured information (such as an entity relationship extracted from the text), etc., and present the result to the user in an appropriate form according to the task type.

[0027] In the case where there are multiple conversation task results, all the conversation task results are integrated to obtain the final conversation task result. Specifically, through natural language processing technology, it is necessary to collect the output results of all models, use NLP technology to extract the key information in each result, and fuse the extracted key information, which can be achieved in various ways, such as weighted average, voting mechanism, sequence combination, etc. The choice of the specific method depends on the task type, the characteristics of the model output, and the integration goal.

[0028] Through the above technical solutions, obtaining the user's conversation information and determining the conversation task type can more accurately understand the user's needs and intentions. And selecting a matching model from the preset model library can ensure that the most suitable model is used to process specific conversation tasks, thereby improving the accuracy and efficiency of user conversation task processing. In addition, by processing various types of conversation tasks, such as long text, document, image recognition, etc. tasks, it can meet the needs of different users in different scenarios. At the same time, by integrating the results of multiple models, tasks involving multiple fields or multiple skills can be comprehensively processed, which can not only provide users with more comprehensive and personalized solutions, but also improve the business processing efficiency and user experience.

[0029] In one implementation manner of this embodiment, the conversation task type includes picture task processing, and the picture task processing includes image recognition and image generation, such as Figure 2 shown, Figure 2 schematically shows a flowchart of a picture task processing according to an embodiment of the present application. The method may include the following steps: S210. In response to receiving a picture uploaded by the user, obtain the URL and format type of the picture; S220. Store the URL and format type of the picture into a fixed entity class; S230. When the picture task is processed into image recognition or image generation, call the image AI model in the preset model library to process the picture and obtain the picture processing result, where the image AI model is used to process instructions for image recognition or image generation; S240. Save the picture processing result to a fixed entity class and send it to the user side through the first instruction.

[0030] In response to receiving a picture uploaded by the user, obtain the URL and format type of the picture. The URL is a string used to identify a resource on the network. Specifically, the user selects and uploads a picture file through a front-end interface (such as a web page, mobile application, etc.). The front-end code usually encapsulates the picture file in an HTTP request and sends it to the back-end server through a POST request. After receiving the picture file, the back-end system may choose to temporarily store the picture file in the server's file system. According to the storage location of the picture file, the back-end system generates a URL that can access the picture. If the picture is stored in the local file system, the URL may be a link pointing to the static resource path of the server; if it is stored in a cloud storage service, the URL is the access link provided by the cloud storage service. The format type of the picture can be obtained by checking the MIME type of the file. The MIME type is sent through the Content-Type header in the HTTP request and is used to indicate the media type of the resource. For example, the MIME type of a JPEG picture is image / jpeg, and the MIME type of a PNG picture is image / png.

[0031] Subsequently, store the URL and format type of the picture in a fixed entity class. In this embodiment, the fixed entity class refers to SessionDetail, which is an entity class used to store detailed information related to the user session. This information may include session ID, task ID, user ID, model type, model function type, question content, answer content, picture URL, picture format, etc. That is, store the URL and format type of the picture in SessionDetail.

[0032] In the case where the picture task is processed into image recognition or image generation, the image AI model in the preset model library is called to process the picture, and the picture processing result is obtained. Among them, the image AI model is used to process the instructions of image recognition or image generation. In this embodiment, image recognition refers to analyzing and understanding the input picture to identify information such as objects, scenes, and text in the picture. For example, identifying the license plate number, face, object category, etc. in the picture; image generation refers to generating a new picture according to the input instructions or data. For example, generating an image according to a text description, generating an artistic style picture according to a style transfer algorithm, etc. The image AI model is an AI model for processing image tasks and can be a convolutional neural network (CNN), a generative adversarial network (GAN), etc. based on deep learning. Specifically, when the system receives an image recognition or image generation task, it will select a suitable image AI model from the preset model library, and the selected model will process the input picture (or related instructions) to execute the image recognition or image generation task. After the processing is completed, the model will output the processing result. For the image recognition task, the result may be information such as the identified object category and location; for the image generation task, the result may be the generated new picture.

[0033] Subsequently, the picture processing result is saved to a fixed entity class and transmitted to the user side through the first instruction. In this embodiment, the fixed entity class refers to SessionDetail; the first instruction refers to ChatGPTCreateResponse. That is to say, the picture processing result is saved to the entity class, and this result is transmitted to the user side through ChatGPTCreateResponse. The ChatGPTCreateResponse class is mainly used to encapsulate the result after processing the ChatGPT request, including the session ID, task ID, model type, model function type, processing result (such as the generated text, picture URL, etc.), error information, etc. This class is used as a response entity to construct the response body of the API so as to return the processing result to the user side.

[0034] Through the above process, the efficient processing of the picture task can be realized, the user experience can be improved, the system intelligence level can be enhanced, and it is helpful for data management and analysis.

[0035] In one implementation manner of this embodiment, the session task type includes document task processing, such as Figure 3 shown, Figure 3 Schematically shows a flow diagram of a document task processing according to an embodiment of the present application. The method may include the following steps: S310. In response to receiving a document uploaded by a user, obtain the URL and document type of the document; S320. Store the URL and document type of the document in a fixed entity class; S330. Invoke the file download instruction and the file reading instruction to download and read the document content of the document; S340. Generate a session title and record the session title when the document content is blank; S350. Invoke the document AI model corresponding to the document task processing in the preset model library; S360. Input the document content into the document AI model to obtain the document processing result; S370. Save the document processing result in a fixed entity class and transmit it to the user side through the second instruction.

[0036] In response to receiving a document uploaded by the user, obtain the URL and document type of the document. The URL is a string used to identify a resource on the network. Specifically, the user selects and uploads a document through a front-end interface (such as a web page, a mobile application, etc.). When the document is successfully uploaded, a URL (Uniform Resource Locator) will be generated, and the URL will be used to reference the document in subsequent processing or access. In addition to the URL of the document, it is also necessary to determine the document type. The document type is usually determined by the file extension (such as.txt,.pdf,.jpg, etc.) or the metadata in the file content.

[0037] After obtaining the URL and document type of the document, store the URL and document type of the document in a fixed entity class. In this embodiment, the fixed entity class refers to SessionDetail, which is an entity class used to store detailed information related to the user session. These information may include session ID, task ID, user ID, model type, model function type, question content, answer content, picture URL, picture format, etc. That is to say, store the URL and document type of the document in SessionDetail.

[0038] Call the file download instruction and the file reading instruction to download and read the document content. In this embodiment, the file download instruction refers to HttpSupport.doGetDownload, and the file reading instruction is DocumentUtil.readFile. HttpSupport.doGetDownload is a method call used to download files from a remote server or cloud storage service; DocumentUtil.readFile is a method call used to read the content of the downloaded file. This method may be a general file reading tool method that supports reading of multiple file formats. That is, download the document content through HttpSupport.doGetDownload and use DocumentUtil.readFile to read the document content.

[0039] In the case where the document content is blank, generate a session title and record the session title. That is, when the content of the document is blank, generate a session label through the system to mark the document as a blank document and record the session title for use in subsequent processing flows. Specifically, store the title in a database or cache and associate it with other relevant information of the session (such as session ID, user information, etc.).

[0040] In the preset model library, call the document AI model corresponding to document task processing. In this embodiment, the preset model library can be determined according to the actual situation. The document AI model is an AI model used to process and analyze document data. That is, in the preset model library, based on document task processing, call the AI model for processing text tasks.

[0041] Subsequently, input the document content into the document AI model to obtain the document processing result. Specifically, the document AI model will extract key features or information from the document content. These features may include keywords in the text, the grammatical structure of sentences, objects or scenes in images, etc. Based on the extracted features, the model will conduct in-depth understanding and analysis of the document content. According to the requirements of the task, the model will generate corresponding outputs, which may be a summary, classification label, entity list, relationship diagram, etc., thus obtaining the document processing result.

[0042] Save the document processing result into a fixed entity class and send it to the client through a second instruction. In this embodiment, the fixed entity class refers to SessionDetail; the second instruction refers to ChatGPTCreateResponse. That is to say, the document processing result is saved into the entity class, and this result is sent to the client through ChatGPTCreateResponse. The ChatGPTCreateResponse class is mainly used to encapsulate the result after processing the ChatGPT request, including the session ID, task ID, model type, model function type, processing result (such as the generated text, picture URL, etc.), error information, etc. This class, as a response entity, is used to construct the response body of the API to return the processing result to the client.

[0043] Through the above steps, efficient processing of document tasks can be achieved, improving the document processing efficiency, providing an intelligent document analysis service, supporting flexible task processing, improving data management and utilization, which can not only enhance the user experience but also improve the efficiency of document processing.

[0044] In one implementation manner of this embodiment, the session task type includes long text task processing, such as Figure 4 shown Figure 4 Schematically shows a flowchart of a long text task processing according to an embodiment of the present application. The method may include the following steps: S410. In response to receiving a long text uploaded by the user, obtain the text length of the long text; S420. In the case where the text length is less than a preset text length threshold, perform segmentation processing on the long text through a session task processing system to obtain the segmented text; S430. In a preset model library, call a long text AI model corresponding to the long text task processing; S440. Input the segmented text into the long text AI model to obtain a long text processing result; S450. Save the long text processing result into a fixed entity class and send it to the client through a third instruction.

[0045] In response to receiving a long text uploaded by a user, obtain the text length of the long text, which can be achieved through a session task processing system. The session task processing system is a system capable of processing various session tasks initiated by users (such as text input, picture upload, file upload, etc.). That is to say, the system can check whether the text length exceeds a preset limit to ensure that subsequent processing will not go wrong or experience performance degradation due to the text being too long. The system receives the long text uploaded by the user through the front-end interface. In the background, the system uses the string processing functions provided by the programming language (such as String.length() in Java, len() in Python, etc.) to calculate the text length. And based on the text length, determine whether the length of the text meets the processing standard.

[0046] In the case where the text length is less than the preset text length threshold, through the session task processing system, perform segmented processing on the long text to obtain the segmented text. In this embodiment, the preset text length threshold can be determined according to the actual situation. That is to say, the session task processing system first checks the length of the received long text. If the text length is less than the preset text length threshold, no segmented processing is required and subsequent processing or display can be directly performed; if the text length exceeds the threshold, segmented processing is required. According to the actual requirements and application scenarios, determine an appropriate segmentation strategy. For example, it can be segmented by paragraph, sentence, number of words, or other logical units. Subsequently, according to the determined segmentation strategy, perform segmented processing on the long text. After the segmented processing is completed, the segmented text is obtained. When implementing segmented processing, the string processing functions provided by the programming language can be used. For example, in Java, regular expressions or string splitting methods can be used to segment by paragraph or sentence.

[0047] After obtaining the segmented text, in the preset model library, call the long text AI model corresponding to the long text task processing. The preset model library is a collection containing various AI models. These models have been trained and optimized and can process different types of tasks, such as image recognition, speech recognition, natural language processing, etc. The long text AI model refers to an AI model specifically used for processing long text tasks. That is to say, according to the task type, the system selects an appropriate long text AI model from the preset model library. Once an appropriate model is selected, the system will call this model to process the long text data. The model will analyze, understand, or generate the input long text and output the processing result.

[0048] Subsequently, input the segmented text into the segmented text to obtain the long text processing result. Specifically, the long text is first segmented according to a certain strategy (such as by paragraph, sentence, number of words, etc.). These segmented texts are input into the long text AI model to obtain the long text processing result.

[0049] Save the long text processing result into a fixed entity class and send it to the client through the third instruction. In this embodiment, the fixed entity class refers to SessionDetail; the third instruction refers to ChatGPTCreateResponse. That is to say, the long text processing result is saved into the entity class and sent to the client through ChatGPTCreateResponse. The ChatGPTCreateResponse class is mainly used to encapsulate the result after processing the ChatGPT request, including session ID, task ID, model type, model function type, processing result (such as generated text, picture URL, etc.), error information, etc.

[0050] Through the above steps, the efficient processing of long text tasks can be achieved, improving the processing efficiency, enhancing the text processing ability, improving the user experience, facilitating data management and analysis. It can not only support the automated processing of long text tasks, but also promote the application and development in the field of long text processing.

[0051] In one implementation manner of this embodiment, the method further includes: S510. When the text length is greater than the preset text length threshold, in response to receiving the abnormal text, obtain the abnormal information of the abnormal text; S520. Send out a preset abnormal class through the session task processing system; S530. Enclose the abnormal information into the preset abnormal class and return the abnormal class to the client.

[0052] When the text length is greater than the preset text length threshold, in response to receiving the abnormal text, obtain the abnormal information of the abnormal text. Specifically, the system first checks whether the text length uploaded by the user exceeds the preset text length threshold. If the text length exceeds the preset text length threshold, the system enters the long text processing flow; if not, it is processed according to the normal process. Once the system detects the abnormal text, it is necessary to obtain the abnormal information of the text. The abnormal information may include the type of error, the location of the error, the description of the error, etc. In this embodiment, the abnormal information refers to that the text length exceeds the preset text length threshold.

[0053] Through the session task processing system, a preset exception class is issued. In this embodiment, the preset exception class can be ErrorCode.TEXT_TOO_LONG, etc. The preset exception classes are defined during the system design phase and are used to identify and describe specific exception situations. These exception classes may contain error information, error codes, the location where the exception occurred, and other information. Specifically, when the session task processing system encounters an exception situation, it will select a suitable exception class to throw according to the preset rules or logic. The purpose of throwing the exception class is to notify the caller (which may be another system, module, or service) that an error has occurred and provide some information about the error. The caller can take corresponding error handling measures based on the received exception class information, such as retrying the operation, logging, notifying the user, etc.

[0054] Next, the exception information is encapsulated into the preset exception class and the exception class is returned to the client. Specifically, when the system encounters an exception situation while processing a user request, these exception information need to be organized in a structured way. An exception class can be created, which contains all the necessary information related to the exception. And the exception information is encapsulated into the exception class, and an instance of the exception class needs to be returned to the client. This is usually achieved through an HTTP response, and the HTTP status code is used to indicate the result of the request.

[0055] By introducing an exception handling mechanism when the text length exceeds the preset threshold, encapsulating the exception information and returning it to the client, this method can significantly improve the robustness of the system, enhance the user experience, facilitate error tracking and debugging, and support flexible exception handling strategies.

[0056] In one implementation manner of this embodiment, based on the session information, the session task type of the user is determined, including: S610. Extract keywords from the session information and count the frequency of the keywords in the session information to obtain the task type indicator words in the session information; S620. Based on the task type indicator words, use the preset natural language processing technology to identify the session task type of the user.

[0057] Based on the session information, determine the type of the user's session task. Specifically, extract the keywords in the session information and count the frequency of the keywords in the session information to obtain the task type indicator words in the session information. First, extract the keywords from the session information. Keywords are usually nouns, verbs, adjectives, etc. in a sentence, which can summarize the main content of the sentence or express the user's core intention. Lexical analysis, syntactic analysis, etc. in natural language processing (NLP) technology can be used to extract keywords. Algorithms such as TF-IDF (Term Frequency-Inverse Document Frequency) and TextRank can also be used. These algorithms can evaluate the importance of vocabulary based on factors such as the frequency of occurrence, position, and association with other vocabulary in the document, so as to extract keywords. After extracting the keywords, count the number of times each keyword appears in the session information. Word frequency statistics can help understand the user's attention to the topic and the key vocabulary in the session information. Traverse each vocabulary in the session information and count the number of times each keyword appears. Specifically, a hash table can be used to store the keywords and their corresponding word frequencies, so as to obtain the word frequency of the keywords. After counting the word frequency of the keywords, the task type indicator words in the session can be identified based on these keywords. Task type indicator words refer to those words that can clearly indicate the user's intention or needs, such as "query", "reservation", "purchase", etc. This can be achieved by constructing a task type indicator word dictionary. First, a dictionary containing common task type indicator words needs to be constructed. This dictionary can be constructed based on business requirements, user feedback, or domain knowledge. Then, match the counted keywords with the task type indicator word dictionary. If a keyword matches a certain task type indicator word in the dictionary, then it can be considered that the session contains this type of task. In addition to direct matching, the word frequency and context information of the keywords can also be combined to judge the task type. For example, if a keyword appears frequently in the session and is associated with a specific context (such as asking about price, time, etc.), then it can be more confidently considered that the session contains a task related to this keyword, so as to obtain the task type indicator words in the session information.

[0058] Subsequently, based on the task type indicator, using preset natural language processing techniques, the session task type of the user is identified. Natural language processing techniques can include various means such as part-of-speech tagging, syntactic analysis, and semantic role labeling. For example, in syntactic analysis, the user's actions or behaviors can be determined by analyzing the components of a sentence, such as the subject, predicate, and object; in semantic role labeling, semantic roles such as agent, patient, and instrument in the sentence can be identified to further understand the user's intention. Specifically, first, the task type indicator needs to be extracted from the session information. An indicator refers to some words or phrases with clear directivity that can directly reflect the user's intention or need. For example, in a session in the e-commerce field, common task type indicators may include "purchase", "query", "return", etc. The method of extracting the task type indicator can be based on various methods such as keyword matching, rule matching, or machine learning models. In this example, we have extracted keywords from the session information through previous steps and identified potential task type indicators. After identifying the potential task type indicators, natural language processing techniques are used for context understanding, and these information are combined to identify the user's session task type. It can be based on various methods such as rule matching, template matching, or machine learning models. By training a suitable machine learning model, the task type indicator, context information, etc. are used as features for input, so as to output the user's session task type.

[0059] By extracting keywords from the session information, counting the frequency of the keywords to obtain the task type indicator, and using the preset natural language processing techniques to identify the session task type of the user, the accuracy of session task recognition can be significantly improved, the user experience can be enhanced, the intelligent level of the system can be increased, multiple session task types can be supported, and the data analysis and mining capabilities can be improved.

[0060] In one implementation manner of this embodiment, extracting keywords from the session information and counting the frequency of the keywords in the session information to obtain the task type indicator in the session information includes: S710. Split the session information into multiple phrases through a preset word segmentation algorithm to obtain the keywords in the session information; S720. Traverse each keyword and count the number of times the keyword appears; S730. Use the keywords whose number of times exceeds the preset threshold as candidate task type indicators, and obtain the types of the candidate task type indicators; S740. Perform context analysis on the session information to obtain the context of the session information; S750. Combine the context and the types of the candidate task type indicators to obtain the task type indicator in the session information.

[0061] Extract keywords from the conversation information and count the frequency of the keywords in the conversation information to obtain the task type indicator in the conversation information. First, split the conversation information into multiple phrases through a preset word segmentation algorithm to obtain the keywords in the conversation information. Specifically, first, it is necessary to determine the word segmentation algorithm, which can be rule-based word segmentation, statistics-based word segmentation, and a method that combines rules and statistics. The appropriate algorithm can be selected according to the specific application requirements and the characteristics of the conversation information. Subsequently, use the preset word segmentation algorithm to segment the preprocessed conversation information. The word segmentation algorithm will split the conversation information into multiple phrases according to the lexical information and algorithm rules in the preset dictionary. For example, for the sentence "What's the weather like today?", the word segmentation algorithm may split it into phrases such as "today", "the", "weather", "what", and "like". After word segmentation, it is necessary to extract keywords from the obtained phrases. Keywords are usually some representative or important words that can summarize the main content of the conversation information or express the user's core intention. The method of extracting keywords can be based on word frequency statistics, TF-IDF algorithm, TextRank algorithm, etc. In this embodiment, keywords can be extracted by the method of word frequency statistics to obtain the keywords in the conversation information.

[0062] After obtaining the keywords in the conversation information, traverse each keyword and count the number of times the keyword appears. Specifically, traverse the text after word segmentation. For each word or phrase, check whether it is a keyword. If it is, increment the corresponding counter by 1. After counting, output each keyword and the number of times it appears. This can be presented in various forms such as a list, chart, etc., to facilitate an intuitive understanding of the distribution of each keyword in the text.

[0063] Take the keywords whose frequency exceeds the preset threshold as candidate task type indicators and obtain the types of the candidate task type indicators. In this embodiment, the preset threshold can be determined according to the actual situation. After obtaining the number of times each keyword appears, compare these keywords with the preset threshold. If the number of times a certain keyword appears exceeds the threshold, then consider it as a candidate task type indicator. The candidate indicator may represent the main topic or task type in the text. For each selected candidate task type indicator, perform semantic analysis or category annotation on the indicator. Specifically, the type of the indicator can be obtained through a preset category mapping. A mapping table from keywords to task types can be predefined in advance. For each candidate indicator, the mapping table can be searched to determine its corresponding task type. After obtaining the candidate task type indicators and their types, this information can be applied to various application scenarios. For example, in an intelligent customer service system, the user's intention can be quickly identified based on the task type indicator in the user input, and corresponding services or answers can be provided.

[0064] Next, perform a context analysis on the conversation information to obtain the context of the conversation information. Specifically, first, analyze the sentence structure to identify components such as the subject, predicate, and object, as well as the relationships between them. Subsequently, identify the named entities in the text (such as person names, place names, organization names, etc.) and determine the emotional tendency expressed in the text (such as positive, negative, neutral), which helps to understand the user's emotional state. Finally, identify the intention or purpose of the user's input text, such as querying information, expressing emotions, making requests, etc. Combine the previous conversation records to understand the context and background information of the current conversation, and use a domain knowledge base or common sense knowledge base to provide additional information support for the context analysis. Represent the context information in vector form for subsequent calculation and analysis. Use deep learning models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and Transformers to model and analyze the context to obtain the context of the conversation information. For example, suppose there is a conversation information "What's the weather like in Beijing tomorrow?" When performing context analysis, it can be identified that "Beijing" is a place name entity, "tomorrow" is a time word, "weather" is the topic word, and "What's... like" expresses the intention of asking. Combining this information, it can be inferred that the user wants to query the weather conditions in Beijing tomorrow.

[0065] Combine the context and the type of candidate task type indicators to obtain the task type indicators in the conversation information. That is to say, through context association, semantic feature analysis, etc., combine the context analysis results with the type information of the candidate indicators to comprehensively judge which or which words are most likely to be the keywords indicating the task type. Finally, output the task type indicators in the conversation information.

[0066] Through the above steps, not only the accuracy of task type recognition is improved, but also the intelligence level of the system is enhanced, the user experience is improved, and multiple conversation scenarios and task types are supported. By continuously collecting and analyzing user data, it can also provide decision-making support for enterprises and promote the continuous development of business.

[0067] In one implementation manner of this embodiment, in a preset model library, select at least one matching model according to the conversation task type, including: S810. When the conversation task type is single, select the corresponding matching model from the preset model library according to the conversation task type; S820. When the conversation task type is multiple, regard each conversation task type as multiple sub-task types; S830. For any one of the sub-task types, select a corresponding matching model in the preset model library; S840. Arrange each matching model in a preset order to obtain a matching model combination, and the matching model combination is used to represent the matching model.

[0068] In the case where the session task type is single, according to the session task type, select the corresponding matching model from the preset model library. Specifically, it is necessary to establish a mapping relationship between the task type and the model, which can be achieved through a trie tree. A trie tree is a data structure specifically used to handle string matching problems. In the scenario of mapping between task types and models, if the task type can be represented as a string, then a trie tree can be used to store these task types and their corresponding models. According to the task type of the session, the system finds the corresponding model in the mapping relationship and selects it as the processing model for the current session.

[0069] In the case where the session task type is multiple, each session task type is regarded as multiple subtask types. Specifically, when there are multiple task types in the session, these tasks need to be split into multiple independent subtasks. Each subtask corresponds to a specific task type and can be processed by an independent model. For each subtask, the system needs to select a suitable model from the preset model library for processing.

[0070] For any one of the subtask types, select a corresponding matching model from the preset model library. That is to say, for each subtask, it is necessary to select a suitable model from the preset model library for processing, and find the corresponding model in the mapping relationship between the task type and the model according to the task type of the subtask. Since each subtask is independent, models can be assigned to each subtask in parallel and inference calculations can be performed simultaneously.

[0071] Arrange each matching model in a preset order to obtain a matching model combination. The matching model combination is used to represent the matching models. That is to say, after parallel processing of all subtasks, it is necessary to integrate the results of each subtask to generate a final response. Specifically, it includes operations such as sorting, merging, and filtering of subtask results to ensure the accuracy and consistency of the final response. For example, if the subtasks include querying the weather and querying the route, the system may need to process these two tasks separately first, and then integrate the weather information and route information into a response and return it to the user.

[0072] Processing each session task type as multiple subtask types is an effective strategy, which can improve the flexibility and response ability of the system, and provide a more comprehensive and personalized service experience for users. At the same time, through parallel processing of subtasks and optimization of resource utilization, the system can also achieve higher processing efficiency and performance.

[0073] This application also provides a machine-readable storage medium, on which instructions are stored. These instructions are used to make the machine execute the above-mentioned intelligent session task processing method based on multi-model collaboration.

[0074] The present application also provides an intelligent conversation task processing system based on multi-model collaboration, including: A memory configured to store instructions; and A processor configured to call instructions from the memory and, when executing the instructions, be able to implement the above-mentioned intelligent conversation task processing method based on multi-model collaboration.

[0075] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0076] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.

[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.

[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.

[0079] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0080] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0081] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0082] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0083] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. An intelligent conversation task processing method based on multi-model collaboration, characterized in that, Applied to a session task processing system, including: In response to receiving a session task initiated by a user, obtain the session information of the user; Based on the session information, determine the type of the user's session task; When the session task type is single, select a corresponding matching model from a preset model library according to the session task type; When the session task type is multiple, regard each session task type as multiple subtask types; For any one of the subtask types, select a corresponding matching model in a preset model library; Arrange each of the matching models in a preset order to obtain a matching model combination, and the matching model combination is used to represent the matching models; Input the session information into at least one of the matching models, and output at least one session task result; When there are multiple session task results, integrate all the session task results to obtain a final session task result.

2. The method according to claim 1, wherein The session task type includes picture task processing, and the picture task processing includes image recognition and image generation, including: In response to receiving the picture uploaded by the user, obtain the URL and format type of the picture; Store the URL and the format type of the picture into a fixed entity class; When the picture task processing is the image recognition or the image generation, call an image AI model in a preset model library to process the picture to obtain a picture processing result, where the image AI model is used to process instructions for the image recognition or the image generation; Save the picture processing result into the fixed entity class, and send it to the user side through a first instruction.

3. The method according to claim 1, characterized in that, The session task type includes document task processing, including: In response to receiving the document uploaded by the user, obtain the URL and document type of the document; Store the URL and the document type of the document into a fixed entity class; Call a file download instruction and a file reading instruction to download and read the document content of the document; When the document content is blank, generate a session title and record the session title; In a preset model library, call a document AI model corresponding to the document task processing; Input the document content into the document AI model to obtain a document processing result; Save the document processing result into a fixed entity class, and send it to the user side through a second instruction.

4. The method according to claim 1, wherein The session task type includes long text task processing, including: In response to receiving the long text uploaded by the user, obtain the text length of the long text; When the text length is less than a preset text length threshold, perform segmentation processing on the long text through the session task processing system to obtain segmented text; In a preset model library, call a long text AI model corresponding to the long text task processing; Input the segmented text into the long text AI model to obtain a long text processing result; Save the long text processing result into a fixed entity class, and send it to the user side through a third instruction.

5. The method according to claim 4, wherein The method further includes: When the text length is greater than a preset text length threshold, in response to receiving abnormal text, obtain the abnormal information of the abnormal text; Through the session task processing system, issue a preset abnormal class; Package the abnormal information into a preset abnormal class and return the abnormal class to the client.

6. The method according to claim 1, wherein Determining the session task type of the user based on the session information includes: Extract keywords from the session information and count the frequency of occurrence of the keywords in the session information to obtain task type indicator words in the session information; Based on the task type indicator words, use preset natural language processing techniques to identify the session task type of the user.

7. The method according to claim 6, wherein The extracting keywords from the session information and counting the frequency of occurrence of the keywords in the session information to obtain task type indicator words in the session information includes: Through a preset word segmentation algorithm, split the session information into multiple phrases to obtain keywords in the session information; Traverse each keyword and count the number of times the keyword appears; Use the keywords whose number of occurrences exceeds a preset threshold as candidate task type indicator words and obtain the types of the candidate task type indicator words; Conduct a context analysis on the session information to obtain the context of the session information; Combine the context and the types of the candidate task type indicator words to obtain task type indicator words in the session information.

8. A machine-readable storage medium, characterized in that, Instructions are stored on the machine-readable storage medium, and the instructions are used to cause the machine to execute the intelligent session task processing method based on multi-model collaboration according to any one of claims 1 to 7.

9. An intelligent conversation task processing system based on multi-model collaboration, characterized in that, Including: A memory configured to store instructions; And A processor configured to call the instructions from the memory and capable of implementing the intelligent session task processing method based on multi-model collaboration according to any one of claims 1 to 7 when executing the instructions.

Citation Information

Patent Citations

  • Method and device for decomposing and scheduling basic model task with enhanced thinking map prompt

    CN117370638A

  • Question and answer method, device and equipment based on intelligent agent and medium

    CN120162404A

Cited By

  • MaaS multi-model release management method and system and storage medium

    CN120872616A

  • Multi-model dynamic collaborative interpretable strategy generation and evaluation method and system, electronic equipment and storage medium

    CN122088669A