Session processing method and device, computer equipment and program product
By analyzing and splitting the user sessions in a multi-dimensional way, combining the model resource pool information, selecting the most matching AI model and server for processing, the problem of unreasonable AI model selection and server resource scheduling is solved, and the session results that meet users' expectations are achieved.
Patent Information
- Application Number
- CN202510740479.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In AI model selection and server resource scheduling, the prior art cannot reasonably schedule the AI model and select servers, resulting in the inability to provide session processing results that meet users' expectations.
By performing multi-dimensional analysis of user sessions, obtaining multi-dimensional features and session complexity, filtering the initial model and candidate servers using the model resource pool, splitting the session based on the number of models and session complexity, selecting the most matching target model and server for processing, and integrating the results.
Based on server resource load balancing, it is realized to rationally utilize model capabilities and server resources to provide session results that best meet users' expectations, avoiding the problem of unreasonable server scheduling and model use.
Smart Images

Figure CN120256591A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a conversation processing method, apparatus, computer device, and program product. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, various AI models have emerged and are widely used in natural language processing, knowledge question answering, task processing and other fields. However, since there may be one or more servers for AI models to choose and use, and different AI models have different areas of expertise, when a user initiates a session request, how to reasonably schedule the AI model and select the server that provides service resources for the AI model for session processing has become a technical issue worthy of attention. Summary of the invention
[0003] The embodiments of the present disclosure at least provide a session processing method, apparatus, computer device, and program product.
[0004] In a first aspect, an embodiment of the present disclosure provides a session processing method, including: Using the conversation analysis model, the user conversation is analyzed in multiple dimensions to obtain the multi-dimensional features and conversation complexity corresponding to the user conversation; According to the multi-dimensional features, searching for each initial model matching the user session and each candidate server that can be used by the initial model from the latest model resource pool; wherein the model resource pool is constructed according to the model capability information of each session model and the resource information of the initial server supporting the session model; When the number of the initial models and the session complexity indicate that the user session requires multi-model processing, split the user session into multiple sub-sessions, and screen out a target model matching each sub-session and a target server required to be used by the target model from the initial model and the candidate servers; Each of the target servers is scheduled to run the target model to process each of the sub-sessions, obtain a sub-session result corresponding to each of the sub-sessions, and determine a session result corresponding to the user session according to each of the sub-session results.
[0005] In a possible implementation, the multiple dimensions include at least a conversation domain dimension, a conversation question type dimension, and a conversation question expertise dimension.
[0006] In a possible implementation, the model resource pool is constructed according to the following steps: For any conversation model, collect the resource information of each initial server that supports the conversation model in real time, and determine the availability score of the conversation model on each initial server according to the resource information, the reserved resource requirements, the resource usage requirements corresponding to the conversation model, and the historical resource usage information; Obtain the model capability information corresponding to the conversation model from the model capability label library; the model capability information includes the capability indication labels and the proficiency levels of various capabilities possessed by the conversation model; Construct the model resource pool according to the availability score of each conversation model on each initial server and the model capability information.
[0007] In a possible implementation manner, the step of finding each initial model and each candidate server available for the initial model that match the user conversation from the latest model resource pool according to the multi-dimensional features includes: Determine the first weighted weight of each dimension feature in the multi-dimensional features; Determine the first matching degree between each conversation model on each initial server and the user conversation according to the availability score of each conversation model on each initial server and the model capability information in the latest model resource pool, and the first weighted weight of each dimension feature; Filter out each initial model and each candidate server available for the initial model according to the first matching degree.
[0008] In a possible implementation manner, the step of filtering out each initial model and each candidate server available for the initial model according to the first matching degree includes: Determine the second weighted weight of each conversation model relative to the initiating user according to the user attribute information of the initiating user of the user conversation and the weight mapping relationship between the user attributes and the conversation model; Use the second weighted weight of each conversation model to perform weighted processing on the first matching degree between the conversation model on each initial server and the user conversation to obtain the second matching degree; Filter out each initial model and each candidate server available for the initial model according to the second matching degree.
[0009] In a possible implementation manner, the step of filtering out each initial model and each candidate server available for the initial model according to the first matching degree includes: Determine the expected success rate of each conversation model on each initial server for the user conversation according to the historical conversations of each conversation model on each initial server. Determine whether the user session is a new session initiated by the initiating user; If not, determine the model switching loss of each of the session models on each of the initial servers for the user session according to the correlation degree between the user session and the historical session of the initiating user, and the historical session model and historical server used in the historical session; Filter out each of the initial models and each candidate server available for the initial model according to the model switching loss, the expected success rate, and the first matching degree.
[0010] In a possible implementation manner, the determining the session result corresponding to the user session according to each of the sub-session results includes: Determine the model capability information of the target model corresponding to each sub-session in the session domain to which the sub-session belongs; Determine the third weighted weight of the sub-session result corresponding to each sub-session according to the model capability information in the session domain and the importance of the sub-session relative to the user session; Use a text integration model to perform multi-round session integration on each of the sub-session results according to the third weighted weight of each of the sub-session results to obtain the session result.
[0011] In a possible implementation manner, after determining the session result corresponding to the user session, it further includes: Obtain the initial feedback data of the session result; the initial feedback data includes interaction data and usage data for the session result; Determine the feedback score of the session result according to the initial feedback data, and combine the feedback score and the time stamp into intermediate feedback data, and store it in the feedback data pool; For each of the session models, according to the time stamp of each intermediate feedback data in the feedback data pool and a preset time range, obtain the first target feedback data of the session result of the session model in different session domains from the feedback data pool; Update the model capability information of the session model according to the first target feedback data of the session result of the session model in different session domains.
[0012] In a possible implementation manner, the updating the model capability information of the session model according to the first target feedback data of the session result of the session model in different session domains includes: For any session domain, determine the first capability score of the session model in the session domain according to the first target feedback data of the session result of the session model in the session domain; Determine the second ability score of the session model in the session domain according to each first ability score of the session model in the session domain within the update time range; Update the model ability information of the session model according to the second ability scores of the session model in each of the session domains.
[0013] In a possible implementation manner, the method further includes: Obtain second target feedback data corresponding to the processed user sessions of different session models in different session domains from the feedback data pool; Update the weight mapping relationship between the user attributes and the session model according to the second target feedback data and the user attribute information of the initiating user of the processed user session.
[0014] In a second aspect, an embodiment of the present disclosure further provides a session processing apparatus, including: A parsing module, configured to perform multi-dimensional parsing on a user session by using a session analysis model to obtain multi-dimensional features and session complexity corresponding to the user session; A searching module, configured to search for each initial model matching the user session and each candidate server available for the initial model from the latest model resource pool according to the multi-dimensional features; wherein, the model resource pool is constructed according to the model ability information of each session model and the resource information of the initial server supporting the session model; A splitting module, configured to split the user session into multiple sub-sessions when the number of the initial models and the session complexity indicate that the user session needs multi-model processing, and screen out a target model matching each sub-session and a target server required by the target model from the initial models and the candidate servers; A determining module, configured to schedule each of the target servers to run the target model to process each sub-session respectively, obtain sub-session results corresponding to each sub-session, and determine a session result corresponding to the user session according to each sub-session result.
[0015] In a third aspect, an alternative implementation manner of the present disclosure further provides a computer device, including a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the steps in the first aspect or any possible implementation manner in the first aspect are executed.
[0016] In a fourth aspect, an optional implementation of the present disclosure also provides a computer program product, including a computer program which, when run, implements the steps in the first aspect or any possible implementation manner in the first aspect.
[0017] The session processing method, apparatus, computer device, and program product provided by the embodiments of the present disclosure can obtain the dimensional features of a user session from different perspectives and reasonably estimate the complexity of the session by performing multi-dimensional parsing on the user session. By using the multi-dimensional features, as well as the resource information and model capability information of the initial servers supporting the session model, it is possible to initially screen out various initial models that can be used for session processing and candidate servers that can be used by the initial models on the basis of ensuring the load balance of server resources. Splitting the user session based on the number of models and session complexity can achieve fine-grained division of complex sessions to obtain refined sub-sessions when supported by the session model. Then, by separately screening out the target models and target servers that best match each sub-session from the initial models, it is possible to rationally utilize the model capabilities of each target model and server resources, avoiding unreasonable problems in server scheduling and model usage. Finally, by integrating the results of each sub-session, a session result that best matches the user session can be obtained. In the entire solution, by parsing and splitting the user session and combining the resource information and model capability information of each session model, the session model and the initial server are scheduled, so that the capabilities of each available model and server resources can be rationally utilized to process the session, thereby obtaining a session result that best meets the user's expectations on the basis of ensuring relatively balanced server load.
[0018] For the description of the effects of the above session processing apparatus, computer device, and computer program product, refer to the description of the above session processing method, which will not be elaborated here.
[0019] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically described below in conjunction with the accompanying drawings. Description of the Drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. The accompanying drawings are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 The figure shows a flowchart of a session processing method provided by an embodiment of the present disclosure; Figure 2 The figure shows a specific implementation flowchart of a session processing method provided by an embodiment of the present disclosure; Figure 3 The figure shows a schematic diagram of a session processing apparatus provided by an embodiment of the present disclosure; Figure 4 The figure shows a schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some, but not all, of the embodiments of the present disclosure. The components of the embodiments of the present disclosure described and illustrated herein generally may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure that is claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0023] In addition, the terms "first", "second", etc. in the description, claims, and the above-mentioned drawings of the embodiments of the present disclosure are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way may be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein.
[0024] As used herein, "a plurality of" or "several" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0025] It has been found that with the continuous development of AI large model technology, large models of various scales and capabilities have emerged, such as some common large models like the 7 billion (7B) parameter large model, 32B parameter large model, 70B parameter large model, 617B parameter large model, 1.56-bit dynamic quantization large model, q4 standard quantization model, etc. In the initial stage of the application of AI large models, when users have a need to use AI models, the commonly adopted method is to manually select the most suitable target AI from a large number of AI models, and then input the session information into the target AI to obtain the required session result. However, this method relies on the user's understanding of the model, and there are often problems where the model selected by the user cannot provide accurate answers. Subsequently, some inference scheduling systems for multi-model environments emerged. Considering that the current common AI large model deployment schemes mainly include two methods: full Graphics Processing Unit (GPU) inference and Central Processing Unit (CPU) / GPU hybrid inference, these systems usually schedule models to handle requests based on the priority of the inference engine and the device status when receiving an external request from the user. However, this model scheduling strategy is too simple, which will lead to unreasonable scheduling of both the model server and the session model, and cannot provide users with response results that meet their expectations.
[0026] Based on the above research, the present disclosure provides a session processing method, apparatus, computer device, and program product. By performing multi-dimensional analysis on user sessions, dimensional features of user sessions from different perspectives can be obtained, and a reasonable estimation of session complexity can be achieved. Using the multi-dimensional features, as well as the resource information and model capability information of the initial servers supporting the session models, it is possible to preliminarily screen out each initial model that can be used for session processing and the candidate servers that can be used by the initial models on the basis of ensuring the load balance of server resources. Splitting the user session based on the number of models and session complexity can achieve a fine-grained division of complex sessions and obtain refined sub-sessions when supported by session models. Then, by separately screening out the target models and target servers that best match each sub-session from the initial models, a reasonable utilization of the model capabilities of each target model and server resources can be achieved, avoiding unreasonable problems in server scheduling and model usage. Finally, by integrating the results of each sub-session, a session result that best matches the user session can be obtained. In the entire solution, by analyzing and splitting the user session and combining the resource information and model capability information of each session model, the session models and initial servers are scheduled, and the capabilities of each available model and server resources can be reasonably utilized to process the session, thereby obtaining a session result that best meets the user's expectations on the basis of ensuring relatively balanced server load.
[0027] Regarding the defects existing in the above solutions, they are all the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should be the contributions made by the inventors to the present disclosure during the process of the present disclosure.
[0028] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0029] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0030] To facilitate the understanding of this embodiment, a session processing method disclosed in the embodiments of the present disclosure will be introduced in detail first. The execution subject of the session processing method provided in the embodiments of the present disclosure is generally a terminal device or other processing device with certain computing capabilities. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a personal digital assistant device (PDA), a handheld device, a computer device, etc. In some possible implementation manners, the session processing method may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0031] Next, the session processing method provided in the embodiments of the present disclosure will be described by taking the execution subject as the server as an example.
[0032] As Figure 1 shown, it is a flowchart of a session processing method provided in the embodiments of the present disclosure, which may include the following steps: S101: Use a session analysis model to perform multi-dimensional parsing on a user session to obtain multi-dimensional features and session complexity corresponding to the user session.
[0033] Here, the user session is a question to be answered submitted by the user. The session analysis model can be a lightweight analysis model, which is used to perform multi-dimensional parsing and complexity analysis on the user session to obtain multi-dimensional features and session complexity of the user session. Exemplarily, the session analysis model can be a model fine-tuned based on a distillation model, which can be used to identify and classify user sessions, so as to realize disassembling user sessions into information required for subsequent steps.
[0034] The multi-dimensions can at least include a session domain dimension, a session question type dimension, and a session question professional degree dimension. The multi-dimensional features can specifically include session features under each dimension. Among them, the dimension features under the session domain dimension are used to characterize the domain attribute of the user session, such as the technical field, the management field, the creative field, the document production field, the drawing field, the scientific field, the engineering field, the official document writing field, the solution generation field, the news field, the knowledge retrieval field, etc. The dimension features under the session question type dimension are used to characterize the question type of the user session, such as the analysis type, the decision type, the question and answer type, the creation type, etc. The dimension features under the session question professional degree dimension are used to characterize the professional degree of the user session, such as the entry level, the primary level, the intermediate level, the advanced level, the professional level, etc.
[0035] The dimensional features under each dimension can specifically be the dimensional feature vectors output by the conversation analysis model, and the multi-dimensional features can be the multi-dimensional feature vectors output by the conversation analysis model. The conversation complexity is used to characterize the computational complexity of the user conversation. The higher the computational complexity, the higher the processing difficulty of the user conversation, and the higher the requirement for the capabilities of the conversation model to be used.
[0036] In specific implementation, the server can receive the user conversation submitted by the user, and then schedule the conversation analysis model to perform multi-dimensional parsing and complexity analysis on the user conversation to obtain multi-dimensional features and conversation complexity. For example, based on the analysis result of the conversation analysis model, a multi-dimensional feature vector can be obtained. This vector contains the core feature information of the problem and is in a standardized format for subsequent processing. For example: [Domain: Technical field: Confidence 0.9; Sub-domain: AI architecture field: Confidence 0.8; Problem type: Analysis type: Confidence 0.75; Conversation complexity: Medium, Confidence 0.6; Professional level: Advanced: Confidence 0.8; Computational amount: Medium: Confidence 0.5]. This multi-dimensional feature vector will be an important basis for subsequent model selection and task decomposition.
[0037] Optionally, the conversation processing method provided in this embodiment of the present disclosure can be executed by an agent set on the server.
[0038] S102: According to the multi-dimensional features, find each initial model that matches the user conversation and each candidate server available for the initial model from the latest model resource pool; where the model resource pool is constructed according to the model capability information of each conversation model and the resource information of the initial servers that support the conversation models.
[0039] Here, the server can call multiple AI conversation models. Different conversation models have different processing capabilities, and the deployment schemes of the conversation models can include two main methods: full GPU inference and CPU / GPU hybrid inference. The initial server can be understood as the server that supports the operation of the conversation model. One conversation model can have one or more initial servers that support its operation, and one server can support the operation of one or more conversation models. When there are multiple initial servers that support a conversation model, when it is determined that the conversation model needs to be used, one server can be selected from the multiple initial servers as the server required for using the conversation model this time.
[0040] The initial model is a part of the conversation models selected from the conversation models included in the model resource pool for this user conversation, and the candidate server is a part of the servers selected from the initial servers that can be used to run the initial model for this user conversation.
[0041] The model resource pool may include at least one session model and at least one initial server that currently supports the operation of the session model, and each of the included session models and initial servers is constructed based on the resource information of all the initial servers that support the session model and the model capability information of the session model. The resource information is used to characterize the operating state data of the initial servers that support the session model, and the operating state data may include, but is not limited to, information such as the CPU usage rate, GPU load, memory usage, and network bandwidth occupancy of the initial servers. The GPU load may specifically include information such as computing load and video memory occupancy.
[0042] To facilitate the collection of the resource information of each session model, it can be implemented using a distributed monitoring probe network. In this network, lightweight monitoring agents can be included for the servers that support the operation of each session model, and real-time collection of the resource information can be achieved based on these agents. Optionally, in order to improve the real-time performance of resource information collection, the distributed monitoring probe network can be implemented using a low-latency message queue mechanism. Among them, the interval of information collection can be a preset interval, and the preset interval can be set according to experience. The embodiments of the present disclosure do not make specific limitations. For example, the preset interval can be 0.5 seconds, 1 second, etc.
[0043] The model capability information may include the basic parameter information of the session model, the capability indication tags, and the proficiency levels of the model capabilities under each capability indication tag. Among them, the basic parameter information may include the number of parameters of the model, the quantization type of the model, performance indicators (such as concurrency ability, number of tokens per second, thinking, answering, and outputting tokens), etc. The basic parameter information can also be collected in real time using a distributed monitoring probe network. A session model may include at least one capability indication tag, and the proficiency level can be represented by a normalized score of 0-1. For example, the capability indication tags of a certain session model may include [main field: computer field, proficiency level 0.9; sub-field: operating system, proficiency level 0.9; sub-field: AI architecture, proficiency level 0.85; sub-field: compilation principle, proficiency level 0.95]. Among them, the capability indication tags and proficiency levels of each session model can be dynamically updated, and the update process will be introduced in detail later. Exemplarily, the server side can maintain a dynamically updated model capability tag library, which records the capability indication tags of each session model and the proficiency levels of the capabilities under these tags.
[0044] In specific implementation, the server can determine the currently available session models and the initial servers that these session models can use based on the resource information of each initial server supporting each session model collected during implementation, the basic parameter information of the session models, and the ability indication tags and ability proficiency levels of each session model recorded in the model ability tag library. Then, it can construct a model resource pool based on these session models and the initial servers that can be used. For example, for each session model, it can determine the initial servers whose resource information meets the set conditions based on the resource information of each initial server supporting the session model. And when the computing power of the session model meets the set ability of the session model, it can use this session model and the initial servers that meet the conditions as the session models and the initial servers supporting the session models in the model resource pool. Since the resource information of each initial server supporting the session model, the ability indication tags of the session model, and the ability proficiency levels are all updatable information, the session models and the initial servers supporting the session models included in the model resource pool are also dynamically changing, and the number of session models and the number of initial servers supporting the session models included in the model resource pool are also dynamically changing. For example, the number of session models in the model resource pool can be 0 to N, where N is the number of session models that the server can schedule, and the number of initial servers supporting the operation of each session model can be 1 to M, and M is the maximum number of initial servers supporting the operation of the session model. When the number of session models and initial servers in the model resource pool is 0, after waiting for a set time, it can be determined again whether there are session models and initial servers in the latest model resource pool. If so, it can screen the initial models and candidate servers for these session models and initial servers; if not, it can return to the step of waiting for the set time until the number of return times reaches the preset number, and feedback to the user the indication information that there are currently no available resources for session processing.
[0045] The initial model is a session model in the model resource pool that matches at least one dimensional feature in the multi-dimensional features. The candidate server is a server in each initial server indicated by the model resource pool that supports the operation of the initial model and matches at least one dimensional feature in the multi-dimensional features. For example, multi-dimensional matching calculations can be performed on the multi-dimensional feature vector and the model capability label to obtain the calculation results under each capability indication label, and the calculation results can be weighted using the proficiency level of the capabilities under the capability indication label to obtain the matching degree between the session model and the user session. The session model with a matching degree greater than the set matching degree is used as the initial model that matches the user session. Alternatively, in the order from high to low matching degree, the various session models are sorted according to the matching degree, and the session model whose sorting order meets the set order is used as the initial model. When determining the initial model, it is also possible to select a candidate server available for the initial model from the initial servers with the goal of load balancing each server and supporting the processing of user sessions based on the resource information of the initial servers that support the operation of the initial model and the matching situation of the multi-dimensional features.
[0046] Exemplarily, after obtaining the multi-dimensional features, it is possible to determine each session model included in the latest model resource pool and the initial server corresponding to each session model, and then determine the first server whose resource information can meet the session complexity according to the resource information of the initial server corresponding to each session model. According to the model capability information of the session models supported by each first server and the multi-dimensional features, calculate the matching degree between these session models and the user session respectively, and use the session model with a matching degree that meets the set matching degree as the initial model, and use the first server corresponding to the initial model as the candidate server. Here, the number of initial models selected and the number of candidate servers corresponding to each initial model can both include at least one.
[0047] In one embodiment, the model resource pool can be constructed according to the following steps 1 to 3: Step 1: For any session model, obtain the resource information of each initial server that supports the session model, and determine the availability score of the session model on each initial server according to the resource information, reserved resource requirements, resource usage requirements corresponding to the session model, and historical resource usage information.
[0048] Here, the reserved resource requirements are used to indicate various resources that need to be reserved on the server when running the session model. The resource usage requirements are used to indicate various resources required for the operation of the session model. The historical resource usage information is used to indicate the resource usage of the session model during the processing of historical sessions. Based on the historical resource usage information, information such as the historical load trend and historical resource usage trend of the session model can be determined. The availability score is used to indicate the level of usability of the session model on the initial server. The higher the availability score, the stronger the usability of the session model on the initial server and the better the model usage effect; the lower the availability score, the worse the usability of the session model on the initial server and the worse the model usage effect.
[0049] In specific implementation, the server can collect in real time the resource information of each initial server that supports the operation of each callable session model based on the distributed monitoring probe network. For example, collect the computing load, video memory occupancy, memory usage, network bandwidth occupancy, CPU usage rate counted by core, etc. of each initial server that supports the operation of the session model. For each session model, the availability score of the session model on each initial server can be calculated according to the comparison information between the resource information of each initial server that supports the operation of the session model and the resource usage requirements corresponding to the session model, whether the remaining resources after allocating the required resources for the session model on the initial server meet the reserved resource requirements, and whether the resource information meets the historical load trend characterized by the historical resource usage information. Exemplarily, a preset calculation formula can be used to calculate the availability score of the session model on each initial server according to the resource information, reserved resource requirements, resource usage requirements corresponding to the session model, and historical resource usage information.
[0050] For example, if a session model that requires 8GB of video memory has only 6GB of video memory remaining on a certain initial server, the availability score of the session model on this initial server will be significantly reduced.
[0051] Step 2: Obtain the model capability information corresponding to the session model from the model capability tag library; the model capability information includes the capability indication tags and the degree of proficiency of various capabilities possessed by the session model.
[0052] In specific implementation, for each callable session model, the capability indication tags and the degree of proficiency of various capabilities possessed by the session model can be obtained from the current model capability tag library.
[0053] Step 3: Construct a model resource pool according to the availability score of each session model on each initial server and the model capability information.
[0054] In specific implementation, each session model and the initial server supporting the session model can be added to the model resource pool, and each session model is associated with its availability score, model capability information, basic parameter information, resource information, etc. on each initial server. Alternatively, based on the availability scores of each session model on each initial server, each session model and initial server with an availability score greater than the set score can be screened out, and a model resource pool can be constructed based on these session models and initial servers. Among them, each session model in the model resource pool is at least associated with the following information: basic parameter information (such as the number of parameters, quantization type, etc.), set of capability labels and proficiency in capabilities, current availability score on each initial server, and resource information.
[0055] It can be understood that the model resource pool is dynamically updated and supports hot plugging of models. The session models and initial servers included therein will be updated in real time as the resource information of the initial servers supporting the session models changes, so as to provide comprehensive candidate models and servers for model and server scheduling. When a new model available for scheduling is deployed, the model resource pool will also be automatically updated to add the new model to the model resource pool for scheduling use.
[0056] In one embodiment, for the step of finding the initial model and candidate servers in S102, it can be implemented according to the following steps: S102-1: Determine the first weighted weight of each dimension feature in the multi-dimensional features.
[0057] In specific implementation, each dimension feature can be set with a default weighted weight, and the default weighted weight can be directly used as the first weighted weight.
[0058] Alternatively, for different user sessions, different weighted weights can be set for different dimension features according to the session problem type and session domain. For example, for user sessions of technical analysis type, the default weight of the dimension features under the dimension of session problem professionalism can be increased, and the default weights of the remaining dimension features can be kept unchanged; for user sessions of creative creation type, the default weight of the dimension features under the dimension of session problem professionalism can be decreased, and the default weight of the dimension features under the session domain dimension can be increased.
[0059] Exemplarily, according to the session problem type and session domain indicated by the multi-dimensional features of the user session, the adjustment method of the default weight of each dimension feature can be determined, and the adjustment can be made according to the adjustment method to obtain the first weighted weight of each dimension feature.
[0060] S102-2: Determine the first matching degree between each session model and the user session on each initial server according to the availability score and model capability information of each session model in the latest model resource pool on each initial server, and the first weighted weight of each dimensional feature.
[0061] During specific implementation, for each dimensional feature in the multi-dimensional features, the cosine similarity between the ability indication label in the model capability information of the session model and this dimensional feature can be calculated respectively, and the cosine similarity can be weighted according to the first weighted weight of this dimensional feature and the proficiency degree of the ability under the ability indication label, to obtain the weighted similarity between the ability indication label and this dimensional feature. The ability indication label with the highest weighted similarity is used as the ability indication label matching this dimensional feature. Furthermore, the first matching degree between the session model and the user session on each initial server can be determined according to the weighted similarity of the ability indication label matched by the session model and each dimensional feature and the availability score of the session model on each initial server. For example, the availability score can be used to weight the mean value of the weighted similarities of the ability indication labels matched by the session model and each dimensional feature to obtain the first matching degree. Or, the maximum similarity can be determined from the weighted similarities of the ability indication labels matched by the session model and each dimensional feature, and the availability score can be used to weight this maximum similarity to obtain the first matching degree.
[0062] Or, based on a preset matching formula, the availability score, ability indication label, proficiency degree of the ability, and the first weighted weight of each dimensional feature of the session model on each initial server in the latest model resource pool can be substituted into the matching formula to obtain the first matching degree between the session model and the user session on each initial server.
[0063] S102-3: Screen out each initial model and each candidate server available for the initial model according to the first matching degree.
[0064] Exemplarily, the session model and the initial server with the first matching degree greater than the set matching degree can be used as the screened initial model and candidate server. Or, the first matching degrees of each session model on each initial server can be sorted in descending order of the matching degree, and the initial model and candidate server can be screened by using the size relationship between the sorting order and the set number of times.
[0065] In one embodiment, the above S102-3 can also be implemented according to the following steps: S102-3-1: Determine the second weighted weight of each session model relative to the initiating user according to the user attribute information of the initiating user of the user session and the weight mapping relationship between the user attribute and the session model.
[0066] Here, the initiating user is the user who initiates the user session, and the user attribute information is used to indicate the user group to which the initiating user belongs. The user group can include, for example, software developers, algorithm engineers, product managers, cartographers, ordinary users, and so on. Among them, the user attribute information can be actively submitted by the initiating user to the server, or can be obtained by the server through clustering analysis of the historical interaction data of the initiating user under the condition of user authorization. The historical interaction data can include, but is not limited to, the historical session records of the initiating user, the user's preference model, the scoring information of the user for different session models, etc.
[0067] For example, after the initiating user logs in, the server can, under the condition of the initiating user's authorization, obtain the historical interaction data of the initiating user within a set time range in the past and perform clustering analysis on the historical interaction data, so as to classify the initiating user into a predefined group type. For example, when the historical interaction data of the initiating user indicates that the user has initiated a large number of historical sessions related to software programming, the initiating user can be classified into the software developer type.
[0068] The weight mapping relationship between user attributes and session models is used to indicate the preference weights of user groups under different attributes for different session models. The higher the preference weight, the higher the preference degree, satisfaction, and usage degree of the user group for the session model. For example, the evaluation of a certain professional field model by the R & D engineer group is significantly higher than that of the general model, and this preference will be quantified as a preference weight and stored in the weight mapping relationship between the R & D engineer group and the professional field model. The weight mapping relationship between user attributes and session models can be determined according to the satisfaction and usage degree of each user with the session model of the user attributes. The server can maintain the weight mapping relationship between user attributes and session models through a multi-dimensional matrix.
[0069] Exemplarily, after obtaining the first matching degree, the latest weight mapping relationship between user attributes and session models can be obtained. According to this weight mapping relationship, the preference weights associated with the user attribute information of the initiating user for each session model are determined, and this preference weight is used as the second weighted weight of each session model relative to the initiating user.
[0070] S102-3-2: Use the second weighted weight of each session model to perform weighted processing on the first matching degree between the session model and the user session on each initial server to obtain the second matching degree.
[0071] Exemplarily, for any session model, the second weighted weight of this session model can be multiplied by the first matching degree between this session model and the user session on each initial server respectively to obtain the second matching degree of this session model on each initial server.
[0072] S102-3-3: Screen out each initial model and each candidate server available for the initial model according to the second matching degree.
[0073] Exemplarily, the session model and the initial server with a second matching degree greater than the set matching degree can be used as the screened initial model and candidate server. Alternatively, the second matching degrees of each session model on each initial server can be sorted in descending order of the matching degree, and the initial model and candidate server can be screened by using the magnitude relationship between the sorting order and the set number of times.
[0074] In another embodiment, the above S102-3 can also be implemented according to the following steps A to D: Step A: Determine the expected success rate of the session model for the user session on each initial server according to the historical sessions of each session model on each initial server.
[0075] Here, the historical session data is used to indicate each historical user session processed by using the session model when running the session model on the initial server, and the processing result of the historical user session. The processing result is used to indicate whether the historical user session is successfully answered. The expected success rate is used to indicate the probability that the session model can successfully process the user session and output a valid session result on the initial server.
[0076] In specific implementation, for any session model and any initial server that supports the operation of the session model, the target historical session similar to the current user session can be determined from the historical sessions of the session model on the initial server, and the expected success rate of the session model for the user session on the initial server can be calculated by using the Bayesian probability model according to the computational complexity and processing result of each target historical session and the computational complexity of the current user session.
[0077] Step B: Determine whether the user session is a new session initiated by the initiating user.
[0078] In specific implementation, after the server obtains the user session, it can also determine whether the user session is a continuation of the historical session of the initiating user and / or a newly established session by the user, so as to determine whether the user session is a new session.
[0079] For example, in the case where there is no historical session for the current user session, the user session is determined to be a new session. In the case where there is a historical session, it can be determined whether it is a new session based on the correlation degree between each historical session and the user session. For example, the semantic similarity between the user session and each historical session can be calculated, and different fourth weighting weights can be set for each historical session according to the chronological order of the initiation time of the historical sessions. Using the fourth weighting weights, a weighted sum process is performed on each semantic similarity to obtain the correlation degree between the user session and the historical sessions. In the case where the correlation degree is greater than the set correlation degree, it is determined that the user session is not a new session initiated by the initiating user; on the contrary, in the case where the correlation degree is not greater than the set correlation degree, the user session is determined to be a new session.
[0080] In the case where the user session is a new session, the initial model and candidate servers can be directly filtered according to the expected success rate and the first matching degree. For example, the session model and the initial server that simultaneously meet the conditions that the expected success rate is greater than the preset success rate and the first matching degree is greater than the set matching degree can be used as the initial model and candidate servers. Or, the session model and the initial server whose product of the expected success rate of the session model on the initial server and the first matching degree of the session model on the initial server is greater than the set value can be used as the initial model and candidate servers. Or, the first matching degree can also be used to determine the second matching degree according to the above S102-3-1~S102-3-3, and the initial model and candidate servers can be filtered according to the second matching degree and the expected success rate.
[0081] In the case where the user session does not belong to a new session, the following step C can be executed instead.
[0082] Step C: If not, then according to the correlation degree between the user session and the historical session of the initiating user, and the historical session model and historical server used in the historical session, determine the model switching loss of each session model on each initial server for the user session.
[0083] Here, the model switching loss of the session model on the initial server for the user session is used to characterize the possible loss information when switching to the initial server to run the session model to process the user session.
[0084] Exemplarily, in the case where the user session is not a new session, the server can perform a consistency evaluation on the user session. Among them, the evaluation metrics include, but are not limited to, the coherence of the knowledge system of the session model for the user session, the requirements of the user session for answer consistency, and the context understanding ability of the session model. And, the server can determine the model switching loss of switching to different initial servers to run different session models for session processing according to the correlation between the user session and the historical sessions of the initiating user, the historical session model and the historical server used in the latest historical session, and the consistency evaluation result. The historical session model used in the latest historical session can be a certain model or a model combination of multiple session models.
[0085] Step D: Screen out each initial model and each candidate server available for the initial model according to the model switching loss, the expected success rate, and the first matching degree.
[0086] For example, the session model and the initial server with a model switching loss less than the set loss, an expected success rate greater than the preset success rate, and a first matching degree greater than the set matching degree can be used as the initial model and the candidate server. Or, the session model and the initial server with the product of the model switching loss, the expected success rate, and the first matching degree greater than the set value can be used as the initial model and the candidate server. Or, the first matching degree can also be used to determine the second matching degree according to the above S102-3-1~S102-3-3, and the initial model and the candidate server can be screened out according to the second matching degree, the model switching loss, and the expected success rate.
[0087] S103: In the case where the number of initial models and the session complexity indicate that the user session requires multi-model processing, split the user session into multiple sub-sessions, and screen out the target model matching each sub-session and the target server required for the target model from the initial models and candidate servers.
[0088] Here, one sub-session can correspond to at least one target model, and one target model corresponds to one target server. The target server is a server that needs to run the target model for processing the user session this time. When there are multiple target models corresponding to the sub-session, the sub-session results output by different target models for the same sub-session can be preliminarily integrated when integrating the sub-session results subsequently to obtain the integrated sub-session result. After the preliminary integration is completed, the sub-session results corresponding to each sub-session (when there is an integrated sub-session result, use this integrated sub-session result) can be secondarily integrated to obtain the final session result.
[0089] In specific implementation, when the number of models is 1, it indicates that there is only one model available for use at present. At this time, it can be determined that the user session will not be split. If there is only one available candidate server for this unique initial model at this time, the candidate server can be used to run the initial model and process the user session to obtain a session result. If there are multiple available candidate servers for this unique initial model at this time, the candidate server with the most sufficient resource information can be used as the target server, and the target server can be used to run the initial model and process the user session to obtain a session result.
[0090] When the number of models is greater than 1, it is determined whether to split the user session based on whether the initial servers available for each initial model completely overlap. If the initial servers available for each initial model completely overlap, since it is difficult for one initial server to satisfy the operation of multiple models simultaneously, it can be determined that the user session will not be split. Then, the initial model with the highest matching degree between the model capability information and the user session can be used as the target model, and the candidate server with the richest resources can be selected as the target server from the candidate servers that support the operation of the target model. Then, the target server is called to run the target model to process the user session and obtain a session result.
[0091] If the initial servers available to each initial model do not completely overlap, it is possible to determine whether to split the session based on the session complexity. For example, determine whether to split the session based on whether the session complexity is greater than a preset complexity. If not, the initial model with the highest matching degree between the model capability information and the user session can be used as the target model, and among the candidate servers that support the operation of the target model, the candidate server with the richest resources can be selected as the target server. Then, call the target server to run the target model to process the user session and obtain the session result. If so, it is possible to determine a method that can use multi-model collaborative processing. Furthermore, based on the decomposition algorithm of the knowledge graph, the user session can be split into multiple relatively independent sub-sessions according to the multi-dimensional features of the user session and the model capability information of each initial model. For example, for a technical evaluation session, it may be decomposed into sub-sessions such as "technical principle analysis", "application scenario evaluation", and "advantage and disadvantage comparison". Among them, when the server splits the sub-sessions and matches the target models, it can use a task planning algorithm to specify the target model, execution order, data interaction method, etc. for each sub-session to ensure that the dependency relationships between the individual sub-tasks are correctly processed while optimizing the overall execution efficiency. Further, the server can select the target model that matches the sub-session from the initial models based on the multi-dimensional features of each sub-session and the model capability information of each initial model, and select the candidate server with the richest resources among the candidate servers that support the operation of the target model as the target server. Moreover, the server can establish a session task dependency graph to clearly mark the sequence and data dependency relationships between the individual sub-sessions. Then, the server can start parallel session scheduling, using a directed acyclic graph to manage the individual sub-sessions, supporting the maximum parallel execution that conforms to the dependency relationships. Each sub-session is encapsulated into a standard request packet, including task description, context information, output requirements, etc. The system distributes these sub-session requests to the target servers corresponding to the selected target models through an asynchronous message queue and monitors the execution progress in real time.
[0092] S104: Schedule each target server to run the target model to process each sub-session respectively, obtain the sub-session result corresponding to each sub-session, and determine the session result corresponding to the user session according to the results of each sub-session.
[0093] Here, the session result is used to indicate the answer result of the user session, and this result is used to feedback to the initiating user.
[0094] During specific implementation, the server can schedule each target server to run the corresponding target model according to the execution order and data interaction method, and use each running target model to process the matching sub-sessions respectively to obtain the session result corresponding to each sub-session. Then, the server can schedule the text integration model to perform session integration and optimization on the results of each sub-session to obtain the final session result.
[0095] In one embodiment, for the step of determining the session result in S104, it can be implemented according to the following steps P1 to P3: P1: Determine the model ability information of the target model corresponding to each sub-session in the session domain to which the sub-session belongs.
[0096] During specific implementation, each sub-session also belongs to a session domain. For each sub-session, the model ability information of the target model used to process the sub-session in this session domain can be determined according to the session domain to which the sub-session belongs.
[0097] P2: Determine the third weighted weight of the sub-session result corresponding to each sub-session according to the model ability information in the session domain and the importance of the sub-session relative to the user session.
[0098] Here, the importance of the sub-session relative to the user session can be determined when the server splits the user session.
[0099] During specific implementation, the product of the proficiency level in the model ability information of the target model in the session domain and the importance of the sub-session relative to the user session can be determined, and this product is used as the third weighted weight of the sub-session result corresponding to the sub-session.
[0100] Alternatively, the user group indicated by the user attribute information can also be obtained, as well as the preference degree for the answer results in the session domain to which the sub-session belongs. Then, the product of this preference degree, the proficiency level in the model ability information of the target model in the session domain, and the importance of the sub-session relative to the user session is determined, and the product of the three is used as the third weighted weight of the sub-session result corresponding to the sub-session.
[0101] P3: Use the text integration model to perform multi-round session integration on the results of each sub-session according to the third weighted weight of each sub-session result to obtain the session result.
[0102] Here, the text integration model is a pre-trained text processing model, and this model has text integration and text optimization capabilities.
[0103] In specific implementation, the server can schedule a pre-trained text integration model, and input each sub-session result and its corresponding third weighted weight into the text integration model. The text integration model performs multi-round refinement and coherent session integration on each sub-session result, and outputs an overall session result. It should be noted here that the session integration process of the present disclosure is not a simple text splicing, but an intelligent fusion process. The text integration model will generate a coherent and complete session result through multi-round refinement by combining each sub-session result according to the third weighted weight. Moreover, during the integration process, the text integration model will pay special attention to the logical coherence of the result content, the consistency of professional terms, and the unity of the overall writing style, and selectively add appropriate transition sentences as needed to ensure that the final session result not only integrates the content of each sub-session result but also has readability.
[0104] In this way, by performing multi-dimensional analysis on the user session, the dimensional features of the user session from different perspectives can be obtained and a reasonable estimation of the session complexity can be achieved. Using the multi-dimensional features, as well as the resource information and model ability information of the initial servers supporting the session model, it is possible to initially screen out each initial model that can be used for session processing and the candidate servers that can be used by the initial models on the basis of ensuring the load balance of the server resources. Based on the number of models and the session complexity, the user session can be split to achieve a fine-grained division of complex sessions and obtain refined sub-sessions with the support of the session model. Then, by separately screening out the target models and target servers that best match each sub-session from the initial models, the reasonable utilization of the model capabilities of each target model and the server resources can be achieved, and the problems of unreasonable server scheduling and model use can be avoided. Finally, by integrating the results of each sub-session, a session result that best matches the user session can be obtained. In the entire solution, by analyzing and splitting the user session and combining the resource information and model ability information of each session model, the session model and the initial servers are scheduled, and the capabilities of each available model and the server resources can be reasonably utilized to process the session, so as to obtain the session result that best meets the user's expectations on the basis of ensuring the relative balance of the server load.
[0105] In one embodiment, after determining the session result corresponding to the user session, the model ability information of the session model can also be updated according to the following steps T1 to T4: T1: Obtain the initial feedback data of the session result; the initial feedback data includes interaction data and usage data for the session result.
[0106] Here, the server can adopt a multi-channel user feedback collection mechanism to capture both explicit and implicit feedback data simultaneously, and update the model ability information based on this data. Specifically, the explicit feedback data can include interaction data, such as the scoring data and review operation data initiated by the user for the session result. The review operation data can include like / does not meet expectations, text and image evaluation data, etc. The implicit feedback data can include usage data, such as the reading duration of the answer, the scroll heat map for the session result, the copy behavior, and the re-questioning data for the session result, the secondary usage data for the session result, etc. The secondary usage data includes data such as whether it is copied, shared, or cited.
[0107] Exemplarily, after the session result is fed back to the initiating user, various interaction data and various usage data of the initiating user for the session result can be obtained, and these data are all used as the initial feedback data.
[0108] T2: According to the initial feedback data, determine the feedback score of the session result, and combine the feedback score and the timestamp into the intermediate feedback data, and store it in the feedback data pool.
[0109] Here, a feedback data pool can be set for each session model. The feedback data pool is used to store the intermediate feedback data of the session model for each user session. The intermediate feedback data includes the feedback score of the session result and the timestamp corresponding to the determination time of the feedback score.
[0110] In specific implementation, the server can perform standardization processing on the initial feedback data, and convert the feedback score corresponding to the initial feedback data according to the standardization processing result. The user session, the session result, the feedback score of the session result, the timestamp corresponding to the determination time of the feedback score, and the session domain corresponding to the session result are associated and stored in the feedback data pool corresponding to the session model as the intermediate feedback data.
[0111] T3: For each session model, according to the timestamps of the intermediate feedback data in the feedback data pool and the preset time range, obtain the first target feedback data of the session result of the session model in different session domains from the feedback data pool.
[0112] Here, the preset time range can be set according to experience. For example, the preset time range can be the past three days, the past day, the past week, the past 12 hours, etc.
[0113] During specific implementation, for each session model, according to the timestamps of the intermediate feedback data and the current time, from each piece of intermediate feedback data included in the feedback data pool corresponding to the session model, the target feedback data that meets the preset time range can be filtered out, and according to the session fields in each piece of target feedback data, the first target feedback data of the session results of the session model in different session fields can be determined from the target feedback data. Here, for any session model, the above-mentioned target feedback data corresponding to the session model can be the feedback data obtained after processing different sessions by running the session model using different initial servers.
[0114] T4: Update the model capability information of the session model according to the first target feedback data of the session results of the session model in different session fields.
[0115] Exemplarily, for each session field, according to the feedback scores in each piece of the first target feedback data of the session model in this session field, the comprehensive feedback score of the session model in this session field can be determined. For example, the comprehensive feedback score can be determined according to the mean / variance / extreme value / standard deviation, etc. of the feedback scores in each piece of the first target feedback data. Then, when the comprehensive feedback score is greater than the set feedback score, the comprehensive feedback score is used as the new proficiency level of the session model in this session field. When the comprehensive feedback score is not greater than the set feedback score, it is determined that the session model is not proficient in this session field, and the ability indication label and proficiency level of the session model in this session field are cancelled; or, the proficiency level of the session model in this session field can be updated to the default value, and when the consecutive number of times that the proficiency level of the session model in this session field is the default value reaches the preset number of times, the ability indication label and proficiency level of the session model in this session field are cancelled.
[0116] In one embodiment, the above T4 can also be implemented according to the following steps: T4-1: For any session field, according to the first target feedback data of the session results of the session model in the session field, determine the first ability score of the session model in the session field.
[0117] During specific implementation, for any session field, according to the feedback score in the first target feedback data of the session model in this session field, using a time decay function, determine the decay score corresponding to the feedback score. Then, according to the mean of the decay scores corresponding to the feedback scores in each piece of the first target feedback data, obtain the first ability score of the session model in this session field.
[0118] Among them, the decay score corresponding to any feedback score can be determined according to the following formula (1): ; (Formula 1) Among them, represents the attenuation score, represents the feedback score, is an adjustable attenuation coefficient, which can be determined according to the knowledge update speed in different conversation domains. For example, in the rapidly developing AI field, the value will be set relatively large, and the older feedback scores will decay rapidly; while in the basic mathematics field, the value is relatively small, and the attenuation of historical feedback scores will not be very fast, ensuring the lasting influence of historical conversation results. t represents the current time, represents the timestamp in the first target feedback data to which the feedback score belongs.
[0119] T4-2: Determine the second ability score of the conversation model in the conversation domain according to each first ability score of the conversation model within the update time range in the conversation domain.
[0120] Here, the size of the update time range can be specifically set according to experience, and the embodiments of the present disclosure do not specifically limit it. For example, a finer time range can be two weeks, half a month, one month, etc. However, the time length of the set update time range needs to be greater than the time length of the above preset time range.
[0121] In specific implementation, each to-be-used ability score can be screened from each first ability score according to the time corresponding to each first ability score of the conversation model in the conversation domain, the current time, and the update time range. Then, according to each to-be-used ability score, the second ability score of the conversation model in the conversation domain is determined. Among them, the time corresponding to the first ability score can be the determination time of the first ability score, or the determination time of the feedback score corresponding to the first ability score.
[0122] Exemplarily, the second ability score can be determined according to the following formula two: ; (Formula two) Among them, represents the second ability score, represents the to-be-used ability score closest to the current time, represents except the mean value of each to-be-used ability score other than. represents a preset stability coefficient, usually set between 0.7 and 0.9, and can be selectively set according to different conversation domains .
[0123] T4-3: Update the model ability information of the conversation model according to the second ability scores of the conversation model in each conversation domain.
[0124] Exemplarily, for each session domain, the comprehensive score of the session model in this session domain can be determined according to the respective second ability scores of the session model in this session domain. For example, the comprehensive score can be determined according to the mean / variance / extreme value / standard deviation, etc. of the respective second ability scores. Then, when the comprehensive score is greater than the preset threshold, the comprehensive score is used as the new ability proficiency level of the session model in this session domain. When the comprehensive score is not greater than the preset threshold, it is determined that the session model is not proficient in this session domain, and the ability indication label and ability proficiency level of the session model in this session domain are cancelled; or, the ability proficiency level of the session model in this session domain can be updated to the default value, and when the consecutive number of times that the ability proficiency level of the session model in this session domain is the default value reaches the preset number of times, the ability indication label and ability proficiency level of the session model in this session domain are cancelled.
[0125] In this way, based on the cumulative evaluation results within the update time range, the model ability information is updated regularly, and the update process adopts a progressive method to avoid drastic changes in ability evaluation caused by short-term fluctuations.
[0126] In one embodiment, the embodiments of the present disclosure can not only update the model ability information, but also update the weight mapping relationship between the user attributes and the session model. Specifically, the weight mapping relationship can be updated according to the following steps: Obtain the second target feedback data corresponding to the processed user sessions of different session models in different session domains from the feedback data pool; update the weight mapping relationship between the user attributes and the session model according to the second target feedback data and the user attribute information of the initiating user of the processed user session.
[0127] In specific implementation, for any conversation model, the second target feedback data corresponding to the processed user conversations of the conversation model in different conversation domains can be regularly obtained from the feedback data pool according to a set update period. Among them, the processed user conversations of the conversation model in different conversation domains can include the user conversations in different domains processed by each initiating user using the conversation model. For example, for any conversation domain and any conversation model, the second target feedback data corresponding to the processed user conversations of the conversation model in this conversation domain can be obtained. Then, according to the user attribute information of the initiating users of each processed user conversation and the feedback scores in the second target feedback data, the preference degrees of the initiating users with various user attribute information for the conversation model in different conversation domains can be determined. For example, for each conversation domain of any conversation model, the initiating users with the same user attribute information can be divided according to the user attribute information of different initiating users. Then, for each user attribute information, according to the average value of the feedback scores of the initiating users with this user attribute information for the processed user conversations in this conversation domain, the preference degree of the user group with this user attribute information for the conversation model in this conversation domain can be determined. After that, according to the user groups under different user attribute information, for the preference degrees of each conversation model in different conversation domains, the weight mapping relationship between different user attributes and different conversation models can be updated. For example, for any user attribute and any conversation model, according to the average value of the preference degrees of the user group with this user attribute for the conversation model in different domains, the model preference degree of the user group with this user attribute for the conversation model can be determined. The model preference degree is converted into a preference weight, and the weight mapping relationship between this user attribute and this conversation model is updated using this preference weight. For example, an incremental learning method can be used to perform weighted fusion of the new preference weight and the historical preference weight to achieve the update of the weight mapping relationship.
[0128] In this way, by maintaining a multi-dimensional matrix to record the preference degrees of different user groups (such as developers, researchers, students, etc.) for each conversation model. The collaborative filtering algorithm is used in the optimization process to update the preference degrees. Moreover, group feature re-clustering can be performed regularly to adapt to the dynamic changes of user group features. For example, if it is found that a certain user group begins to differentiate into new usage patterns, the server can automatically adjust the group division to ensure that the preference matrix can accurately reflect the actual usage scenario. Through this dynamic optimization mechanism, the accuracy of model scheduling can be continuously improved, and more personalized services can be provided for different user groups.
[0129] To facilitate the understanding of the embodiments of the present disclosure, the conversation processing method of the present disclosure will be described below with a specific embodiment: Suppose a senior software architect submits a user session: "Please analyze the respective performance bottlenecks of Service A and Service B in a high-concurrency microservices architecture and give optimization suggestions."
[0130] First, the server receives the user session, determines that the initiating user belongs to the "architect" user group, and historical data shows that the initiating user prefers in-depth technical analysis and often involves topics related to distributed systems. The session analysis model can quickly parse the user session, generate a feature vector: [Domain: Distributed Architecture: 0.95; Sub-domain: Microservices: 0.9; Complexity: Advanced: 0.8; Analysis Dimension: Multidimensional: 0.85], and detects that this is a new session, so there is no need to consider the historical context.
[0131] In the model resource pool, there are currently three session models available: General Architecture Model (with an expertise level of 0.88 in the field of distributed systems); Microservices Specialization Model (with an expertise level of 0.92 in the field of microservices); Performance Optimization Model (with an expertise level of 0.85 in the field of system performance analysis).
[0132] Each of these three session models has a server supporting its operation, and the monitoring probes show that the server resources supporting these three session models are sufficient, with the CPU usage rate below 50% and the GPU load being moderate. These models are used as the initial models, and the servers supporting the operation of the initial models are used as candidate servers.
[0133] Entering the intelligent scheduling decision-making stage, the server notices that the user session involves multiple professional dimensions (microservices architecture, performance analysis, optimization suggestions), and decides to adopt a multi-model collaboration strategy. Through matching degree calculation and combined with the weight mapping relationship of the architect group, the system formulates a collaboration plan: The Microservices Specialization Model is responsible for architecture comparison and analysis; the Performance Optimization Model is responsible for bottleneck analysis; the General Architecture Model is responsible for integrating suggestions; and the candidate servers of each session model are directly used as target servers.
[0134] In the multi-model collaborative processing stage, the user session is decomposed into three sub-sessions: "Comparison of the architectural principles of Service A and Service B"; "Analysis of the performance bottlenecks of the two architectures in high-concurrency scenarios"; "Integration of optimization suggestions based on the actual scenario".
[0135] Schedule the target servers corresponding to the three session models to run the three session models respectively, and use the three session models to process their respective sub-sessions in parallel. For example, schedule the target server corresponding to the microservice specialization model to run the microservice specialization model to focus on analyzing the implementation differences between the two architectures in terms of service discovery, load balancing, circuit breaker degradation, etc.; schedule the target server corresponding to the performance optimization model to run the performance optimization model to focus on evaluating performance metrics such as network proxy overhead, memory occupancy, and CPU consumption; schedule the target server corresponding to the general architecture model to run the general architecture model to generate targeted optimization suggestions based on the outputs of the previous two models and in combination with the actual application scenario. Use the text integration model to perform multiple rounds of session integration on the results of each sub-session to obtain the session results.
[0136] Finally, enter the dynamic optimization and feedback learning stage. The initiating user gave a "very satisfied" evaluation of the session results and stated that the suggestions were very helpful for the actual project. The server can record this positive feedback and update the proficiency of each session model's capabilities through a time decay function: The proficiency of the microservice specialization model in the "microservice architecture" field increased from 0.92 to 0.925; the proficiency of the performance optimization model in the "performance analysis" field increased from 0.85 to 0.855; the proficiency of the general architecture model in the distributed system field increased from 0.88 to 0.885.
[0137] Update the weight mapping relationship between the architect group and these three session models: This feedback also affects the weight mapping relationship between the architect group and the three session models. The server determines that the satisfaction of the architect group with the multi-model collaboration solution is extremely high. Therefore, when dealing with similar complex technical problems, it will be more inclined to adopt the multi-model collaboration strategy.
[0138] In this way, the entire process forms a closed loop. Obtain the feature vector through pre-analysis, and then perform intelligent scheduling on the session models in the model resource pool. The multi-model collaboration processing after scheduling generates session results, and the user feedback feeds back to the model ability evaluation system, ultimately realizing the continuous optimization of the system.
[0139] As Figure 2 shown, the following is a specific implementation flowchart of a session processing method provided by an embodiment of the present disclosure, which may include the following steps: S201: Use the session analysis model to perform multi-dimensional parsing on the user session to obtain the multi-dimensional features and session complexity corresponding to the user session.
[0140] S202: Determine the first weighted weight of each dimension feature in the multi-dimensional feature, and determine the first matching degree between each session model on each initial server and the user session according to the availability score and model ability information of each session model in the latest model resource pool on each initial server and the first weighted weight of each dimension feature.
[0141] S203: Determine the second weighted weight of each session model relative to the initiating user according to the user attribute information of the initiating user of the user session and the weight mapping relationship between the user attribute and the session model; and use the second weighted weight of each session model to perform weighted processing on the first matching degree between the session model on each initial server and the user session to obtain the second matching degree.
[0142] S204: Determine the expected success rate of each session model on each initial server for the user session according to the historical session of each session model.
[0143] S205: In the case where the user session is a new session, determine the model switching loss of each session model on each initial server for the user session according to the correlation degree between the user session and the historical session of the initiating user, and the historical session model and historical server used by the historical session.
[0144] S206: Screen out each initial model and each candidate server available for the initial model according to the model switching loss, expected success rate, and second matching degree.
[0145] S207: In the case where the number of initial models and the session complexity indicate that multi-model processing is required, split the user session into multiple sub-sessions, and screen out the target model matching each sub-session and the target server required by the target model from the initial models and candidate servers.
[0146] S208: Schedule each target server to run the target model to process each sub-session respectively, and obtain the sub-session result corresponding to each sub-session.
[0147] S209: Determine the model ability information of the target model corresponding to each sub-session in the session domain to which the sub-session belongs, and determine the third weighted weight of the sub-session result corresponding to each sub-session according to the model ability information in the session domain and the importance of the sub-session relative to the user session.
[0148] S210: Use the text integration model to perform multi-round session integration on each sub-session result according to the third weighted weight of each sub-session result to obtain the session result.
[0149] S211: Update the model ability information of the session model and the weight mapping relationship between the user attribute and the session model according to the session result.
[0150] Here, for the specific implementation process of S211, reference can be made to the specific steps of T1 to T4 and the specific steps of updating the weight mapping relationship above, which will not be elaborated here.
[0151] Regarding the specific implementation processes of S201 to S211 above, reference can be made to the above embodiments, which will not be elaborated here either.
[0152] In this way, the present disclosure realizes an adaptive scheduling system with a full-link closed loop, constructs a complete link from user session analysis, model selection, collaborative processing to feedback optimization, and continuously improves the system performance. Through a multi-dimensional intelligent scheduling decision-making mechanism, multiple dimensions such as model professional capabilities, server resource status, user feedback, and group characteristics are incorporated into the decision-making consideration to improve the scheduling accuracy of servers and models. At the same time, based on the time-decay dynamic ability evaluation, a time-decay function is introduced to process user feedback, making the model ability evaluation more in line with the actual performance changes. Consider the guarantee of session-level model consistency. Through session state tracking and consistency evaluation, the user experience fluctuations caused by model switching are reduced. Utilize group-characteristic-driven personalized services, construct a preference matrix based on user group characteristics, and implement a differentiated model scheduling strategy.
[0153] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0154] Based on the same inventive concept, a session processing apparatus corresponding to the session processing method is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the apparatus in the embodiments of the present disclosure is similar to the above session processing method of the embodiments of the present disclosure, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0155] As Figure 3 shown, it is a schematic diagram of a session processing apparatus provided by an embodiment of the present disclosure, including: A parsing module 301, configured to use a session analysis model to perform multi-dimensional parsing on a user session to obtain multi-dimensional features and session complexity corresponding to the user session; A searching module 302, configured to search, according to the multi-dimensional features, for each initial model matching the user session and each candidate server available for the initial model from the latest model resource pool; wherein, the model resource pool is constructed according to the model ability information of each session model and the resource information of the initial server supporting the session model; The splitting module 303 is configured to split the user session into multiple sub - sessions when the number of the initial models and the session complexity indicate that the user session requires multi - model processing, and screen out a target model matching each sub - session and a target server required for using the target model from the initial models and the candidate servers; The determining module 304 is configured to schedule each of the target servers to run the target model to process each sub - session respectively, obtain a sub - session result corresponding to each sub - session, and determine a session result corresponding to the user session according to each sub - session result.
[0156] In a possible implementation manner, the multi - dimensions at least include a session domain dimension, a session problem type dimension, and a session problem professionalism dimension.
[0157] In a possible implementation manner, the apparatus further includes a constructing module 305, configured to construct the model resource pool according to the following steps: For any session model, collect resource information of each initial server supporting the session model in real time, and determine an availability score of the session model on each initial server according to the resource information, reserved resource requirements, resource usage requirements corresponding to the session model, and historical resource usage information; Obtain model capability information corresponding to the session model from a model capability tag library; the model capability information includes capability indication tags and capability proficiency degrees of various capabilities possessed by the session model; Construct the model resource pool according to the availability score of each session model on each initial server and the model capability information.
[0158] In a possible implementation manner, when the searching module 302 searches for each initial model matching the user session and each candidate server available for the initial model from the latest model resource pool according to the multi - dimensional features, it is configured to: Determine a first weighting weight of each dimension feature in the multi - dimensional features; Determine a first matching degree between each session model on each initial server and the user session according to the availability score of each session model on each initial server and the model capability information in the latest model resource pool and the first weighting weight of each dimension feature; Screen out each initial model and each candidate server available for the initial model according to the first matching degree.
[0159] In a possible implementation, when the lookup module 302 filters out each of the initial models and each candidate server available for the initial model according to the first matching degree, it is used for: Determine the second weighted weight of each session model relative to the initiating user according to the user attribute information of the initiating user of the user session and the weight mapping relationship between the user attributes and the session models; Use the second weighted weight of each session model to perform a weighted process on the first matching degree between the session model and the user session on each initial server to obtain a second matching degree; Filter out each of the initial models and each candidate server available for the initial model according to the second matching degree.
[0160] In a possible implementation, when the lookup module 302 filters out each of the initial models and each candidate server available for the initial model according to the first matching degree, it is used for: Determine the expected success rate of each session model for the user session on each initial server according to the historical sessions of each session model on each initial server; Determine whether the user session is a new session initiated by the initiating user; If not, determine the model switching loss of each session model for the user session on each initial server according to the correlation between the user session and the historical session of the initiating user, and the historical session model and historical server used in the historical session; Filter out each of the initial models and each candidate server available for the initial model according to the model switching loss, the expected success rate, and the first matching degree.
[0161] In a possible implementation, when the determination module 304 determines the session result corresponding to the user session according to each sub-session result, it is used for: Determine the model capability information of the target model corresponding to each sub-session in the session domain to which the sub-session belongs; Determine the third weighted weight of the sub-session result corresponding to each sub-session according to the model capability information in the session domain and the importance of the sub-session relative to the user session; Use a text integration model to perform multi-round session integration on each sub-session result according to the third weighted weight of each sub-session result to obtain the session result.
[0162] In a possible implementation, the device further includes a first update module 306, which, after determining the session result corresponding to the user session, is configured to: Obtain the initial feedback data of the session result; the initial feedback data includes interaction data and usage data for the session result; According to the initial feedback data, determine the feedback score of the session result, and combine the feedback score and the timestamp into intermediate feedback data, and store it in the feedback data pool; For each session model, according to the timestamps of the intermediate feedback data in the feedback data pool and a preset time range, obtain from the feedback data pool the first target feedback data of the session results of the session model in different session domains; Update the model capability information of the session model according to the first target feedback data of the session results of the session model in different session domains.
[0163] In a possible implementation, when the first update module 306 updates the model capability information of the session model according to the first target feedback data of the session results of the session model in different session domains, it is configured to: For any session domain, according to the first target feedback data of the session result of the session model in the session domain, determine the first capability score of the session model in the session domain; According to the respective first capability scores of the session model within the update time range in the session domain, determine the second capability score of the session model in the session domain; Update the model capability information of the session model according to the second capability scores of the session model in each session domain.
[0164] In a possible implementation, the device further includes a second update module 307, which is configured to: Obtain from the feedback data pool the second target feedback data corresponding to the processed user sessions of different session models in different session domains; Update the weight mapping relationship between the user attributes and the session model according to the second target feedback data and the user attribute information of the initiating user of the processed user session.
[0165] The description of the processing flow of each module in the device and the interaction flow between the modules can refer to the relevant descriptions in the above method embodiments, and will not be elaborated here.
[0166] Based on the same technical concept, an embodiment of the present application further provides a computer device. Refer to Figure 4As shown in the figure, it is a schematic structural diagram of a computer device provided by an embodiment of the present application, including: A processor 401, a memory 402, and a bus 403. Among them, the memory 402 stores machine-readable instructions executable by the processor 401. The processor 401 is used to execute the machine-readable instructions stored in the memory 402. When the machine-readable instructions are executed by the processor 401, the processor 401 performs the following steps: S101: Using a session analysis model, perform multi-dimensional parsing on the user session to obtain multi-dimensional features and session complexity corresponding to the user session; S102: According to the multi-dimensional features, search in the latest model resource pool for each initial model that matches the user session and each candidate server available for the initial model; where the model resource pool is constructed based on the model ability information of each session model and the resource information of the initial servers that support the session models; S103: In the case where the number of initial models and the session complexity indicate that the user session requires multi-model processing, split the user session into multiple sub-sessions, and screen out the target model that matches each sub-session and the target server required by the target model from the initial models and candidate servers; and S104: Schedule each target server to run the target model to process each sub-session respectively, obtain the sub-session results corresponding to each sub-session, and determine the session result corresponding to the user session according to each sub-session result.
[0167] The above-mentioned memory 402 includes an internal memory 4021 and an external memory 4022; here, the internal memory 4021 is also called the main memory, which is used to temporarily store the operation data in the processor 401 and the data exchanged with the external memory 4022 such as the hard disk. The processor 401 exchanges data with the external memory 4022 through the internal memory 4021. When the computer device is running, the processor 401 communicates with the memory 402 through the bus 403, so that the processor 401 executes the execution instructions mentioned in the above method embodiments.
[0168] An embodiment of the present disclosure also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the session processing method described in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0169] An embodiment of the present disclosure also provides a computer program product. The computer product carries program code. The instructions included in the program code can be used to execute the steps of the software update method described in the above method embodiments. For details, refer to the above method embodiments and will not be elaborated here.
[0170] The computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0171] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0172] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0173] In addition, in each embodiment of the present disclosure, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0174] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0175] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the personal's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the personal's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the device for processing personal information, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0176] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A session processing method, characterized in that, Including: Using a conversation analysis model, perform multi-dimensional parsing on the user conversation to obtain the multi-dimensional features and conversation complexity corresponding to the user conversation; According to the multi-dimensional features, in the latest model resource pool, find each initial model that matches the user conversation and each candidate server available for the initial model; wherein, the model resource pool is constructed according to the model capability information of each conversation model and the resource information of the initial servers supporting the conversation model; In the case where the number of the initial models and the conversation complexity indicate that the user conversation requires multi-model processing, split the user conversation into multiple sub-conversations, and screen out the target model that matches each sub-conversation and the target server required by the target model from the initial models and the candidate servers; Schedule each of the target servers to run the target model to process each sub-conversation respectively, obtain the sub-conversation result corresponding to each sub-conversation, and determine the conversation result corresponding to the user conversation according to each sub-conversation result.
2. The method according to claim 1, wherein The multi-dimension at least includes a conversation domain dimension, a conversation question type dimension, and a conversation question professional degree dimension.
3. The method according to claim 1, characterized in that, The model resource pool is constructed according to the following steps: For any conversation model, collect the resource information of each initial server supporting the conversation model in real time, and determine the availability score of the conversation model on each initial server according to the resource information, reserved resource requirements, the resource usage requirements corresponding to the conversation model, and historical resource usage information; Obtain the model capability information corresponding to the conversation model from the model capability tag library; the model capability information includes the capability indication tags and capability proficiency levels of various capabilities possessed by the conversation model; Construct the model resource pool according to the availability score of each conversation model on each initial server and the model capability information.
4. The method according to claim 3, characterized in that The step of finding each initial model that matches the user conversation and each candidate server available for the initial model from the latest model resource pool according to the multi-dimensional features includes: Determine the first weighted weight of each dimension feature in the multi-dimensional features; According to the availability score of each conversation model on each initial server in the latest model resource pool, the model capability information, and the first weighted weight of each dimension feature, determine the first matching degree between each conversation model on each initial server and the user conversation; According to the first matching degree, screen out each initial model and each candidate server available for the initial model.
5. The method according to claim 4, wherein The step of screening out each initial model and each candidate server available for the initial model according to the first matching degree includes: According to the user attribute information of the initiating user of the user conversation and the weight mapping relationship between the user attribute and the conversation model, determine the second weighted weight of each conversation model relative to the initiating user; Using the second weighted weight of each of the session models, perform a weighted process on the first matching degree between the session models and the user session on each of the initial servers to obtain a second matching degree; According to the second matching degree, screen out each of the initial models and each candidate server available for the initial models.
6. The method according to claim 4 or 5, characterized in that, The screening out each of the initial models and each candidate server available for the initial models according to the first matching degree includes: According to the historical sessions of each session model on each of the initial servers, determine the expected success rate of the session model for the user session on each of the initial servers; Determine whether the user session is a new session initiated by the initiating user; If not, then according to the correlation between the user session and the historical sessions of the initiating user, and the historical session models and historical servers used in the historical sessions, determine the model switching loss of each session model for the user session on each of the initial servers; According to the model switching loss, the expected success rate and the first matching degree, screen out each of the initial models and each candidate server available for the initial models.
7. The method according to claim 1, characterized in that, The determining the session result corresponding to the user session according to each of the sub-session results includes: Determine the model ability information of the target model corresponding to each sub-session in the session domain to which the sub-session belongs; According to the model ability information in the session domain and the importance of the sub-session relative to the user session, determine the third weighted weight of the sub-session result corresponding to each sub-session; Using a text integration model, perform multi-round session integration on each of the sub-session results according to the third weighted weights of each of the sub-session results to obtain the session result.
8. The method according to claim 1, wherein After determining the session result corresponding to the user session, it further includes: Obtain the initial feedback data of the session result; the initial feedback data includes interaction data and usage data for the session result; According to the initial feedback data, determine the feedback score of the session result, and combine the feedback score and the timestamp into intermediate feedback data, and store it in the feedback data pool; For each of the session models, according to the timestamps of the intermediate feedback data in the feedback data pool and a preset time range, obtain the first target feedback data of the session results of the session model in different session domains from the feedback data pool; According to the first target feedback data of the session results of the session model in different session domains, update the model ability information of the session model.
9. The method according to claim 8, characterized in that, The updating the model ability information of the session model according to the first target feedback data of the session results of the session model in different session domains includes: For any session domain, according to the first target feedback data of the session result of the session model in the session domain, determine the first ability score of the session model in the session domain; According to each of the first ability scores of the session model in the session domain within the update time range, determine the second ability score of the session model in the session domain; Update the model capability information of the session model according to the second capability scores of the session model in each of the session domains.
10. The method according to claim 5, characterized in that, The method further includes: Obtain, from the feedback data pool, second target feedback data corresponding to processed user sessions of different session models in different session domains; Update the weight mapping relationship between the user attributes and the session model according to the second target feedback data and the user attribute information of the initiating user of the processed user session.
11. A session processing device, characterized in that, Includes: A parsing module, configured to use a session analysis model to perform multi-dimensional parsing on a user session to obtain multi-dimensional features and session complexity corresponding to the user session; A searching module, configured to search, according to the multi-dimensional features, for each initial model and each candidate server available for the initial model that match the user session from the latest model resource pool; wherein the model resource pool is constructed according to the model capability information of each session model and the resource information of the initial servers supporting the session model; A splitting module, configured to, in a case where the number of the initial models and the session complexity indicate that the user session requires multi-model processing, split the user session into multiple sub-sessions, and screen out, from the initial models and the candidate servers, target models that match each of the sub-sessions and target servers required by the target models; A determining module, configured to schedule each of the target servers to respectively run the target models to process each of the sub-sessions to obtain sub-session results corresponding to each of the sub-sessions, and determine a session result corresponding to the user session according to the sub-session results.
12. A computer device, characterized in that, Includes: A processor and a memory, where the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor executes the steps of the session processing method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is run on a computer device, the computer device executes the steps of the session processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Message distribution method, device and system for cloud system
CN106790610A
Conversational agent learning model service selection to address a client service request
CN110162606A
Session record searching method and device, electronic equipment and storage medium
CN111897943A
Information feedback method and device, terminal and storage medium
CN113157876A
Session recommendation method and device, electronic equipment and storage medium
CN115618079A