Consultation method and device based on large model
By employing a large-model-based consultation approach, multiple paths are generated through embedded computation. This approach enables label-free quality assessment and sample conditional routing, along with security and process verification and correction. It addresses the issues of scarce labeled data, difficult path selection, and insufficient security in generating consultation answers within vertical domains, thereby achieving high-quality and secure consultation answers.
Patent Information
- Application Number
- CN202511343046.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing technologies face challenges in generating consultation responses in vertical domains, including scarce labeled data, difficulty in selecting generation paths, insufficient security and controllability, and limited domain adaptability, making it difficult to generate high-quality, secure, and controllable consultation responses.
A large-model-based consultation approach is adopted, which generates multiple paths through embedded computation, performs label-free quality assessment, performs sample conditional routing, performs security and process verification and correction, and performs vertical strategy mapping to generate an optimized response.
It enables the generation of high-quality, secure, and controllable consultation responses without the need for labeled data, improving the adaptability and interactive effects of large models in vertical fields.
Smart Images

Figure CN120832405A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, in particular to a consulting method and device based on a large model. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, large language models (LLM for short) have been widely applied in various vertical fields, especially in professional fields such as emotional counseling, predictive analysis services, and mental health assessment. Among them, emotional counseling based on artificial intelligence technology refers to using natural language processing, machine learning, and deep learning technologies to simulate or assist the process of emotional counseling; this service aims to generate corresponding counseling answers through an intelligent system for user input counseling information, realize communication with users, and provide emotional support, psychological counseling, or suggestions for users. When the artificial intelligence system interacts with the user, it needs to generate high-quality, reliable, and scenario-process-compliant responses to the counseling information to output the required counseling answers to the user, achieving a better interaction effect.
[0003] However, the counseling answers generated by the existing technology face many challenges and technical bottlenecks in vertical field applications. First, the existing technology cannot generate counseling answers with good interaction effects without relying on high-quality labeled data in the vertical field; and even if high-quality labeled data can be obtained, it cannot ensure the effectiveness of counseling answer optimization and selection. Second, for different types of counseling information, the counseling answers generated by the existing technology have significant differences in effectiveness, making it difficult to achieve optimal path (i.e., model, prompt, or strategy) selection for generating counseling answers in different scenarios, affecting the quality of the generated counseling answers. And the counseling answers generated according to the existing technology may have problems of insufficient security and controllability, which is poor in practicality for vertical fields such as emotional counseling that require high security and controllability. In addition, the existing technology cannot guarantee the quality of the counseling answers generated for different vertical fields, affecting user experience.
[0004] Therefore, overcoming the defects of the existing technology is a problem that needs to be solved in this technical field. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a consulting method and device based on a large model.
[0006] The present application adopts the following technical solutions: In a first aspect, the present application provides a consulting method based on a large model, obtaining counseling information input by a user; generating a plurality of paths using the counseling information based on embedding calculation; performing unlabeled quality evaluation on the plurality of paths; based on the results of the unlabeled quality evaluation, sample-conditioned routing is performed on the plurality of paths to determine an optimal candidate; The optimal candidate is subjected to security and process verification and correction to obtain a final output; The final output is mapped to an optimized response through vertical strategy to generate a consultation answer.
[0007] Further, the embedding-based calculation uses the consultation information to generate a plurality of paths; the unlabeled quality evaluation of the plurality of paths comprises: A candidate path pool is constructed using the consultation information to obtain a plurality of candidate outputs to generate a plurality of paths corresponding to the consultation information; A combined embedding of each of the candidate outputs is generated; using the combined embedding, a local covariance matrix is estimated in the neighborhood of the consultation information to obtain a whitening transformation matrix; The combined embedding and the whitening transformation matrix are used to solve the consensus embedding of the consultation information, and the final quality indicator of the candidate output is determined based on the consensus embedding to complete the unlabeled quality evaluation.
[0008] Further, the combined embedding and the whitening transformation matrix are used to solve the consensus embedding of the consultation information, and the final quality indicator of the candidate output is determined based on the consensus embedding, comprising: Using the combined embedding, the whitening transformation matrix and the consensus embedding, the near neighbor closeness of the consultation information is determined based on Mahalanobis metric; The density-weighted pairwise Mahalanobis distance is calculated in the neighborhood of the consultation information, and a consistency graph is constructed; the page ranking result or feature vector center degree of the consistency graph is determined as the consistency graph center degree; The candidate output is subjected to high-priority security and medium-priority process checking to determine a rule consistency factor; The near neighbor closeness, the consistency graph center degree and the rule consistency factor are fused based on weighted geometric mean and normalized to obtain the final quality indicator of the candidate output.
[0009] Further, based on the results of the unlabeled quality evaluation, sample-conditioned routing is performed on the plurality of paths to determine an optimal candidate, comprising: The K-nearest neighbor set of the consultation information is retrieved from the history library; The local covariance matrix, the whitening transformation matrix and the consensus embedding are determined according to the K-nearest neighbor set, and the density-weighted pairwise Mahalanobis distance is calculated in the K-nearest neighbor set; For all candidate outputs of the consultation information, the candidate output with the maximum final quality indicator is determined as the optimal candidate to complete the sample-conditioned routing.
[0010] Further, the determining the optimal candidate from all candidate outputs of the consultation information according to the maximum final quality index further comprises: When the maximum final quality index is less than a routing threshold, determining the alternative path as the optimal candidate according to a fallback priority.
[0011] Further, the performing security and process verification and correction on the optimal candidate to obtain a final output comprises: Performing taboo topic detection on the optimal candidate by keyword matching, and performing capability boundary detection on the optimal candidate based on professional advice identification to perform security rule checking; When the taboo topic detection and / or the capability boundary detection hit, the optimal candidate fails the security rule checking, and determining a capability boundary declaration as the final output; When the optimal candidate passes the security rule checking, performing process rule checking on the optimal candidate to determine the final output.
[0012] Further, the performing process rule checking on the optimal candidate to determine the final output when the optimal candidate passes the security rule checking comprises: When the optimal candidate has a dialogue round number less than a round number threshold, and the optimal candidate belongs to a suggestion type output, the optimal candidate fails the process rule checking, and determining non-suggestion content as the final output; wherein the non-suggestion content comprises in-depth exploration output or empathy support output; When the optimal candidate is repeated with a historical output, the optimal candidate fails the process rule checking, and determining a candidate output with a second maximum final quality index as the final output; When the optimal candidate fails the process rule checking, determining the optimal candidate as the final output.
[0013] Further, the generating an optimized response from the final output by vertical strategy mapping comprises: Constructing a strategy label system for a specific vertical field; Determining a corresponding strategy label of the final output in the strategy label system; and determining a structured description guide according to the strategy label; Fine-tuning the structured description guide according to user history, emotional state and / or question type to obtain an optimized response.
[0014] In a second aspect, the present application further provides a consultation device based on a large model, which is used to implement the consultation method based on a large model in the first aspect, and the device comprises: At least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the large model-based consulting method of the first aspect.
[0015] In a third aspect, the present application further provides a non-volatile computer storage medium, which stores computer executable instructions executed by one or more processors to complete the large model-based consulting method of the first aspect.
[0016] The present application constructs an output quality evaluation mechanism without labeling, generates multiple candidate paths through embedded calculation, and automatically evaluates the quality of candidate output, without relying on high-quality labeled data, so as to ensure the quality optimization and selection effect of the consulting answer. Through sample conditioning dynamic routing, the optimal candidate is selected based on quality evaluation, so as to realize the optimal path selection in different scenarios; and the optimal candidate is subjected to safety and process checking and correction to obtain the final output, thereby improving the safety and controllability of the large model output in the vertical field. Through vertical policy mapping of the final output, the optimized response is generated, so as to ensure the consulting answer quality in different vertical fields, and break through the limitation of the adaptation ability of the large model output in the vertical field. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Figure 1 is a flowchart of a large model-based consulting method provided by an embodiment of the present application; Figure 2 is a general flowchart of a large model-based consulting method provided by an embodiment of the present application; Figure 3 is a flowchart of step 10 provided by an embodiment of the present application; Figure 4 is a flowchart of a labeling-free quality evaluation provided by an embodiment of the present application; Figure 5 is a flowchart of step 103 provided by an embodiment of the present application; Figure 6 is a flowchart of step 20 provided by an embodiment of the present application; Figure 7is a sample conditioning routing flowchart provided by an embodiment of the present application; Figure 8 is a flowchart of step 30 provided by an embodiment of the present application; Figure 9 is a schematic diagram of a hierarchical rule engine provided by an embodiment of the present application; Figure 10 is a flowchart of step 303 provided by an embodiment of the present application; Figure 11 is a flowchart of step 40 provided by an embodiment of the present application; Figure 12 is a schematic diagram of the architecture of a large model-based consultation device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0020] In the description of the present application, the terms "inner", "outer", "longitudinal", "transverse", "upper", "lower", "top", "bottom", and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and are not required to be constructed and operated in a specific orientation, therefore should not be understood as a limitation on the present application.
[0021] In the present application, the terms "first", "second", and the like are only for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second", and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0022] In the present application, unless otherwise specified and limited, the term "connection" should be understood broadly, for example, "connection" can be fixed connection, or detachable connection, or integral; can be directly connected, or indirectly connected through intermediate medium. In addition, the term "coupling" can be an electrical connection mode for realizing signal transmission.
[0023] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as there is no conflict.
[0024] Embodiment 1: The application of the large model output of the prior art in the vertical field has problems such as a lack of labeled data, difficulty in selecting a generation path (i.e., a model, a prompt, or a strategy), insufficient safety and controllability, and limited field adaptation capability, as follows.
[0025] (1) Lack of labeled data: Many output quality optimization and selection methods for generating answers rely on a large amount of high-quality labeled data (e.g., preference pairs, reward model labels, and correct response examples). However, in vertical fields, especially in professional fields such as emotional counseling and predictive analysis services, it is difficult to obtain such labeled data. On the one hand, it requires labeling personnel with professional knowledge, which is high in labor cost. Even if the cost is not considered, the labeling consistency is difficult to guarantee due to the strong individualization and contextualization of user expression, and the use of such labeling may affect the training process and effect. On the other hand, there are ethical and legal risks involved in sensitive user privacy content; the data distribution is long-tailed, with extremely rare data in some scenarios; that is, most of the data in the labeled data is relatively small in quantity, but the types are diverse, and only a small number of scenarios have a large amount of data, resulting in a long tail of overall data distribution, with a large number of rare data points in the tail; some scenario data is extremely rare, resulting in a very low proportion of these scenarios in the data set. The above factors cause the prior art to face the "data starvation" dilemma when generating counseling answers using large models, making it difficult to fully develop the potential of the model.
[0026] (2) Difficulty in selecting a generation path (i.e., a model, a prompt, or a strategy): In actual application scenarios of generating counseling answers, engineers usually have access to multiple models and prompt schemes, which have significant differences in effectiveness for different types of counseling information. For example, a certain model performs well in emotional counseling but is weak in relationship analysis; a certain prompt is good at short answers to short questions but performs poorly in handling long rants; and different decoding strategies (e.g., temperature, top-k, or top-p) also affect the output style and stability. The prior art often uses static single-path selection methods based on global average performance, or trains a supervised router based on labeled data and a simple integration method (e.g., voting or averaging) to solve the problem of difficulty in selecting a generation path, but these methods ignore sample differences, rely on labeling, and / or have high computational costs, making it difficult to achieve sample-level optimal path selection and severely affecting the interaction effect.
[0027] (3) Insufficient safety and controllability: Emotional counseling and other scenarios have high requirements for safety and controllability. Improper counseling answers may cause psychological harm or compliance risks. The prior art often uses pure generative methods, which have the following defects: risk of taboo topics and boundary crossing; lack of understanding of the dialogue process, improper timing and strategy selection; single and repetitive response style; lack of explainability and auditability, making it difficult to control quality.
[0028] (4) Limited field adaptation capability: the professional terms and expressions of vertical fields have specialities, and the corresponding tag and strategy system have great differences. The quality of the consultation answers generated by the general model used in the existing technology in the vertical field is uneven, and the adaptation to the specific vertical field is poor, and the consultation answers with good interaction effect cannot be generated. And the existing adaptation method (such as fine-tuning or migration) still relies on annotation to realize, and cannot fundamentally solve the problem of data scarcity.
[0029] In recent years, based on artificial intelligence technology, unsupervised and weakly supervised methods have made some progress in model output selection and quality evaluation, but most of them focus on pattern discovery or specific tasks, and it is difficult to directly solve the sample-level multi-path output selection in professional fields such as emotional counseling, prediction analysis services and mental health evaluation, resulting in the inability to generate consultation answers with good effect.
[0030] And the traditional model integration method (such as Bagging, Boosting or Stacking) is mostly suitable for supervised scenarios. Recently, unsupervised routing technology in the prior art has developed, but it is mostly limited to specific tasks or relies on additional priori. Weak supervision uses incomplete, inaccurate or noisy supervision information to train the model, alleviating the lack of annotation; but the existing methods mostly focus on learning ontology, rather than multi-path output selection and optimization under complete no label. Therefore, the prior art still cannot generate consultation answers with good interaction effect without relying on high-quality annotation data in vertical fields.
[0031] In order to solve the problems in the prior art and meet the output optimization needs of vertical fields, especially professional fields such as emotional counseling, a consultation method based on large model is needed, which can automatically evaluate and select the optimal output path without annotation data; dynamically adjust the routing strategy according to the characteristics of the input sample (i.e. consultation information), realize personalized output optimization; have perfect security mechanism and process rules to ensure the safety and consistency of the consultation answers; flexibly adapt to different vertical fields and strategy systems; and have explainability and controllability, which is convenient for artificial supervision and quality management.
[0032] In order to solve the above problems, as shown in Figure 1 The embodiment of the present application provides a consultation method based on a large model, which comprises: Step 10: obtaining the consultation information input by the user; generating a plurality of paths based on the consultation information using the embedding calculation; and performing annotation-free quality evaluation on the plurality of paths.
[0033] The embodiment of the present application takes the actual application scene of emotional counseling as an example: wherein the consultation information input by the user is the text and other consultation information input into the dialogue box by the user in the dialogue process of emotional counseling based on artificial intelligence technology.
[0034] The embodiment of the application converts the consultation information into an embedding vector representation, generates multiple paths of the embedding vector representation in parallel based on the embedding calculation, and evaluates the generated paths in a weakly supervised and unlabeled manner based on embedding technology to obtain a final quality indicator of each path. The specific way of generating multiple paths according to the consultation information is selected by a person skilled in the art according to the specific use scenario, which is not limited here. The path of the embodiment of the application is a model, a prompt or a strategy involved in the output of the large model generated for the consultation information. In order to facilitate description, the output will be used to refer to the path hereinafter.
[0035] Step 20: Based on the results of the unlabeled quality evaluation, sample conditional routing is performed on the multiple paths to determine an optimal candidate.
[0036] Sample conditional routing predicts or selects one or more paths that are most suitable for processing the consultation information according to the characteristics (such as content, semantics, etc.) of the consultation information.
[0037] Step 30: Security and process verification and correction are performed on the optimal candidate to obtain a final output.
[0038] For the optimal candidate, security and process verification and correction are used to ensure its security and controllability. When the optimal candidate fails the security and process verification, it is corrected or other pre-set output that passes the security and process verification is used as the final output. A specific example will be given hereinafter.
[0039] Step 40: The final output is mapped to an optimized response by vertical strategy mapping to generate a consultation answer.
[0040] The abstract strategy (i.e., the final output) is mapped to a specific vertical field optimized response, and then a specific output (i.e., a consultation answer) is generated according to the guidance, strategy or prompt of the optimized response to respond to the consultation information input by the user. The specific implementation way of generating the consultation answer according to the optimized response is selected by a person skilled in the art according to the specific use scenario, which is not limited here.
[0041] The application constructs an output quality evaluation mechanism without labeling, generates multiple candidate paths through embedded calculation, and automatically evaluates the quality of the candidate output without relying on high-quality labeled data, so as to ensure the quality optimization and selection effect of the consultation answer. Through sample conditioning dynamic routing, the optimal candidate is selected based on quality evaluation, so as to realize the optimal path selection in different scenes; and the optimal candidate is subjected to safety and process checking and correction to obtain the final output, thereby improving the safety and controllability of the large model output in the vertical field. Through vertical field strategy mapping on the final output, the optimized response is generated, so as to ensure the consultation answer quality in different vertical fields, and break through the limitation of the adaptation ability of the large model output in the vertical field.
[0042] The application aims to provide a large model-based consultation method based on sample-level multi-model routing and hierarchical rule coordination in a weakly supervised manner without labeling, so as to dynamically select the optimal generation path at the sample level without labeled data, and ensure safety and process consistency through rule coordination.
[0043] In an embodiment, the large model-based consultation method of the application can be realized through four core modules: an unlabeled quality evaluation module, a sample conditioning routing module, an output verification and conversation rule management module, and a vertical field strategy mapping and response generation module.
[0044] As shown in Figure 2 , it is a whole process schematic diagram of the large model-based consultation method of the embodiment of the application, and the specific description will be given as follows: In an embodiment, in order to describe the processing flow of the unlabeled quality evaluation module, as shown in Figure 3 , the step 10 comprises: Step 101: using the consultation information to construct a candidate path pool to obtain multiple candidate outputs, so as to generate multiple paths corresponding to the consultation information.
[0045] As shown in Figure 4 , in an embodiment, for the consultation information , a candidate path pool is constructed ; each receives the consultation information to output a response or a strategy label. The candidate path pool comprises a plurality of pre-training and / or adaptation models, prompt templates and decoding strategy combinations.
[0046] Step 102: generating a combined embedding of each of the candidate outputs; using the combined embedding, estimating a local covariance matrix in the neighborhood of the consultation information to obtain a whitening transformation matrix.
[0047] The combined embedding of each is formed, and the consultation information The local covariance matrix is estimated from the neighborhood of , and the whitening transformation matrix is obtained using the local covariance matrix. Map input and output combinations to dimensional embedding space, and obtain the combined embedding ;in, Represented by the pre-trained model The corresponding embedding mapping function (i.e., encoder) converts the input sequence (such as or spliced pairs ) is mapped to dimensional vector representation.
[0048] To avoid the scale and noise sensitivity of the Euclidean distance, we first construct the local covariance in the local neighborhood retrieved in the input embedding space: ; in, for the identity matrix, is the regularization coefficient (used to ensure numerical stability), is the embedding space dimension, Indicates covariance calculation, Indicates consulting information The K nearest neighbor set of Indicates the number of candidate outputs generated in step 101.
[0049] Different from the prior art that directly loads the global sample covariance, the embodiment of the present invention consults the information Construct covariances in a subset of , thereby preventing the local covariance matrix from degenerating or becoming ill-conditioned, and thus achieving greater sensitivity to the local structural features of the input and accurately reflecting the consulting information The intrinsic distribution of surrounding samples has stronger numerical robustness to noise and outliers; and because the metric after whitening transformation focuses on consulting information The local Mahalanobis distance is used instead of the global sample space, so it can improve the discrimination of subsequent quality scores.
[0050] And using the local covariance matrix Define the whitening transformation matrix .
[0051] Step 103: Use the combined embedding and the whitening transformation matrix to solve the consensus embedding of the consultation information, and determine the final quality index of the candidate output based on the consensus embedding to complete the label-free quality assessment.
[0052] In the whitening space, solve the consulting information according to the following formula Consensus embedding : ; wherein, is Huber loss function, which has robustness and can suppress the influence of outliers; denotes dimensional real number space.
[0053] In an embodiment, to illustrate the process of determining the final quality indicator of the candidate output based on consensus embedding, as shown in FIG. 3, the step 103 comprises: Figure 5 Step 1031: determining the neighbor closeness of the consultation information based on Mahalanobis distance using the combined embedding, the whitening transformation matrix and the consensus embedding.
[0054] The embodiments of the present application define three types of complementary quality factors based on consensus embedding: neighbor closeness, consistency graph centrality and rule consistency factor.
[0055] The expression of neighbor closeness is as follows: ; wherein, squares of Euclidean norms of corresponding variables.
[0056] Step 1032: calculating density-weighted pairwise Mahalanobis distance in the neighborhood of the consultation information to construct a consistency graph; determining the PageRank (i.e., webpage ranking result) or eigenvector centrality of the consistency graph as the consistency graph centrality.
[0057] In the neighborhood of the consultation information , calculate density-weighted pairwise Mahalanobis distance to construct a consistency graph , and take PageRank (i.e., webpage ranking result) or eigenvector centrality as the consistency graph centrality corresponding to the consultation information . The specific way of determining the PageRank or eigenvector centrality of the consistency graph is selected by those skilled in the art according to the specific use scenario, which is not limited here.
[0058] Step 1033: performing high-priority safety and medium-priority process checks on the candidate output to determine the rule consistency factor.
[0059] Performing high-priority safety and medium-priority process checks on the output , the corresponding expression is as follows: ; wherein, a rule violation penalty weight coefficient, to control the penalty strength of rule violation; a hierarchical rule set, including safety rules and process rules, which is determined by a person skilled in the art according to a specific use scenario, and a specific example of the hierarchical rule set will be given below; a rule violation penalty function, to calculate a penalty value according to the matching degree of the output content and the corresponding hierarchical rule.
[0060] Step 1034: Based on the weighted geometric mean and normalized fusion of the neighbor closeness, the consistent graph centrality and the rule consistency factor, the final quality indicator of the candidate output is obtained.
[0061] The expression of the final quality indicator (i.e., the path quality score) is as follows: the neighbor closeness, the exponential weight corresponding to the neighbor closeness; the consistent graph centrality, the exponential weight corresponding to the consistent graph centrality; the rule consistency factor, the exponential weight corresponding to the rule consistency factor.
[0062] The above formula adopts a weighted geometric mean and normalized fusion architecture, which is different from the traditional linear weighted normalization. In the fusion of the neighbor closeness, the consistency centrality and the rule consistency factor, the present application freely adjusts the relative contribution of the three by using the exponential weight, and takes into account the characteristics of each quality factor; through denominator normalization, it ensures , which conforms to the probability distribution property and is convenient for subsequent routing decision-making; the rule consistency factor is internalized into the quality scoring system for the first time, realizing seamless coupling of model scoring and rule checking.
[0063] The above design is different from the scheme that only relies on global Euclidean distance and triple identity, and the specific differences are as follows: local Mahalanobis metric is used instead of global Euclidean distance; consensus embedding is constructed; the consistent graph centrality is used to capture the consistent information of the cross-path structure; and the rule consistency is directly internalized into the quality score.
[0064] The embodiment of the present application automatically evaluates the quality of the candidate output through the latent variable graph model and embedding calculation, evaluates the output quality score of different candidate outputs without labeled data by constructing a latent variable graph model, and evaluates the output quality in combination with three types of complementary quality factors.
[0065] In one embodiment, in order to illustrate the processing flow of the sample conditional routing module, as shown in FIG. 7, the processing flow of the sample conditional routing module includes the following steps:Figure 6 As shown, the step 20 includes: Step 201: Retrieve the K nearest neighbor set of the consultation information from the history database.
[0066] like Figure 7 As shown, the embodiment of the present invention designs a sample conditional dynamic routing algorithm to select the best candidate output for each consultation information. Among them, the history library is used for the consultation information received historically. The specific implementation method of the history library is selected by those skilled in the art according to the specific usage scenario and is not limited here. In an optional embodiment, the consultation information Input pre-trained model Get vector representation ,use In the historical database retrieval, the K nearest neighbor set is searched through the nearest neighbor, and the neighborhood is weighted with the kernel weight.
[0067] Step 202: Determine the local covariance matrix, whitening transformation matrix and consensus embedding according to the K-nearest neighbor set, and calculate the density-weighted pairwise Mahalanobis distance within the K-nearest neighbor set.
[0068] The specific implementation method of calculating the density-weighted pairwise Mahalanobis distance within the K nearest neighbor set is selected by those skilled in the art according to the specific usage scenario.
[0069] Step 203: For all candidate outputs of the consultation information, determine the candidate output with the largest final quality index as the optimal candidate to complete sample conditional routing.
[0070] The embodiment of the present invention generates the optimal candidate based on dynamic routing and abstention. Dynamic routing is implemented with the final quality index as the sample-level quality score, according to the following formula: ; in, represents the best candidate, Represents the final quality indicator.
[0071] In one embodiment, to illustrate the situation of abandonment, step 203 further includes: when the maximum final quality indicator is less than the routing threshold, determining the alternative path as the optimal candidate according to the fallback priority.
[0072] The maximum final quality index refers to the maximum final quality index among multiple candidate outputs generated for the same consultation information; the routing threshold is determined by those skilled in the art according to specific usage scenarios and is not limited here.
[0073] when When the fallback (historical consistency / majority consistency) and the safe fallback are executed, the alternative path is determined as the optimal candidate. The specific implementation of the alternative path is determined by the person skilled in the art according to the specific use scenario.
[0074] The fallback priority is: safe fallback > A-class boundary management > highest historical consistency strategy. In an embodiment, the corresponding descriptions of the "safe fallback", "A-class boundary management" and "highest historical consistency strategy" can be pre-set as the output, that is, the corresponding alternative path. When the maximum final quality index is less than the routing threshold, the safety of the system is primarily ensured, and the "safe fallback" alternative path is taken as the optimal candidate. Secondly, if there is no "safe fallback" alternative path, boundary management is performed, for example, by taking the "A-class boundary management" alternative path as the optimal candidate, it is declared that the capability boundary of the output content is limited. Finally, if there is no "A-class boundary management" alternative path, the historical optimal candidate is taken as the optimal candidate generated this time according to the highest historical consistency strategy.
[0075] In an embodiment, in order to illustrate the hierarchical logical rules, as shown in Figure 8 the step 30 comprises: Step 301: taboo topic detection is performed on the optimal candidate by keyword matching, and capability boundary detection is performed on the optimal candidate based on professional advice identification, so as to perform safety rule checking.
[0076] As shown in Figure 9 , the embodiment of the application establishes a hierarchical logical rule engine to check and correct the candidate output with high-priority safety rules and medium-priority process rules. The corresponding rule execution process is: safety rule checking is performed first, and then process rule checking is performed; if neither is triggered, the candidate is kept.
[0077] The high-priority safety rule checking includes: (1) taboo topic detection; (2) capability boundary detection. The specific implementation of the taboo topic detection by keyword matching and the capability boundary detection based on professional advice identification is selected by the person skilled in the art according to the specific use scenario, which is not limited here.
[0078] Step 302: when the taboo topic detection and / or the capability boundary detection hit, the optimal candidate fails to pass the safety rule checking, and the capability boundary declaration is determined as the final output.
[0079] The taboo topic detection is: keyword and semantic similarity double-channel detection, and once it hits, the "capability boundary declaration" is forced to be output. The capability boundary detection is: when the intent exceeds the service range (for example, medical, legal or investment, etc.), the "capability boundary declaration" is forced to be output. In an embodiment, as shown in Figure 9As shown, a strategic labeling system can be developed in advance for each vertical field, and the capability boundary declaration can be numbered as A5.2. This allows the user to know that their input has exceeded the capability boundary of the model and that the model cannot provide the expected answer when the final output is subsequently displayed to the user.
[0080] Step 303: When the optimal candidate passes the security rule check, a process rule check is performed on the optimal candidate to determine a final output.
[0081] Medium-priority process rules: Avoid premature advice in early rounds, prioritize empathy or exploration; avoid strategy duplication.
[0082] In one embodiment, if Figure 10 As shown, step 303 includes: Step 3031: When the number of dialogue turns of the optimal candidate is lower than the turn threshold and the optimal candidate belongs to the suggestion type output, the optimal candidate fails to pass the process rule check and the non-suggestion content is determined as the final output; wherein, the non-suggestion content includes in-depth exploration output or empathy support output.
[0083] Among them, the turn threshold, the specific content of the non-suggested content in-depth exploration output, and the specific content of the empathy support output are all determined by those skilled in the art based on the specific usage scenario. Follow the following rules: avoid providing suggestions too early when the number of conversation turns is small, and give priority to empathy support and / or in-depth exploration. Figure 9 As shown, in one embodiment, the round threshold can be 3. Suggestion output refers to outputting a description that provides suggestions based on the consultation information. In-depth exploration output refers to outputting a description that further analyzes and explores the consultation information, allowing the user to perceive that the model is conducting in-depth exploration of the consultation information. Empathy and support output refers to outputting a description that is based on the consultation information.
[0084] Step 3032: When the optimal candidate is repeated with the historical output, the optimal candidate fails the process rule check, and the candidate output with the second largest final quality index is determined as the final output.
[0085] Among them, historical output refers to the output of the previous round in the conversation with the same user. It is carried out according to the following rules: detect strategy duplication and temporarily reduce its priority, select the suboptimal strategy, and improve the richness of interaction. Figure 9 As shown, for example, if the optimal candidate is repeated in the historical output, that is, the outputs in the last two rounds are repeated, the priority of the optimal candidate obtained in step 302 is lowered, and according to the ranking of the final quality indicators corresponding to the consultation information in step 203, the candidate output with the second highest final quality indicator is determined as the final output, thereby modifying the output strategy for this time.
[0086] Step 3033: When the optimal candidate fails the process rule check, the optimal candidate is determined as the final output.
[0087] The embodiment of the present invention introduces a management mechanism that combines output verification with session rules. The layered rule engine verifies and corrects the security and process consistency of candidate outputs to ensure the security, professionalism and coherence of the final response.
[0088] In one embodiment, in order to illustrate the process of vertical policy mapping and response generation, as shown in FIG. Figure 11 As shown, the step 40 includes: Step 401: Build a strategic labeling system for specific vertical fields.
[0089] By building a strategy label system, we can optimize the strategy system and implementation for vertical scenarios.
[0090] The following is a specific example of a policy labeling system: policy labels A1 to A6 and their definitions, applicable scenarios, implementation guidance, and example expressions.
[0091] “A1: Consultation initiation and setup (A1.1 Opening greeting, A1.2 Self-introduction, A1.3 Consultation question acquisition, A1.4 In-depth exploration, A1.5 Energy connection statement); A2: Interpretation and Analysis (A2.1 Interpretation of the Essence of Relationships, A2.2 Identification of Obstacles, A2.3 Analyzing Others' Psychological Behaviors, A2.4 User Perception, A2.5 Future Prediction); A3: Interactive response (A3.1 direct answer, A3.2 provide suggestions); A4: Emotional support (A4.1 Empathy and support, A4.2 Positive empowerment); A5: Boundary Management (A5.1 Credibility Protection, A5.2 Capability Boundary Declaration, A5.3 Payment Guidance); A6: General Politeness (A6.1 Standard Politeness); Output Mapping: Find implementation guidance and generate structured response points for the final strategy.” Step 402: Determine the policy tag corresponding to the final output in the policy tag system; and determine a structured description guide according to the policy tag.
[0092] For example, the policy tag is "A5.2 Capability Boundary Declaration". According to the final obtained policy tag, the corresponding description information is searched and a structured description guide is generated. The specific implementation method of this process is selected by those skilled in the art according to the specific usage scenario.
[0093] Step 403: details of the structured description guidance are fine-tuned according to the user history, emotional state and / or question type, to obtain an optimized response.
[0094] In an embodiment, the model uses a set of corresponding parameters to represent the user history, emotional state or question type of the user in the current dialogue, and then adjusts the structured description guidance in detail according to whether the parameters exceed a certain value, such as adding soothing statements in the structured description guidance, so as to realize the mapping of the abstract strategy to the specific vertical field response implementation.
[0095] In order to verify the effectiveness of the large model-based counseling method of the embodiments of the present application, the embodiments of the present application are evaluated on multiple vertical field data sets. Among them, the vertical field data sets used include: emotional counseling data set, prediction analysis service data set and mental health evaluation data set.
[0096] The embodiments of the present application provide a specific example of a set of experimental results based on a commonly used evaluation index system, as follows: The output quality (i.e., the success rate of optimization) of the prior art method and the large model-based counseling method of the embodiments of the present application in multiple vertical fields is shown in Table 1 below.
[0097] Table 1 Output quality comparison table
[0098] Among them, the baseline method in the prior art in Table 1 includes: randomly selecting (Random Selection, abbreviated as Selection) to generate a path; using a single path with the best global performance, i.e., the best single path (Best Single Path, abbreviated as BSP); majority voting (Majority Voting, abbreviated as MV); supervised routing based on labeled training, i.e., supervised router (Supervised Router, abbreviated as SR); global unsupervised routing (SMOOTHIE-Global, abbreviated as SMOOTHIE-G). As can be seen from the data in the above table, the output quality of the large model-based counseling method of the embodiments of the present application in the fields of emotional counseling, prediction analysis service and mental health is higher than that of the prior art method, and the corresponding average success rate and standard deviation reflect the robustness of the large model-based counseling method of the embodiments of the present application. In each data set, the highest output quality is obtained, which is an average of 2.8 percentage points higher than the supervised router and 5.5 percentage points higher than the unsupervised baseline, and has higher stability.
[0099] The safety indicators (Safety Rate) of the prior art method and the large model-based counseling method of the embodiments of the present application in multiple vertical fields are shown in Table 2 below.
[0100] Table 2 Safety performance evaluation table
[0101] The comparison results of the user experience scores (UX Score) of the prior art method and the large model-based consultation method of the embodiments of the present application in multiple vertical fields are shown in Table 3.
[0102] Table 3 User experience evaluation table
[0103] The system performance index results of the large model-based consultation method of the embodiments of the present application in multiple vertical fields are shown in Table 4.
[0104] Table 4 System performance index table
[0105] In order to verify the large model-based consultation method of the embodiments of the present application, a specific example of ablation experiment results is provided as shown in Table 5.
[0106] Table 5 Ablation experiment results
[0107] The large model-based consultation method of the embodiments of the present application is particularly suitable for large model-based consultation tasks in vertical fields such as emotional consultation, prediction analysis service and mental health assessment, to realize response optimization and output control. It can significantly improve the output quality and system security while reducing the labeling cost. The experimental results show that compared with the traditional fixed model and single path method, the present application improves the output quality (optimization success rate) by 8-15%, reduces the labeling cost by more than 80%, and realizes 100% security guarantee through the logic rule mechanism.
[0108] As shown in Figure 12 , it is an architecture schematic diagram of the large model-based consultation device of the embodiments of the present application. The large model-based consultation device of the present embodiment includes one or more processors 21 and memories 22. Among them, Figure 12 The processor 21 is taken as an example.
[0109] The processor 21 and the memory 22 can be connected through a bus or other means, Figure 12 The connection through the bus is taken as an example.
[0110] The memory 22, as a non-volatile computer readable storage medium, can be used to store non-volatile software programs and non-volatile computer executable programs, such as the large model-based consulting method in the embodiments. The processor 21 executes the large model-based consulting method by running the non-volatile software programs and instructions stored in the memory 22.
[0111] The memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 22 can optionally include a memory disposed remotely with respect to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0112] The program instructions / modules are stored in the memory 22, and when executed by the one or more processors 21, perform the large model-based consulting method in the above-mentioned embodiments, for example, perform each step of the large model-based consulting method described above.
[0113] It is worth noting that the information interaction, execution process, etc. between the modules and units in the above-mentioned apparatus and system are based on the same concept as the processing method embodiments of the present application, and the specific content can be referred to the description in the method embodiments of the present application, which will not be repeated here.
[0114] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be instructed by a program to relevant hardware, and the program can be stored in a computer readable storage medium, which can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0115] The above only describes the preferred embodiments of the present application and should not be used to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A large model-based consulting method, characterized by, The method comprises: obtaining user input consultation information; generating multiple paths using the consultation information based on embedding calculation; performing unlabeled quality evaluation on the multiple paths; based on the results of the unlabeled quality evaluation, sample conditional routing is performed on the multiple paths to determine the optimal candidate; safety and process checking and correction are performed on the optimal candidate to obtain the final output; the final output is mapped to generate an optimized response through vertical strategy mapping to generate a consultation answer.
2. The large model-based consulting method of claim 1, wherein, The method comprises: using the consultation information to construct a candidate path pool to obtain multiple candidate outputs to generate multiple paths corresponding to the consultation information; generate a combined embedding for each candidate output; using the combined embedding, estimate the local covariance matrix in the neighborhood of the consultation information to obtain the whitening transformation matrix; using the combined embedding and the whitening transformation matrix, the consensus embedding of the consultation information is solved, and the final quality indicator of the candidate output is determined based on the consensus embedding to complete the unlabeled quality evaluation.
3. The large model-based consulting method of claim 2, wherein, The method comprises: using the combined embedding, the whitening transformation matrix and the consensus embedding, the near neighbor closeness of the consultation information is determined based on Mahalanobis metric; calculate the density-weighted pairwise Mahalanobis distance in the neighborhood of the consultation information to construct a consistency graph; the page ranking result or feature vector center degree of the consistency graph is determined as the consistency graph center degree; high-priority safety and medium-priority process checking are performed on the candidate output to determine the rule consistency factor; based on weighted geometric mean and normalized fusion, the near neighbor closeness, the consistency graph center degree and the rule consistency factor are fused to obtain the final quality indicator of the candidate output.
4. The large model-based consulting method of claim 1, wherein, The method comprises: retrieve the K-nearest neighbor set of the consultation information in the history library; determine the local covariance matrix, the whitening transformation matrix and the consensus embedding according to the K-nearest neighbor set, and calculate the density-weighted pairwise Mahalanobis distance in the K-nearest neighbor set; for all candidate outputs of the consultation information, the candidate output with the maximum final quality indicator is determined as the optimal candidate to complete sample conditional routing.
5. The large model-based advice sample conditioned quality score method of claim 4, wherein, The method further comprises: when the maximum final quality indicator is less than the routing threshold, the alternative path is determined as the optimal candidate according to the fallback priority.
6. The large model based consulting method of claim 1, wherein, The method comprises: taboo topic detection is performed on the optimal candidate by keyword matching, and ability boundary detection is performed on the optimal candidate based on professional advice identification to perform safety rule checking; when the taboo topic detection and / or ability boundary detection hit, the optimal candidate fails the safety rule check, and the ability boundary declaration is determined as the final output; when the optimal candidate passes the safety rule check, process rule check is performed on the optimal candidate to determine the final output.
7. The large model-based consulting method of claim 6, wherein, The method comprises: when the number of dialogue rounds of the optimal candidate is lower than the round threshold, and the optimal candidate belongs to the suggestion class output, the optimal candidate fails the process rule check, and the non-suggestion content is determined as the final output; wherein the non-suggestion content includes in-depth exploration output or empathy support output; when the optimal candidate is repeated with historical outputs, the optimal candidate fails the process rule check, and a final output is determined as a candidate output with a second largest final quality indicator; when the optimal candidate fails the process rule check, the optimal candidate is determined as the final output.
8. The large model-based consulting method according to any one of claims 1-7, characterized in that, The method comprises: constructing a strategy tag system for a specific vertical field; determining a corresponding strategy tag of the final output in the strategy tag system; determining a structured description guide according to the strategy tag; optimizing the structured description guide according to user history, emotional state and / or question type to obtain an optimized response.
9. A non-transitory computer storage medium, comprising: The computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors to complete the consultation method based on the large model in any one of claims 1-8.
10. A large model-based consultation device, characterized by, comprise: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the consultation method based on the large model in any one of claims 1-8.
Citation Information
Patent Citations
Quality evaluation model training method, multi-round dialogue quality evaluation method and multi-round dialogue quality evaluation device
CN117556005A
Dimension reduction method of annotated data, anomaly detection method, device, system and equipment
CN118230320A
Vector database retrieval method and device based on large model, terminal and medium
CN118312594A
system
JP2025055140A