Large language model optimization method and device based on multi-modal feedback and reinforcement learning

By obtaining reply information and page feedback from the large language model, and using the target clustering algorithm to filter and train samples, the problem of inability to capture user needs in training large language model is solved, achieving more accurate model output and improving user experience.

CN120386849AActive Publication Date: 2025-07-29HAIYAN COUNTY NANBEIHU MEDICAL ARTIFICIAL INTELLIGENCE RES INST +1

Patent Information

Application Number
CN202510885917.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

During the training process, existing large language models cannot fully capture the user's complex attitudes and diverse needs for model output, resulting in the output responses that are inconsistent with the user's actual needs, poor accuracy and poor user experience.

Method used

By obtaining the reply information set output by the large language model and the page feedback set collected by the device, the target clustering algorithm is used to remove abnormal reply information, determine the satisfactory information set, generate the feedback data set, and filter the positive and negative samples according to the quality of multiple rounds of reply, and train the large language model.

Benefits of technology

It improves the output accuracy of large language models, enhances user experience, and ensures that the model output is more in line with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386849A_ABST
    Figure CN120386849A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large language model optimization method and device based on multi-modal feedback and reinforcement learning. A specific embodiment of the method comprises the steps of obtaining a reply information set and a page feedback set; removing the abnormal reply information to obtain a normal reply information set; determining satisfaction information corresponding to each piece of normal reply information; generating a first feedback data set; screening out a target reply information set; for each piece of target reply information, executing a data generation step: taking the corresponding initial reply information as an anchor sample, taking the reply content of which the corresponding reply quality is higher than that of the anchor sample as a positive sample, and taking the reply content of which the corresponding reply quality is lower than that of the anchor sample as a negative sample; generating second feedback data; and carrying out model training on the large language model. According to the embodiment, through the multi-modal information fed back by the page and the performance condition of multi-round output of the large language model, the large language model can be efficiently trained, and the large language model with more accurate output is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method and apparatus for optimizing large language models based on multi-modal feedback and reinforcement learning. Background Art

[0002] Currently, with the advent of the intelligent era, large language models empower various fields to improve industrial production capacity and work efficiency. For the training of large language models, the commonly adopted method is to directly train an initial large language model through a pre-acquired data set to obtain a large language model.

[0003] However, when using the above method to train large language models, the following technical problems often exist: The large language models trained by conventional methods cannot comprehensively capture the complex attitudes and diverse needs of users towards the model output, resulting in the responses output by the large language models not matching the content of the actual needs of users, with poor accuracy of the large language models and poor user experience. Summary of the Invention

[0004] This section of the present disclosure is used to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. This section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present disclosure propose a method and apparatus for optimizing large language models based on multi-modal feedback and reinforcement learning to solve one or more of the technical problems mentioned in the above background art section.

[0006] In a first aspect, some embodiments of the present disclosure provide a method for optimizing a large language model based on multimodal feedback and reinforcement learning, including: obtaining a set of response information output by the large language model and a set of page feedback collected by a device; using a target clustering algorithm to remove abnormal response information in the above set of response information to obtain a set of normal response information; determining satisfactory information corresponding to each normal response information according to the page feedback subset corresponding to the above set of normal response information to obtain a set of satisfactory information; generating a first feedback data set according to the above set of satisfactory information and the set of normal response information; screening out target response information with multi-round responses in the corresponding response content from the above set of normal response information to obtain a set of target response information; for each target response information, performing a data generation step: determining the initial response information corresponding to the above target response information as an anchor sample, determining response content with a response quality higher than the anchor sample in the multi-round responses as positive samples, and determining response content with a response quality lower than the anchor sample in the multi-round responses as negative samples; generating second feedback data for the above anchor sample, positive sample set, and negative sample set; training the above large language model according to the above first feedback data set and the second feedback data set to obtain a trained large language model.

[0007] In a second aspect, some embodiments of the present disclosure provide an apparatus for optimizing a large language model based on multimodal feedback and reinforcement learning, including: an obtaining unit configured to obtain a set of response information output by the large language model and a set of page feedback collected by a device; a removing unit configured to use a target clustering algorithm to remove abnormal response information in the above set of response information to obtain a set of normal response information; a determining unit configured to determine satisfactory information corresponding to each normal response information according to the page feedback subset corresponding to the above set of normal response information to obtain a set of satisfactory information; a generating unit configured to generate a first feedback data set according to the above set of satisfactory information and the set of normal response information; a screening unit configured to screen out target response information with multi-round responses in the corresponding response content from the above set of normal response information to obtain a set of target response information; an execution unit configured to, for each target response information, perform a data generation step: determining the initial response information corresponding to the above target response information as an anchor sample, determining response content with a response quality higher than the anchor sample in the multi-round responses as positive samples, and determining response content with a response quality lower than the anchor sample in the multi-round responses as negative samples; generating second feedback data for the above anchor sample, positive sample set, and negative sample set; a training unit configured to train the above large language model according to the above first feedback data set and the second feedback data set to obtain a trained large language model.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0010] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the method for optimizing a large language model based on multi-modal feedback and reinforcement learning according to some embodiments of the present disclosure, through the multi-modal information of page feedback and the performance of the large language model in multiple rounds of output, the large language model can be efficiently trained to obtain a large language model with more accurate output. Specifically, the reason for the inaccurate output of the relevant large language model is that: the large language model trained conventionally cannot comprehensively capture the complex attitudes and diverse needs of users towards the model output, resulting in the reply output by the large language model not matching the content of the actual needs of users, with poor accuracy of the large language model and poor user experience. Based on this, in the method for optimizing a large language model based on multi-modal feedback and reinforcement learning according to some embodiments of the present disclosure, first, a reply information set output by the large language model and a page feedback set collected by the device are obtained. Here, by obtaining the reply information set, it serves as the data basis for subsequent model training of the large language model. The page feedback set can reflect the effective feedback of users on the reply content during the reply process. By collecting the page feedback set, the subsequent large language model can comprehensively consider the page feedback situation and achieve precise training of the corresponding model. Then, using the target clustering algorithm, the abnormal reply information in the above-mentioned reply information set can be efficiently and accurately removed to obtain a normal reply information set. Next, according to the page feedback subset corresponding to the above-mentioned normal reply information set, through various situations of page feedback, the satisfaction information corresponding to each normal reply information can be accurately determined to obtain a satisfaction information set. Through the satisfaction information, the subsequent large language model can know whether the reply content output during training meets the needs of users, so as to achieve precise training of the large language model. Then, according to the above-mentioned satisfaction information set and normal reply information set, an accurate first feedback data set can be generated for subsequent training of the large language model to generate a more precise model. Further, target reply information with multiple rounds of replies in the corresponding reply content is screened out from the above-mentioned normal reply information set to obtain a target reply information set, so as to facilitate the key screening of data for the performance of the large language model in multiple rounds of replies, so that the subsequent large language model can reduce the number of rounds and improve the output accuracy during multiple rounds of replies. Furthermore, for each target reply information, a data generation step is executed: First step, the initial reply information corresponding to the above-mentioned target reply information is determined as the anchor sample, the reply content with a reply quality higher than the anchor sample in multiple rounds of replies is determined as the positive sample, and the reply content with a reply quality lower than the anchor sample in multiple rounds of replies is determined as the negative sample, so as to facilitate the subsequent large language model to learn the reply situation in multiple rounds of replies. Second step, generate the second feedback data of the above-mentioned anchor sample, positive sample set and negative sample set as training samples, so that the subsequent large language model can improve the output accuracy under the condition of reducing the number of rounds of replies and continuously learn. Finally, according to the above-mentioned first feedback data set and second feedback data set, the above-mentioned large language model is trained to obtain a trained large language model.In summary, by considering the multi-modal feature information in terms of page feedback and the output accuracy in terms of multi-round responses, the large language model can be continuously improved with respect to the feedback content and multi-round response situations, thereby enhancing the output accuracy of the large language model. Description of the Drawings

[0011] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flowchart of some embodiments of a method for optimizing a large language model based on multi-modal feedback and reinforcement learning according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of an apparatus for optimizing a large language model based on multi-modal feedback and reinforcement learning according to the present disclosure; Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Description of the Embodiments

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0014] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0015] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions performed by these devices, modules, or units or their interdependent relationships.

[0016] It should be noted that the modifiers "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0019] Referring to Figure 1 , a flowchart 100 of some embodiments of a large language model optimization method based on multimodal feedback and reinforcement learning according to the present disclosure is shown. The large language model optimization method based on multimodal feedback and reinforcement learning includes the following steps: Step 101, obtain a reply information set output by the large language model and a page feedback set collected by the device.

[0020] In some embodiments, the execution subject (e.g., an electronic device) of the above large language model optimization method based on multimodal feedback and reinforcement learning can obtain the reply information set output by the large language model and the page feedback set collected by the device in a wired or wireless manner. Among them, the large language model (LLM) can be a large-scale pre-trained model constructed based on a deep neural network, which learns the statistical laws and semantic associations of language through a large amount of text data and has capabilities such as text generation, understanding, and reasoning. The reply information can be the reply result after the large language model and the user answer. The reply information set can be a massive set of reply contents output by the large language model within a target time period. For example, the target time period can be half a year. The number of the reply information set can be massive. The large language model can be a common large language model on the market that supports answering in various fields. The page feedback can be a page operation by the user to give quality feedback on the output content of the large language model. In practice, the page feedback can be empty information. That is, it represents that the user has not given any page feedback on the reply content. The page feedback can also be, but is not limited to, one of the following: page like, page content collection, page content copy, page content negative review, page dwell time. The page feedback set can be obtained by using refined event listening technology in the client-side buried point design. There is a one-to-one correspondence between the page feedback in the page feedback set and the reply information in the reply information set. The large language model has a corresponding model operation page. The model operation page supports the input of questions and page-related operations on the model reply content. The reply information can include, but is not limited to, at least one of the following: reply user information, reply time, reply question, reply content.

[0021] Step 102, use a target clustering algorithm to remove abnormal reply information in the above reply information set to obtain a normal reply information set.

[0022] In some embodiments, the above-mentioned execution entity may use a target clustering algorithm to remove abnormal response information from the above-mentioned response information set, so as to obtain a normal response information set. Among them, the target clustering algorithm may be the K-Means algorithm, or may also be hierarchical clustering (Hierarchical Methods). The abnormal response information may be response information that does not proceed normally in the response process. That is, the abnormal response information may be response information of a malicious feedback attack. The abnormal response information is noise response information. The normal response information is response information that proceeds normally in the response process. That is, the normal response information may be response information that is not a malicious feedback attack.

[0023] As an example, first, use the K-Means algorithm to perform clustering processing on each response information in the response information set according to various indicators corresponding to the response information, so as to use the response information whose sum of distances from each cluster center is higher than the target value as abnormal response information, and obtain an abnormal response information set. Then, remove the abnormal response information set from the response information set to obtain a normal response information set.

[0024] In some optional implementation manners of some embodiments, the above-mentioned execution entity may use a target clustering algorithm to remove abnormal response information from the above-mentioned response information set, so as to obtain a normal response information set, including the following steps: The first step is to use the above-mentioned target clustering algorithm to perform clustering processing on the above-mentioned response information set to obtain a response information cluster set. Among them, the number of clusters corresponding to the response information clusters may be preset. Each response information cluster has corresponding cluster label information. The cluster label information includes: multiple labels. Each label may characterize the information category corresponding to each response information in the response information cluster. For example, the cluster label information may include: "female", "high school", "campus problem type".

[0025] As an example, the above-mentioned execution entity may use features such as "problem type" and "user group" as clustering features, and use the above-mentioned target clustering algorithm to perform clustering processing on the above-mentioned response information set to obtain a response information cluster set.

[0026] Step 2: Determine the between-cluster variance between each of the response information clusters in the above response information cluster set and the within-cluster variance corresponding to each response information cluster. Among them, the between-cluster variance can characterize the degree of separation between different response information clusters. In practice, the between-cluster variance can reflect the degree of difference between the cluster centers of different response information clusters and the global mean. In practice, the larger the between-cluster variance, the higher the degree of inter-cluster separation and the better the clustering effect. The within-cluster variance can characterize the degree of dispersion between the data points within the same cluster and the cluster center, and can reflect the tightness of the data within the cluster. In practice, the smaller the corresponding value of the within-cluster variance, the more concentrated the data points within the cluster and the better the clustering effect.

[0027] Step 3: According to the above between-cluster variance and the obtained within-cluster variance set, adaptively adjust each response information cluster to obtain an adjusted response information cluster set.

[0028] As an example, the above-mentioned execution entity can try to remove the marginal response information for each response information cluster to determine the between-cluster variance and the within-cluster variance set, so as to adjust the between-cluster variance to the maximum value and the minimum value of each within-cluster variance in an equilibrium state, and obtain an adjusted response information cluster set. Among them, the marginal response information can be the response information in the response information cluster whose distance from the cluster center is higher than a predetermined value.

[0029] Step 4: According to the above response information cluster set and the above adjusted response information cluster set, screen out the initial abnormal response information set from the above response information set. Among them, the initial abnormal response information can be the response information that is preliminarily determined to possibly have abnormal problems. The abnormal problem can be an abnormal response. The abnormal response can be an abnormal response.

[0030] As an example, the above-mentioned execution entity compares the response information clusters in the response information cluster set with the response information clusters in the adjusted response information cluster set to obtain a set of differential response information as the initial abnormal response information set.

[0031] Step 5: For each initial abnormal response information, perform the following determination steps: Sub-step 1: Obtain the target response content corresponding to the above initial abnormal response information. Among them, the target response content can be the actual response content of the initial abnormal response information.

[0032] Sub-step 2: Determine the response content type and content validity corresponding to the above target response content. Among them, the response content type can be the content type corresponding to the response content. The content validity can be the proportion of the normal content in the response content. The normal content can be non-sensitive and non-violating content. In practice, through semantic extraction of the response content, it can be determined whether the content is a normal response content, and through word extraction, it can be determined whether there is sensitive content in the response content.

[0033] Sub-step 3: In response to determining that the above content validity does not meet the target validity requirement and / or the response content type belongs to the set of predetermined content types, determine the above initial abnormal response information as abnormal response information. Among them, each predetermined content type in the set of predetermined content types can be a content type indicating that the corresponding response content is sensitive content. In practice, the predetermined content type can indicate that there is sensitivity in the corresponding response question or the response content. For example, the sensitive content is illegal content. The target validity requirement can be that the numerical value corresponding to the content validity is higher than the target value. The content validity can be information in numerical form. The higher the content validity, the more likely the response content in the corresponding response information is a content for normal response.

[0034] Step 6: Remove the obtained set of abnormal response information from the above set of response information to obtain a set of normal response information.

[0035] Step 103: Determine the satisfaction information corresponding to each normal response information according to the page feedback subset corresponding to the above set of normal response information to obtain a set of satisfaction information.

[0036] In some embodiments, the above execution subject can determine the satisfaction information corresponding to each normal response information according to the page feedback subset corresponding to the above set of normal response information to obtain a set of satisfaction information. Among them, there is a one-to-one correspondence between the normal response information in the set of normal response information and the page feedback in the page feedback subset. The satisfaction information can indicate the satisfaction degree of the user with the response content output by the large language model in the normal response information. In practice, the satisfaction information can be information in numerical form or in the form of a label. The higher the corresponding numerical value, the more satisfied the user is with the response content output by the large language model in the normal response information.

[0037] As an example, for each normal response information, the above execution subject can numericalize each page feedback in the page feedback subset to obtain a page feedback numerical subset. Then, input the page feedback numerical subset into a linearly regression model with pre-trained parameter weights to obtain the satisfaction information.

[0038] In some alternative implementations of some embodiments, the page feedback in the above page feedback set includes at least one of the following: an explicit feedback information set after an operation on the model operation page, an implicit feedback information set after an operation on the model operation page. Among them, the explicit feedback information can be the feedback content under an explicit operation. The explicit operation can be a page operation that is relatively obvious on the model operation page. The implicit feedback information can be the feedback information that can be summarized only through relevant operations implicitly. In practice, for the reply information, the user can perform multiple explicit operations on the model operation page to obtain the explicit feedback information set. The explicit feedback information includes at least one of the following: like feedback, comment feedback. Among them, the like feedback represents the like processing of the reply information on the model operation page. The like feedback can include: the like behavior and the like timestamp. The comment feedback can represent the comment processing of the reply information on the model operation page. The comment feedback can include: the comment content and the comment timestamp. The comment content can be the comment content after preprocessing. That is, when the user finishes inputting and submits, text preprocessing is immediately performed. The text preprocessing can include: removing special characters, performing word segmentation, etc. The processed comment data is stored together with meta-information such as the user ID and the question ID to obtain the comment content. The implicit feedback information includes at least one of the following: page stay duration, context switching frequency, number of answer modifications. The page stay duration can be the duration that the user stays on the model operation page to view the reply content. The context switching frequency can be counted by monitoring the number of jumps between different question pages or pages related to the same question, and at the same time, using the browser's history API to analyze the path of the user's context switching to obtain it. The number of answer modifications can be that when the user edits the generated answer each time, the background counter automatically accumulates, and records the time and content of each modification.

[0039] Optionally, the above execution subject can determine the satisfaction information corresponding to each normal reply information according to the page feedback subset corresponding to the above normal reply information set, including the following steps: The first step is to determine the target page feedback corresponding to the above normal reply information.

[0040] The second step is to determine the comment sentiment tendency information according to the comment feedback included in the above target page feedback. Among them, the comment sentiment tendency information can be the sentiment tendency of the user corresponding to the comment content. In practice, the comment sentiment tendency information can be one of the following: positive sentiment tendency, neutral sentiment tendency, negative sentiment tendency.

[0041] As an example, the above-mentioned execution entity can input the comment content in the comment feedback into a pre-trained sentiment analysis model to obtain comment sentiment tendency information. In a specific scenario, the sentiment analysis model can be a part of the large language model, that is, the large language model supports sentiment analysis processing. The sentiment analysis model can also be a deep learning model independent of the large language model. The sentiment analysis model can be a common model for sentiment analysis.

[0042] In the third step, the above-mentioned comment sentiment tendency information and the above-mentioned like feedback are quantized to obtain a sentiment tendency value and a like value. Among them, information quantization can be to numerically transform the information.

[0043] In practice, a positive sentiment tendency can be converted into the numerical value "1", a neutral sentiment tendency can be converted into the numerical value "0", and a negative sentiment tendency can be converted into the numerical value "-1". In response to the like content in the like feedback being that a like has been given, the like value corresponding to the like feedback is set to the numerical value "1", and in response to the like content in the like feedback being that no like has been given, the like value corresponding to the like feedback is set to the numerical value "-1".

[0044] In the fourth step, obtain the first signal coefficient and the first decay factor corresponding to the sentiment tendency value at the current time, the second signal coefficient and the second decay factor corresponding to the above-mentioned like value, the third signal coefficient and the third decay factor corresponding to the page stay duration, the fourth signal coefficient and the fourth decay factor corresponding to the context switching frequency, and the fifth signal coefficient and the fifth decay factor corresponding to the number of answer modifications. Among them, each signal coefficient and each decay factor are periodically updated by the gradient descent method. They are initialized according to the importance of the feedback signal type. For example, for explicit like feedback, since it directly reflects the user's recognition of the answer, the signal coefficient can be set to a relatively high value. For implicit context switching frequency, the signal coefficient value is relatively low. The decay factor is optimized through training a large amount of historical data using the gradient descent algorithm to ensure that the time decay factor can accurately reflect the timeliness of user behavior. The signal coefficient can represent the importance degree of each corresponding feedback factor. The decay factor can represent the degree of change in the importance degree of the corresponding content as the corresponding feedback factor changes over time.

[0045] In the fifth step, generate the first dynamic weight corresponding to the sentiment tendency value according to the above-mentioned first signal coefficient, the above-mentioned first decay factor, and the feedback time corresponding to the comment feedback.

[0046] First, determine the time difference between the feedback time corresponding to the comment feedback and the current time. Then, multiply the time difference by the first decay factor to obtain a multiplication value. Next, use the opposite value of the multiplication value as the exponent and the numerical value e as the base to generate an exponential value. Finally, multiply the exponential value by the first signal coefficient to obtain the first dynamic weight.

[0047] Step 6: Generate a second dynamic weight corresponding to the above-mentioned like value according to the above-mentioned second signal coefficient, the above-mentioned second attenuation factor, and the feedback time corresponding to the like feedback. The generation method of the second dynamic weight can refer to the generation method of the first dynamic weight. Details are not described herein again.

[0048] Step 7: Generate a third dynamic weight corresponding to the above-mentioned page stay duration according to the above-mentioned third signal coefficient, the above-mentioned third attenuation factor, and the feedback time corresponding to the page stay duration. Details are not described herein again.

[0049] Step 8: Generate a fourth dynamic weight corresponding to the above-mentioned context switching frequency according to the above-mentioned fourth signal coefficient, the above-mentioned fourth attenuation factor, and the feedback time corresponding to the context switching frequency. Details are not described herein again.

[0050] Step 9: Generate a fifth dynamic weight corresponding to the above-mentioned number of answer modifications according to the above-mentioned fifth signal coefficient, the above-mentioned fifth attenuation factor, and the feedback time corresponding to the number of answer modifications. Details are not described herein again.

[0051] Step 10: Generate satisfaction information according to the above-mentioned sentiment tendency value, the above-mentioned like value, the target page stay duration included in the above-mentioned target page feedback, the target context switching frequency, the target number of answer modifications, the above-mentioned first dynamic weight, the above-mentioned second dynamic weight, the above-mentioned third dynamic weight, the above-mentioned fourth dynamic weight, and the fifth dynamic weight.

[0052] As an example, the above-mentioned execution entity may multiply the sentiment tendency value by the first dynamic weight to obtain a first multiplication value. Multiply the like value by the second dynamic weight to obtain a second multiplication value. Multiply the target page stay duration by the third dynamic weight to obtain a third multiplication value. Multiply the target context switching frequency by the fourth dynamic weight to obtain a fourth multiplication value. Multiply the target number of answer modifications by the fifth dynamic weight to obtain a fifth multiplication value. Finally, add the first multiplication value, the second multiplication value, the third multiplication value, the fourth multiplication value, and the fifth multiplication value to obtain the above-mentioned satisfaction information.

[0053] Step 104: Generate a first feedback data set according to the above-mentioned satisfaction information set and the normal reply information set.

[0054] In some embodiments, the above-mentioned execution entity may generate a first feedback data set according to the above-mentioned satisfaction information set and the normal reply information set. Each normal reply information has a corresponding first feedback data.

[0055] As an example, the above-mentioned execution entity can use the satisfaction information as labels, the normal reply information as training samples, and combine the labels and training samples one by one to generate first feedback data and obtain a first feedback data set.

[0056] Step 105: Screen out target reply information with multi-round replies in the corresponding reply content from the above-mentioned normal reply information set to obtain a target reply information set.

[0057] In some embodiments, the above-mentioned execution entity can screen out target reply information with multi-round replies in the corresponding reply content from the above-mentioned normal reply information set to obtain a target reply information set. Among them, multi-round replies can be the reply process in which the user conducts multi-round Q&A for a question (i.e., the reply question). For example, for the content output by the large language model that is not what the user needs, the user will continuously refine the question asked to let the large language model output. This process will result in multi-round replies.

[0058] As an example, the above-mentioned execution entity can screen out target reply information with multi-round replies in the corresponding reply content from the above-mentioned normal reply information set according to the number of questions asked and the number of reply contents to obtain a target reply information set.

[0059] Step 106: For each target reply information, execute the data generation step: Step 1061: Determine the initial reply information corresponding to the above-mentioned target reply information as the anchor sample, the reply content with a reply quality higher than the anchor sample in the multi-round replies as the positive sample, and the reply content with a reply quality lower than the anchor sample in the multi-round replies as the negative sample.

[0060] In some embodiments, the above-mentioned execution entity can determine the initial reply information corresponding to the above-mentioned target reply information as the anchor sample, the reply content with a reply quality higher than the anchor sample in the multi-round replies as the positive sample, and the reply content with a reply quality lower than the anchor sample in the multi-round replies as the negative sample. Among them, the initial reply information may include: the question first proposed by the user and the reply content first replied by the large language model. The reply quality can be the degree of compliance of the content output by the large language model with the content required by the user.

[0061] In some optional implementation manners of some embodiments, the above-mentioned execution entity can determine the initial reply information corresponding to the above-mentioned target reply information as the anchor sample, the reply content with a reply quality higher than the anchor sample in the multi-round replies as the positive sample, and the reply content with a reply quality lower than the anchor sample in the multi-round replies as the negative sample, including the following steps: Step 1: Extract the response content sequence and the Q&A question sequence for each response stage from the above target response information. Among them, the answer content in the answer content sequence has a corresponding response stage. The response content can be the response result of the large language model. The Q&A question can be the question input by the user on the model operation page in each response stage. There is a unique corresponding Q&A question for the response content.

[0062] Step 2: Determine the target page feedback corresponding to the above target response information.

[0063] Step 3: For each response content in the above response content sequence, extract the feedback sub - information corresponding to the above response content from the above target page feedback. Among them, there are interface feedbacks (i.e., feedback sub - information) for each response stage in multiple response stages in the target page feedback. For example, the multiple response stages include: the first response stage, the second response stage, and the third response stage. The target page feedback includes: the feedback sub - information corresponding to the first response stage (which can include explicit feedback and implicit feedback), the feedback sub - information corresponding to the second response stage, and the feedback sub - information corresponding to the third response stage.

[0064] Step 4: Extract the question semantic information of each Q&A question in the above Q&A question sequence and the question association semantic information between every two adjacent Q&A questions to obtain the question semantic information sequence and the question association semantic information sequence. Among them, the question semantic information can be information in vector form representing the semantic content of the question features corresponding to the Q&A question. The question association semantic information can be the association relationship between the question semantic information corresponding to two adjacent Q&A questions. The question association semantic information can also be information in vector form.

[0065] Step 5: Cross - fuse the above question semantic information sequence and question association semantic information sequence to obtain the semantic information sequence.

[0066] As an example, the above execution entity can cross - fuse the above question semantic information sequence and question association semantic information sequence in the order of rounds to obtain the semantic information sequence.

[0067] Step 6: Perform feature dimensionality reduction on the concatenated semantic information corresponding to the above semantic information sequence in an adaptive scale reduction format to obtain feature semantic dimensionality reduction information. Among them, the later the sequence position of the semantic information in the semantic information sequence, the higher the proportion of the feature content in the above feature semantic dimensionality reduction information. Among them, the adaptive scale reduction format can be that the scale reduction of the semantic content corresponding to the time distance closest to the current time in the concatenated semantic information is more delicate, that is, the features extracted from the semantic content corresponding to the time distance closest to the current time are more detailed and more important. These contents are also the key contents that the subsequent large language model needs to learn. The feature semantic dimensionality reduction information can be information representing the feature semantic content in the form of a vector with a preset dimension. The concatenated semantic information can be information in the form of a vector obtained by sequentially concatenating each semantic information in the semantic information sequence.

[0068] As an example, first, the above execution entity can use at least one convolution kernel of different scales to perform downsampling convolution processing on the concatenated semantic information to obtain at least one downsampling convolution information. Then, perform vector dimension complementation on each downsampling convolution information in the above at least one downsampling convolution information (for example, perform vector complementation according to the target value) to obtain at least one complemented vector with the same vector dimension. Finally, perform horizontal vector concatenation on the at least one complemented vector to obtain the feature semantic dimensionality reduction information.

[0069] Step 7: Use a generative model to generate a comprehensive question corresponding to the above feature semantic dimensionality reduction information. Among them, the generative model can be a deep learning model for generating summary questions. The comprehensive question can be a question without redundant content. That is, it is a comprehensive question obtained by summarizing the questions in each round. In practice, the generative model can be a part of the large language model or a deep learning model independent of the large language model. For a deep learning model independent of the large language model, the generative model can be a generative model based on Transformer.

[0070] As an example, the above execution entity can directly input the feature semantic dimensionality reduction information into the generative model to obtain a comprehensive question.

[0071] Step 8: Generate high-quality response content corresponding to the above comprehensive question. Among them, the high-quality response content can be response content with very high response accuracy.

[0072] As an example, first, the above execution entity can input the comprehensive question into at least one large language model of different model types to obtain at least one response content. Then, summarize the content of each response content in the at least one response content to obtain high-quality response content.

[0073] In the ninth step, generate the high-quality content semantic information corresponding to the above high-quality response content, and generate the response content semantic sequence corresponding to the response content sequence. Among them, the high-quality content semantic information can be information representing the content semantic information corresponding to the response content in vector form. There is a one-to-one correspondence between the response content semantics in the response content semantic sequence and the response content in the response content sequence. The response content semantics can be information representing the content semantic information in vector form. As an example, the above-mentioned execution entity can use the content semantic extraction model to implement the generation of high-quality content semantic information and the response content semantic sequence.

[0074] In the tenth step, determine the semantic similarity information between each response content semantic information in the response content semantic sequence and the above high-quality content semantic information, and obtain the semantic similarity information sequence. Among them, the semantic similarity information can represent the similarity of the semantic content between the response content semantic information and the high-quality content semantic information. The semantic similarity information can be a value between 0 and 1, and the higher the value, the more similar the semantic content between the two. In practice, the semantic similarity information can be the cosine similarity.

[0075] In the eleventh step, according to the above semantic similarity information sequence, use the initial response information corresponding to the above target response information as the anchor sample, the response content with a response quality higher than the anchor sample in multiple rounds of responses as the positive sample, and the response content with a response quality lower than the anchor sample in multiple rounds of responses as the negative sample.

[0076] As an example, first, screen out at least one first response content from the response content sequence whose corresponding semantic similarity information is higher than the semantic similarity of the anchor sample. Combine each first response content with the first label to generate a positive sample, and obtain at least one positive sample. Among them, the first label can represent that the response quality is higher than the response quality of the anchor sample. Then, screen out at least one second response content from the response content sequence whose corresponding semantic similarity information is lower than or equal to the semantic similarity of the anchor sample. Combine each second response content with the second label to generate a negative sample, and obtain at least one negative sample. Among them, the second label can represent that the response quality is not higher than the response quality of the anchor sample.

[0077] Optionally, the above-mentioned execution entity can perform feature dimensionality reduction on the spliced semantic information corresponding to the above semantic information sequence in an adaptive scale dimensionality reduction format to obtain the feature semantic dimensionality reduction information, including the following steps: Step 1: Obtain different semantic information partitioning methods. Among them, different semantic information partitioning methods correspond to different semantic information partitioning nodes. The semantic information partitioning method can be a method of partitioning semantic information by information interception. For example, the semantic information partitioning method can be a partitioning method according to the information length of "1:2:3", or a partitioning method according to rounds. The partitioning methods corresponding to each semantic information partitioning method are different. Different semantic information partitioning methods can be manually set on the relevant page.

[0078] Step 2: For each semantic information partitioning method, perform the following second generation steps: Sub-step 1: According to the above semantic information partitioning method, divide the above concatenated semantic information into a semantic information sequence.

[0079] Sub-step 2: Determine the downsampling convolutional block corresponding to each semantic information in the above semantic information sequence to obtain a downsampling convolutional block sequence. Among them, the downsampling convolutional block is a convolutional block used for downsampling processing. Among them, the matrix dimension corresponding to the downsampling convolutional block with a more backward sequence position is smaller. For example, the downsampling convolutional block sequence includes: the first downsampling convolutional block, the second downsampling convolutional block, and the third downsampling convolutional block. The dimension of the convolutional kernel corresponding to the first downsampling convolutional block is larger than the dimension of the convolutional kernel corresponding to the second downsampling convolutional block. The dimension of the convolutional kernel corresponding to the second downsampling convolutional block is larger than the dimension of the convolutional kernel corresponding to the third downsampling convolutional block. In practice, the dimensions of the convolutional kernels corresponding to each downsampling convolutional block can be set in an arithmetic progression.

[0080] Sub-step 3: For each semantic information, use the corresponding downsampling convolutional block to perform feature dimensionality reduction on the above semantic information to obtain semantic dimensionality reduction information.

[0081] Sub-step 4: Sequentially concatenate the semantic dimensionality reduction information sequences to obtain a semantic concatenated semantic information.

[0082] Step 3: Perform semantic fusion on the obtained semantic concatenated semantic information to obtain feature semantic dimensionality reduction information.

[0083] As an example, the above execution entity can supplement and concatenate the content of each semantic concatenated semantic information to obtain a semantic concatenated information. Input the semantic concatenated information into the fully connected layer to output the feature semantic dimensionality reduction information.

[0084] Step 1062: Generate the second feedback data of the above anchor samples, positive sample set, and negative sample set.

[0085] In some embodiments, the above execution entity can generate the second feedback data of the above anchor samples, positive sample set, and negative sample set.

[0086] As an example, for the anchor sample, there is a corresponding user input problem for each sample in the positive sample set and the negative sample set. The anchor sample, the positive sample set, the negative sample set, and the corresponding user input problems can be combined in the form of key-value pairs to obtain the second feedback data. For example, the second feedback data can be {anchor sample: problem 1, positive sample 1: problem 2, positive sample 2: problem 3, negative sample 1: problem 4, negative sample 2: problem 5}.

[0087] Step 107, according to the above first feedback data set and second feedback data set, perform model training on the above large language model to obtain a trained large language model.

[0088] In some embodiments, the above execution subject may perform model training on the above large language model according to the above first feedback data set and second feedback data set to obtain a trained large language model.

[0089] As an example, the above execution subject may use the first feedback data set and the second feedback data set as the training data set, and use a training method based on the gradient descent method to perform model training on the above large language model to obtain a trained large language model.

[0090] In some optional implementation manners of some embodiments, the above execution subject may perform model training on the above large language model according to the above first feedback data set and second feedback data set to obtain a trained large language model, including the following steps: First step, obtain the cluster label information set corresponding to the obtained abnormal reply information cluster set. Among them, there is a one-to-one correspondence between the abnormal reply information clusters in the abnormal reply information cluster set and the cluster label information in the cluster label information set. There is a cluster correspondence between the abnormal reply information clusters in the abnormal reply information cluster set and the reply information clusters in the reply information cluster set. That is, the abnormal reply information cluster can be the cluster after removing the normal reply information from the reply information cluster. The cluster label information can be the label information corresponding to the reply information cluster.

[0091] Second step, for each abnormal reply information, perform the first generation step: Sub-step 1, obtain the abnormal reply information cluster corresponding to the above abnormal reply information.

[0092] Sub-step 2: According to the cluster label information corresponding to the above abnormal reply information cluster, obtain the corresponding knowledge graph as the first knowledge graph. Among them, the knowledge graph can be a knowledge graph of general knowledge related to the cluster label information. In practice, the cluster label information includes multiple labels. The multiple labels can include: feature labels, scenario labels, and object labels. The corresponding knowledge graph can be a graph of basic knowledge related to the multiple labels. For example, for the cluster label information including: "boys", "campus", "basketball skills", the corresponding knowledge graph can be a graph related to "boys", "campus", and "basketball skills". In practice, the knowledge graph corresponding to the cluster label information can be the knowledge graph output by a related model. The main nodes in the knowledge graph are each label and the important associated content corresponding to the label, and the edges are the association relationships between the labels.

[0093] Sub-step 3: According to the key knowledge nodes and key knowledge edges in the above first knowledge graph, adjust the above abnormal reply information multiple times to obtain an adjusted abnormal reply information group. Among them, the key knowledge node can be the graph node corresponding to the label or the node with the largest number of edge connections. The key knowledge edge can be the edge corresponding to the key knowledge node.

[0094] In practice, according to the general knowledge and general knowledge connections corresponding to the key knowledge nodes and key knowledge edges, the content of the abnormal reply information can be adjusted for errors to obtain an adjusted abnormal reply information group. For example, if the key knowledge node and key knowledge edge are "boys often play basketball on campus", the content in the abnormal reply information can be changed to "boys cannot play basketball on campus" to obtain an adjusted abnormal reply information.

[0095] Third step: In response to the total number of information corresponding to the abnormal reply information set and the adjusted abnormal reply information group set reaching the target data volume, fuse the above abnormal reply information set and the above adjusted abnormal reply information group set to obtain an abnormal fusion reply information set. The target data volume can be a pre-set data volume.

[0096] Fourth step: For each abnormal fusion reply information, generate third feedback data according to the corresponding content validity, reply content type, and the above abnormal fusion reply information.

[0097] As an example, the above execution entity can generate a triple corresponding to the content validity, reply content type, and the above abnormal fusion reply information as the third feedback data.

[0098] Fifth step: According to the above first feedback data set, the above second feedback data set, and the third feedback data set, train the above large language model to obtain a trained large language model.

[0099] As an example, the above-mentioned execution entity can use the first feedback data set as the first training set, the second feedback data set as the second training set, and the third feedback data set as the third training set, and train the above-mentioned large language model by means of backpropagation to obtain the trained large language model.

[0100] In some optional implementation manners of some embodiments, after step 107, the steps further include: First step, in response to receiving the target question input by the target object on the model operation page, decompose the above-mentioned target question into multiple semantic units, and determine the corresponding reply information cluster of the above-mentioned target question as the target reply information cluster. Among them, the target object may be a user who queries the reply content corresponding to the target question. Among them, the multiple semantic units may be semantic units corresponding to sub-questions in the target question. The semantic unit may be information representing the semantic content of the sub-question in the form of a vector.

[0101] Second step, determine the corresponding second knowledge graph according to the cluster label information corresponding to the above-mentioned target reply information cluster. Details are not described herein again.

[0102] Third step, according to the above-mentioned second knowledge graph, determine the semantic knowledge concept and semantic knowledge relationship corresponding to each semantic unit in the above-mentioned multiple semantic units.

[0103] As an example, the above-mentioned execution entity can query the semantic knowledge concept and semantic knowledge relationship of multiple keywords corresponding to each semantic unit from the second knowledge graph by means of keyword query. The semantic knowledge relationship may be an association relationship between semantic knowledge concepts.

[0104] Fourth step, generate triples for the above-mentioned semantic knowledge concept, the above-mentioned semantic knowledge relationship, and the target question.

[0105] Fifth step, generate prompt information indicating generating reply information, an evidence chain, and a confidence level according to the triples. Among them, the evidence chain may be evidence information representing the basis for judging the correctness of the reply information in a chain form. The judgment basis may be the basis content for correct judgment logic. The confidence level may be a value between 0 and 1. The confidence level may represent the correct probability of each judgment basis.

[0106] Sixth step, input the above-mentioned prompt information into the above-mentioned trained large language model to obtain optimized reply information, the reply evidence chain corresponding to the above-mentioned optimized reply information, and the confidence levels of the respective basis contents corresponding to the above-mentioned reply evidence chain.

[0107] Seventh step, display the above-mentioned optimized reply information, the above-mentioned reply evidence chain, and each confidence level on the above-mentioned model operation page.

[0108] In some alternative implementations of some embodiments, the above-mentioned large language model includes: a first large language sub-model for explicit feedback information and a second large language sub-model for implicit feedback information. Among them, the first large language sub-model can be a large language model that replies to the reply question during the first reply of the reply content to improve the satisfaction of the reply content and enables the target user to give explicit feedback as much as possible. In practice, the first large language sub-model can be a temporal neural network model. The model input corresponding to the first large language sub-model is the reply question, and the outputs are the reply content and the satisfaction information. The satisfaction information can be the satisfaction of the reply content. For example, the first large language sub-model can be a language model based on the Transformer model. The second large language sub-model can be a language model that improves the accuracy of the reply content of the large language model according to the content feedback by the user during multiple rounds of replies. The training purpose of the second large language sub-model is to accurately understand the content feedback by the user in a multi-round dialogue scenario, and on the basis of understanding the user's question-asking habits corresponding to the target user, reduce the number of reply rounds to output accurate reply content. That is, the first large language sub-model is a large language model for outputting accurate reply content to improve user satisfaction in the first-round dialogue scenario. The second large language sub-model is a large language model that reduces the number of reply rounds for replies in a multi-round dialogue scenario. For example, the second large language model can also be a language model based on the Transformer model.

[0109] Optionally, the above-mentioned execution entity can train the above-mentioned large language model according to the above-mentioned first feedback data set and second feedback data set to obtain a trained large language model, including the following steps: First step, query the feedback data subset corresponding to the above-mentioned target object from the above-mentioned first feedback data set as the target feedback data subset. Among them, the target feedback data subset is a feedback data set generated based on the question-asking behavior corresponding to the target object. The target feedback data subset can be an empty data set. In the case of an empty data set, it can indicate that the target object uses the large language model to ask questions for the first time.

[0110] Second step, in response to determining that the data volume corresponding to the above-mentioned target feedback data subset does not reach the target data volume and receiving the target object's support for obtaining page information, obtain user operation behavior information from the operation page or operation application corresponding to the above-mentioned target object. Among them, the page information can be the information of a web page or the application operation information in an operation application. In practice, the user operation behavior information can be obtained by means of crawling or local acquisition. Among them, the user operation behavior information is the behavior trajectory information of the user's page operation behavior.

[0111] Step 3: Generate the above user feature information set according to the above user operation behavior information. Among them, the user feature information can represent the feature preference information of the target object under the target feature.

[0112] As an example, the above execution entity can determine the user feature information set by means of word frequency statistics.

[0113] Step 4: Crawl the question information set for which the above target object asks questions from the web page or related applications. For example, the web page can be a search page. The related applications can be shopping applications or short video applications.

[0114] Step 5: Generate a page key content information set corresponding to the page click information set according to the page click information set corresponding to the question information set. Among them, the page click information can be the page information of the target object's web page click or application click for the question information. The page key content information can be the main content of the clicked page. There is a one-to-one correspondence between the question information in the question information set and at least one page click information in the page click information set. That is, for the question information, multiple pages may be clicked.

[0115] As an example, for each page click information in the page click information set, the above execution entity can extract the page key content information from the page content according to the page content corresponding to the page click information and the above user feature information set. Among them, the page key content information can be the key content in the page related to the user feature information set.

[0116] As an example, the above execution entity can match the page local semantic information with the user feature information set to screen out at least one page local content closely related to at least one user feature information from the page content, and obtain at least one page local content. Then, summarize at least one page local content to obtain the page key content information.

[0117] Step 6: Generate a page reply information set for the page key content information set and the question information set. Among them, the page reply information includes: the question information and the corresponding at least one page key content information.

[0118] Step 7: For each page reply information, perform the following feedback data generation steps: Sub-step 1: Determine at least one page key content information and question information corresponding to the above page reply information, and use them as at least one target page key content information and target question information respectively.

[0119] Sub-step 2: Generate the problem feature semantic information corresponding to the above target problem information, and generate at least one page key content semantic information corresponding to at least one page key content information. Among them, the problem feature semantic information can be information in the form of a vector representing the problem semantic content corresponding to the target problem information. Each page key content information has a corresponding page key content semantic information. The page key content semantic information can be information in the form of a vector representing the page key content semantics.

[0120] Sub-step 3: Input the above at least one page key content semantic information into the same page content extraction layer based on the multi-layer residual layer to generate the page key content common semantic information. Among them, the page key content common semantic information can be the same semantic content corresponding to each page key content semantic information in the at least one page key content semantic information. The page key content common semantic information can be information in the form of a vector representing the common page content of at least one page key content information.

[0121] Sub-step 4: Input the above problem feature semantic information and the above page key content common semantic information into the generative layer based on the Transformer model to obtain the page reply content.

[0122] Sub-step 5: Generate the page feedback data corresponding to the above page reply information according to the above page reply content, the above page reply information, and the click information.

[0123] The eighth step: Combine the obtained page feedback data set and the target feedback data subset to obtain the initial explicit feedback data set corresponding to the above target object.

[0124] The ninth step: In response to determining that the data volume corresponding to the above target feedback data subset reaches the target data volume, the above execution entity can directly determine the target feedback data subset as the explicit feedback data set.

[0125] The tenth step: In response to determining that the data volume corresponding to the above initial explicit feedback data set reaches the target data volume, the above execution entity can directly determine the initial explicit feedback data set as the explicit feedback data set.

[0126] The eleventh step: In response to receiving that the target object does not support page information acquisition or the data volume corresponding to the initial explicit feedback data set does not reach the target data volume, determine the object group corresponding to the above target object according to the user feature information set. Among them, the characteristics corresponding to the object group are similar to the characteristics corresponding to the target object.

[0127] In the twelfth step, combine the pre-stored feedback data set corresponding to the above object group, the target feedback data subset, or the initial explicit feedback data set at the data end to obtain a combined feedback data set as the explicit feedback data set. The feedback data set corresponding to the object group can be a data set under explicit feedback.

[0128] In the thirteenth step, set the first data set weight corresponding to the above initial explicit feedback data set, the second data set weight corresponding to the feedback data set, and the third data set weight corresponding to the target feedback data subset. The third data set weight is higher than the first data set weight. The first data set weight is higher than the second data set weight. For example, the third data set weight is 1. The first data set weight is 0.7. The second data set weight is 0.4.

[0129] In the fourteenth step, according to the above explicit feedback data set, the corresponding data set weights, and the second feedback data set, use the backpropagation method to train the above first large language sub-model and second large language sub-model to obtain a trained large language model.

[0130] As an example, the above execution entity can use the backpropagation method to train the above first large language sub-model according to the above explicit feedback data set and the corresponding data set weights to obtain a trained first large language sub-model. Then, for each feedback data in the second feedback data set, use the trained first large language model to generate a response content set corresponding to the problem set of each feedback data. Use the second feedback data set as the input data and the response content set as the output label to train the second large language sub-model to obtain a trained second large language sub-model. Finally, replace the first large language sub-model and the second large language sub-model in the large language model with the trained first large language sub-model and the trained second large language sub-model respectively to obtain a trained large language model.

[0131] It should be noted that the above "first step - fourteenth step", as one of the inventive points of the present disclosure, solves at least one other technical problem: "The large language model cannot achieve accurate first-round responses to answering questions, and during the multi-round answering process, there are often situations where there are ambiguities in understanding questions, resulting in the object asking multiple questions and the large language model answering multiple times. This way not only wastes a lot of computing resources but also leads to a poor user experience", and also solves the problem that "the large language model cannot output customized content according to the preferences of the target object, resulting in the output content of the large language model not being what the target object wants, and the situation of content output ambiguity occurs". Based on this, in the present disclosure, first, a dataset corresponding to the target object is extracted from the first feedback dataset. Then, when it is determined that the data volume of the dataset is small and the subsequent first large language sub-model cannot fully learn the relevant feature information of the response preferences corresponding to the target object, with the permission of the target object, the explicit response information is filled through the historical page browsing record and the application browsing record, so as to obtain more initial explicit feedback datasets matching the above target object. When the initial explicit feedback dataset has not reached the target data volume, the target object corresponding object group is determined according to the user feature information set to further supplement the explicit response information. Here, through the same page content extraction layer and the generative layer based on the Transformer model, accurate generation of page response content can be achieved. In addition, by setting the weights of each dataset, the importance of the data corresponding to the true and false datasets can be judged, so that the subsequent first large language sub-model can achieve more accurate model training. Based on this, through the above explicit feedback dataset, the corresponding weights of each dataset, and the second feedback dataset, the large language model can be accurately trained by using the backpropagation method.

[0132] In some optional implementation manners of some embodiments, the above execution subject can use the backpropagation method to perform model training on the above first large language sub-model and the second large language sub-model according to the above explicit feedback dataset, the corresponding weights of each dataset, and the second feedback dataset, to obtain the trained large language model, including the following steps: First step, according to the above explicit feedback dataset and the weights of each dataset, use the backpropagation training method to perform model training on the first large language sub-model to obtain the trained first large language sub-model. Here, during the training process of the first large language sub-model, for the datasets corresponding to the weights of each dataset, selective learning of feature semantic information is performed according to the size of the dataset weights (that is, the higher the dataset weight, the more the first large language sub-model focuses on learning. That is, the lower the dataset weight, the more superficially the first large language sub-model learns).

[0133] Step 2: Extract the second feedback data from the second feedback dataset as the extracted feedback data.

[0134] Step 3: For the extracted feedback data, perform the following training steps: Sub-step 1: Determine the feedback question information set under multiple rounds of responses corresponding to the above-mentioned extracted feedback data.

[0135] Sub-step 2: For each round of response in the multiple rounds of responses, perform the following question set steps: The first sub-step: Determine the feedback question information corresponding to this round of response and the historical feedback question information set before this round of response. Among them, the feedback question information can be the questions raised by the target object during this round of response. The historical feedback question information set can be each question set raised by the target object during the response process before this round of response.

[0136] The second sub-step: Combine the feedback question information set and the historical feedback question information set to obtain the combined question information.

[0137] Sub-step 3: Use the above-trained first large language sub-model to generate the response content set and the satisfaction information set corresponding to the obtained comprehensive question information set, as the feedback response content set and the feedback satisfaction information set respectively. Among them, each feedback response content has a corresponding feedback satisfaction information. Each comprehensive question information has a corresponding feedback response content and feedback satisfaction information.

[0138] Sub-step 4: For each round of response in the multiple rounds of responses, combine the response content corresponding to the above-mentioned extracted feedback data corresponding to this round of response with the corresponding feedback response content to obtain a content group.

[0139] Sub-step 5: Generate multiple content loss differences corresponding to the multiple rounds of responses according to the obtained multiple content groups. Among them, the content loss difference can be used to represent the difference between the feedback response content in the content group and the response content corresponding to the second feedback data. Each content group has a corresponding content loss difference. The higher the content loss difference, the smaller the difference between the corresponding feedback response content and the response content corresponding to the second feedback data.

[0140] As an example, first, generate the content feature semantic information corresponding to each content in the content group to obtain a content feature semantic information group. Then, input the above content feature semantic information group into the cross-entropy loss function to obtain the content loss difference.

[0141] Sub-step 6: For the positive sample set, negative sample set, and anchor sample corresponding to the second feedback data, randomly combine the negative samples, positive samples, and anchor samples to obtain a contrastive learning sample set, where each contrastive learning sample includes: positive sample, negative sample, and anchor sample.

[0142] Sub-step 7: Using a contrastive loss function, generate the contrastive loss information corresponding to each contrastive learning sample in the contrastive learning sample set, obtaining a contrastive loss information set. For example, the contrastive loss function can be Triplet loss (triplet loss function).

[0143] Sub-step 8: Perform weighted summation processing on the contrastive loss information set and multiple content loss differences to obtain weighted loss information.

[0144] Sub-step 9: In response to determining that the loss information sequence corresponding to the weighted loss information tends to the target loss value, determine the second largest language sub-model as the trained second largest language sub-model.

[0145] Sub-step 10: Replace the first largest language sub-model and the second largest language sub-model in the large language model with the trained first largest language sub-model and the trained second largest language sub-model respectively.

[0146] Fourth step: In response to determining that the loss information sequence corresponding to the weighted loss information does not tend to the target loss value, based on the weighted loss information, use the training method of backpropagation to train the second largest language sub-model, obtaining an initial second largest language sub-model.

[0147] Fifth step: Use the initial second largest language sub-model as the second largest language model, and the feedback data extracted again as the extracted feedback data, and continue to execute the above training steps.

[0148] It should be noted that "the above first step - fifth step" is another inventive point of the present disclosure, which solves another technical problem of "how to achieve precise training of the first largest language sub-model and the second largest language sub-model to ensure that the trained second largest language sub-model can quickly reply with precise reply content in the case of fewer reply rounds". Based on this, in the present disclosure, first, on the basis of training through a large number of the above-mentioned explicit feedback data sets for the target object and the corresponding weights of each data set, precise reply to the problem of a single round can be achieved. On this basis, the first largest language sub-model may not be good at precisely replying to the reply content based on the adjustment of the problem of the target object during the multi-round reply process. On this basis, for each round of reply, through the first largest language model, the generation of the reply content label in a single round is realized to generate the content loss information for each round of reply. In addition, through the setting of positive samples, negative samples and anchor samples, the construction of contrastive learning samples can be realized, so that the second largest language sub-model can quickly master the quality differences of different reply contents between rounds, and through the contrastive loss information, the second largest language sub-model can achieve precise reply in the case of multi-rounds. Thus, the large language model can be accurately obtained.

[0149] In some alternative implementations of some embodiments, the large language model further includes: a response evidence chain and confidence generation model. The response evidence chain and confidence generation model can also be a temporal neural network model based on the Transformer model. The response evidence chain and confidence generation model can generate the response evidence chain and confidence based on prompt words.

[0150] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the large language model optimization method based on multimodal feedback and reinforcement learning in some embodiments of the present disclosure, through the multimodal information of page feedback and the performance of the large language model in multiple rounds of output, the large language model can be efficiently trained to obtain a large language model with more accurate output. Specifically, the reason for the inaccurate output of the relevant large language model is that: the large language model trained conventionally cannot comprehensively capture the complex attitudes and diverse needs of users towards the model output, resulting in the reply output by the large language model not matching the content of the actual needs of users, the accuracy of the large language model being poor and the user experience being poor. Based on this, in some embodiments of the present disclosure, the large language model optimization method based on multimodal feedback and reinforcement learning first obtains the reply information set output by the large language model and the page feedback set collected by the device. Here, by obtaining the reply information set, it serves as the data basis for subsequent model training of the large language model. The page feedback set can reflect the effective feedback of users on the reply content during the reply process. By collecting the page feedback set, the subsequent large language model can comprehensively consider the page feedback situation and achieve precise training of the corresponding model. Then, using the target clustering algorithm, the abnormal reply information in the above-mentioned reply information set can be efficiently and accurately removed to obtain a normal reply information set. Next, according to the page feedback subset corresponding to the above-mentioned normal reply information set, through various situations of page feedback, the satisfaction information corresponding to each normal reply information can be accurately determined to obtain a satisfaction information set. Through the satisfaction information, the subsequent large language model can know whether the reply content output during training meets the needs of users, so as to achieve precise training of the large language model. Then, according to the above-mentioned satisfaction information set and normal reply information set, an accurate first feedback data set can be generated for subsequent training of the large language model to generate a model with more accurate output. Further, target reply information with multiple rounds of replies in the corresponding reply content is selected from the above-mentioned normal reply information set to obtain a target reply information set, so as to facilitate the key screening of data for the performance of the large language model in multiple rounds of replies, so that the subsequent large language model can reduce the number of rounds and improve the output accuracy during multiple rounds of replies. Furthermore, for each target reply information, a data generation step is executed: First step, the initial reply information corresponding to the above-mentioned target reply information is determined as the anchor sample, the reply content with a reply quality higher than the anchor sample in multiple rounds of replies is determined as the positive sample, and the reply content with a reply quality lower than the anchor sample in multiple rounds of replies is determined as the negative sample, so as to facilitate the subsequent large language model to learn the reply situation in multiple rounds of replies. Second step, generate the second feedback data of the above-mentioned anchor sample, positive sample set and negative sample set as training samples, so that the subsequent large language model can improve the output accuracy under the condition of reducing the number of rounds of replies and continuously learn. Finally, according to the above-mentioned first feedback data set and second feedback data set, the above-mentioned large language model is trained to obtain a trained large language model.In summary, by considering the multimodal feature information in terms of page feedback and the output accuracy in terms of multi-round responses, the large language model can be continuously improved in response to the feedback content and multi-round response situations, thereby enhancing the output accuracy of the large language model.

[0151] For further reference Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an optimization device for a large language model based on multimodal feedback and reinforcement learning. These device embodiments correspond to Figure 1 the method embodiments shown, and the optimization device for the large language model based on multimodal feedback and reinforcement learning can be specifically applied to various electronic devices.

[0152] As shown in Figure 2 , an optimization device 200 for a large language model based on multimodal feedback and reinforcement learning includes: an acquisition unit 201, a removal unit 202, a determination unit 203, a generation unit 204, a screening unit 205, an execution unit 206, and a training unit 207. Among them, the acquisition unit 201 is configured to acquire a response information set output by the large language model and a page feedback set collected by the device; the removal unit 202 is configured to use a target clustering algorithm to remove abnormal response information in the above response information set to obtain a normal response information set; the determination unit 203 is configured to determine satisfaction information corresponding to each normal response information according to a page feedback subset corresponding to the above normal response information set to obtain a satisfaction information set; the generation unit 204 is configured to generate a first feedback data set according to the above satisfaction information set and normal response information set; the screening unit 205 is configured to screen out target response information with multi-round responses in the corresponding response content from the above normal response information set to obtain a target response information set; the execution unit 206 is configured to, for each target response information, execute a data generation step: determine the initial response information corresponding to the above target response information as an anchor sample, determine response content with a response quality higher than the anchor sample in multi-round responses as positive samples, and determine response content with a response quality lower than the anchor sample in multi-round responses as negative samples; generate second feedback data of the above anchor sample, positive sample set, and negative sample set; the training unit 207 is configured to perform model training on the above large language model according to the above first feedback data set and second feedback data set to obtain a trained large language model.

[0153] It can be understood that the units described in the optimization device 200 for the large language model based on multimodal feedback and reinforcement learning correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the optimization device 200 for the large language model based on multimodal feedback and reinforcement learning and the units included therein, and will not be elaborated herein.

[0154] Reference is made below to Figure 3 , which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for use in implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0155] As Figure 3 shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0156] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 3 Each block shown in

[0157] may represent a device or, as needed, multiple devices.

[0158] It should be noted that, in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0159] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0160] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a set of response information output by a large language model and a set of page feedback collected by the device; use a target clustering algorithm to remove abnormal response information from the set of response information to obtain a set of normal response information; determine satisfaction information corresponding to each normal response information according to a page feedback subset corresponding to the set of normal response information to obtain a set of satisfaction information; generate a first feedback data set according to the set of satisfaction information and the set of normal response information; screen out target response information with multi-round responses in the corresponding response content from the set of normal response information to obtain a set of target response information; for each target response information, execute a data generation step: determine the initial response information corresponding to the target response information as an anchor sample, determine response content with a response quality higher than the anchor sample in the multi-round responses as positive samples, and determine response content with a response quality lower than the anchor sample in the multi-round responses as negative samples; generate second feedback data for the anchor sample, the set of positive samples, and the set of negative samples; and perform model training on the large language model according to the first feedback data set and the second feedback data set to obtain a trained large language model.

[0161] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider using the Internet).

[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0163] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a removal unit, a determination unit, a generation unit, a screening unit, an execution unit, and a training unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit that acquires the response information set output by the large language model and the page feedback set collected by the device".

[0164] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0165] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An optimization method for large language models based on multimodal feedback and reinforcement learning, characterized in that, Including: Obtaining a set of reply information output by a large language model and a set of page feedbacks collected by a device; Using a target clustering algorithm to remove abnormal reply information from the set of reply information, obtaining a set of normal reply information; Determining, according to a page feedback subset corresponding to the set of normal reply information, satisfaction information corresponding to each normal reply information, obtaining a set of satisfaction information; Generating a first feedback data set according to the set of satisfaction information and the set of normal reply information; Screening out target reply information with multiple rounds of replies in the corresponding reply content from the set of normal reply information, obtaining a set of target reply information; For each target reply information, perform a data generation step: Determining the initial reply information corresponding to the target reply information as an anchor sample, reply content with a reply quality higher than the anchor sample in multiple rounds of replies as positive samples, and reply content with a reply quality lower than the anchor sample in multiple rounds of replies as negative samples; Generating second feedback data for the anchor sample, the set of positive samples, and the set of negative samples; Performing model training on the large language model according to the first feedback data set and the second feedback data set, obtaining a trained large language model.

2. The method according to claim 1, wherein The step of using a target clustering algorithm to remove abnormal reply information from the set of reply information, obtaining a set of normal reply information, includes: Using the target clustering algorithm to perform clustering processing on the set of reply information, obtaining a set of reply information clusters; Determining the between-cluster variance between each reply information cluster in the set of reply information clusters and the within-cluster variance corresponding to each reply information cluster; Performing adaptive adjustment on each reply information cluster according to the between-cluster variance and the obtained set of within-cluster variances, obtaining an adjusted set of reply information clusters; Screening out an initial set of abnormal reply information from the set of reply information according to the set of reply information clusters and the adjusted set of reply information clusters; For each initial abnormal reply information, perform the following determination steps: Obtaining the target reply content corresponding to the initial abnormal reply information; Determining the reply content type and content validity corresponding to the target reply content; In response to determining that the content validity does not meet the target validity requirement and / or the reply content type belongs to a predetermined set of content types, determining the initial abnormal reply information as abnormal reply information; Removing the obtained set of abnormal reply information from the set of reply information, obtaining a set of normal reply information.

3. The method according to claim 2, wherein The step of performing model training on the large language model according to the first feedback data set and the second feedback data set, obtaining a trained large language model, includes: Obtaining a set of cluster label information corresponding to the obtained set of abnormal reply information clusters; For each abnormal reply information, perform a first generation step: Obtaining the abnormal reply information cluster corresponding to the abnormal reply information; Obtaining a corresponding knowledge graph according to the cluster label information corresponding to the abnormal reply information cluster, as a first knowledge graph; Performing multiple adjustments on the abnormal reply information according to the key knowledge nodes and key knowledge edges in the first knowledge graph, obtaining a set of adjusted abnormal reply information; In response to the abnormal reply information set and the total number of information corresponding to the adjusted abnormal reply information group set reaching the target data volume, fuse the abnormal reply information set and the adjusted abnormal reply information group set to obtain an abnormal fusion reply information set; For each abnormal fusion reply information, generate third feedback data according to the corresponding content validity, reply content type, and the abnormal fusion reply information; Perform model training on the large language model according to the first feedback data set, the second feedback data set, and the third feedback data set to obtain a trained large language model.

4. The method according to claim 1, characterized in that, The page feedback in the page feedback set includes at least one of the following: an explicit feedback information set after an operation on the model operation page, an implicit feedback information set after an operation on the model operation page. The explicit feedback information includes at least one of the following: like feedback, comment feedback. The implicit feedback information includes at least one of the following: page dwell time, context switching frequency, answer modification times; And Determining the satisfaction information corresponding to each normal reply information according to the page feedback subset corresponding to the normal reply information set includes: Determine the target page feedback corresponding to the normal reply information; Determine the comment sentiment tendency information according to the comment feedback included in the target page feedback; Quantify the comment sentiment tendency information and the corresponding like feedback to obtain a sentiment tendency value and a like value; Obtain the first signal coefficient and the first attenuation factor corresponding to the sentiment tendency value, the second signal coefficient and the second attenuation factor corresponding to the like value, the third signal coefficient and the third attenuation factor corresponding to the page dwell time, the fourth signal coefficient and the fourth attenuation factor corresponding to the context switching frequency, and the fifth signal coefficient and the fifth attenuation factor corresponding to the answer modification times at the current time. Among them, each signal coefficient and each attenuation factor are periodically updated by the gradient descent method; Generate a first dynamic weight corresponding to the sentiment tendency value according to the first signal coefficient, the first attenuation factor, and the feedback time corresponding to the comment feedback; Generate a second dynamic weight corresponding to the like value according to the second signal coefficient, the second attenuation factor, and the feedback time corresponding to the like feedback; Generate a third dynamic weight corresponding to the page dwell time according to the third signal coefficient, the third attenuation factor, and the feedback time corresponding to the page dwell time; Generate a fourth dynamic weight corresponding to the context switching frequency according to the fourth signal coefficient, the fourth attenuation factor, and the feedback time corresponding to the context switching frequency; Generate a fifth dynamic weight corresponding to the answer modification times according to the fifth signal coefficient, the fifth attenuation factor, and the feedback time corresponding to the answer modification times; Generate satisfaction information according to the sentiment tendency value, the like value, the target page dwell time included in the target page feedback, the target context switching frequency, the target answer modification times, the first dynamic weight, the second dynamic weight, the third dynamic weight, the fourth dynamic weight, and the fifth dynamic weight.

5. The method according to claim 1, characterized in that, Determining the initial reply information corresponding to the target reply information as an anchor sample, and determining the reply content with a reply quality higher than that of the anchor sample in multiple rounds of replies as a positive sample, and determining the reply content with a reply quality lower than that of the anchor sample in multiple rounds of replies as a negative sample, includes: Extracting the reply content sequence and the question and answer question sequence in each round of reply stage from the target reply information; Determining the target page feedback corresponding to the target reply information; For each reply content in the reply content sequence, extracting the feedback sub-information corresponding to the reply content from the target page feedback; Extracting the question semantic information of each question and answer question in the question and answer question sequence and the question association semantic information between every two adjacent question and answer questions, to obtain a question semantic information sequence and a question association semantic information sequence; Performing cross-fusion on the question semantic information sequence and the question association semantic information sequence to obtain a semantic information sequence; Performing feature reduction on the concatenated semantic information corresponding to the semantic information sequence in an adaptive scale reduction format to obtain feature semantic reduction information, where the later the sequence position of the semantic information in the semantic information sequence, the higher the proportion of the feature content in the feature semantic reduction information; Using a generative model to generate a comprehensive question corresponding to the feature semantic reduction information; Generating a high-quality reply content corresponding to the comprehensive question; Generating the high-quality content semantic information corresponding to the high-quality reply content, and generating a reply content semantic sequence corresponding to the reply content sequence; Determining the semantic similarity information between each reply content semantic information in the reply content semantic sequence and the high-quality content semantic information, to obtain a semantic similarity information sequence; According to the semantic similarity information sequence, using the initial reply information corresponding to the target reply information as an anchor sample, the reply content with a reply quality higher than that of the anchor sample in multiple rounds of replies as a positive sample, and the reply content with a reply quality lower than that of the anchor sample in multiple rounds of replies as a negative sample.

6. The method according to claim 5, wherein The performing feature reduction on the concatenated semantic information corresponding to the semantic information sequence in an adaptive scale reduction format to obtain feature semantic reduction information includes: Obtaining different semantic information division methods, where different semantic information division methods correspond to different semantic information division nodes; For each semantic information division method, performing the following second generation steps: Dividing the concatenated semantic information into a semantic information sequence according to the semantic information division method; Determining the downsampling convolutional block corresponding to each semantic information in the semantic information sequence, to obtain a downsampling convolutional block sequence, where the later the sequence position of the downsampling convolutional block, the smaller the matrix dimension corresponding to the downsampling convolutional block; For each semantic information, using the corresponding downsampling convolutional block to perform feature reduction on the semantic information to obtain semantic reduction information; Sequentially splicing the semantic reduction information sequences to obtain semantic concatenated semantic information; Performing semantic fusion on the obtained semantic concatenated semantic information to obtain feature semantic reduction information.

7. The method according to claim 1, characterized in that The method further includes: In response to receiving the target question input by the target object, decomposing the target question into multiple semantic units, and determining the response information cluster corresponding to the target question as the target response information cluster; Determining a corresponding second knowledge graph according to the cluster label information corresponding to the target response information cluster; Determining, according to the second knowledge graph, the semantic knowledge concept and semantic knowledge relationship corresponding to each semantic unit in the multiple semantic units; Generating triples for the semantic knowledge concept, the semantic knowledge relationship, and the target question; Generating a prompt message indicating that response information, an evidence chain, and a confidence level are generated according to the triples; Inputting the prompt message into the trained large language model to obtain optimized response information, the response evidence chain corresponding to the optimized response information, and the confidence levels of the respective basis contents corresponding to the response evidence chain; Displaying the optimized response information, the response evidence chain, and the respective confidence levels on the model operation page.

8. An optimization device for large language models based on multi-modal feedback and reinforcement learning, characterized in that, Including: An acquisition unit configured to acquire a response information set output by the large language model and a page feedback set collected by the device; A removal unit configured to use a target clustering algorithm to remove abnormal response information in the response information set to obtain a normal response information set; A determination unit configured to determine the satisfaction information corresponding to each normal response information according to the page feedback subset corresponding to the normal response information set to obtain a satisfaction information set; A generation unit configured to generate a first feedback data set according to the satisfaction information set and the normal response information set; A screening unit configured to screen out target response information with multi-round responses in the corresponding response content from the normal response information set to obtain a target response information set; An execution unit configured to, for each target response information, execute a data generation step: determining the initial response information corresponding to the target response information as an anchor sample, determining the response content with a response quality higher than the anchor sample in the multi-round responses as a positive sample, and determining the response content with a response quality lower than the anchor sample in the multi-round responses as a negative sample; Generating second feedback data for the anchor sample, the positive sample set, and the negative sample set; A training unit configured to perform model training on the large language model according to the first feedback data set and the second feedback data set to obtain a trained large language model.

9. An electronic device, characterized in that, Including: One or more processors; A storage device on which one or more programs are stored, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, The program, when executed by the processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-round dialogue method and system based on machine learning

    CN112836028A

  • System for providing personalized reply for user based on multi-modal large model

    CN119337995A

  • Generative search data processing method and device, equipment and storage medium

    CN119514612A

  • Domain question and answer large model training and question and answer method, related equipment and program product

    CN119961422A

  • System and method for dynamic online search result generation

    US20190205761A1

Cited By

  • Code generation task reply method and device, medium and electronic equipment

    CN120994173A