Large language model parameter tuning method, evaluation method and device
By tuning the parameters of the large language model, using the business results and user interest information in the training sample, the problem of inefficiency in the existing large language model when evaluating the business results in the generative task is solved, and more efficient and stable evaluation performance is achieved.
Patent Information
- Application Number
- CN202510436569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing large language models are difficult to effectively evaluate business results in generative tasks, especially when business updates and iterations are fast.
Through the large language model parameter tuning method, the parameters of the first largest language model are adjusted using the business results, labeled results and user interest tendencies in the training sample to improve its evaluation performance.
It improves the evaluation performance of the large language model, makes it more adaptable to business changes, and improves the accuracy and stability of the evaluation results.
Smart Images

Figure CN119938964A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular, to a large language model parameter tuning method and device, and a large language model evaluation method and device. Background Art
[0002] In the field of artificial intelligence, large models usually refer to machine learning models with a large number of parameters (such as tens of billions or hundreds of billions of parameters) and complex structures. Through training on massive amounts of data, these models can complete a variety of complex tasks. At present, many businesses are developed by integrating large model technology, especially large language models (LLM) and multimodal technology to meet various project requirements, such as generative tasks that generate corresponding business content based on the subject in the query image. For each business, large models can also be used to evaluate whether the business has achieved the expected results.
[0003] However, the business results generated by such generative tasks are highly uncertain and complex, and business updates and iterations are very fast. The large models currently used for evaluation are difficult to effectively evaluate such businesses. Summary of the invention
[0004] The embodiments of this specification provide a large language model parameter tuning and large language model evaluation solution, which can improve the evaluation performance of the large language model through parameter tuning.
[0005] In a first aspect, an embodiment of the present specification provides a large language model parameter tuning method, wherein a first large language model is used to evaluate the execution performance of a target business, and the target business is used to generate a business result for a subject in a user's query image based on a second large language model; the method comprises: obtaining a training sample, the training sample comprising the business result, a labeling result and a first information indicating the user's interest tendency, the labeling result being used to label an evaluation label of the business result; the first large language model performs an evaluation based on the business result to obtain an evaluation result for the training sample, the evaluation result and the labeling result comprising a correlation between the business result and the first information; and adjusting the parameters of the first large language model based on the evaluation result and the labeling result.
[0006] In some embodiments, the training samples include query images.
[0007] In some embodiments, the evaluation results and the annotation results also include at least one of the following evaluation indicators: the correlation between the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; and the information content of the business result.
[0008] In some embodiments, adjusting the parameters of the first large language model based on the evaluation result and the annotation result includes: determining whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; when it is determined that the training sample is an offset sample, upsampling the offset sample to obtain an amplified offset sample; and adjusting the parameters of the first large language model based on the amplified offset sample.
[0009] In some embodiments, the method further includes: determining whether the training sample is a shifted sample based on the difference between the evaluation result and the annotation result; when it is determined that the training sample is a shifted sample, based on the subjectivity of the shifted sample and each of the evaluation indicators, adjusting the prompt words of the first large language model, and performing evaluation again based on the business results to obtain a second evaluation result for the training sample; and adjusting the parameters of the first large language model based on the second evaluation result and the annotation result.
[0010] In some embodiments, determining whether the training sample is a shifted sample based on the difference between the evaluation result and the annotation result includes: in response to the consistency between the evaluation indicator in the evaluation result and the evaluation indicator in the annotation result being lower than a preset threshold, determining the training sample as a shifted sample.
[0011] In some embodiments, obtaining a training sample includes: obtaining a first business result and a first annotation result generated by the target business for a first candidate image, the first business result including location information of at least one subject obtained based on recognition of the first candidate image, and the first annotation result being used to annotate a location tag of the at least one subject; screening the first candidate image based on the location information in the first business result and the location tag in the first annotation result to determine a first image; obtaining the business result and annotation result generated by the target business for the first image as a training sample.
[0012] In some embodiments, the obtaining of training samples includes: obtaining the category of at least one subject in the second candidate image; screening the second candidate image according to the first category requirements to determine the second image; and obtaining the business results and annotation results generated by the target business for the second image as training samples.
[0013] In some embodiments, the method further includes: obtaining actual business data when the target business is running, the actual business data including at least one of the following data: a query image, a first information indicating the user's interest tendency; determining sample offset data based on the difference between the actual business data and the training sample; and constructing a new training sample based on the sample offset data.
[0014] In a second aspect, an embodiment of the present specification provides a large language model evaluation method, the method comprising: obtaining an execution result of a target business, the execution result being generated by the target business based on a second large language model for a subject in a user's query image; evaluating the execution result by a first large language model to obtain evaluation data for the execution result, wherein the first large language model is a large language model obtained after parameter tuning based on any of the large language model parameter tuning methods described in the first aspect.
[0015] In some embodiments, the method further includes: obtaining an intermediate execution result and category annotation data of the target business, the intermediate execution result being the subject category in the query image obtained by identifying the target business, and the category annotation data being used to label the category label of the subject of the query image; and using a third language model to evaluate the intermediate execution result according to the category annotation data to obtain the accuracy of the intermediate execution result.
[0016] In some embodiments, obtaining the intermediate execution results and category annotation data of the target business includes: obtaining an evaluation benchmark data set, wherein the evaluation benchmark data set contains query images of multiple users and category annotation data corresponding to the query images; and generating, by the target business, an intermediate execution result for the subject in the query image in the evaluation benchmark data set.
[0017] In some embodiments, obtaining an evaluation benchmark data set includes: obtaining historical business data, the historical business data including a user's query image; filtering the historical business data according to data filtering requirements to obtain target business data; and annotating the query image in the target business data to obtain category annotation data.
[0018] In some embodiments, filtering the historical business data according to data screening requirements to obtain target business data includes: obtaining the main category of the query image in the historical business data; and filtering the query image in the historical business data according to a second category requirement to obtain the target business data.
[0019] In some embodiments, obtaining the evaluation benchmark data set includes: obtaining business data from a third-party platform, the business data including query images in actual user usage scenarios; and annotating the query images in the business data to obtain category annotation data.
[0020] In some embodiments, the method further includes: determining whether the query image corresponding to the intermediate execution result is an offset image based on the accuracy of the intermediate execution result; and if it is determined that the query image is an offset image, constructing a new evaluation benchmark data set based on the offset image.
[0021] In a third aspect, an embodiment of the present specification provides a large language model parameter tuning device, wherein a first large language model is used to evaluate the execution performance of a target business, and the target business is used to generate a business result for a subject in a user's query image based on a second large language model; the device comprises: a data acquisition module, configured to acquire training samples, the training samples comprising the business results, annotation results and first information indicating the user's interest tendency, the annotation results being used to annotate an evaluation label of the business results; a business evaluation module, configured to perform an evaluation based on the business results by the first large language model to obtain an evaluation result for the training sample, the evaluation result and the annotation result comprising the correlation between the business result and the first information; a parameter adjustment module, configured to adjust the parameters of the first large language model based on the evaluation result and the annotation result.
[0022] In a fourth aspect, an embodiment of the present specification provides a large language model evaluation device, the device comprising: a result acquisition module, configured to obtain an execution result of a target business, the execution result being generated by the target business based on a second large language model for a subject in a user's query image; a performance evaluation module, configured to evaluate the execution result using a first large language model to obtain evaluation data for the execution result, the first large language model being a large language model obtained after parameter tuning based on any of the large language model parameter tuning methods described in the first aspect.
[0023] In a fifth aspect, an embodiment of the present specification provides a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the implementation modes of the first aspect and the second aspect is implemented.
[0024] In the scheme provided in the above-mentioned embodiments of the present specification, an evaluation of image-based content generation tasks can be implemented, wherein the business results of the training samples are evaluated to obtain the evaluation results, and then the parameters of the first language model are tuned according to the evaluation results and the annotation results, so as to help the first language model adapt to different business changes, and then the generation performance of the second language model under different conditions can be tested, thereby improving the evaluation effect of the first language model; in addition, the correlation between the business results and the first information used to indicate the user's interest tendency is also taken into account during the evaluation, so as to ensure that the evaluation results are consistent with the user's subjective feelings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 It is a schematic diagram of a large language model parameter tuning process in an embodiment of this specification; Figure 2 It is a schematic diagram of a process of implementing a target service in an embodiment of this specification; Figure 3 It is a business result of a target business in the embodiment of this specification; Figure 4 is a flow chart of a large language model parameter tuning method in an embodiment of this specification; Figure 5 is a schematic diagram of the degree of consistency in the embodiments of this specification; Figure 6 is a schematic diagram of an acquisition and processing path of an offset sample in an embodiment of this specification; Figure 7 It is a flowchart of a large language model evaluation method in an embodiment of this specification; Figure 8 is a schematic diagram of the construction process of the evaluation benchmark data set in the embodiments of this specification; Fig. 9 It is a structural diagram of a large language model parameter tuning device in an embodiment of this specification; Fig.10 It is a structural diagram of a large language model evaluation device in an embodiment of this specification. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0028] As mentioned above, with the rapid development of large language model technology, the use of large language models for content generation in related businesses has been widely used in various industries. In order to test whether the business related to the content generated by the large language model has achieved the expected results, the generated content can be evaluated. However, the generated content of the large language model is often divergent and unstable, especially when the content is generated based on the subject in the image. Due to the richness of visual information in the image, the diversity of scenes in the image, and the multimodal fusion of visual information and text information that may be involved, the evaluation of such business results needs to comprehensively consider various uncertainties and complex factors. There are currently methods for manual evaluation and methods for evaluation using open source large language models.
[0029] However, the manual evaluation method has the following problems: First, the efficiency of manual evaluation is very low and the cost is extremely high. Second, manual evaluation is highly subjective because the results are easily affected by personal experience, knowledge, emotions and preferences. The evaluation method using open source large language models has the following problems: First, the accuracy of open source large models is low and the evaluation effect cannot be guaranteed. Manual intervention is still required in the future. Second, the rapid development of large model technology has led to a fast iteration speed of the target business itself, while the open source large models and the benchmark data sets they use are difficult to quickly iterate the evaluation system in line with the business model update.
[0030] In order to solve the above technical problems, the embodiments of this specification provide a large language model parameter tuning method and a large language model evaluation method. Figure 1 FIG. 1 is a schematic diagram showing a large language model parameter tuning process according to an embodiment of this specification. Figure 1As shown, in some embodiments, a training sample may be first obtained, and the training sample includes: business results of the target business, annotation results, and first information for indicating the user's interest tendency. Then, according to the indication of the prompt word, the first language model is used to evaluate based on the business results to obtain the evaluation results for the training sample, and the evaluation indicators in the evaluation results and the annotation results include the correlation between the business results and the first information. Then, based on the evaluation results and the annotation results, the parameters of the first language model are adjusted. In another embodiment, the training sample also includes: a query image. Among them, the query image and the first information indicating the user's interest tendency are input into the target business, and the business results generated by the target business based on the second language model for the subject in the user's query image can be obtained, and the annotation results can be obtained by annotating the business results.
[0031] The following first describes the target business to be evaluated.
[0032] In the embodiment of this specification, the target service is used to generate service results for the subject in the user's query image based on the second largest language model, so as to satisfy the user's curiosity or inspire the user's inspiration and make personalized recommendations to the user. For example, when the user takes a photo of the surrounding environment to obtain a query image and upload it to the target service, the target service can identify the query image to obtain at least one subject, and obtain the first information indicating the user's interest tendency, and then determine the target subject in the at least one subject based on the first information, and then generate a service result for the user based on the target subject through the second largest language model.
[0033] Among them, the business result can be the recommended content used to meet the user's query needs for the subject. Specifically, the business result can be generated content in different modes, for example, it can be one or more of text mode, image mode, video mode, and audio mode. Exemplarily, when the query image is an image obtained by the user taking a photo of the desktop, the teacup, paper, pen, laptop, tea leaves in the teacup, and the user's exposed pants on the desktop in the image can all be considered as the subject. The generated business result can be generated content in text mode, such as: "What kind of tea is suitable for this teacup?".
[0034] In practice, the first information may be a topic or service that the user is interested in, specifically a topic or online service related to people's livelihood and government affairs, commodity information, cultural and technological fields, historical and geographical knowledge, and current affairs information. For example, the first information may be a football match that the user is concerned about, a traditional culture that the user is interested in, a housing fund withdrawal service, etc. In the case where there are multiple identified subjects, the target subject can be determined from the user's interest tendency indicated by the first information, that is, the target subject is the subject that the user is interested in. For example, still taking the desktop object as an example, when there are multiple subjects on the desktop, such as a teacup, paper, pen, laptop, tea leaves in the teacup, and the user's exposed pants, when the user's first information indicates that the user is interested in Chinese traditional culture, because teacups and tea leaves are not only daily necessities, but also have semantic associations with Chinese traditional culture, and carry rich cultural connotations and symbolic meanings, the determined target subject may be teacups and tea leaves.
[0035] The prompt words (Prompt) constructed based on the subject and the user can be input into the second largest language model so that the second largest language model generates corresponding business results. Among them, the prompt words contain information about the subject and can also contain information related to the user, so that the second largest language model generates recommended content that the user may be interested in, and the recommended content is semantically related to the subject. The recommended content can be text information, goods, services or other resources recommended to the user in a personalized manner. For example, when the subject in the query image is a computer host, the prompt words constructed by the user can be "Based on the subject, raise questions that the user may be interested in", and the business results generated by the second largest language model based on the prompt words can be text information such as "How can brand A computer host improve work efficiency?" "How to design the cooling system of brand A computer host?" When the user clicks on the text information, the text information can be input into the second largest language model or other large language models to obtain more relevant text or picture content generated by the large language model to answer related questions and satisfy the user's curiosity.
[0036] The embodiments of this specification do not limit the method of obtaining the user's first information. For example, the user's first information can be obtained by analyzing the user's historical behavior data, or by analyzing the attribute information uploaded by the user. For example, if the historical behavior data is the user's bill record, running shoes, knee pads, sports drinks and other items appear in the bill record many times, the user's first information can be analyzed to be a topic word related to sports and fitness, indicating the user's interest in sports and fitness.
[0037] As an implementation method, in order to accurately obtain the first information of the user, the associated information of the user may be obtained first, and then the first information of the user may be generated based on the associated information through the second largest language model.
[0038] The user's associated information is user personal information related to the user, such as the user's occupation, the user's crowd portrait, the user's historical behavior, etc. Exemplarily, the user's associated information may include at least one of the following information: user attribute information, user historical behavior data, user location information, and user current time information.
[0039] Specifically, the user's attribute information is used to describe the user's basic situation, interest preferences, behavioral habits, such as education level, social role, exercise habits, etc., in order to better model the user portrait; the user's historical behavior data is used to describe the user's online behavior, for example, it may include the user's billing records, records of using mini-programs, search history, and historical records of visiting specific points on the client at a specific time, from which the user's continuous behavior in different time and space scenarios can be obtained to better analyze the user's behavior patterns; the user's location information can be the geographical location information of the user when taking or uploading the query image, or it can be the shooting location information carried by the query image itself, in order to better predict the behavior pattern through the location information; the user's current time information can be the time when the user took the query image, or it can be the time when the user uploaded the query image to the client, in order to better predict the user's behavior pattern through time information.
[0040] The second largest language model used here is an artificial intelligence model based on deep learning, which is specifically used to understand and generate natural language. It can be obtained by fine-tuning the open source large language model, or it can be trained so that the second largest language model can learn the deep connection between the user's associated information and the first information through supervised training.
[0041] In one embodiment, in response to the dynamic needs of users, the second largest language model used may be a spatial temporal-LLM (ST-LLM for short), which is obtained by training or fine-tuning the large language model using large-scale spatiotemporal data (associated information). The large language model has certain spatiotemporal prediction capabilities and can capture the specific needs of users in different scenarios. The generated recommended content can better meet the current needs of users.
[0042] In practice, the prompt information can be constructed based on the associated information, and the prompt information can be input into the spatiotemporal large language model to obtain the first information of the user output by the spatiotemporal large language model. Exemplarily, the constructed prompt information can be: based on the user's portrait information {p} and recent purchase history {h}, combined with the current time {t} and location information {l}, predict what kind of service {f} the user needs, thereby obtaining the service prediction {f} output by the model.
[0043] Next, combine Figure 2The implementation process of the target business is described in more detail.
[0044] like Figure 2 As shown, the target business can be a recommendation business based on a spatiotemporal large language model, and the business mainly includes four modules: an input module, which is used to obtain the input data required in the execution of the target business, such as the user's historical behavior data and attribute information and other related information as well as the user's query image; a spatiotemporal guidance module, which is used to fine-tune the spatiotemporal large language model; a preference discovery module, which consists of step 1, step 2, step 3 and step 4, where the sequence number of the step is only used to identify different steps, and does not limit the execution order of the steps.
[0045] Wherein, in step 1: a scene graph is generated based on the query image, the scene graph includes at least one identified subject and the relationship between the subjects, the scene graph can be represented by a set of triples, each triple is used to represent the relationship between two subjects, illustratively, the subjects included in the scene graph are: a teacup, paper, a pen, a laptop, tea leaves in a teacup, and pants, and an arrow indicates the relationship between two subjects, for example, tea leaves are in the teacup and the pen is on the paper.
[0046] In step 2: a knowledge graph is generated based on the user's attribute information and the scene graph. The knowledge graph can also be represented by a set of triples. Each triple represents the association relationship between a subject and a vocabulary. Multiple triples form a multi-hop chain relationship. Specifically, it can be based on a preset corpus to retrieve association information and / or context information of at least one subject, and then through a large model, based on the association information, at least one subject and context information, generate vocabulary that has an association relationship with the subject. The preset corpus can be various open knowledge bases or private corpora. Vocabulary can be regarded as knowledge enhancement based on the scene graph to better utilize large language models as user agents to simulate user preferences and interests. Figure 2 Take the knowledge graph in as an example, where the vocabulary in the blue module is expanded from the subject, and the vocabulary in the yellow module is expanded from the associated information.
[0047] In step 3: the first information is generated based on the user's attribute information and historical behavior data through the pre-trained spatiotemporal language model. Figure 2 As shown in step 3 of the preference discovery module, the associated information input into the spatiotemporal large language model may include the attribute information of the user and the historical behavior data of the user. After the associated information is input into the spatiotemporal large language model, the spatiotemporal large language model may model the user's interests and predict the first information indicating the user's interest tendency.
[0048] In step 4: according to the first information, search and sort in the knowledge graph to determine the target subject and the vocabulary associated with the target subject. Specifically, based on the first information, the target subject and the vocabulary associated with the target subject are determined in at least one subject and the vocabulary associated with the subject. In practice, the semantic similarity between the first information and each subject-vocabulary pair (i.e., the subject and the vocabulary associated with the subject) can be calculated, and then one or several subject-vocabulary pairs with the highest similarity are determined as the target subject and the vocabulary associated with the target subject, so as to find the target subject and vocabulary closest to the user's interest tendency. Exemplarily, in order to further construct a personalized knowledge graph, hop-by-hop semantic indexing can be performed on the knowledge graph to effectively retrieve relevant information and efficiently capture first-order neighbors with clear semantic associations, thereby serving as a new basis for knowledge retrieval, thereby enhancing the knowledge matching ability of user interests. For example, a pre-trained language model (e.g., a generative encoder) can be introduced to semantically encode the first information and the triples in the knowledge graph, respectively, to obtain the first semantic vector and the second semantic vector, thereby capturing their potential semantic information. In the knowledge graph, since the relationships between multiple subjects and words are connected through a chain relationship consisting of multiple triples, the semantic similarity between the first semantic vector corresponding to the first information and the second semantic vector corresponding to the triple can be calculated hop by hop.
[0049] When calculating hop by hop, the similarity between the first semantic vector and the second semantic vector corresponding to the first hop can be calculated first. If the similarity is higher than the set threshold, the triplet is considered to be highly correlated and retained; otherwise, continue to propagate backward along the graph structure and recursively calculate the similarity of the adjacent triplet of the next hop to ensure the scalability of the retrieval range. For example, calculate the similarity between the first semantic vector and the second semantic vector corresponding to the second hop, and determine whether it is higher than the set threshold. Repeat this cycle until the entire knowledge graph is traversed.
[0050] Finally, after completing the hop-by-hop indexing, the retained associations are screened and sorted according to semantic similarity. For example, the top five associations are selected, and the subjects and words associated with the associations are used as the target subjects and the words that have an association relationship with the target subjects for subsequent personalized recommendations or knowledge enhancement tasks.
[0051] like Figure 2As shown in the reordering of step 4, by comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association between tea and beverages, the similarity score of the association between tea and beverages is 0.9, and by comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association between tea and cultural experience, the similarity score of the association between tea and cultural experience is 0.7. The higher the similarity score, the closer the two are semantically. It can be seen that the user has a high interest in beverages. Question generation based on this can meet the needs of users. If only the target image is analyzed without combining the user's interests, it is difficult to guide the concept of beverages. Through the above method, structured knowledge can be effectively extracted from complex visual content, and personalized knowledge graphs can be generated according to user preferences, thereby improving the accuracy and relevance of information retrieval and recommendation.
[0052] The personalized recommendation module is used to generate at least one recommended content, i.e., a business result, for the user based on the target subject and the vocabulary associated with the target subject using the spatiotemporal large language model. In practice, the prompt word input large language model can be constructed based on the target subject and the vocabulary associated with the target subject, and the large language model can output the corresponding recommended content. The prompt word can include the user's attribute information and the user's historical behavior data to provide the large model with rich background information, and can also include the user's location information and the user's current time information to assist the large model in capturing the user's complex and real dynamic needs that change with time and space according to the actual time and geographical environment. Exemplary recommended content can be "What kind of tea is suitable for this teacup?", "What are the types and materials of traditional Chinese teacups?" and "What are the types of tea in China?", so as to arouse the user's interest.
[0053] like Figure 3 As shown, Figure 3 An exemplary target business result provided in this specification is shown, and the business result is displayed on the display interface of the user terminal. The display interface includes a text section and a view section, wherein the text content in the text section is the business result generated by the second largest language model, and the content in the view section is the image of the subject in the query image. Exemplarily, the generation process of the business result can be: the user takes an image of sausages drying on the balcony and uploads it to the target business for query. After obtaining the query image, the target business identifies the subject "sausage" from the query image, and then constructs a prompt word based on the sausage "Based on the sausage in the image, raise questions that the user may be interested in" and inputs it into the second largest language model, and the second largest language model generates: "What are the different versions of sausages in the world?" "How are sausages made?" "Why has sausage become one of the classic delicacies" and so on.
[0054] In the process of gradually realizing the needs of the target business, powerful indicator data is needed to evaluate whether the target business has achieved the expected results. The embodiment of this specification divides the evaluation of the execution performance of the target business into two modules, among which the understanding module mainly ensures the business rationality of the business results, and the intention module mainly ensures the accuracy of the subject category identification in the view section. In the understanding module, the first language model is used to evaluate the business results of the target business. The detailed process of tuning the parameters of the first language model is described below.
[0055] Figure 4 A flow chart of a large language model parameter tuning method according to an embodiment of the present disclosure is shown. The parameter tuning process can be performed by any device, platform or device cluster with computing and processing capabilities, including steps S401-S403 as shown below.
[0056] like Figure 4 As shown, in step S401, a training sample is obtained.
[0057] Among them, each training sample includes a business result and a corresponding annotation result, as well as a first information for indicating the interest tendency of the user in the training sample. The business result is the result of the output of the above-mentioned target business, and the annotation result is used to annotate the evaluation label of the business result. The annotation result can be the result of manually scoring and annotating the business result according to the evaluation criteria. It can be understood that the evaluation criteria can also be used as the evaluation criteria of the first language model to evaluate the business results and obtain the evaluation results. The first information can be information generated during the implementation of the target business described in the aforementioned embodiment.
[0058] In this embodiment, the evaluation indicators of the evaluation results and the annotation results may include the relevance of the business results to the first information, so as to evaluate whether the content of the target business output is consistent with the user's personal intention. By paying attention to this evaluation indicator, the situation where the user is not interested in the business results can be avoided.
[0059] In different embodiments, different evaluation criteria can be used. In one example, in order to avoid erroneous knowledge information and illegal guidance in business results, three business-related evaluation indicators of relevance, correctness, and logic can be set in the evaluation criteria. In another example, a series of standards and criteria can be used to constrain the evaluation of the model to comprehensively consider various uncertainties and complex factors and reduce evaluation bias. In order to evaluate the effectiveness of the target business in generating personalized recommended content, the evaluation criteria not only need to capture the relevance of the recommended content to the visual query, but also need to measure its fit with the user's interests. The evaluation results and annotation results may also include at least one of the following evaluation indicators: the relevance of the semantic subject in the business result to the subject in the query image; the correctness of the business result; the logic of the business result; the information content of the business result; the relevance of the business result to the first information. It can be understood that the specific details of the evaluation indicators are related to the content of the business results.
[0060] Exemplarily, in the case where the business result is recommended content in text mode (such as a question), the correlation between the semantic subject in the business result and the subject in the query image can be considered as image-text correlation, which is used to evaluate whether the output recommended content is related to the subject of the visual query; the correctness of the business result, that is, the correctness of the recommended content, is used to evaluate whether the output recommended content contains misleading information that violates compliance; the logic of the business result, that is, the logic of the question, is used to evaluate whether the output question is logical and easy to understand; the information content of the business result can be considered as the attractiveness of the question, which is used to evaluate whether the output question provides new perspectives or involves content that is not well known to most people. Generally speaking, the greater the information content of the question, the more attractive the question is; the correlation of the business result with the first information, that is, the personalization of the question, is used to evaluate whether the output question is highly relevant to the user's personal intentions and interests.
[0061] In different embodiments, the scoring criteria for the evaluation indicators may be different. In one embodiment, three gears of high, medium, and low (1 / 0 / -1) may be set for manual evaluation, wherein the higher the gear or the higher the score, the better the execution performance of the target business. In addition, considering that subjective evaluation is easily affected by personal experience, knowledge, emotion, and preference, in order to reduce the deviation that may be caused by subjectivity, it is also possible to combine objective data or use the opinions of multiple evaluators for comprehensive analysis. For example, during the evaluation process, multiple groups of annotation personnel with similar education can be used for evaluation to ensure the validity of the results. For example, when the business result is a problem, the annotation result can be "diversity of problem angles: 1 point, logic of problem: 1 point, personalization of problem: -1 point, information content of problem: 1 point, correctness of content: 1 point, relevance of pictures and texts: 1 point".
[0062] In practice, training samples can be obtained from a pre-prepared training sample set or from other specified locations.
[0063] In one embodiment, in order to improve the parameter tuning effect of the first language model and reduce the annotation cost, based on the idea of active learning, this embodiment also constructs a high-quality training sample set to assist the continuous fine-tuning and iteration of the first language model. The most valuable data is selected in the training sample set to improve the performance and learning efficiency of the model, and reduce the demand for annotated data. The following examples illustrate several methods of obtaining training samples in the training sample sets. It can be understood that the methods of obtaining training samples in the following examples can be used separately or in combination.
[0064] In one example, in order to meet the data validity of the training sample set, the data can be strictly filtered. When obtaining the training sample, the first business result and the first annotation result generated by the target business for the first alternative image can be obtained, and the first business result includes the location information of at least one subject obtained based on the first alternative image recognition, and the first annotation result is used to annotate the location label of at least one subject; based on the location information in the first business result and the location label in the first annotation result, the first alternative image is screened to determine the first image; and the business result and the annotation result generated by the target business for the first image are obtained as training samples.
[0065] Considering that the business results of the target business are generated based on the subject in the user's query image, if the subject in the query image is detected incorrectly, the subject is unclear, or the user's query intention is unclear, it will affect the quality of the generated business results. Evaluating such business results will not reflect the actual execution performance of the target business, and such business results can be considered invalid data. In this example, the above invalid data can be filtered out by verifying the accuracy of the position of the subject detected in the target business.
[0066] Specifically, the first candidate image can be a query image in the historical business data of the target business. When the target business identifies the subject in the first candidate image, a first business result is generated. The first business result includes the location information of at least one subject obtained based on the first candidate image recognition, such as the location box where the subject is located. By manually annotating the first candidate image, a first annotation result, i.e., the ground truth (GT for short), can be obtained. Then, the degree of regional overlap between the location box and the annotation box can be measured by IoU (Intersection over Union) or other screening strategies. If the IoU value is greater than a certain threshold (such as 0.5), the detection is considered successful, and the first candidate image is retained and determined as the first image. Otherwise, it means that the target business has failed to successfully detect the subject, and the first candidate image is filtered out to ensure that samples with unclear subjects and unclear intentions do not flow into the training samples.
[0067] In another example, in order to meet the data diversity of the training sample set, when obtaining training samples, the category of at least one subject in the second candidate image can be obtained; according to the requirements of the first category, the second candidate images are screened to determine the second image; and the business results and annotation results generated by the target business for the second image are obtained as training samples.
[0068] Among them, the second candidate image can be an image collected from historical business data, various online platforms or manual shooting, and the subject recognition of the second candidate image can obtain the categories of different subjects in the second candidate image. Exemplarily, the category of the subject in the embodiment of this specification can refer to the secondary category. When the primary category is electronic products, the secondary category can be mobile phones, televisions, headphones, computers and other categories. In other embodiments, the category of the subject can also refer to categories of other levels. The first category requirement can be a requirement for the richness of various categories in the category, for example, it can be a requirement for the proportion of different categories. The second candidate image that meets the first category requirement is determined as the second image, and the business results generated by the target business based on the second image are manually marked to obtain the corresponding annotation results, thereby obtaining a training sample. By sampling different subject categories, the category richness is guaranteed so that the training samples provide enough differentiated information, thereby ensuring the comprehensive evaluation performance of the first language model.
[0069] In another example, considering that the business iteration speed of the target business is very fast, in order to enable the first language model to adjust parameters in time with business changes, the actual business data when the target business is running can also be obtained. The actual business data includes at least one of the following data: a query image, a first information indicating the user's interest tendency; based on the difference between the actual business data and the training sample, determining the sample offset data; based on the sample offset data, constructing a new training sample.
[0070] In practice, the actual business data generated by the target business during runtime includes query images collected by different users in different scenarios, as well as the first information of different users. Due to the rapid development of current society, many new things are generated every day, which makes the categories of subjects in the query images and the interests of users constantly updated. By comparing the differences between the actual business data and the training samples in the training sample set, the sample offset data after the business change can be determined. The sample offset data contains query images and / or first information that have never appeared in the training samples. New training samples are constructed through the sample offset data, so that the training sample set can be quickly updated with the business iteration of the target business, thereby improving the evaluation performance of the first language model for newly emerging business data.
[0071] Next, in step S402, the first language model is evaluated based on the business results to obtain evaluation results for the training samples.
[0072] In practice, a prompt word can be constructed and input into the first language model to prompt the first language model to evaluate the business results and obtain the evaluation results. The prompt word refers to the information input into the first language model, which can include information of different modes such as text and images. The purpose is to guide the first language model to evaluate the business results according to the evaluation criteria or evaluation requirements in the prompt word, and generate corresponding evaluation results. The first language model can be an open source large model or a pre-trained large model. In different embodiments, the first language model can be a large model of different specific types or with different neural network structures, and the prompt word can also be a different specific prompt word. This specification does not limit this.
[0073] In different embodiments, the specific scoring criteria indicated by the prompt word may be different. In one embodiment, when the business result is a question generated for the subject in the query image, the evaluation indicator in the scoring criteria may be the logic of the business result, and the prompt word may be "Please score the logic of question q, 5 points: correct logic, precise and fluent expression; 4 points: basically correct logic, correct expression; 3 points: basically correct logic; 2 points: there are loopholes in the logic and the result is partially incorrect; 1 point: there are obvious loopholes in the logic and some content is misunderstood; 0 points: false content".
[0074] The evaluation indicators in the scoring criteria may also be the relevance between the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; the information content of the business result and the relevance of the business result to the first information. When the business result is a question generated for the subject in the query image, the prompt word A may be "the content of the question is: q, Please score the two related items only for the semantic subject in the question (only the semantic subject is considered) and the subject in the picture. The score range is [-1, 0, 1]. The specific meaning of each range is -1: The semantic subject of the question (only the semantic subject is considered) is irrelevant to the subject in the figure. 0: The subject of the question (only the semantic subject is considered) is too broad. 1: The subject of the question (only the semantic subject is considered) is clear and consistent with the subject in the figure. Please score according to the score level rules. Note that you only need to output the score and put it in the first element of the list; Please rate the correctness of the question content. The score range is [-1, 0, 1]. The specific meaning of each range is: -1: The content is wrong or the content has incorrect guidance costs. 0: The question is normal, but the knowledge is unclear. 1: The question has a certain level of knowledge. Please score according to the score level rules. Note that you only need to output the score and put it in the second element of the list; Please rate the logic of the question. The score range is [-1, 1]. The specific meaning of each range is: -1: The question description is illogical and difficult to understand; or there are references to names in the question. 1: The logic of the question is clear. Please score according to the score level rules. Note that you only need to output the score and put it in the third element of the list; Please rate the information content of the question. The score range is [-1, 0, 1]. The specific meaning of each range is: -1: The answer to the question is too simple (it can be seen by the naked eye and basically does not require any thinking); 0: The question is mediocre and the answer is known to most people (common sense but requires some thought); 1: The question raises new ideas or involves content that is not well known to most people. Please score according to the score level rules. Note that you only need to output the score and put it in the fourth element of the list; User's personal interest keyword (i.e. the first information): 'usr_interest'. Please rate the relevance between the question and any item of user interest. The score range is [-1, 1]. The specific meaning of each range is: -1: Not related to the user's personal interests; 1: It is relevant to the user's personal interests and has a sense of freshness and desire to explore. Please score according to the score level rules. Note that you only need to output the score and put it in the fifth element of the list; Finally, output the result list. Note that only the list format is output. If other problems occur, output [0, 0, 0, 0].
[0075] In the above embodiment, the prompt word includes the business result and the first information. In other embodiments, the prompt word may also include a query image to assist the first language model to better perform the evaluation. After the prompt word is input into the first language model, the obtained evaluation result may include a score for each evaluation indicator.
[0076] It is understandable that business results are not limited to questions, but can also be knowledge content, emotion-related content, etc. For other types of business results besides questions, prompt words can be generated in a manner similar to the above example, which will not be elaborated in this embodiment.
[0077] Next, in step S403, the parameters of the first language model are adjusted based on the evaluation results and the annotation results.
[0078] In practice, there will be differences between the evaluation results output by the largest language model and the manually annotated results. For example, for a certain evaluation indicator, the score in the evaluation result is low, while the score in the annotation result is high. By adjusting the parameters of the largest language model, this difference can be gradually narrowed, making the evaluation results closer and closer to the annotation results, thereby continuously improving the evaluation performance of the largest language model. For example, the loss function can be calculated based on the difference between the evaluation results and the annotation results, and then the parameters of the largest language model can be adjusted with the goal of minimizing the loss value of the loss function. When adjusting the parameters, you can fine-tune some of the parameters, you can adjust all the parameters, or you can optimize the prompt words through P-Tuning (Prompt Tuning).
[0079] In one embodiment, for the part where the evaluation result predicted by the first language model is significantly different from the manual annotation result, expert knowledge may be introduced for secondary calibration and used as a difficult sample (Hard Case) to strengthen the training of the first language model.
[0080] In one example, it may be possible to first determine whether a training sample is an offset sample based on the difference between an evaluation result and an annotation result, then, if the training sample is determined to be an offset sample, upsample the offset sample to obtain an amplified offset sample, and finally, based on the amplified offset sample, adjust the parameters of the first language model.
[0081] In practice, the difference between the evaluation results and the annotation results can be measured first. Different embodiments can use different ways to measure. For example, a loss function can be used to calculate the difference between the two. The larger the loss value of the loss function, the greater the difference. When the loss value exceeds the preset threshold A, the training sample is determined as an offset sample. Alternatively, the similarity between the two can be calculated. The lower the similarity, the greater the difference. When the similarity is lower than the preset threshold B, the training sample is determined as an offset sample.
[0082] As an implementation method, in response to the consistency of the evaluation indicators in the evaluation results and the evaluation indicators in the annotation results being lower than a preset threshold, the training sample can be determined as a shifted sample. For example, the Pearson correlation coefficient, Spearman rank correlation coefficient or Cohen's Kappa coefficient can be used to calculate the consistency of the evaluation results and the annotation results. The exact match can also be used to calculate the consistency between the two, that is, to calculate the matching degree of the content and format in the evaluation results and the annotation results. If the consistency between them is lower than the preset threshold, it means that there is a large difference between the evaluation results and the annotation results, and it is determined as a shifted sample. Exemplary, Figure 5 A schematic diagram of the degree of consistency is shown. For Exact match, the closer the score is to 1, the higher the degree of consistency. For the Kappa coefficient, when the score is greater than , it is considered that the two are highly consistent.
[0083] After determining the offset samples, expert experience can be used to determine the reasons for the huge differences in the offset samples and take corresponding measures. For example, if it is caused by errors in manual labeling results, the offset samples can be relabeled. If it is caused by model prediction errors, the offset samples can be used as hard cases to recalibrate the model.
[0084] In this example, when the offset sample is a hard case, the offset sample can be upsampled to obtain an amplified offset sample. This embodiment does not limit the specific upsampling method. For example, the offset sample can be copied by simple oversampling to increase its number in the training sample set, or a synthetic minority class oversampling technique can be used to generate new samples by interpolation between training samples, thereby amplifying the offset sample and increasing its weight in the training sample set. Finally, step S402 and step S403 can be performed again according to the amplified offset sample to adjust the parameters of the first language model. The adjusted model can achieve better evaluation results when facing hard cases.
[0085] In another example, when it is determined that the training sample is an offset sample, based on the offset sample and the subjectivity of each evaluation indicator, the prompt words of the first language model can be adjusted, and the evaluation can be performed again based on the business results to obtain a second evaluation result for the training sample. Finally, based on the second evaluation result and the labeling result, the parameters of the first language model can be adjusted.
[0086] In actual implementation, since the scores of various evaluation indicators in the annotation results are greatly affected by the subjectivity of the annotators, the evaluation results predicted by the model are quite different from the annotation results. For such situations, this embodiment considers adjusting the prompt words of the input model according to the subjectivity of each evaluation indicator to make it closer to people's subjective thinking, thereby narrowing the difference between the annotation results and the evaluation results and improving the evaluation effect of the model.
[0087] For example, when the annotator scores different evaluation indicators, the order of the degree of subjectivity of each evaluation indicator in the evaluation logic of the annotator can be: attractiveness>logicity>content correctness>subject relevance>interest relevance. For example, the prompt word can be added with "Please score the above evaluation indicators subjectively according to the subjective order of attractiveness>logicity>content correctness>subject relevance>interest relevance", or the scoring standard of the prompt word can be directly optimized to make it closer to people's subjective thinking. For example, the prompt word B obtained by optimizing the prompt word A in the above embodiment can be "The content of the question is: q. Regarding the relevance of the question to the main picture, please score according to the following steps, and finally only output the score: 1. Identify the semantic subject from the question text, that is, the main object of emphasis or attention.
[0088] 2. Observe and identify the subject in the image and identify the main object shown in the image.
[0089] 3. Compare the semantic subject in the question with the subject in the image: If they are not relevant please rate -1.
[0090] If the semantic subject of the question is too broad to be matched, please score 0.
[0091] If the semantic subject is clear and the subject is the same as in the figure, please score 1.
[0092] If there is any uncertainty, just give it a score of 0 and output the score in a list. The score should be placed in the first element of the list.
[0093] Regarding the correctness of the questions, please analyze the content of the following questions in detail, and finally only output the scores: 1. Make sure you fully understand the question.
[0094] 2. Determine whether there are factual errors in the content: If there are obvious errors that may lead to wrong conclusions, -1 point will be given.
[0095] 3. Determine whether the question is knowledge-based: If the question is not clearly knowledge-based, give 0 points.
[0096] Give 1 point if the question stimulates intellectual discussion and exploration.
[0097] 4. If there is any uncertainty, just give it a score of 0. Output the score in a list format, with the score being the second element of the list.
[0098] Please rate the logic of the question. Please analyze the content of the following questions in detail and only output the score in the end: 1. Make sure you fully understand the question.
[0099] 2. Determine whether there are factual errors in the content: If there are obvious errors that may lead to wrong conclusions, -1 point will be given.
[0100] 3. Determine whether the question is knowledge-based: If the question is not clearly knowledge-based, give 0 points.
[0101] Give 1 point if the question stimulates intellectual discussion and exploration.
[0102] 4. If there is any uncertainty, just give it a score of 0. Based on the analysis, the final score should be placed in the third element of the list.
[0103] Please rate the attractiveness of the question, specify the rating criteria, and finally output only the rating: -1 means the answer to the question is obvious and can be obtained without thinking; 0 means the answer to the question is common knowledge but requires a short thought; 1 means the question introduces new ideas or involves knowledge that is not widely known.
[0104] Problem Analysis: Is the problem simple: Is the answer obvious? Level of common sense: Is the issue understood by most people? Novelty: Does the question present new ideas or unusual knowledge? Select the score: According to the analysis of the question, select the corresponding score. If there is no confirmation, just give it a score of 0. The score should be placed in the fourth element of the list; According to the user's personal interests, only the ratings need to be output in the end. Input: User's personal interest keyword: 'usr_interest' Questions to be graded: q Output: Output the association score as the fifth element of the list.
[0105] Execution steps: 1. Extract the user’s personal interest keywords.
[0106] 2. Get the question text to be evaluated.
[0107] 3. Analyze the relevance of questions to user interests: Find occurrences of the keyword of interest in the question.
[0108] Determine whether the question content is likely to generate freshness or exploration desire for users.
[0109] 4. Scoring will be based on the following criteria: -1: The question is not relevant to the user's interest.
[0110] 1: The question is related to the user’s interests and may stimulate exploration.
[0111] 5. If there is no confirmation, just give it a score of 0. Store the score as the fifth element of the list and output it. Note: Finally, the result list is output. Note that only the list format is output. If other problems occur, output [0, 0, 0, 0]" The adjusted prompt words are closer to the human evaluation logic, so that the evaluation results of the model are closer to the human subjective feelings. Then, based on the adjusted prompt words, step S402 and step S403 are performed again until the difference between the annotation result and the evaluation result reaches the preset requirement.
[0112] In other embodiments, the offset sample may also be determined from the evaluation of the first language model on other test samples other than the training sample, such as Figure 6 As shown, Figure 6 Another path for obtaining and processing offset samples is shown. For the manually annotated benchmark data set, in the first stage of training, it can be sampled to obtain a test sample set 1, and the training sample set 1 is screened by IoU and other methods. The specific screening method is mentioned in the previous text and will not be repeated here. The training samples in the training sample set 1 are used to tune the parameters of the first language model to obtain the first language model of version V1. The first language model of version V1 can be used to evaluate the business results of the test samples in the test sample set 1 to obtain the evaluation results corresponding to the test samples, which are compared with the annotation results and the samples with large differences between the two are determined as offset samples. In the second stage of training, the benchmark data set can be sampled to obtain test sample set 2, and a new training sample set 2 can be constructed by using the offset samples and training sample set 1. Similarly, the training samples in training sample set 2 are used to tune the parameters of the first largest language model of version V1 to obtain the first largest language model of version V2. The first largest language model of version V2 can be used to evaluate the business results of the test samples in the test sample set 2 to obtain the evaluation results corresponding to the test samples, which are compared with the annotation results and the samples with large differences between the two are determined as new offset samples. This process is repeated until the adjusted first largest language model meets the business needs.
[0113] The large language model parameter tuning method provided in the embodiment of this specification can, on the one hand, automatically evaluate the business results generated by the target business through the first large language model, which greatly improves the evaluation efficiency of the business results, reduces the labor cost and time cost consumed, and is more objective compared to manual evaluation. On the other hand, by adjusting the parameters of the first large language model using the evaluation results and annotation results of the training samples, the accuracy and stability of the first large language model during evaluation are improved, and the first large language model can be adapted to different business changes, thereby being able to test the generation performance of the second large language model under different iterative versions. In addition, the correlation between the business results and the first information used to indicate the user's interest tendency is also considered during the evaluation, which can ensure that the evaluation results are consistent with the user's subjective feelings.
[0114] The following is an explanation of the evaluation process of the large language model after parameter tuning. Figure 7 A flow chart of a large language model evaluation method according to an embodiment of the present disclosure is shown. The method can be executed by any device, platform or device cluster with computing and processing capabilities, including steps S701-S702 as shown below.
[0115] like Figure 7 As shown, in step S701, the execution result of the target business is obtained.
[0116] The execution result is generated by the target business based on the second largest language model for the subject in the user's query image. For a more detailed explanation of the second largest language model, the target business and its execution results, please refer to the previous description of the second largest language model, the target business and its business results, which will not be repeated here.
[0117] In step S702, the execution result is evaluated by the first large language model to obtain evaluation data for the execution result.
[0118] The first large language model is a large language model obtained after parameter tuning by any method of the aforementioned embodiment. In practice, prompt words can be constructed according to the execution result, the first information, and the query image and incorporated into the first large language model to obtain output evaluation data. The evaluation indicators in the evaluation data can include the relevance of the execution result and the first information, and can also include the relevance of the semantic subject in the execution result and the subject in the query image, the correctness of the execution result, the logic of the execution result, and the information content of the execution result.
[0119] In the solution provided in the above embodiment, an evaluation of image-based content generation tasks can be implemented, wherein the first language model used can adapt to changes in the execution results of the second language model when facing different businesses through parameter tuning, and can test the generation performance of the second language model under different conditions.
[0120] In one embodiment, this specification also provides an evaluation method for the intent module, which can be used to evaluate the intent module. Figure 2 Specifically, the intermediate execution result and category annotation data of the target business can be obtained. The intermediate execution result is the subject category in the query image obtained by the target business recognition. The category annotation data is used to annotate the category label of the subject of the query image. Then, the third language model evaluates the intermediate execution result according to the category annotation data to obtain the accuracy of the intermediate execution result.
[0121] In practice, during the execution of the target business, before generating the execution result based on the first language model, it is necessary to detect and identify the subject in the query image to obtain the location box and subject category of the subject, that is, the intermediate execution result. The accuracy of the intermediate execution result affects the actual performance of the target business.
[0122] Among them, the evaluation of the accuracy of the subject's position is relatively objective. The IoU method can be used to compare the degree of overlap between the manually annotated subject's annotation box and the location box in the intermediate business result. If the IoU value is greater than a certain threshold (such as 0.8), the detection is considered accurate. The evaluation of the subject category can be evaluated through the third language model. For example, a prompt word can be constructed, and the content of the prompt word can be "Please compare the category label in the category annotation data with the subject category of the intermediate execution result, and output a score of 0-10". The higher the score, the more accurate the intermediate execution result.
[0123] Considering that there are many subject categories in actual application scenarios, and different subject categories have different recognition difficulties, as an implementation method, automatic evaluation coverage can be gradually performed according to the difficulty of the attributes of the subject categories. For example, four evaluation modes are illustrated as follows: a. For categories with urgent needs and subjective evaluation, manual evaluation can be used. For example, categories related to medical care and safety can be considered as urgent needs, and categories affected by personal experience can be considered as subjective categories, and manual evaluation can be used to ensure its accuracy.
[0124] b. For objective evaluation categories that do not require ambiguity, such as medicines, the names, categories and other attributes of the medicines are very certain and there is no ambiguity. Pre-programmed code logic can be used for evaluation, and manual review and evaluation can be added to reduce the amount of manual involvement.
[0125] c. For ambiguous objective evaluation categories, such as wine categories, there are many types of wine and their appearances are very similar, which will cause some deviations in the evaluation. You can use a third-largest language model with unimodal input for evaluation. For example, you can construct prompt words of text content and input them into the third-largest language model for evaluation.
[0126] d. For ambiguous objective evaluation categories and subjective evaluation categories, that is, categories that are easily influenced by personal experience, such as animals, plants, ingredients, beauty products and other categories with complicated aliases or common names, the third language model with multimodal input can be used for evaluation. When constructing prompt words, multimodal prompt words can be constructed based on the query image, the location box of the subject, the subject category and the category label, so as to inquire the model about the matching degree of the subject category and the category label, so as to combine multimodal data for more accurate evaluation.
[0127] Specifically, an evaluation mode can be preset for each category. During the evaluation process, a single or a combination of multiple evaluation methods can be used to output the accuracy index according to the complexity of the category attributes. In practice, by continuously optimizing the code scripts and model prompts, the overall evaluation accuracy can reach more than 95% after manual sampling review and verification.
[0128] In some embodiments, in order to better evaluate different versions of the target business, an evaluation benchmark data set for the target business can be pre-constructed, and the evaluation benchmark data set contains query images of multiple users and category annotation data corresponding to the query images. When obtaining the intermediate execution results and category annotation data of the target business, the evaluation benchmark data set can be obtained first, and then the target business generates the intermediate execution results for the subject in the query image in the evaluation benchmark data set. For example, the query image and category annotation data of a test sample in the evaluation benchmark data set are obtained, and the target business obtains the subject category based on the query image recognition. In this way, different versions of the target business can be evaluated based on a unified standard evaluation benchmark data set, which can better determine the pros and cons of the execution performance of different versions of the target business.
[0129] Combine the following Figure 8 , the construction process of the benchmark dataset is explained. Figure 8 The full-graph data in the data may include historical business data, business data of a third-party platform, and return data of the target business. After obtaining the full-graph data, it may be returned or dropped into a table to be stored in the database.
[0130] Exemplarily, the construction of the benchmark dataset can be divided into three stages: Phase 1: Historical data reuse and processing In one embodiment, historical business data may be obtained first, the historical business data including the user's query image, and then the historical business data may be filtered according to data filtering requirements to obtain target business data, and then the query image in the target business data may be annotated to obtain category annotation data.
[0131] In practice, the historical business data may be the historical data of other image-based query businesses similar to the target business. The historical business data is cleaned according to the data screening requirements to remove invalid data. The invalid data may be duplicate data, incomplete data, or erroneous data. Different embodiments may adopt different cleaning methods according to the data screening requirements. For example, Figure 8As shown, the large model can be directly used to clean the historical business data in the database. Then, the position of the subject in the query image can be manually annotated, that is, the box selection and pre-marking can be performed to obtain the annotation box, and the subject category in the query image can be manually annotated to obtain the category annotation data for the secondary category. The category annotation data is randomly sampled to obtain the evaluation benchmark data set available to the intent module. The evaluation benchmark data set can include the annotation box and secondary category of the subject. The category annotation data can also include the necessary attributes of the subject. The necessary attributes can include the brand, category, and first-level category of the subject to ensure that the category name of the subject is not ambiguous. For example, the necessary attributes of an apple can be plant, fruit, etc., to distinguish it from the necessary attributes of an Apple phone, which is a mobile phone or an electronic product. The necessary attributes can be automatically annotated by a trained large model first. In order to improve the accuracy of the annotation, such as Figure 8 As shown in Figure 1, a variety of different large models (large model 1, large model 2, large model 3 and large model 4) can be used to annotate the subject in the same query image, and then the answers output by different large models are combined as reference answers, and manual annotation is performed based on the reference answers to obtain richer category annotation data, that is, Figure 8 The detailed attributes of the standard categories in the dataset can be used as a richer benchmark dataset. The model will also output corresponding richer category evaluation data based on the richer benchmark dataset. The above datasets can be manually reviewed to ensure their accuracy.
[0132] As an implementation method, when performing data cleaning, the main category of the query image in the historical business data may be obtained, and the query image in the historical business data may be screened according to the second category requirement to obtain the target business data.
[0133] The second category requirement may be a requirement for the richness of various categories in the category, for example, it may be a requirement for the proportion of different categories or the number of different categories. By screening different main categories, the category richness is ensured so that the benchmark data set can provide enough differentiated information, thereby being able to evaluate the comprehensive performance of the target business.
[0134] During this phase, historical data from other businesses are reused. After being cleaned, labeled, and processed, these data are persistently stored. During this phase, existing data resources can be utilized to quickly build an initial data set when data resources are insufficient.
[0135] In other embodiments, the historical data may also include historical business data of the target business.
[0136] Phase 2: Data crawling and collection
[0137] In one embodiment, business data of a third-party platform may be obtained, where the business data includes query images in actual user usage scenarios, and then the query images in the business data are annotated to obtain category annotation data.
[0138] In practice, third-party platforms can be travel, social, video sharing or other types of websites or applications. Data close to the actual user usage scenarios can be obtained through data crawling to ensure that the data set can better fit the usage form of the target business. Through third-party platforms, rich and diverse data samples can be obtained, thereby improving the quality and representativeness of the data set. The specific annotation process can be seen in the first stage, which will not be repeated here.
[0139] Phase 3: Link backflow and directional enhancement
[0140] In the third stage, data can be reflowed from the actual business data of the target business and targeted for reinforcement construction. Similar to the first stage, the large model can be used to clean the data for the randomly sampled category results and their proportions to ensure that the data meets the preset category requirements, and then the cleaned data can be reflowed into the evaluation benchmark data set. This step helps to supplement and improve the data set to ensure that it covers a wider range of situations.
[0141] In one embodiment, the directional reinforcement construction can be performed based on the Hard Case, i.e., the offset image, in the evaluation benchmark data set. Specifically, based on the accuracy of the intermediate execution result, it can be determined whether the query image corresponding to the intermediate execution result is an offset image. When it is determined that the query image is an offset image, a new evaluation benchmark data set is constructed based on the offset image.
[0142] Exemplarily, when the accuracy of the intermediate execution result is lower than a preset threshold, it means that the subject category and the category annotation data in the query image obtained by the target business identification are very different, and the evaluation performance of the model on the subject category needs to be strengthened. The query image corresponding to the intermediate execution result can be determined as an offset image, and then the offset image is subjected to directed enhancement construction. For example, the offset image can be upsampled to obtain an expanded offset image and added to the evaluation benchmark data set. Alternatively, data samples for specific scenarios can be generated based on the scenes in the offset image to enhance the diversity and complexity of the data set, thereby improving the model's understanding and generalization capabilities.
[0143] In addition, for the understanding module, the above process can also be used to construct a corresponding evaluation benchmark data set to evaluate the execution performance of different versions of the target business based on a unified standard evaluation benchmark data set. The difference lies in the annotation part of the query image. For example, after determining the query image of the evaluation benchmark data set, the target business can generate a business result based on the query image and the first information, and then manually annotate the various evaluation indicators of the business result to obtain an evaluation label. Finally, the evaluation benchmark data set is constructed based on the business results, the annotation results, the query image and the first information.
[0144] Fig. 9 : is a schematic diagram of the structure of the large language model parameter tuning device in the embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. Among them, the first large language model is used to evaluate the execution performance of the target business, and the target business is used to generate business results for the subject in the user's query image based on the second large language model; the device includes:
[0145] The data acquisition module 901 is configured to acquire a training sample, wherein the training sample includes a business result, a labeling result, and first information indicating the user's interest tendency, wherein the labeling result is used to label an evaluation label of the business result;
[0146] The business evaluation module 902 is configured to evaluate the business result based on the first language model to obtain the evaluation result for the training sample, and the evaluation result and the annotation result include the correlation between the business result and the first information;
[0147] The parameter adjustment module 903 is configured to adjust the parameters of the first language model based on the evaluation results and the annotation results.
[0148] In one implementation, the training samples include query images.
[0149] In one embodiment, the evaluation results and annotation results also include at least one of the following evaluation indicators: the relevance of the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; and the information content of the business result.
[0150] In one embodiment, the business evaluation module 902 is specifically configured to determine whether a training sample is an offset sample based on the difference between the evaluation result and the annotation result; when it is determined that the training sample is an offset sample, upsample the offset sample to obtain an amplified offset sample; and adjust the parameters of the first language model based on the amplified offset sample.
[0151] In one embodiment, the device also includes: a prompt word optimization module (not shown in the figure), which is configured to determine whether the training sample is a shifted sample based on the difference between the evaluation result and the annotation result; when the training sample is determined to be a shifted sample, based on the subjectivity of the shifted sample and each evaluation indicator, by adjusting the prompt word of the first language model, and performing evaluation again based on the business result, to obtain a second evaluation result for the training sample; based on the second evaluation result and the annotation result, the parameters of the first language model are adjusted.
[0152] In one embodiment, when the business evaluation module 902 or the prompt word optimization module determines whether a training sample is an offset sample based on the difference between the evaluation result and the annotation result, it is specifically configured to determine the training sample as an offset sample in response to the consistency between the evaluation index in the evaluation result and the evaluation index in the annotation result being lower than a preset threshold.
[0153] In one embodiment, the data acquisition module 901 is specifically configured to: obtain a first business result and a first annotation result generated by the target business for the first candidate image, the first business result including location information of at least one subject obtained based on the first candidate image recognition, and the first annotation result being used to annotate the location label of at least one subject; based on the location information in the first business result and the location label in the first annotation result, screen the first candidate image to determine the first image; obtain the business result and annotation result generated by the target business for the first image as training samples.
[0154] In one embodiment, the data acquisition module 901 is specifically configured to: obtain the category of at least one subject in the second candidate image; screen the second candidate image according to the first category requirements to determine the second image; obtain the business results and annotation results generated by the target business for the second image as training samples.
[0155] In one embodiment, the data acquisition module 901 is also configured to: acquire actual business data when the target business is running, the actual business data including at least one of the following data: a query image, a first information indicating the user's interest tendency; based on the difference between the actual business data and the training sample, determine the sample offset data; based on the sample offset data, construct a new training sample.
[0156] Fig.10 Schematic diagram of the structure of the large language model evaluation device in the embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes:
[0157] A result acquisition module 101 is configured to acquire an execution result of a target service, where the execution result is generated by the target service based on the second largest language model for a subject in a user's query image;
[0158] The performance evaluation module 102 is configured to evaluate the execution result by using a first large language model to obtain evaluation data for the execution result. The first large language model is a large language model obtained after parameter tuning by any large language model parameter tuning method.
[0159] In one embodiment, the device also includes a category evaluation module (not shown in the figure), which is configured to obtain an intermediate execution result and category annotation data of the target business, the intermediate execution result is a subject category in the query image obtained by identifying the target business, and the category annotation data is used to mark the category label of the subject of the query image; the intermediate execution result is evaluated by the third language model according to the category annotation data to obtain the accuracy of the intermediate execution result.
[0160] In one embodiment, when the category evaluation module obtains the intermediate execution results and category annotation data of the target business, it is specifically configured to obtain an evaluation benchmark data set, where the evaluation benchmark data set contains query images of multiple users and category annotation data corresponding to the query images; the target business generates an intermediate execution result for the subject in the query image in the evaluation benchmark data set.
[0161] In one embodiment, the category evaluation module is specifically configured to obtain historical business data when obtaining an evaluation benchmark data set, and the historical business data includes a user's query image; filter the historical business data according to data screening requirements to obtain target business data; and annotate the query image in the target business data to obtain category annotation data.
[0162] In one embodiment, the category evaluation module filters the historical business data according to the data screening requirements to obtain the target business data, and is specifically configured to obtain the main category of the query image in the historical business data; according to the second category requirement, the query image in the historical business data is filtered to obtain the target business data.
[0163] In one embodiment, when obtaining the evaluation benchmark data set, the category evaluation module is specifically configured to obtain business data of a third-party platform, where the business data includes query images in actual user usage scenarios; and annotate the query images in the business data to obtain category annotation data.
[0164] In one embodiment, the category evaluation module is also configured to determine whether the query image corresponding to the intermediate execution result is an offset image based on the accuracy of the intermediate execution result; if it is determined that the query image is an offset image, a new evaluation benchmark data set is constructed based on the offset image.
[0165] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the following Figure 4 and Figure 7 Describe the method.
[0166] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the following is implemented: Figure 4 and Figure 7 Describe the method.
[0167] The embodiments of the present specification also provide a computer program product, including a computer program / instruction, which is executed by a processor to implement the following Figure 4 and Figure 7 Describe the steps of the method.
[0168] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0169] In some cases, the actions or steps described in the claims may be performed in a different order than in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0170] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not used to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A method for tuning parameters of a large language model, wherein: The first language model is used to evaluate the execution performance of a target service, and the target service is used to generate a service result for a subject in a user's query image based on the second language model; the method comprises: Acquire a training sample, wherein the training sample includes the business result, a marking result, and first information indicating the interest tendency of the user, wherein the marking result is used to mark an evaluation label of the business result; The first language model is used to perform an evaluation based on the business result to obtain an evaluation result for the training sample, wherein the evaluation result and the annotation result include a correlation between the business result and the first information; Based on the evaluation result and the annotation result, parameters of the first language model are adjusted.
2. The method according to claim 1, wherein: The training samples include the query image.
3. The method according to claim 1, wherein: The evaluation result and the annotation result also include at least one of the following evaluation indicators: the correlation between the semantic subject in the business result and the subject in the query image; The correctness of the business results; the logic of the business results; the information content of the business results.
4. The method according to claim 1, wherein: The adjusting the parameters of the first language model based on the evaluation result and the annotation result includes: Based on the difference between the evaluation result and the labeling result, determining whether the training sample is a shifted sample; In the case where it is determined that the training sample is an offset sample, upsampling the offset sample to obtain an amplified offset sample; Based on the amplified offset sample, the parameters of the first language model are adjusted.
5. The method according to claim 1, wherein: The method further comprises: Based on the difference between the evaluation result and the labeling result, determining whether the training sample is a shifted sample; In the case where it is determined that the training sample is a shifted sample, based on the shifted sample and the subjectivity of each of the evaluation indicators, by adjusting the prompt words of the first language model, and performing evaluation again based on the business result, a second evaluation result for the training sample is obtained; Based on the second evaluation result and the annotation result, parameters of the first language model are adjusted.
6. The method according to claim 4 or 5, wherein: The determining whether the training sample is a shifted sample based on the difference between the evaluation result and the labeling result includes: In response to the consistency between the evaluation index in the evaluation result and the evaluation index in the annotation result being lower than a preset threshold, the training sample is determined as a shifted sample.
7. The method according to claim 1, wherein: The obtaining of training samples comprises: Acquire a first business result and a first annotation result generated by the target business for the first candidate image, wherein the first business result includes location information of at least one subject obtained based on recognition of the first candidate image, and the first annotation result is used to annotate a location tag of the at least one subject; Based on the location information in the first business result and the location tag in the first annotation result, the first candidate images are screened to determine a first image; A business result and a labeling result generated by the target business on the first image are obtained as training samples.
8. The method according to claim 1, wherein: The obtaining of training samples comprises: Obtaining a category of at least one subject in the second candidate image; According to the first category requirement, the second candidate images are screened to determine the second image; A business result and a labeling result generated by the target business on the second image are obtained as training samples.
9. The method according to claim 1, wherein: The method further comprises: Acquire actual business data when the target business is running, the actual business data including at least one of the following data: a query image, first information indicating the interest tendency of the user; Determining sample offset data based on the difference between the actual business data and the training sample; Based on the sample offset data, a new training sample is constructed.
10. A large language model evaluation method, the method comprising: Acquire an execution result of the target business, where the execution result is generated by the target business based on the second largest language model for a subject in a query image of a user; The execution result is evaluated by a first large language model to obtain evaluation data for the execution result, wherein the first large language model is a large language model obtained after parameter tuning based on any one of the methods described in claims 1-9.
11. The method according to claim 10, wherein: The method further comprises: Acquire an intermediate execution result and category annotation data of the target business, wherein the intermediate execution result is a subject category in the query image obtained by the target business identification, and the category annotation data is used to annotate a category label of the subject of the query image; The intermediate execution result is evaluated by the third language model according to the category annotation data to obtain the accuracy of the intermediate execution result.
12. The method according to claim 11, wherein: The obtaining of the intermediate execution result and category annotation data of the target business includes: Acquire an evaluation benchmark data set, wherein the evaluation benchmark data set includes query images of multiple users and class annotation data corresponding to the query images; The target business generates an intermediate execution result for the subject in the query image in the evaluation benchmark data set.
13. The method according to claim 12, wherein: The obtaining of the evaluation benchmark data set includes: Acquire historical business data, where the historical business data includes a user's query image; According to the data screening requirements, the historical business data is screened to obtain target business data; The query image in the target business data is labeled to obtain category annotation data.
14. The method according to claim 13, wherein: The step of screening the historical business data according to the data screening requirements to obtain target business data includes: Obtaining the subject category of the query image in the historical business data; According to the second category requirement, the query image in the historical business data is screened to obtain the target business data.
15. The method according to claim 12, wherein: The obtaining of the evaluation benchmark data set includes: Acquire business data of a third-party platform, wherein the business data includes query images in actual usage scenarios of users; The query image in the business data is labeled to obtain category annotation data.
16. The method according to claim 12, wherein: The method further comprises: Based on the accuracy of the intermediate execution result, determining whether the query image corresponding to the intermediate execution result is an offset image; When it is determined that the query image is an offset image, a new evaluation benchmark data set is constructed based on the offset image.
17. A large language model parameter tuning device, wherein: The first language model is used to evaluate the execution performance of a target service, and the target service is used to generate a service result for a subject in a user's query image based on the second language model; the device comprises: A data acquisition module is configured to acquire a training sample, wherein the training sample includes the business result, a marking result and first information indicating the user's interest tendency, wherein the marking result is used to mark an evaluation label of the business result; A business evaluation module is configured to evaluate the business result based on the first language model to obtain an evaluation result for the training sample, wherein the evaluation result and the annotation result include a correlation between the business result and the first information; The parameter adjustment module is configured to adjust the parameters of the first language model based on the evaluation result and the annotation result.
18. A large language model evaluation device, the device comprising: A result acquisition module is configured to acquire an execution result of a target service, wherein the execution result is generated by the target service based on a second language model for a subject in a user's query image; The performance evaluation module is configured to evaluate the execution result by a first large language model to obtain evaluation data for the execution result, wherein the first large language model is a large language model obtained after parameter tuning based on any one of the methods described in claims 1-9.
19. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 16 is implemented.
Citation Information
Patent Citations
Model training method and device
CN113344201A
Model evaluation method and device, computer storage medium and electronic equipment
CN117909700A
Training sample generation method and device, electronic equipment and storage medium
CN118673334A
Answer quality evaluation method, related device, equipment and storage medium
CN119599021A
Evaluation method and device for retrieval enhancement generation application, equipment and medium
CN119719275A