Large Language Model Parameter Tuning Method, Evaluation Method and Device
By obtaining training samples and user interest information, and using the first largest language model for parameter tuning, the problems of uncertainty in the evaluation and rapid iteration of large language models are solved, and efficient and personalized evaluation and recommendation effects are achieved.
Patent Information
- Application Number
- CN202510436569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing large language models have high evaluation uncertainty in generative tasks and are difficult to adapt to rapid business iteration, low manual evaluation efficiency and high cost, and low accuracy of open source models is difficult to meet the needs of personalized evaluation.
By obtaining training samples, including business results, labeling results and user interest information, use the first largest language model for evaluation, adjust its parameters to improve evaluation performance, and capture user interests in combination with the space-time language model to build personalized recommended content.
The parameter tuning of the large language model is realized, the relevance and accuracy of the evaluation results are improved, the relevance and accuracy of the evaluation results are adapted to business changes, the subjectivity and cost of manual evaluation are reduced, and the personalization and logic of the generated content is improved.
Smart Images

Figure CN119938964B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology. Specifically, they relate to a method and apparatus for optimizing large language model parameters, as well as a method and apparatus for evaluating large language models. Background Art
[0002] In the field of artificial intelligence, large models usually refer to machine learning models with a large number of parameters (such as tens of billions or hundreds of billions of parameters) and complex structures. Through training on massive amounts of data, these models can complete a variety of complex tasks. Currently, many businesses are developed by integrating large model technologies, especially large language models (LLMs) and multimodal technologies, to meet various project requirements. For example, for generative tasks such as generating corresponding business content based on the subject in a query image. For each business, large models can also be used to evaluate whether the business meets the expected results.
[0003] However, the business results generated by such generative tasks have great uncertainty and complexity, and the business is updated and iterated quickly. The current large models used for evaluation are difficult to effectively evaluate such businesses. Summary of the Invention
[0004] The embodiments of this specification provide large language model parameter tuning and large language model evaluation solutions, which can improve the evaluation performance of large language models through parameter tuning.
[0005] In a first aspect, the embodiments of this specification provide a method for optimizing large language model parameters. Among them, a first large language model is used to evaluate the execution performance of a target business, and the target business is used to generate a business result based on a second large language model for the subject in a user's query image. The method includes: obtaining training samples, where the training samples include the business result, an annotation result, and first information used to indicate the user's interest tendency. The annotation result is used to annotate the evaluation label of the business result; the first large language model evaluates based on the business result to obtain an evaluation result for the training sample. The evaluation result and the annotation result include the correlation between the business result and the first information; based on the evaluation result and the annotation result, the parameters of the first large language model are adjusted.
[0006] In some embodiments, the training samples include query images.
[0007] In some embodiments, the evaluation result and the annotation result further include at least one of the following evaluation metrics: the correlation between the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; the information content of the business result.
[0008] In some embodiments, adjusting the parameters of the first large language model based on the evaluation result and the annotation result includes: determining whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; in the case of determining that the training sample is an offset sample, oversampling the offset sample to obtain an amplified offset sample; and adjusting the parameters of the first large language model based on the amplified offset sample.
[0009] In some embodiments, the method further includes: determining whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; in the case of determining that the training sample is an offset sample, based on the subjectivity of the offset sample and each evaluation metric, by adjusting the prompt words of the first large language model, evaluating again based on the business result to obtain a second evaluation result for the training sample; and adjusting the parameters of the first large language model based on the second evaluation result and the annotation result.
[0010] In some embodiments, determining whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result includes: in response to the degree of consistency between the evaluation metrics in the evaluation result and the evaluation metrics in the annotation result being lower than a preset threshold, determining the training sample as an offset sample.
[0011] In some embodiments, obtaining the training sample includes: obtaining a first business result and a first annotation result generated by the target business for a first alternative image, where the first business result includes position information of at least one subject recognized based on the first alternative image, and the first annotation result is used to label the position tags of the at least one subject; screening the first alternative image based on the position information in the first business result and the position tags in the first annotation result to determine a first image; and obtaining the business result and the annotation result generated by the target business for the first image as the training sample.
[0012] In some embodiments, obtaining the training sample includes: obtaining the categories of at least one subject in a second alternative image; screening the second alternative image according to the first category requirement to determine a second image; and obtaining the business result and the annotation result generated by the target business for the second image as the training sample.
[0013] In some embodiments, the method further includes: obtaining actual service data during the operation of the target service, where the actual service data includes at least one of the following data: a query image, and first information for indicating the user's interest tendency; determining sample offset data based on the difference between the actual service data and the training samples; and constructing new training samples based on the sample offset data.
[0014] In a second aspect, an embodiment of this specification provides a large language model evaluation method, where the method includes: obtaining an execution result of a target service, where the execution result is generated by the target service for the subject in a query image of a user based on a second large language model; and evaluating the execution result by a first large language model to obtain evaluation data for the execution result, where the first large language model is a large language model obtained by performing parameter tuning based on any one of the large language model parameter tuning methods in the first aspect.
[0015] In some embodiments, the method further includes: obtaining an intermediate execution result of the target service and class annotation data, where the intermediate execution result is the subject category of the query image identified by the target service, and the class annotation data is used to label the class tag of the subject of the query image; and evaluating the intermediate execution result by a third large language model according to the class annotation data to obtain the accuracy of the intermediate execution result.
[0016] In some embodiments, the obtaining of the intermediate execution result of the target service and the class annotation data includes: obtaining an evaluation benchmark data set, where the evaluation benchmark data set contains query images of multiple users and the class annotation data corresponding to the query images; and generating an intermediate execution result for the subject in the query images in the evaluation benchmark data set by the target service.
[0017] In some embodiments, the obtaining of the evaluation benchmark data set includes: obtaining historical service data, where the historical service data contains query images of users; screening the historical service data according to data screening requirements to obtain target service data; and labeling the query images in the target service data to obtain class annotation data.
[0018] In some embodiments, the screening of the historical service data according to data screening requirements to obtain target service data includes: obtaining the subject category of the query images in the historical service data; and screening the query images in the historical service data according to a second category requirement to obtain target service data.
[0019] In some embodiments, the obtaining of the evaluation benchmark dataset includes: obtaining service data of a third-party platform, where the service data includes query images in the actual usage scenarios of users; and annotating the query images in the service data to obtain class annotation data.
[0020] In some embodiments, the method further includes: determining whether the query image corresponding to the intermediate execution result is an offset image based on the accuracy of the intermediate execution result; and constructing a new evaluation benchmark dataset based on the offset image when it is determined that the query image is an offset image.
[0021] In a third aspect, an embodiment of this specification provides a large language model parameter tuning device. A first large language model is used to evaluate the execution performance of a target service, and the target service is used to generate a service result for the main body in a user's query image based on a second large language model. The device includes: a data acquisition module configured to acquire training samples, where the training samples include the service result, an annotation result, and first information for indicating the user's interest tendency, and the annotation result is used to annotate the evaluation label of the service result; a service evaluation module configured to evaluate the service result by the first large language model to obtain an evaluation result for the training samples, where the evaluation result and the annotation result include the relevance between the service result and the first information; and a parameter adjustment module configured to adjust the parameters of the first large language model based on the evaluation result and the annotation result.
[0022] In a fourth aspect, an embodiment of this specification provides a large language model evaluation device. The device includes: a result acquisition module configured to acquire the execution result of a target service, where the execution result is generated by the target service for the main body in a user's query image based on a second large language model; and a performance evaluation module configured to evaluate the execution result by a first large language model to obtain evaluation data for the execution result, and the first large language model is a large language model obtained by tuning parameters according to any one of the large language model parameter tuning methods in the first aspect.
[0023] In a fifth aspect, an embodiment of this specification provides a computing device including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method described in any implementation manner of the first aspect and the second aspect is implemented.
[0024] In the solution provided by the above embodiments of this specification, it is possible to implement the evaluation of image-based content generation tasks. Among them, by evaluating the business results of training samples, an evaluation result is obtained. Then, based on the evaluation result and the annotation result, the parameters of the first large language model are tuned to help the first large language model adapt to different business changes, and further, the generation performance of the second large language model under different conditions can be tested, improving the evaluation effect of the first large language model. In addition, when evaluating, the relevance between the business result and the first information used to indicate the user's interest tendency is also considered, ensuring that the evaluation result is consistent with the user's subjective feeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the multiple embodiments disclosed in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a schematic diagram of the parameter tuning process of a large language model in an embodiment of this specification;
[0027] Figure 2 It is a schematic diagram of the implementation process of a target business in an embodiment of this specification;
[0028] Figure 3 It is the business result of a target business in an embodiment of this specification;
[0029] Figure 4 It is a flowchart of the method for tuning the parameters of a large language model in an embodiment of this specification;
[0030] Figure 5 It is a schematic diagram of the scores of the consistency degree in an embodiment of this specification;
[0031] Figure 6 It is a schematic diagram of the acquisition and processing path of an offset sample in an embodiment of this specification;
[0032] Figure 7 It is a schematic diagram of the process of the method for evaluating a large language model in an embodiment of this specification;
[0033] Figure 8 It is a schematic diagram of the construction process of an evaluation benchmark data set in an embodiment of this specification;
[0034] Figure 9 It is a schematic diagram of the structure of the device for tuning the parameters of a large language model in an embodiment of this specification;
[0035] Figure 10It is a schematic structural diagram of the large language model evaluation device in the embodiments of this specification. Detailed implementation manners
[0036] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0037] As mentioned above, with the rapid development of large language model technology, the use of large language models for content generation in related businesses has achieved extensive applications in various industries. In order to verify whether the businesses related to the content generated by large language models meet the expected effects, the generated content can be evaluated. However, the generated content of large language models often has divergence and instability. Especially when generating content based on the main body in an image, due to factors such as the richness of visual information in the image, the diversity of scenes in the image, and the possible multi-modal fusion of visual information and text information, the evaluation of the results of such businesses requires comprehensive consideration of a large number of uncertainties and complex factors. Currently, there are methods of manual evaluation and methods of using open-source large language models for evaluation.
[0038] However, the method of manual evaluation has the following problems: First, the efficiency of relying on manual methods for evaluation is extremely low and the cost is extremely high. Second, manual evaluation is highly subjective because its results are easily affected by personal experience, knowledge, emotions, and preferences. The method of using open-source large language models for evaluation has the following problems: First, the accuracy of open-source large models is low, and the evaluation effect cannot be guaranteed, and manual intervention is still required later. Second, the large model technology develops rapidly, resulting in a relatively fast iteration speed of the target business itself. However, it is difficult for open-source large models and the benchmark datasets they use to quickly iterate the evaluation system in coordination with the model updates of the business.
[0039] To solve the above technical problems, the embodiments of this specification provide a large language model parameter tuning method and a large language model evaluation method. Figure 1 It shows a schematic diagram of a large language model parameter tuning process according to the embodiments of this specification. As Figure 1As shown, in some embodiments, training samples may be obtained first. The training samples include: the business results of the target business, annotation results, and first information indicating the user's interest tendency. Then, according to the instructions of the prompt words, the first large language model evaluates based on the business results to obtain an evaluation result for the training samples. The evaluation metrics in the evaluation result and the annotation result include the relevance between the business result and the first information. Then, based on the evaluation result and the annotation result, the parameters of the first large language model are adjusted. In another embodiment, the training samples further include: query images. Among them, by inputting the query images and the first information indicating the user's interest tendency into the target business, the business results generated by the target business for the subjects in the user's query images based on the second large language model can be obtained, and the annotation results can be obtained by annotating the business results.
[0040] First, the target business to be evaluated will be described below.
[0041] In the embodiments of this specification, the target business is used to generate business results for the subjects in the user's query images based on the second large language model to satisfy the user's curiosity or inspire the user's inspiration and provide personalized recommendations to the user. For example, when the user takes a photo of the surrounding environment to obtain a query image and uploads it to the target business, the target business can identify at least one subject from the query image, obtain the first information indicating the user's interest tendency, then determine the target subject among the at least one subject based on the first information, and then generate the business results for the user based on the target subject through the second large language model.
[0042] Among them, the business result can be recommended content for satisfying the user's query needs for the subject. Specifically, the business result can be generated content in different modalities. For example, it can be one or several of the text modality, image modality, video modality, and audio modality. Exemplarily, when the query image is an image obtained by the user taking a photo of the desktop, the tea cup, paper, pen, laptop computer, tea leaves in the tea cup, and the trousers shown by the user on the desktop in the image can all be considered as subjects. The generated business result can be generated content in the text modality, such as: "What kind of tea is suitable for this tea cup?"
[0043] In practice, the first information can be topics or services that the user is interested in. Specifically, it can be topics or online services related to people's livelihood and government affairs, commodity information, culture and technology fields, historical and geographical knowledge, and current affairs news. Exemplarily, the first information can be a football game that the user cares about, traditional culture that the user is interested in, the provident fund withdrawal service, etc. In the case where there are multiple identified entities, the target entity can be determined from them according to the user interest tendency indicated by the first information, that is, the target entity is the entity that the user is interested in. For example, still taking the desktop objects as an example, in the case where there are multiple entities such as a teacup, paper, a pen, a laptop, the tea leaves in the teacup, and the user's exposed pants on the desktop, when the user's first information indicates that the user has an interest in Chinese traditional culture, since the teacup and the tea leaves are not only daily necessities but also semantically related to Chinese traditional culture and carry rich cultural connotations and symbolic meanings, the determined target entities can be the teacup and the tea leaves.
[0044] The prompt constructed based on the entity and the user can be input into the second large language model so that the second large language model generates corresponding business results. Among them, the prompt contains the information of the entity and can also contain information related to the user, so that the second large language model generates recommended content that the user may be interested in, and the recommended content is semantically related to the entity. The recommended content can be personalized text information, commodities, services, or other resources recommended to the user. For example, when the entity in the query image is a computer host, the prompt constructed by the user can be "Based on the entity, propose questions that the user may be interested in". The business results generated by the second large language model according to the prompt can be text information such as "How can the computer host of brand A improve work efficiency?" and "How is the cooling system of the computer host of brand A designed?". When the user clicks on this text information, this text information can be input into the second large language model or other large language models to obtain more relevant text or picture content generated by the large language model to answer relevant questions and satisfy the user's curiosity.
[0045] The embodiments of this specification do not limit the acquisition method of the user's first information. For example, it can be analyzed based on the user's historical behavior data to obtain the user's first information, or it can be analyzed based on the attribute information uploaded by the user himself to obtain the user's first information. Taking the user's bill record as an example of historical behavior data, items such as running shoes, knee pads, and sports drinks appear multiple times in the bill record. It can be analyzed that the user's first information is a topic word related to sports and fitness, indicating the user's interest tendency in sports and fitness.
[0046] As an implementation method, in order to accurately obtain the user's first information, it can be to first obtain the user's associated information, and based on the associated information, generate the user's first information through the second large language model.
[0047] Among them, the associated information of the user is the user's personal information related to the user. For example, the user's occupation, the user's population portrait, the user's historical behavior, etc. Exemplarily, the associated information of the user may include at least one of the following information: the user's attribute information, the user's historical behavior data, the user's location information, and the user's current time information.
[0048] Specifically, the user's attribute information is used to describe the user's basic situation, interest preferences, behavior habits, etc., such as educational level, social role, exercise habits, etc., to better model the user portrait; the user's historical behavior data is used to describe the user's online behavior. For example, it may include the user's bill records, records of using mini-programs, search history records, and historical records of accessing specific points of the client at a specific time, so as to obtain the continuous behavior of the user in different space-time scenarios to better analyze the user's behavior pattern; the user's location information can be the geographical location information of the user when taking or uploading the query image, or the shooting location information carried by the query image itself, so as to better predict the behavior pattern through the location information; the user's current time information can be the time when the user takes the query image, or the time when the user uploads the query image to the client, so as to better predict the user's behavior pattern through the time information.
[0049] The second large language model used here is an artificial intelligence model based on deep learning, which is specifically used to understand and generate natural language. It can be obtained by fine-tuning an open-source large language model, or can be trained, so that the second large language model learns the deep association between the user's associated information and the first information through supervised training.
[0050] In one embodiment, for the dynamic needs of the user, the second large language model used can be a spatio-temporal large language model (Spatial Temporal-LLM, abbreviated as ST-LLM), which is trained or fine-tuned using a large amount of spatio-temporal data (associated information). It has a certain spatio-temporal prediction ability and can capture the specific needs of the user in different scenarios. The recommended content generated is more in line with the user's current needs.
[0051] In practice, it can be to construct prompt information based on the associated information, input the prompt information into the spatio-temporal large language model, and obtain the first information of the user output by the spatio-temporal large language model. Exemplarily, the constructed prompt information can be: According to the user's portrait information {p} and recent purchase history {h}, combined with the current time {t} and location information {l}, predict what kind of service {f} the user needs, so as to obtain the service prediction {f} output by the model.
[0052] Then, combined with Figure 2Describe the implementation process of the target business in more detail.
[0053] As Figure 2 shown, the target business can be a recommendation business based on a spatio-temporal large language model. The business mainly includes four modules: an input module for obtaining the input data required during the execution of the target business, such as associated information like the user's historical behavior data and attribute information, as well as the user's query image; a spatio-temporal guidance module for fine-tuning the spatio-temporal large language model; a preference discovery module, which consists of Step 1, Step 2, Step 3, and Step 4. Here, the serial numbers of the steps are only used to identify different steps and do not limit the execution order of the steps.
[0054] Among them, in Step 1: Generate a scene graph based on the query image. The scene graph contains at least one recognized subject and the relationships between the subjects. The scene graph can be represented by a set of triples, and each triple is used to represent the relationship between two subjects. Exemplarily, the subjects included in this scene graph are: a teacup, a piece of paper, a pen, a laptop, the tea leaves in the teacup, and a pair of trousers. The arrows indicate the relationships between two subjects. For example, the tea leaves are in the teacup, and the pen is on the paper.
[0055] In Step 2: Generate a knowledge graph based on the user's attribute information and the scene graph. The knowledge graph can also be represented by a set of triples, and each triple represents the association relationship between a subject and a vocabulary. Multiple triples form a multi-hop chain relationship. Specifically, it can be based on a preset corpus to retrieve the associated information and / or the context information of at least one subject, and then through a large model, based on the associated information, at least one subject, and the context information, generate the vocabulary that has an association relationship with the subject. The preset corpus can be various open knowledge bases or private corpora. The vocabulary can be regarded as a knowledge enhancement based on the scene graph to better utilize the large language model as a user agent to simulate the user's preferences and interests. Taking the Figure 2 knowledge graph in it as an example, among them, the vocabulary in the blue module is expanded from the subject, and the vocabulary in the yellow module is expanded from the associated information.
[0056] In Step 3: Generate the first information according to the user's attribute information and historical behavior data through a pre-trained spatio-temporal large language model. As Figure 2 shown in Step 3 of the preference discovery module in
[0057] In step 4: Retrieve and rank in the knowledge graph according to the first information to determine the target entity and the vocabulary associated with the target entity. Specifically, based on the first information, among at least one entity and the vocabulary associated with the entity, determine the target entity and the vocabulary associated with the target entity. In practice, it can be to calculate the semantic similarity between the first information and each entity-vocabulary pair (i.e., the entity and the vocabulary associated with the entity), and then determine one or several entity-vocabulary pairs with the highest similarity as the target entity and the vocabulary associated with the target entity, so as to find the target entity and vocabulary that are closest to the user's interest tendency. Exemplarily, in order to further construct a personalized knowledge graph, a hop-by-hop semantic indexing can be performed on the knowledge graph to effectively retrieve relevant information and efficiently capture the first-order neighbors with clear semantic associations, which serves as a new basis for knowledge retrieval, thereby enhancing the knowledge matching ability of user interests. For example, a pre-trained language model (such as a certain generative encoder) can be introduced to semantically encode the first information and the triples in the knowledge graph respectively to obtain the first semantic vector and the second semantic vector, so as to capture their potential semantic information. In the knowledge graph, since the relationships between multiple entities and vocabularies are connected by a chain relationship composed of multiple triples, the semantic similarity between the first semantic vector corresponding to the first information and the second semantic vector corresponding to the triples can be calculated hop by hop.
[0058] When calculating hop by hop, it can be to first calculate the similarity between the first semantic vector and the second semantic vector corresponding to the first hop. If the similarity is higher than the set threshold, the triple is regarded as highly relevant and retained; otherwise, continue to propagate backward along the graph structure and recursively calculate the similarity of the next-hop adjacent triples to ensure the expandability of the retrieval scope. For example, calculate the similarity between the first semantic vector and the second semantic vector corresponding to the second hop, and determine whether it is higher than the set threshold, and so on until the entire knowledge graph is traversed.
[0059] Finally, after completing the hop-by-hop indexing, filter the retained association relationships and rank the filtered association relationships according to the semantic similarity. For example, select the top 5 association relationships, and use the entities and vocabularies associated with the association relationships as the target entity and the vocabulary associated with the target entity for subsequent personalized recommendation or knowledge enhancement tasks.
[0060] Such as Figure 2As shown by the reordering of step 4, by comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association relationship between tea and beverages, the similarity score of the association relationship from tea to beverages is 0.9. By comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association relationship between tea and cultural experience, the similarity score of the association relationship from tea to cultural experience is 0.7. The higher the similarity score, the closer the two are semantically. It can be seen that the user has a high interest in beverages. Generating questions based on this can meet the user's needs. If only analyzing based on the target image without combining the user's interests, it is very difficult to lead to the concept of beverages. Through the above method, structured knowledge can be effectively extracted from complex visual content, and a personalized knowledge graph can be generated according to the user's preferences, thereby improving the accuracy and relevance of information retrieval and recommendation.
[0061] The personalized recommendation module is used to use the spatio-temporal large language model to generate at least one piece of recommendation content for the user, that is, the business result, based on the target entity and the vocabulary associated with the target entity. In practice, based on the target entity and the vocabulary associated with the target entity, a prompt can be constructed and input into the large language model, and the large language model can output the corresponding recommendation content. The prompt can include the user's attribute information and the user's historical behavior data to provide rich background information to the large model. The user's location information and the user's current time information can also be included to assist the large model in capturing the complex and real dynamic needs of the user changing with time and space. Exemplary recommendation content can be "What kind of tea is suitable for this teacup?", "What are the types and materials of traditional Chinese teacups?", and "What are the types of tea in China?", thus arousing the user's interest.
[0062] As Figure 3 shown, Figure 3 shows a business result of an exemplary target business provided in this specification, which is displayed on the display interface of the user terminal. The display interface includes a text section and a view section. Among them, the text content in the text section is the business result generated by the second large language model, and the content in the view section is the image of the subject in the query image. Exemplarily, the generation process of this business result can be: the user takes a picture of the sausage drying on the balcony and uploads it to the target business for query. After the target business obtains the query image, it identifies the subject "sausage" from the query image, and then constructs a prompt "Based on the sausage in the image, propose questions that the user may be interested in" and inputs it into the second large language model. The second large language model generates: "What are the different versions of sausage in the world?", "How is sausage made?", "Why has sausage become one of the classic foods?", etc.
[0063] In the process of gradually implementing the requirements of the target business, powerful indicator data is needed to evaluate whether the target business achieves the expected results. The embodiments of this specification divide the evaluation of the execution performance of the target business into two modules. Among them, the understanding module mainly ensures the business rationality of the business results, and the intention module mainly ensures the accuracy of the identification of the main categories in the view section. In the understanding module, the first large language model is used to evaluate the business results of the target business. The following describes the detailed process of tuning the parameters of this first large language model.
[0064] Figure 4 FIG. shows a schematic flowchart of a method for tuning the parameters of a large language model according to an embodiment of the present disclosure. This parameter tuning process can be executed by any device, platform or cluster of devices with computing and processing capabilities, including steps S401 - S403 shown below.
[0065] As Figure 4 shown, in step S401, training samples are obtained.
[0066] Among them, each training sample includes a business result and a corresponding annotation result, as well as first information for indicating the user's interest tendency in the training sample. The business result is the result output by the above-mentioned target business, and the annotation result is used to annotate the evaluation label of the business result. The annotation result can be the result of manually scoring and annotating the business result according to the evaluation criteria. It can be understood that this evaluation criteria can also be used as the evaluation criteria of the first large language model to evaluate the business result to obtain an evaluation result. The first information can be the information generated during the implementation process of the target business described in the foregoing embodiments.
[0067] In this embodiment, the evaluation indicators of the evaluation result and the annotation result can both include the relevance between the business result and the first information, so as to evaluate whether the content output by the target business fits the user's personal intention. By attaching importance to this evaluation indicator, the situation where the user is not interested in the business result can be avoided.
[0068] In different embodiments, different evaluation criteria can be used. In one example, in order to avoid incorrect knowledge information and illegal guidance in business results, three evaluation indicators strongly related to the business, namely relevance, correctness, and logic, can be set in the evaluation criteria. In another example, a series of standards and guidelines can be used to constrain the evaluation of the model, so as to comprehensively consider various uncertainties and complex factors, and reduce evaluation bias. Among them, in order to evaluate the effectiveness of the target business in generating personalized recommendation content, the evaluation criteria not only need to capture the relevance between the recommended content and the visual query, but also need to measure its degree of fit with the user's interests. The evaluation results and the annotation results can also include at least one of the following evaluation indicators: the relevance between the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; the information content of the business result; the relevance between the business result and the first information. It can be understood that the specific details of the evaluation indicators are related to the content of the business result.
[0069] Exemplarily, in the case where the business result is a recommended content (such as a question) in text modality, the relevance between the semantic subject in the business result and the subject in the query image can be regarded as the relevance between text and image, which is used to evaluate whether the output recommended content is relevant to the subject of the visual query; the correctness of the business result is the correctness of the recommended content, which is used to evaluate whether the output recommended content contains incorrect guidance that violates compliance; the logic of the business result is the logic of the question, which is used to evaluate whether the output question is logical and easy to understand; the information content of the business result can be regarded as the attractiveness of the question, which is used to evaluate whether the output question provides new viewpoints or involves content not well-known to most people. Generally speaking, the greater the information content of the question, the stronger the attractiveness of the question; the relevance between the business result and the first information, that is, the personalization of the question, is used to evaluate whether the output question is highly relevant to the user's personal intentions and interests.
[0070] In different embodiments, the scoring criteria for the evaluation indicators can be different. In one embodiment, three levels of high, medium, and low (1 / 0 / -1) can be set for manual evaluation. Among them, the higher the level or the higher the score, the better the execution performance of the target business. In addition, considering that subjective evaluation is easily affected by personal experience, knowledge, emotion, and preference, in order to reduce the bias that may be brought by subjectivity, objective data can also be combined or the opinions of multiple evaluators can be used for comprehensive analysis. For example, in the evaluation process, multiple groups of annotators with similar education levels can be used for evaluation to ensure the effectiveness of the results. For example, when the business result is a question, the annotation result can be "Question angle diversity: 1 point, question logic: 1 point, question personalization: -1 point, question information content: 1 point, content correctness: 1 point, text-image relevance: 1 point".
[0071] In practice, the training samples can be obtained from a pre-prepared training sample set or from other specified locations.
[0072] In one embodiment, in order to improve the parameter tuning effect of the first large language model and reduce the annotation cost, based on the idea of active learning, this embodiment also constructs a high-quality training sample set to assist the continuous fine-tuning and iteration of the first large language model. The most valuable data is selected from the training sample set to improve the model's performance and learning efficiency, and reduce the demand for annotated data. The following will exemplify several ways to obtain training samples in the training sample set. It can be understood that the following exemplified ways to obtain training samples can be used separately or in combination.
[0073] In one example, in order to meet the data validity of the training sample set, the data can be strictly filtered. When obtaining training samples, the first business result and the first annotation result generated by the target business for the first alternative image can be obtained. The first business result includes the position information of at least one subject recognized based on the first alternative image, and the first annotation result is used to label the position tags of at least one subject; based on the position information in the first business result and the position tags in the first annotation result, the first alternative image is screened to determine the first image; the business result and the annotation result generated by the target business for the first image are obtained as training samples.
[0074] Considering that the business result of the target business is generated based on the subject in the user's query image, if the subject in the query image is detected incorrectly, the subject is not clear, or the user's query intention is not clear, it will affect the generation quality of the business result, and evaluating such business results cannot reflect the true execution performance of the target business. Such business results can be considered invalid data. In this example, the above invalid data can be filtered by verifying the position accuracy of the subject detected in the target business.
[0075] Specifically, the first alternative image can be a query image in the historical business data of the target business. When the target business identifies the subject in the first alternative image, a first business result will be generated. The first business result includes the position information of at least one subject identified based on the first alternative image. For example, the position box where the subject is located. By manually annotating the first alternative image, a first annotation result can be obtained, that is, the true annotation box (Ground Truth, abbreviated as GT). Then, the degree of regional overlap between the position box and the annotation box can be measured by IoU (Intersection over Union) or other screening strategies. If the IoU value is greater than a certain threshold (such as 0.5), it is considered that the detection is successful, and the first alternative image is retained and determined as the first image. Otherwise, it means that the target business fails to successfully detect the subject, and the first alternative image is filtered out to ensure that samples with unclear subjects and unclear intentions do not flow into the training samples.
[0076] In another example, to meet the data diversity of the training sample set, when obtaining training samples, it can be to obtain the categories of at least one subject in the second alternative image; according to the first category requirement, screen the second alternative image to determine the second image; obtain the business result and annotation result generated by the target business for the second image as training samples.
[0077] Among them, the second alternative image can be an image collected from historical business data, various online platforms, or manual shooting, etc. By identifying the subjects in the second alternative image, the categories of different subjects in the second alternative image can be obtained. Exemplarily, the category of the subject in the embodiments of this specification can refer to the secondary category. When the primary category is electronic products, the secondary categories can be mobile phones, televisions, headphones, computers, etc. In other embodiments, the category of the subject can also refer to other levels of categories. The first category requirement can be the requirement for the richness of various categories in the category. For example, it can be the requirement for the proportion of different categories. The second alternative image that meets the first category requirement is determined as the second image, and the business result generated by the target business based on this second image is manually labeled to obtain the corresponding annotation result, thereby obtaining training samples. By sampling different subject categories, the category richness is ensured, so that the training samples provide sufficient differential information, and further ensure the comprehensive evaluation performance of the first large language model.
[0078] In another example, considering that the business iteration speed of the target business is very fast, in order to enable the first large language model to adjust its parameters in a timely manner as the business changes, it is also possible to obtain the actual business data during the operation of the target business. The actual business data includes at least one of the following data: query images, and first information indicating the user's interest tendency; determining sample offset data based on the difference between the actual business data and the training samples; constructing new training samples based on the sample offset data.
[0079] In practice, the actual business data generated during the operation of the target business includes query images collected by different users in different scenarios, as well as the first information of different users. Since the current social development is changing rapidly and many new things are emerging every day, the categories of the subjects in the query images and the interests of users are constantly updated. By comparing the differences between the actual business data and the training samples in the training sample set, the sample offset data after the business change can be determined. The sample offset data contains query images and / or first information that did not appear in the training samples. By constructing new training samples based on the sample offset data, the training sample set can be quickly updated as the business iterates of the target business, thereby improving the evaluation performance of the first large language model for newly emerging business data.
[0080] Next, in step S402, the first large language model evaluates based on the business results to obtain an evaluation result for the training samples.
[0081] In practice, a prompt can be constructed and input into the first large language model to prompt the first large language model to evaluate the business results and obtain an evaluation result. A prompt refers to the information input into the first large language model, which can include different modalities of information such as text and images. The purpose is to guide the first large language model to evaluate the business results according to the evaluation criteria or requirements in the prompt and generate corresponding evaluation results. The first large language model can be an open-source large model or a pre-trained large model. In different embodiments, the first large language model can be different specific types or large models with different neural network structures, and the prompt can also be different specific prompts. This specification does not limit this.
[0082] In different embodiments, the specific scoring criteria indicated by the prompt can be different. In one embodiment, when the business result is a question generated for the subject in the query image, the evaluation index in the scoring criteria can be the logic of the business result. The prompt can be "Please score the logic of question q. 5 points: correct logic, precise and fluent expression; 4 points: basically correct logic, correct expression; 3 points: basically correct logic; 2 points: logical loopholes, incorrect results in part; 1 point: obvious logical loopholes, misunderstanding in part of the content; 0 points: false content".
[0083] The evaluation indicators in the scoring criteria can also be the relevance between the semantic subject in the business result and the subject in the query image; the correctness of the business result; the logic of the business result; the information content of the business result, and the relevance between the business result and the first information. When the business result is a question generated for the subject in the query image, the prompt A can be "The question content is: q,
[0084] Please only score two related items for the semantic subject in the question (only consider the semantic subject) and the subject in the figure. The score range is [-1, 0, 1]. The specific meanings of each range are
[0085] -1: The semantic subject of the question (only consider the semantic subject) has no relation to the subject in the figure.
[0086] 0: The corresponding subject of the question (only consider the semantic subject) is too broad.
[0087] 1: The subject of the question (only consider the semantic subject) is clearly pointed and consistent with the subject in the figure.
[0088] Please score according to the score range rules. Note that only output the score and put it in the first element of the list;
[0089] Please score the content correctness of the question. The score range is [-1, 0, 1]. The specific meanings of each range are
[0090] -1: The content is wrong or the content has a misleading cost.
[0091] 0: The question is normal, but the knowledge is not clear.
[0092] 1: The question has a certain amount of knowledge.
[0093] Please score according to the score range rules. Note that only output the score and put it in the second element of the list;
[0094] Please score the logic of the question. The score range is [-1, 1]. The specific meanings of each range are
[0095] -1: The question description has no logic and is not easy to understand; or there are pronoun nouns in the question.
[0096] 1: The question logic is clear.
[0097] Please score according to the score range rules. Note that only output the score and put it in the third element of the list;
[0098] Please score the information content of the question. The score range is [-1, 0, 1]. The specific meanings of each range are
[0099] -1: The answer to the question is too simple (visible to the naked eye, hardly any thinking required).
[0100] 0: The question is mediocre and the answer is known to most people (common sense but requires some thinking).
[0101] 1: The question presents a new perspective or involves content not well-known to most people
[0102] Please score according to the score range rules. Note that only the score needs to be output and placed in the fourth element of the list;
[0103] User's personal interest keyword (i.e., the first information): 'usr_interest'. Please score the relevance between any item of the question and the user's interest. The score range is [-1, 1]. The specific meanings of each score range are
[0104] -1: Irrelevant to the user's personal interest;
[0105] 1: Relevant to the user's personal interest and has a sense of freshness and exploration desire
[0106] Please score according to the score range rules. Note that only the score needs to be output and placed in the fifth element of the list;
[0107] Finally, output the result list. Note that only the list format needs to be output. If there are other problems, output [0, 0, 0, 0, 0].
[0108] In the above embodiments, the prompt words include the business result and the first information. In other embodiments, the prompt words can also include the query image to assist the first large language model in better evaluation. After inputting the prompt words into the first large language model, the evaluation results can include the scores for various evaluation indicators.
[0109] It can be understood that the business result is not limited to the question, but can also be knowledge content, emotion-related content, etc. For other types of business results other than the question, the prompt words can be generated in a similar way as in the above examples, and this embodiment will not elaborate further.
[0110] Next, in step S403, based on the evaluation result and the annotation result, the parameters of the first large language model are adjusted.
[0111] In practice, there will be differences between the evaluation results output by the first large language model and the annotation results manually marked. For example, for a certain evaluation metric, the score in the evaluation result is low, while the score in the annotation result is high. By adjusting the parameters of the first large language model, this difference can be gradually reduced, making the evaluation result closer and closer to the annotation result, thereby continuously improving the evaluation performance of the first large language model. For example, the loss function can be calculated based on the difference between the evaluation result and the annotation result, and then the parameters of the first large language model can be adjusted with the goal of minimizing the loss value of the loss function. When adjusting the parameters, it can be a fine-tuning of some of the parameters, or an adjustment of all the parameters, or the optimization of the prompt words through the way of P-Tuning (Prompt Tuning).
[0112] In one embodiment, for the part where the difference between the evaluation result predicted by the first large language model and the manual annotation result is large, expert knowledge can also be introduced for secondary calibration and used as difficult samples (Hard Case) to strengthen the training of the first large language model.
[0113] In one example, it can be to first determine whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result, then in the case where the training sample is determined to be an offset sample, upsample the offset sample to obtain an amplified offset sample, and finally adjust the parameters of the first large language model based on the amplified offset sample.
[0114] In practice, the difference between the evaluation result and the annotation result can be measured first, and different methods can be used to measure it in different embodiments. For example, the loss function can be used to calculate the difference between the two. The larger the loss value of the loss function, the greater the difference. When the loss value exceeds the preset threshold A, the training sample is determined to be an offset sample; it can also be to calculate the similarity between the two. The lower the similarity, the greater the difference. When the similarity is lower than the preset threshold B, the training sample is determined to be an offset sample.
[0115] As an implementation method, it can be to determine the training sample as an offset sample in response to the consistency degree between the evaluation metrics in the evaluation result and the evaluation metrics in the annotation result being lower than the preset threshold. For example, the Pearson correlation coefficient, Spearman rank correlation coefficient, or Cohen's Kappa coefficient can be used to calculate the consistency degree of each evaluation metric in the evaluation result and the annotation result, and the exact match degree can also be used to calculate the consistency degree between the two, that is, calculate the matching degree of the content and format in the evaluation result and the annotation result. If the consistency degree between them is lower than the preset threshold, it means that there is a large difference between the evaluation result and the annotation result, and it is determined as an offset sample. Exemplarily, Figure 5A schematic diagram showing a degree of consistency is presented. For an exact match, the closer the score is to 1, the higher the degree of consistency. For the Kappa coefficient, when the score is greater than a certain value, it is considered that there is a high degree of consistency between the two.
[0116] After determining the offset samples, expert experience can be involved to judge the reasons for the huge differences in these offset samples and take corresponding measures. For example, if it is caused by incorrect manual annotation results, the offset samples can be re-annotated. If it is caused by incorrect model prediction, the offset samples can be used as hard cases to recalibrate the model.
[0117] In this example, when the offset samples are hard cases, the offset samples can be upsampled to obtain amplified offset samples. This embodiment does not limit the specific upsampling method. For example, the offset samples can be replicated using the simple oversampling method to increase their quantity in the training sample set, or the synthetic minority over-sampling technique (SMOTE) can be used to generate new samples by interpolating between the training samples, thereby achieving the amplification of the offset samples and increasing their weights in the training sample set. Finally, steps S402 and S403 can be executed again according to the amplified offset samples to adjust the parameters of the first large language model. The adjusted model can achieve better evaluation results when facing hard cases.
[0118] In another example, when it is determined that the training samples are offset samples, based on the subjectivity of the offset samples and each evaluation metric, by adjusting the prompt words of the first large language model, and then evaluating again based on the business results to obtain the second evaluation result for the training samples. Finally, based on the second evaluation result and the annotation result, the parameters of the first large language model are adjusted.
[0119] In actual implementation, since the scores of each evaluation metric in the annotation result are greatly affected by the subjectivity of the annotators, resulting in a large difference between the evaluation results predicted by the model and the annotation result. In view of this situation, this embodiment considers adjusting the prompt words input to the model according to the subjectivity of each evaluation metric to make it closer to human subjective thinking, thereby narrowing the difference between the annotation result and the evaluation result and improving the evaluation effect of the model.
[0120] Exemplarily, when the annotator is scoring for different evaluation metrics, the ranking of the subjectivity degrees of the various evaluation metrics in the annotator's evaluation logic can be: attractiveness > logic > content correctness > subject relevance > interest relevance. For example, it can be stated in the prompt "Please subjectively score the above evaluation metrics according to the subjectivity order of attractiveness > logic > content correctness > subject relevance > interest relevance", or the scoring criteria of the prompt can be directly optimized to be closer to human subjective thinking. For example, the prompt B obtained by optimizing the prompt A in the above embodiment can be "The question content is: q. Regarding the relevance between the question and the main figure, please score according to the following steps and only output the score in the end:
[0121] 1. Identify the semantic subject from the question text, that is, the main object emphasized or concerned.
[0122] 2. Observe and identify the subject in the image, and clarify the main object shown in the image.
[0123] 3. Compare the semantic subject in the question with the subject in the image:
[0124] If they are not relevant, score -1.
[0125] If the semantic subject of the question is too broad to match, score 0.
[0126] If the semantic subject is clear and the same as the subject in the figure, score 1.
[0127] If there is uncertainty, directly score 0, and output the scoring result in the form of a list, with the score placed in the first element of the list.
[0128] Regarding the question correctness, please analyze the content of the following question in detail and only output the score in the end:
[0129] 1. Ensure a full understanding of the question content.
[0130] 2. Judge whether there are factual errors in the content:
[0131] If there are obvious errors and may lead to wrong conclusions, give -1 point.
[0132] 3. Judge whether the question has knowledgeability:
[0133] If the question does not have clear knowledgeability, give 0 points.
[0134] If the question can stimulate knowledge discussion and exploration, give 1 point.
[0135] 4. If there is uncertainty, simply assign a score of 0, and output the scoring results in a list. The score should be placed in the second element of the list.
[0136] Please rate the logic of the question. Analyze the content of the following question in detail and finally only output the score:
[0137] 1. Ensure a full understanding of the question content.
[0138] 2. Determine whether there are factual errors in the content:
[0139] If there are obvious errors and they may lead to wrong conclusions, give -1 point.
[0140] 3. Determine whether the question has knowledge content:
[0141] If the question does not have clear knowledge content, give 0 points.
[0142] If the question can stimulate knowledge discussion and exploration, give 1 point.
[0143] 4. If there is uncertainty, simply assign a score of 0. Based on the analysis, obtain the final score, which should be placed in the third element of the list.
[0144] Please rate the attractiveness of the question, clarify the scoring criteria, and finally only output the score:
[0145] -1 means the answer to the question is obvious and can be obtained without thinking;
[0146] 0 means the answer to the question is common sense but requires brief thinking;
[0147] 1 means the question introduces new viewpoints or involves little-known knowledge.
[0148] Question analysis: Is the question simple? Is the answer obvious?
[0149] Degree of common sense: Is the question known to most people?
[0150] Novelty: Does the question put forward new viewpoints or uncommon knowledge?
[0151] Select the score: Based on the analysis of the question, select the corresponding score. If there is uncertainty, simply assign a score of 0, and the score should be placed in the fourth element of the list;
[0152] Regarding the user's personal interest, finally only output the score. Input:
[0153] User's personal interest keywords: 'usr_interest'
[0154] Question to be scored: q
[0155] Output: Output the associated score as the fifth element of the list.
[0156] Execution steps:
[0157] 1. Extract the user's personal interest keywords.
[0158] 2. Obtain the text of the question to be evaluated.
[0159] 3. Analyze the relevance between the question and the user's interests:
[0160] Find the occurrences of the interest keywords in the question.
[0161] Judge whether the content of the question may generate a sense of freshness or exploration desire for the user.
[0162] 4. Score according to the following criteria:
[0163] -1: The question has no relevance to the user's interests.
[0164] 1: The question is relevant to the user's interests and may stimulate the desire for exploration.
[0165] 5. If there is uncertainty, directly score 0, store and output the above score as the fifth element of the list
[0166] Note: The final output result is a list. Note that only the list format is output. If there are other problems, output [0, 0, 0, 0, 0] ”
[0167] The adjusted prompt is closer to the human evaluation logic, so that the evaluation result of the model is closer to the human subjective feeling. Then, based on the adjusted prompt, execute step S402 and step S403 again until the difference between the annotation result and the evaluation result reaches the preset requirement.
[0168] In other embodiments, the offset samples can also be determined from the evaluation of other test samples outside the training samples by the first large language model, such as Figure 6 shown Figure 6Another acquisition and processing path for offset samples is shown. For the benchmark dataset that has been manually annotated, during the first-stage training, it is possible to first sample it to obtain the test sample set 1, and filter to obtain the training sample set 1 through methods such as IoU. For the specific filtering method, refer to the previous text and it will not be elaborated here. Use the training samples in the training sample set 1 to optimize the parameters of the first large language model to obtain the first large language model in version V1. The first large language model in version V1 can be used to evaluate the business results in the test samples in the test sample set 1 to obtain the evaluation results corresponding to the test samples, compare them with the annotation results, and determine the samples with large differences between the two as offset samples. During the second-stage training, it is possible to sample the benchmark dataset to obtain the test sample set 2, and construct a new training sample set 2 through the offset samples and the training sample set 1. Similarly, use the training samples in the training sample set 2 to optimize the parameters of the first large language model in version V1 to obtain the first large language model in version V2. The first large language model in version V2 can be used to evaluate the business results in the test samples in the test sample set 2 to obtain the evaluation results corresponding to the test samples, compare them with the annotation results, and determine the samples with large differences between the two as new offset samples. This process is repeated until the adjusted first large language model meets the business requirements.
[0169] For the large language model parameter tuning method provided in the embodiments of this specification, on the one hand, the business results generated by the target business can be automatically evaluated by the first large language model. Compared with manual evaluation, it greatly improves the evaluation efficiency of business results, reduces the consumed labor cost and time cost, and has stronger objectivity. On the other hand, by using the evaluation results and annotation results of the training samples to adjust the parameters of the first large language model, the accuracy and stability of the first large language model during evaluation are improved, and the first large language model can be made to adapt to different business changes. Furthermore, it can test the generation performance of the second large language model under different iterative versions. In addition, the relevance between the business result and the first information used to indicate the user's interest tendency is also considered during evaluation, and it can ensure that the evaluation result is consistent with the user's subjective feeling.
[0170] The evaluation process of the large language model after parameter tuning will be described below. Figure 7 It shows a schematic flowchart of a large language model evaluation method according to an embodiment of the present disclosure. This method can be executed by any device, platform, or device cluster with computing and processing capabilities, and includes steps S701 - S702 as shown below.
[0171] As Figure 7 shown, in step S701, obtain the execution result of the target business.
[0172] Among them, the execution result is generated by the target service based on the second large language model for the subject in the query image of the user. For a more detailed explanation of the second large language model, the target service and its execution result, please refer to the relevant descriptions of the second large language model, the target service and its service result in the previous text, which will not be elaborated here.
[0173] In step S702, the first large language model evaluates the execution result to obtain evaluation data for the execution result.
[0174] Among them, the first large language model is the large language model obtained by optimizing the parameters of any method in the foregoing embodiments. In practice, the prompt words can be constructed according to the execution result, the first information, and the query image and incorporated into the first large language model to obtain the output evaluation data. The evaluation metrics in the evaluation data can include the relevance between the execution result and the first information, and can also include the relevance between the semantic subject in the execution result and the subject in the query image, the correctness of the execution result, the logic of the execution result, and the information content of the execution result.
[0175] In the solution provided by the foregoing embodiments, the evaluation of the image-based content generation task can be realized. Among them, the first large language model used can adapt to the changes in the execution results of the second large language model when facing different services through parameter tuning, and can test the generation performance of the second large language model under different conditions.
[0176] In one embodiment, this specification also provides an evaluation method for the intent module, which can evaluate, for example Figure 2 the accuracy of the subject category recognition in the view section in. Specifically, the intermediate execution result of the target service and the class annotation data can be obtained. The intermediate execution result is the subject category in the query image recognized by the target service, and the class annotation data is used to label the class tags of the subject in the query image. Then, the third large language model evaluates the intermediate execution result according to the class annotation data to obtain the accuracy of the intermediate execution result.
[0177] In practice, during the execution of the target service, before generating the execution result based on the first large language model, it is necessary to detect and recognize the subject in the query image to obtain the position box where the subject is located and the subject category, that is, the intermediate execution result. The accuracy of this intermediate execution result affects the actual performance of the target service.
[0178] Among them, the evaluation of the accuracy of the position of the main body is relatively objective. The degree of overlap between the annotation box where the manually annotated main body is located and the position box in the intermediate service result can be compared by means of IoU. If the IoU value is greater than a certain threshold (such as 0.8), it is considered that the detection is accurate. The evaluation of the main body category can be carried out through the third large language model. For example, a prompt can be constructed, and the content of the prompt can be "Please compare whether the category label in the category annotation data is consistent with the main body category of the intermediate execution result and output a score from 0 to 10". The higher the score, the higher the accuracy of the intermediate execution result.
[0179] Considering that there are a large number of main body categories in the actual application scenario and the recognition difficulty of different main body categories is different, as an implementation method, the automated evaluation coverage can be gradually carried out according to the difficulty level of the attributes of the main body categories. Exemplarily, the following four evaluation modes are exemplified:
[0180] a. For urgent requirements and relatively subjective evaluation categories, manual evaluation can be adopted. For example, for categories related to medical care and safety, they can be considered as urgent requirement categories, and for categories affected by personal experience, they can be considered as relatively subjective categories, and manual evaluation methods can be used to ensure their accuracy.
[0181] b. For categories with objective evaluations without ambiguity, such as drug categories, the attributes such as the name and category of drugs are very certain and there is no ambiguity. The pre-programmed code logic can be used for evaluation, and then manual review and evaluation can be added to reduce the amount of manual participation.
[0182] c. For categories with objective evaluations with ambiguity, such as the wine category, since there are many wine varieties and their appearances are very similar, some deviations will occur during the evaluation. The third large language model with single-modal input can be used for evaluation. For example, the evaluation can be carried out by constructing a text-based prompt and inputting it into the third large language model.
[0183] d. For categories with objective evaluations with ambiguity and subjective evaluation categories, that is, categories that are very easily affected by personal experience, such as categories with complex aliases or common names like animals, plants, food ingredients, and beauty products, the third large language model with multi-modal input can be used for evaluation. When constructing the prompt, a multi-modal prompt can be constructed according to the query image, the position box where the main body is located, the main body category, and the category label, etc., to ask the model about the matching degree between the main body category and the category label, so as to carry out a more accurate evaluation in combination with multi-modal data.
[0184] Specifically, it can be a preset evaluation mode for each category. During the evaluation process, according to the complexity of the attributes of the category, a single or a combination of multiple evaluation methods is adopted to output the precision-recall index. In practice, by continuously optimizing the code script and the prompt words of the model, the accuracy of the overall evaluation can reach more than 95% after verification by manual sampling review.
[0185] In some embodiments, in order to better evaluate different versions of the target service, an evaluation benchmark data set for the target service can be pre-constructed. The evaluation benchmark data set contains query images of multiple users and the corresponding category annotation data of the query images. When obtaining the intermediate execution result and the category annotation data of the target service, it can be to first obtain the evaluation benchmark data set, and then the target service generates the intermediate execution result for the main body in the query images in the evaluation benchmark data set. For example, obtain the query image and the category annotation data of a certain test sample in the evaluation benchmark data set, and the target service identifies the main body category based on the query image. In this way, for different versions of the target service, it can be evaluated based on a unified standard evaluation benchmark data set, which can better determine the advantages and disadvantages of the execution performance of different versions of the target service.
[0186] The following combines Figure 8 , to illustrate the construction process of the evaluation benchmark data set. Among them, Figure 8 The full-map data in can include historical business data, business data of third-party platforms, and feedback data of the target service. After obtaining the full-map data, it can be subjected to feedback processing or table dropping to store it in the database.
[0187] Exemplarily, the construction of the evaluation benchmark data set can be divided into three stages:
[0188] The first stage: Reuse and processing of historical data
[0189] In one embodiment, historical business data can be obtained first. The historical business data contains query images of users. Then, according to the data screening requirements, the historical business data is screened to obtain the target business data. Next, the query images in the target business data are annotated to obtain the category annotation data.
[0190] In practice, the historical business data can be the historical data of other image-based query services similar to the target service. The historical business data is cleaned according to the data screening requirements to remove invalid data. The invalid data can be duplicate data, incomplete data, or incorrect data, etc. Different embodiments can adopt different cleaning methods according to the data screening requirements. For example, it can be as Figure 8As shown, for the historical business data in the database, a large model can be directly used for data cleaning. Then, the position of the subject in the query image can be manually labeled first, that is, a bounding box is obtained by box selection for pre-labeling, and the subject category in the query image is manually labeled to obtain the category annotation data for the secondary category. Random sampling of the category annotation data is performed to obtain the evaluation benchmark dataset available for the intent module. The evaluation benchmark dataset can include the annotation bounding box of the subject and the secondary category. The category annotation data can also include the necessary attributes of the subject. The necessary attributes can include attributes such as the brand, category, and first-level category to which the subject belongs, to ensure that there is no ambiguity in the category name of the subject. For example, the necessary attributes of an apple can be plant, fruit, etc., to distinguish it from an iPhone with necessary attributes of mobile phone, electronic product. The necessary attributes can be automatically labeled by a trained large model first. To improve the accuracy of the labeling, as Figure 8 shown, multiple different large models (Large Model 1, Large Model 2, Large Model 3, and Large Model 4) can be used to label the subject in the same query image, and then the answers output by different large models are combined as the reference answer. Manual labeling is performed based on this reference answer to obtain richer category annotation data, that is Figure 8 the available dataset of the detailed attributes of the standard category in, and this dataset can be used as a richer evaluation benchmark dataset. The model will also output corresponding richer category evaluation data based on this richer evaluation benchmark dataset. The datasets obtained above can all be manually reviewed to ensure their accuracy.
[0191] As an implementation method, when performing data cleaning, it can be to obtain the subject category of the query image in the historical business data, and filter the query images in the historical business data according to the requirements of the second category to obtain the target business data.
[0192] Among them, the requirements of the second category can be the requirements for the richness of various categories in the category. For example, it can be the requirement for the proportion of different categories or the number of different categories. By filtering different subject categories, the category richness is ensured so that the evaluation benchmark dataset can provide sufficient differential information, and then the comprehensive performance of the target business can be evaluated.
[0193] In this stage, the historical data of other businesses is reused. These data are persistently stored after being cleaned, labeled, and processed. In this stage, the existing data resources can be utilized to quickly construct an initial dataset under the condition of insufficient data resources.
[0194] In other embodiments, the historical data can also include the historical business data of the target business.
[0195] Second stage: Data crawling and collection
[0196] In one embodiment, it may be to obtain business data from a third-party platform. The business data includes query images in the actual usage scenario of users, and then annotate the query images in the business data to obtain class annotation data.
[0197] In practice, the third-party platform can be a travel, social, video-sharing, or other type of website or application. Data close to the actual usage scenario of users can be obtained from it through data crawling to ensure that the dataset can better fit the usage form of the target business. Through the third-party platform, rich and diverse data samples can be obtained, thereby improving the quality and representativeness of the dataset. For the specific annotation process, refer to the first stage and will not be elaborated here.
[0198] The third stage: Link reflux and directional enhancement
[0199] In the third stage, data can be refluxed from the actual business data of the target business and subjected to directional enhancement construction. Similar to the first stage, for the randomly sampled category results and their proportions, a large model can be used for data cleaning to ensure that the data meets the preset category requirements, and then the cleaned data is refluxed to the evaluation benchmark dataset. This step helps to supplement and improve the dataset to ensure its coverage of a wider range of situations.
[0200] In one embodiment, the directional enhancement construction can be based on the Hard Case, i.e., the offset image, in the evaluation benchmark dataset. Specifically, based on the accuracy of the intermediate execution result, it can be determined whether the query image corresponding to the intermediate execution result is an offset image. When it is determined that the query image is an offset image, a new evaluation benchmark dataset is constructed based on the offset image.
[0201] Exemplarily, when the accuracy of the intermediate execution result is lower than the preset threshold, it indicates that the difference between the main category in the query image recognized by the target business and the class annotation data is large, and the evaluation performance of the model on the category of this main body needs to be strengthened. The query image corresponding to the intermediate execution result can be determined as an offset image, and then directional enhancement construction is performed on the offset image. For example, the offset image is upsampled, and the expanded offset image is added to the evaluation benchmark dataset. It can also be to generate data samples in a specific scenario based on the scene in the offset image to enhance the diversity and complexity of the dataset, thereby improving the understanding ability and generalization ability of the model.
[0202] In addition, for the understanding module, the above process can also be used to construct the corresponding evaluation benchmark dataset to evaluate the execution performance of different versions of the target service based on the evaluation benchmark dataset with a unified standard. The difference lies in the annotation part of the query image. Exemplarily, after determining the query image of the evaluation benchmark dataset, the target service can generate a service result based on the query image and the first information, and then manually annotate each evaluation index of the service result to obtain an evaluation label. Finally, an evaluation benchmark dataset is constructed according to the service result, the annotation result, the query image, and the first information.
[0203] Figure 9 It is a schematic structural diagram of a large language model parameter tuning device in an embodiment of this specification. This device can be applied to any device, platform, or device cluster with computing and processing capabilities. Among them, the first large language model is used to evaluate the execution performance of the target service, and the target service is used to generate a service result for the main body in the user's query image based on the second large language model. This device includes:
[0204] A data acquisition module 901, configured to acquire training samples, where the training samples include service results, annotation results, and first information for indicating the user's interest tendency, and the annotation results are used to annotate the evaluation labels of the service results;
[0205] A service evaluation module 902, configured to evaluate based on the service result by the first large language model to obtain an evaluation result for the training sample, and the evaluation result and the annotation result include the relevance between the service result and the first information;
[0206] A parameter adjustment module 903, configured to adjust the parameters of the first large language model based on the evaluation result and the annotation result.
[0207] In one implementation, the training samples include query images.
[0208] In one implementation, the evaluation result and the annotation result further include at least one of the following evaluation indexes: the relevance between the semantic main body in the service result and the main body in the query image; the correctness of the service result; the logic of the service result; the information content of the service result.
[0209] In one implementation, the service evaluation module 902 is specifically configured to determine whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; in the case of determining that the training sample is an offset sample, upsample the offset sample to obtain an amplified offset sample; and adjust the parameters of the first large language model based on the amplified offset sample.
[0210] In one embodiment, the apparatus further includes: a prompt optimization module (not shown in the figure), configured to determine whether a training sample is an offset sample based on the difference between the evaluation result and the annotation result; in the case where it is determined that the training sample is an offset sample, based on the offset sample and the subjectivity of each evaluation metric, adjust the prompt of the first large language model, and evaluate again based on the business result to obtain a second evaluation result for the training sample; and adjust the parameters of the first large language model based on the second evaluation result and the annotation result.
[0211] In one embodiment, when the service evaluation module 902 or the prompt optimization module determines whether a training sample is an offset sample based on the difference between the evaluation result and the annotation result, it is specifically configured to determine the training sample as an offset sample in response to the consistency degree between the evaluation metrics in the evaluation result and the evaluation metrics in the annotation result being lower than a preset threshold.
[0212] In one embodiment, the data acquisition module 901 is specifically configured to: acquire a first business result and a first annotation result generated by the target service for the first alternative image, the first business result including the position information of at least one subject recognized based on the first alternative image, and the first annotation result being used to label the position tags of at least one subject; screen the first alternative image based on the position information in the first business result and the position tags in the first annotation result to determine the first image; and acquire the business result and the annotation result generated by the target service for the first image as training samples.
[0213] In one embodiment, the data acquisition module 901 is specifically configured to: acquire the categories of at least one subject in the second alternative image; screen the second alternative image according to the first category requirement to determine the second image; and acquire the business result and the annotation result generated by the target service for the second image as training samples.
[0214] In one embodiment, the data acquisition module 901 is further configured to: acquire the actual business data during the operation of the target service, the actual business data including at least one of the following data: a query image, and first information for indicating the user's interest tendency; determine sample offset data based on the difference between the actual business data and the training samples; and construct new training samples based on the sample offset data.
[0215] Figure 10 It is a schematic structural diagram of a large language model evaluation apparatus in an embodiment of this specification. This apparatus can be applied to any device, platform, or device cluster with computing and processing capabilities. This apparatus includes:
[0216] A result acquisition module 101, configured to acquire the execution result of the target service, where the execution result is generated by the target service based on the second large language model for the subject in the user's query image;
[0217] The performance evaluation module 102 is configured to evaluate the execution result by a first large language model to obtain evaluation data for the execution result. The first large language model is a large language model obtained by tuning the parameters of any large language model parameter tuning method.
[0218] In one implementation, the device further includes a category evaluation module (not shown in the figure), which is configured to obtain the intermediate execution result of the target service and the category annotation data. The intermediate execution result is the main category in the query image identified by the target service, and the category annotation data is used to annotate the category label of the main body of the query image; the third large language model is used to evaluate the intermediate execution result according to the category annotation data to obtain the accuracy of the intermediate execution result.
[0219] In one implementation, when the category evaluation module obtains the intermediate execution result of the target service and the category annotation data, it is specifically configured to obtain an evaluation benchmark data set, which contains query images of multiple users and the category annotation data corresponding to the query images; the target service generates an intermediate execution result for the main body in the query images in the evaluation benchmark data set.
[0220] In one implementation, when the category evaluation module obtains the evaluation benchmark data set, it is specifically configured to obtain historical service data, which contains the query images of users; according to the data screening requirements, screen the historical service data to obtain the target service data; annotate the query images in the target service data to obtain the category annotation data.
[0221] In one implementation, when the category evaluation module screens the historical service data according to the data screening requirements to obtain the target service data, it is specifically configured to obtain the main category of the query images in the historical service data; screen the query images in the historical service data according to the second category requirement to obtain the target service data.
[0222] In one implementation, when the category evaluation module obtains the evaluation benchmark data set, it is specifically configured to obtain the service data of a third-party platform, and the service data contains query images in the actual usage scenarios of users; annotate the query images in the service data to obtain the category annotation data.
[0223] In one implementation, the category evaluation module is further configured to determine whether the query image corresponding to the intermediate execution result is an offset image based on the accuracy of the intermediate execution result; in the case where it is determined that the query image is an offset image, construct a new evaluation benchmark data set based on the offset image.
[0224] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is caused to execute as Figure 4 and Figure 7 the described method.
[0225] The embodiments of this specification also provide a computing device, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method described as Figure 4 and Figure 7 is implemented.
[0226] The embodiments of this specification also provide a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method described as Figure 4 and Figure 7 are implemented.
[0227] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0228] In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0229] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above are only the specific embodiments of the multiple embodiments disclosed in this specification and are not used to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification shall be included within the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A method for optimizing the parameters of a large language model, wherein, The first large language model is used to evaluate the execution performance of the target service, and the target service is used to generate a service result for the subject in the user's query image based on the second large language model; the method includes: Obtain training samples, where the training samples include the service result, an annotation result, and first information used to indicate the user's interest tendency, and the annotation result is used to annotate the evaluation label of the service result; The first large language model evaluates based on the service result to obtain an evaluation result for the training sample, and the evaluation result and the annotation result include the relevance between the service result and the first information; Adjust the parameters of the first large language model based on the evaluation result and the annotation result; The method further includes: Obtain the actual service data during the operation of the target service, and the actual service data includes at least one of the following data: a query image, and first information used to indicate the user's interest tendency; Determine sample offset data based on the difference between the actual service data and the training sample; Construct a new training sample based on the sample offset data.
2. The method according to claim 1, wherein, The query image is included in the training sample.
3. The method according to claim 1, wherein The evaluation result and the annotation result further include at least one of the following evaluation metrics: the relevance between the semantic subject in the service result and the subject in the query image; The correctness of the service result; the logic of the service result; the information content of the service result.
4. The method according to claim 1, wherein, The adjusting the parameters of the first large language model based on the evaluation result and the annotation result includes: Determine whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; In the case where it is determined that the training sample is an offset sample, upsample the offset sample to obtain an amplified offset sample; Adjust the parameters of the first large language model based on the amplified offset sample.
5. The method according to claim 1, wherein The method further includes: Determine whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result; In the case where it is determined that the training sample is an offset sample, based on the offset sample and the subjectivity of each evaluation metric, adjust the prompt words of the first large language model, and evaluate again based on the service result to obtain a second evaluation result for the training sample; Adjust the parameters of the first large language model based on the second evaluation result and the annotation result.
6. The method according to claim 4 or 5, wherein The determining whether the training sample is an offset sample based on the difference between the evaluation result and the annotation result includes: In response to the degree of consistency between the evaluation metrics in the evaluation result and the evaluation metrics in the annotation result being lower than a preset threshold, determine the training sample as an offset sample.
7. The method according to claim 1, wherein The obtaining the training sample includes: Obtain the first service result and the first annotation result generated by the target service for the first alternative image, where the first service result includes the position information of at least one subject recognized based on the first alternative image, and the first annotation result is used to annotate the position label of the at least one subject; Filter the first alternative image based on the location information in the first business result and the location tag in the first annotation result to determine the first image; Obtain the business result and annotation result generated by the target business for the first image as training samples.
8. The method according to claim 1, wherein The obtaining of the training samples includes: Obtain the category of at least one subject in the second alternative image; Filter the second alternative image according to the first category requirement to determine the second image; Obtain the business result and annotation result generated by the target business for the second image as training samples.
9. A method for evaluating a large language model, the method includes: Obtain the execution result of the target business, where the execution result is generated by the target business based on the second large language model for the subject in the query image of the user; Evaluate the execution result by the first large language model to obtain evaluation data for the execution result, where the first large language model is a large language model obtained by parameter tuning according to the method described in any one of claims 1-8.
10. The method according to claim 9, wherein, The method further includes: Obtain the intermediate execution result and class annotation data of the target business, where the intermediate execution result is the category of the subject in the query image identified by the target business, and the class annotation data is used to label the class tag of the subject in the query image; Evaluate the intermediate execution result by the third large language model according to the class annotation data to obtain the accuracy of the intermediate execution result.
11. The method according to claim 10, wherein, The obtaining of the intermediate execution result and class annotation data of the target business includes: Obtain an evaluation benchmark data set, where the evaluation benchmark data set contains query images of multiple users and the class annotation data corresponding to the query images; Generate an intermediate execution result for the subject in the query image in the evaluation benchmark data set by the target business.
12. The method according to claim 11, wherein, The obtaining of the evaluation benchmark data set includes: Obtain historical business data, where the historical business data contains query images of users; Filter the historical business data according to the data screening requirement to obtain target business data; Annotate the query images in the target business data to obtain class annotation data.
13. The method according to claim 12, wherein, The filtering of the historical business data according to the data screening requirement to obtain target business data includes: Obtain the category of the subject in the query image in the historical business data; Filter the query images in the historical business data according to the second category requirement to obtain target business data.
14. The method according to claim 11, wherein, The obtaining of the evaluation benchmark data set includes: Obtain the business data of the third party platform, where the business data contains query images in the actual usage scenario of the user; Annotate the query images in the business data to obtain class annotation data.
15. The method according to claim 11, wherein, The method further includes: Based on the accuracy of the intermediate execution result, determine whether the query image corresponding to the intermediate execution result is an offset image; In the case of determining that the query image is an offset image, construct a new evaluation benchmark data set based on the offset image.
16. An apparatus for tuning parameters of a large language model, wherein, The first large language model is used to evaluate the execution performance of the target service, and the target service is used to generate a service result for the subject in the user's query image based on the second large language model; The device includes: A data acquisition module, configured to acquire training samples, where the training samples include the service result, an annotation result, and first information for indicating the user's interest tendency, and the annotation result is used to annotate the evaluation label of the service result; A service evaluation module, configured to evaluate by the first large language model based on the service result to obtain an evaluation result for the training sample, and the evaluation result and the annotation result include the relevance between the service result and the first information; A parameter adjustment module, configured to adjust the parameters of the first large language model based on the evaluation result and the annotation result; The data acquisition module is further configured to acquire actual service data during the operation of the target service, and the actual service data includes at least one of the following data: a query image, first information for indicating the user's interest tendency, determine sample offset data based on the difference between the actual service data and the training sample, and construct a new training sample based on the sample offset data.
17. A large language model evaluation device, the device includes: A result acquisition module, configured to acquire the execution result of the target service, and the execution result is generated by the target service for the subject in the user's query image based on the second large language model; A performance evaluation module, configured to evaluate the execution result by the first large language model to obtain evaluation data for the execution result, and the first large language model is a large language model obtained by parameter tuning based on any one of the methods recited in claims 1-8.
18. A computing device, comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the method recited in any one of claims 1-15.
Citation Information
Patent Citations
Model evaluation method and device, computer storage medium and electronic equipment
CN117909700A