Label marking method, device, electronic device and readable storage medium
By using the correlation prediction model to score the text pairs of labels and objects, the problems of misrecall and misrecalls in search recalls in the prior art are solved, and the accuracy of the association between labels and objects and the hit rate of search recalls are improved.
Patent Information
- Application Number
- CN202011300604.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-11-18
AI Technical Summary
In the existing search recall method, when the user's search text matches the object identifier in the database, it is easy to have missed recalls and misrecalls, and the accuracy of the association between the label and the object affects the accuracy of the search recall.
By using the label's own text information and the object's identification text, the target object is tagged using a pre-trained correlation prediction model. The specific steps include obtaining the text of the tag to be marked and the identification text of the target object, inputting the text of the composition to the correlation prediction model, obtaining the correlation score, and marking the tag to be marked for the target object when the score is greater than the preset threshold.
Improve the accuracy of label marking and enhance the accuracy of association between labels and objects, thereby effectively improving the hit rate of search recalls and the accuracy of search results.
Smart Images

Figure CN112507066B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a label marking method, device, electronic device, and readable storage medium. Background Art
[0002] Search recall refers to using search information to obtain data that matches the search information from a specified database and return it to the subject that performed the search. For example, a user enters a search text in the search bar of a shopping website to obtain a list of products corresponding to the search text. Existing search recall methods generally use the user's search text to match the object identifier (such as the product name) in the database, which has a high error rate and is prone to missed recalls and false recalls.
[0003] Related technologies have begun to use object labels for search recall to reduce the error rate of search recall. For example, a user enters a search text, and the search text is matched with the label of the object in the database to return the corresponding object. The accuracy of the association between the label and the object affects the accuracy of the search recall, and how to accurately label the target object and then achieve accurate association between the label and the object is a problem that needs to be solved. Summary of the invention
[0004] The embodiments of the present application provide a labeling method, device, electronic device, and readable storage medium for accurately labeling a target object.
[0005] According to a first aspect of an embodiment of the present application, a label marking method is provided, the method comprising:
[0006] Obtain the text of the label to be marked and the identification text of the target object;
[0007] Inputting a text pair consisting of the identification text of the target object and the text of the tag to be marked into a pre-trained relevance prediction model to obtain a relevance score;
[0008] When the relevance score is greater than a preset threshold, the target object is marked with the tag to be marked.
[0009] Optionally, after marking the to-be-marked tag for the target object, the method further includes:
[0010] Add the text of the tag to be marked to the description text of the target object;
[0011] The display terminal is controlled to display the description text of the target object including the text of the tag to be marked.
[0012] Optionally, after marking the to-be-marked tag for the target object, the method further includes:
[0013] Get the search text;
[0014] Matching the search text with texts of multiple tags marked for multiple objects respectively, the multiple objects including a first target retrieval object;
[0015] When the search text matches the text of a tag marked for the object, determining all objects marked by the tag as the first target retrieval object;
[0016] The first target search object is added to a first target search object set, and a first search result is generated according to the first target search object set, wherein the first search result includes all the first target search objects.
[0017] Optionally, after marking the to-be-marked tag for the target object, the method further includes:
[0018] Obtaining all target search objects in the first target search object set;
[0019] Comparing the label of the first target retrieval object with the labels respectively marked for a plurality of objects for similarity, the plurality of objects including the second target retrieval object;
[0020] When the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity, determining the object as a second target retrieval object;
[0021] The second target search object is added to a second target search object set, and a second search result is generated according to the first target search object set and the second target search object set, wherein the second search result includes all the first target search objects and the second target search objects.
[0022] Optionally, the method further comprises:
[0023] Obtaining a target search text, extracting a tag text from the target search text and adding the tag text to a tag text set, wherein the target search text is determined based on at least one of a corresponding search number, a conversion rate, and a search result; or
[0024] Perform entity recognition on the text of each object to obtain the text of the label and add it to the label text set, wherein the text of an object includes at least one of the following: description text of the object, identification text of the object, brand text of the object, and attribute text of the object;
[0025] Obtaining the text of the tag to be marked includes: determining the text of the tag to be marked from the tag text set.
[0026] Optionally, the method further comprises:
[0027] Obtaining multiple search logs of sample objects, each search log including a sample search request input by a user and a sample search result including the sample object output by a search engine;
[0028] In the case where the user has performed a preset operation on the sample object included in the sample search results, generating a positive sample pair according to the search request and the sample object; and / or
[0029] In the case that the user has not performed the preset operation on the sample object included in the sample search results, a negative sample pair is generated according to the search request and the sample object.
[0030] Optionally, the method further comprises:
[0031] Perform entity recognition on the text of the sample object to obtain the entity recognition result;
[0032] Generate a positive sample pair according to the entity recognition result and the sample object; and / or
[0033] A negative sample pair is generated according to other texts except the entity recognition result and the sample object.
[0034] Optionally, the method further comprises:
[0035] Obtaining a basic model, the basic model is used to identify the relevance between two texts, the basic model is obtained by training the second preset model using sample text pairs marked with relevance labels as training samples;
[0036] The basic model is trained using the multiple sample pairs as training samples to obtain the correlation prediction model.
[0037] Optionally, the method further comprises:
[0038] For the text pair, output a labeling prompt to label the target object with a manually labeled label;
[0039] The relevance prediction model is updated according to a text pair consisting of the identification text of the target object and the text of the manually marked label.
[0040] According to a second aspect of an embodiment of the present application, a label marking device is provided, the device comprising:
[0041] A text acquisition module is used to obtain the text of the label to be marked and the identification text of the target object;
[0042] A relevance prediction module, used to input a text pair consisting of the identification text of the target object and the text of the tag to be marked into a pre-trained relevance prediction model to obtain a relevance score;
[0043] The tag marking module is used to mark the target object with the tag to be marked when the relevance score is greater than a preset threshold.
[0044] Optionally, the device further comprises:
[0045] A description text adding module, used for adding the text of the tag to be marked to the description text of the target object;
[0046] The description text display module is used to control the display terminal to display the description text of the target object including the text of the tag to be marked.
[0047] Optionally, the device further comprises:
[0048] A search text acquisition module, used to obtain the search text;
[0049] A text matching module, used to match the search text with texts of multiple tags marked for multiple objects, respectively, wherein the multiple objects include a first target retrieval object;
[0050] A first target retrieval object determination module, configured to determine all objects marked by the tag as first target retrieval objects when the search text matches the text of a tag marked for the object;
[0051] The first search result generating module is used to add the first target search object to a first target search object set, and generate a first search result according to the first target search object set, wherein the first search result includes all the first target search objects.
[0052] Optionally, the device further comprises:
[0053] A tag acquisition module, used to obtain all target search objects in the first target search object set;
[0054] A similarity comparison module, configured to compare the label of the first target search object with the labels respectively marked for a plurality of objects, the plurality of objects including the second target search object;
[0055] A second target retrieval object determination module, configured to determine the object as a second target retrieval object when the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity;
[0056] The second search result generating module is used to add the second target search object to the second target search object set, and generate a second search result according to the first target search object set and the second target search object set, wherein the second search result includes all the first target search objects and the second target search objects.
[0057] Optionally, the device further comprises:
[0058] A first tag text adding module, used to obtain a target search text, extract a tag text from the target search text and add the tag text to a tag text set, wherein the target search text is determined based on at least one of a corresponding search number, a conversion rate, and a search result;
[0059] A second label text adding module is used to perform entity recognition on the text of each object, obtain the label text and add it to the label text set, wherein the text of an object includes at least one of the following: description text of the object, identification text of the object, brand text of the object, and attribute text of the object;
[0060] The label text acquisition module is used to obtain the text of the label to be marked, including: determining the text of the label to be marked from the label text set.
[0061] Optionally, the device further comprises:
[0062] A search log acquisition module, used to obtain multiple search logs of sample objects, each search log includes a sample search request input by a user and a sample search result containing the sample object output by a search engine;
[0063] A first sample pair acquisition module, configured to generate a positive sample pair according to the search request and the sample object when a user has performed a preset operation on the sample object included in the sample search result; and / or
[0064] In the case that the user has not performed the preset operation on the sample object included in the sample search results, a negative sample pair is generated according to the search request and the sample object.
[0065] Optionally, the device further comprises:
[0066] An entity recognition module is used to perform entity recognition on the text of the sample object and obtain entity recognition results;
[0067] A second sample pair acquisition module is used to generate a positive sample pair according to the entity recognition result and the sample object; and / or
[0068] A negative sample pair is generated based on other texts except the entity recognition result and the sample object.
[0069] Optionally, the device further comprises:
[0070] A basic model acquisition module, used to obtain a basic model, wherein the basic model is used to identify the relevance between two texts, and the basic model is obtained by training the second preset model using sample text pairs marked with relevance labels as training samples;
[0071] The basic model training module is used to train the basic model using the multiple sample pairs as training samples to obtain the association prediction model.
[0072] Optionally, the device further comprises:
[0073] A manual labeling module, for outputting a labeling prompt for the text pair, so as to label the target object with a manually labeled label;
[0074] The prediction model updating module is used to update the relevance prediction model according to the text pair consisting of the identification text of the target object and the text of the manually marked label.
[0075] According to a third aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps in the method described in the first aspect of the present application are implemented.
[0076] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executed, implements the steps in the method described in the first aspect of the present application.
[0077] The label marking method provided in the embodiment of the present application is adopted, and the text of the label itself and the identification text of the object are used to predict the association between the label and the object, thereby completing the label marking. The embodiment of the present application effectively improves the accuracy of label marking by utilizing the text information of the label itself, and can be widely used in various search scenarios for object recall. Applying the method provided in the embodiment of the present application to the search scenario can effectively improve the hit rate of target retrieval object recall. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor.
[0079] Figure 1 is a flowchart of the steps of a label marking method provided by an embodiment of the present application;
[0080] Figure 2 This is a flowchart of the steps for online use of the relevance prediction model provided by an embodiment of the present application;
[0081] Figure 3 is a flowchart of the steps of a method for generating a first search result provided by an embodiment of the present application;
[0082] Figure 4 is a flowchart of the steps of a method for generating a second search result provided by an embodiment of the present application;
[0083] Figure 5 It is an example diagram of a user terminal displaying object information;
[0084] Figure 6 It is a structural block diagram of a label marking device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0085] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of them. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.
[0086] As mentioned in the above background technology, in current search scenarios, the user's search text is usually directly matched with the object identifier, which results in a low recall hit rate and a high error rate in the recall results. Taking commodity objects as an example, it is difficult for the commodity identifier (i.e., the commodity name) to fully describe the commodity. If the recall is based only on the commodity name, there is first a problem of missed recall: if the commodity name is inconsistent with the search text, the commodity cannot be recalled; for example, when a user searches for "milk", the commodity name is matched, and the commodity with the commodity name "certain milk" can be recalled, but the commodity with the commodity name "certain cow's milk" cannot be recalled. Second, there is a problem of false recall: the commodity name and the search text are falsely matched, resulting in false recall of the commodity; for example, when a user searches for "milk", the commodity name is matched, and commodities such as "milk-flavored biscuits" will also be recalled.
[0087] If each object in the database is labeled in advance, when the user performs search recall, the user provides the search text, and the search text is matched with the label text of all objects in the database, and the target retrieval object corresponding to the label can be obtained. Because labels have good text scalability and stronger scene compatibility, they can be used to avoid semantic misjudgments caused by multiple meanings of a word or multiple words in the same meaning. Matching the label text of the object with the search text for search recall can effectively improve the recall hit rate.
[0088] It can be seen that the accuracy of the association between labels and objects can determine the hit rate of user search and recall of target objects; and how to label objects determines the accuracy of labeling, that is, the accuracy of the association between labels and objects.
[0089] If we only use the semantic content contained in the object identifier to identify the label, label the object, and use the inherent label for search recall, since the accuracy of labeling is not substantially improved, such an approach will still have a very limited improvement in search accuracy. For example, for the product name: "XX brand anti-caries mint flavor children's toothpaste", according to the text semantics contained in the object name, it is labeled with "XX brand", "anti-caries", "mint flavor", and "children". However, if a user searches for "XX brand for tooth decay", it should be the user's potential target retrieval object (that is, the potential object that meets the user's expectations), but it cannot appear in the user's search results.
[0090] Based on the above considerations, the inventor of the present application introduced the text information of the label itself into the label marking method, scored the correlation between the text pairs composed of the label text and the identification text of the object, and then directly judged whether the label and the object are associated based on the correlation score, thereby performing label marking and further improving the accuracy of label marking.
[0091] The embodiment of the present application uses the text information of the tag itself as the basis for determining the tag marking, and uses the text pair relationship between the tag and the object to convert the tag marking into a binary classification problem. The binary classification model only needs to output 0 or 1 to determine whether to mark. At the same time, the binary classification model has better scalability and can easily add and process new tags, so that the accuracy of tag marking can also be improved.
[0092] The label marking method in this embodiment can use the text pair consisting of the text of the label to be marked and the identification text of the target object to mark the target object with a label whose correlation degree is greater than a preset threshold. When the user enters the search text, the hit rate of the search recall is improved based on the matching of the identification text of the object with the search text and the matching of the label text marked for the object with the search text. That is, when the search results containing the target retrieval object corresponding to the search text are returned to the user, the accuracy of the returned results is guaranteed.
[0093] The label marking method of the present application can be applied to terminals, and the terminals specifically include but are not limited to: smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, car computers, desktop computers, set-top boxes, smart TVs, wearable devices, etc.
[0094] Figure 1 The following is a flowchart showing the steps of a label marking method provided by an embodiment of the present application. Figure 1 As shown, the method includes:
[0095] Step S11: Obtain the text of the label to be marked and the identification text of the target object.
[0096] The specific steps include:
[0097] S11-1. Obtain the identification text of the target object.
[0098] The target object is an object that can be returned to the user client as a search result, and has different meanings depending on the actual search scenario. That is, in a specific search scenario, there is a corresponding target object. For example, in the commodity search scenario of a shopping software app or a shopping webpage, the target object is a commodity; in the book search scenario of a library's book information query system, the target object is a book or literature; in the merchant search scenario of an offline group purchase website or software app, the target object is a merchant.
[0099] The identification text of the target object may be the name text of the target object, or may also include some attribute information of the target object. Accordingly, the identification text of the commodity may be the text information of the commodity name, such as "A brand B model air conditioner", or "A brand B model air conditioner inverter wall-mounted 1.5 horsepower new on the market".
[0100] The identification text of the target object can be obtained from a corresponding database according to the application scenario.
[0101] S11-2, obtaining the text of the tag to be marked.
[0102] In this embodiment, the tag to be marked is used as the second identifier of the target object to mark the target object. The text of the tag can describe the attribute-related information of the target object to a certain extent. Specifically, the tag can be the name of the target object, the attribute of the target object, or the application scenario of the target object.
[0103] Among the common labeling methods currently used, labels are generally extracted from the target object's identification text through semantic analysis, word segmentation, etc. For example, for the target object "canned sugar-free cola", the associated labels that can be extracted are: "canned", "sugar-free", and "cola".
[0104] However, simply extracting text from object names as labels is far from sufficient to adapt to actual search scenarios. If new and effective labels can be mined from more sources, more comprehensive labeling can be achieved.
[0105] In fact, the user's search text often contains implicit label information, such as the user's high-frequency search term "XX brand of tooth decay", which can be used as a label for "XX brand of anti-caries toothpaste". Of course, the text of the object can also be labeled by performing entity recognition on the object, such as performing entity recognition on "whitening skin care products" to obtain the labels "whitening" and "skin care products". In addition, if the knowledge graph between object names is established in advance, the knowledge graph of the object can be used to link sample objects that are similar or identical to the target object, such as "sunglasses", "sunglasses" and "sunglasses" can be linked to each other, and "cow's milk" and "milk" can be linked to each other.
[0106] In view of the above considerations, this embodiment can obtain the text of the label to be marked by using one or more combinations of the three ways of the user's search text, the result of entity recognition of the object's identification text, and the object's knowledge graph. Among them, obtaining the text of the label to be marked includes: determining the text of the label to be marked from the label text set. It is easy to understand that the label text set of this embodiment is a text set that is continuously updated and dynamically changed by absorbing new label texts and eliminating invalid label texts, and it can be an empty set at the beginning.
[0107] Method 1: Use the user's search text to obtain the label text. This includes:
[0108] Obtain a target search text, extract the text of the label from the target search text and add it to a label text set, wherein the target search text is determined based on at least one of the corresponding search times, conversion rate, and search results; determine the text of the label to be marked from the label text set.
[0109] The target search text is the text provided by at least one user in the historical search behavior for search. That is to say, the search text used to mine the tag text can be obtained from the historical search behavior of a specific user, or from the historical search behavior of all users or multiple users in a specified user group. If the tag is only obtained from one or more specific users and the object is successfully marked, the marking relationship is also only applied to the subsequent search recall of the specific one or more users.
[0110] For example, for map software, the label "Mingtang New School" is mined from the historical search texts of users in area C. Assuming that the label "Mingtang New School" subsequently marks a Cantonese restaurant M in area C, in subsequent applications, only users in area C who search for "Mingtang New School" can search for the Cantonese restaurant M; and the label "Mingtang New School" is also mined from the historical search texts of users in area D. Assuming that the label "Mingtang New School" subsequently marks a tea house N in area D, in subsequent applications, only users in area D who search for "Mingtang New School" can search for the tea house N.
[0111] First, the target search text is determined based on the corresponding number of searches, which means that for a certain search text, the number of searches that users have performed using the search text in a historical period or a specific historical period is greater than a preset threshold, indicating that the search text is a high-frequency search text. The label text is extracted from the search text and added to the label text set.
[0112] An example is given accordingly. In a group buying website, the number of searches for the search text "Hainan Coconut Chicken" by users has reached 1,500 times in the past three months, which is greater than the preset search number threshold of 500. The tag "Hainan Coconut Chicken" extracted from the search text is added to the tag text set.
[0113] Second, the target search text is determined based on the corresponding conversion rate, which may specifically include: for a certain search text, the conversion rate of the user after using the search text in a historical period or a specific historical period is higher than a preset threshold, indicating that the search text is an efficient search text, and the label text is extracted from the search text and added to the label text set.
[0114] If the user obtains the desired target retrieval object through the search text and successfully performs the corresponding target action, it means that the search conversion is successful. The search conversion rate is the ratio of the number of searches in which the user performs the corresponding target action to the total number of searches.
[0115] As an example, in a group buying website, the number of historical searches by users using the search text "XX buns" has reached 200. Among them, in a certain search, the user obtains the page results through the search operation, clicks into a certain merchant or product information on the page, and successfully places an order for the product, which means that the search conversion is successful. If the number of successful conversions in these 300 searches is 80, the conversion rate is 40%, which is higher than the preset conversion rate threshold of 20%, indicating that some of the products or merchants pointed to by the search text may be what the user is happy to see, and can be used for labeling, and the label "XX buns" extracted from the search text is added to the label text set.
[0116] Third, the target search text is determined based on the corresponding search results, which may specifically include: for a certain search text, the number of search results obtained by the user using the search text in a historical period or a specific historical period is less than a preset threshold, indicating that the search text is a search text that lacks an associated object, and the label text is extracted from the search text and added to the label text set.
[0117] An example is given accordingly. In a group buying website, a user searches using the search text "walnut candy". There are no corresponding merchants or products in the search results. The number of search results is 0, which is less than the preset threshold of 3. If necessary, the tag "walnut candy" extracted from the search text is added to the tag text set.
[0118] In addition, the target search text is determined according to the corresponding search times, conversion rate, and search results. Specifically, it can include: based on an adaptive multi-threshold algorithm, it is detected that the search times, conversion rate, and search results of the target search text all meet the corresponding thresholds, and the label text is extracted from the search text and added to the label text set.
[0119] For example, in a group buying website, the number of historical searches for the search text "caramel coffee" by users in the past month reached 300, the conversion rate was 15%, and the number of search results was 10. At this time, the thresholds given by the adaptive algorithm are met: the search number threshold is 200, the conversion rate is 10%, and the number of search results does not have a threshold. This means that the search frequency, conversion rate, and search results of the search text "caramel coffee" have met the requirements, and the tag "caramel coffee" in the search text can be added to the tag text set.
[0120] This embodiment can also use a text pair consisting of a target search text that has been successfully searched and converted and an identification text of a target retrieval object that has been converted from the text as a positive sample pair based on the search results of the user.
[0121] The second approach is to use the result of entity recognition on the object's identification text to obtain the label text. It includes:
[0122] Perform entity recognition on the text of each object to obtain the text of the label and add it to the label text set, wherein the text of an object includes at least one of the following: the description text of the object, the identification text of the object, the brand text of the object, and the attribute text of the object; determine the text of the label to be marked from the label text set.
[0123] Entity recognition refers to identifying entities with specific meanings in text, extracting entity data such as names of people, places, institutions, and attributes. One or more label texts can be identified for the text of an object. For example, for the description text of a certain commodity milk, "high calcium skim milk for middle-aged and elderly people", "high calcium", "skim", "middle-aged and elderly people", and "milk" can be identified.
[0124] Approach 3: Use the knowledge graph of the object to obtain the label text. This includes:
[0125] Therefore, in an embodiment of the present application, the knowledge graph of the sample object can also be used to obtain labels, including: pre-establishing the knowledge graph of the sample object through semantic analysis, and then using the sample object knowledge graph to obtain the synonymous entity name of the target object as the label to be marked.
[0126] For example, "potato-potato", "tomato-tomato", "yogurt" and "sour milk" are synonyms, and all of them are added to the label text set of the to-be-labeled label to improve the label text set.
[0127] In this embodiment, a "marked" mark may be added to all tags in the tag text set that have been used to mark objects. When obtaining tags to be marked from the tag text set, tags with the "marked" mark added will no longer be obtained.
[0128] Step S12: input the text pair consisting of the identification text of the target object and the text of the tag to be marked into a pre-trained relevance prediction model to obtain a relevance score.
[0129] Taking into account that label texts from multiple sources are obtained in the aforementioned step S11, and due to the continuous updating of user search logs, the mined labels will also be continuously updated, this embodiment further uses the updated labels as labels to be labeled, and calculates the association score between the text pair consisting of the identification text of the target object and the text of the label to be labeled.
[0130] The accuracy of the correlation score determines the accuracy of the label marking. To achieve a certain level of accuracy in the label marking, it is necessary to pre-train the correlation prediction model using sample pairs. In this embodiment, sample pairs for training the correlation prediction model can be first obtained, and then the preset model can be trained to obtain the correlation prediction model, and finally the trained correlation prediction model can be used to obtain the correlation score.
[0131] Figure 2 is a flowchart of the steps of the online use method of the correlation prediction model provided by an embodiment of the present application, refer to Figure 2 As shown, the following steps may be specifically included:
[0132] S12-11. Obtain sample pairs for training the association prediction model.
[0133] For the correlation prediction model, it is necessary to construct sample data for model training in advance, input multiple sample pairs into the preset model for training, and then obtain a correlation prediction model with certain application value. Therefore, in this embodiment, the correlation prediction model is a model obtained by training the preset model with multiple sample pairs as input values. The sample pairs used to train the correlation prediction model include the following two methods:
[0134] Sample pairs obtained based on the search log of the sample object and sample pairs obtained using the entity recognition results.
[0135] First, obtaining sample pairs according to the search log of the sample object may include the following steps:
[0136] Obtaining multiple search logs of sample objects, each search log including a sample search request input by a user and a sample search result including the sample object output by a search engine;
[0137] In the case where the user has performed a preset operation on the sample object included in the sample search results, generating a positive sample pair according to the search request and the sample object; and / or
[0138] In the case that the user has not performed the preset operation on the sample object included in the sample search results, a negative sample pair is generated according to the search request and the sample object.
[0139] Among them, for the sample objects contained in the sample search results, those on which the user has clicked are positive sample objects, and constitute a positive sample pair with the corresponding search text; those on which the user has never clicked are negative sample objects, and constitute a negative sample pair with the corresponding search text.
[0140] For example, within a period of time, for multiple objects provided by a certain search behavior, the user has clicked on an object E1, then the text F2 of the label to be marked extracted by mining the search text F1 under the search behavior, and the text pair F2-E2 consisting of the object's identification text E2 can be used as a positive sample pair; for multiple objects provided by a certain search behavior, the user has never clicked on an object G1, then the text F2 of the label to be marked extracted by mining the search text F1 under the search behavior, and the text pair F2-G2 consisting of the object's identification text G2 can be used as a negative sample pair.
[0141] Second, the sample pairs obtained using the entity recognition results may specifically include the following steps:
[0142] Perform entity recognition on the text of the sample object to obtain the entity recognition result;
[0143] Generate a positive sample pair according to the entity recognition result and the sample object; and / or
[0144] A negative sample pair is generated according to other texts except the entity recognition result and the sample object.
[0145] In an example of the present embodiment, for a certain commodity object "XX semi-skimmed milk", the entity recognizes that its attribute contains "semi-skimmed", then text pairs such as "semi-skimmed" - "XX semi-skimmed milk", "half-fat" - "XX semi-skimmed milk" can be constructed as positive sample pairs; text pairs such as "skim" - "XX semi-skimmed milk", "whole fat" - "XX semi-skimmed milk", "whole skimmed" - "XX semi-skimmed milk", "non-skimmed" - "XX semi-skimmed milk" can also be constructed according to the entity recognition result as negative sample pairs.
[0146] S12-12, train the preset model to obtain a correlation prediction model. This includes the following two methods:
[0147] The first preset model is used as the original model of the association prediction model, and the first preset model is trained to obtain the association prediction model; or the basic model is used to complete the model training based on the transfer learning method.
[0148] First, the first preset model is used as the original model of the correlation prediction model, and the first preset model is trained to obtain the correlation prediction model. Including:
[0149] The obtained multiple sample pairs are input into the first preset model, and the network weights and thresholds in the first preset model are adaptively adjusted according to the output multiple association scores, so that the association scores obtained between the sample pairs are closer to the actual environment, that is, more accurate, thereby improving the accuracy of label marking.
[0150] Specifically, the first preset model can select a common neural network model, which will not be described in detail here.
[0151] Second: Use the basic model to complete model training based on transfer learning method.
[0152] In the embodiment of the present application, a basic model with prior knowledge can also be obtained based on the transfer learning method as the starting point for fine-tuning the relevance prediction model to identify the relevance between two texts. Specifically, it includes:
[0153] Obtaining a basic model, the basic model is used to identify the relevance between two texts, the basic model is obtained by training the second preset model using sample text pairs marked with relevance labels as training samples;
[0154] The basic model is trained using the multiple sample pairs as training samples to obtain the correlation prediction model.
[0155] The second preset model is trained to obtain the basic model. The basic model only needs to be fine-tuned and iterated 1-2 times to balance the effect of the basic model in the new application scenario and maximize the use of prior knowledge in the basic model. Therefore, it can be said that the correlation prediction model is the basic model after fine-tuning.
[0156] The basic model with prior knowledge can play a complementary role in specific search scenarios that lack a large amount of labeled data. For example, "Chuandong Spicy" is the name of a hot pot restaurant, but the attributes of the hot pot restaurant cannot be read from the two words "Chuandong Spicy". In a certain website or application software, if the basic model with prior knowledge is used in advance to associate the text pair "Chuandong Spicy" and "hot pot restaurant", when the user searches for "hot pot restaurant", the search results will include the object "Chuandong Spicy".
[0157] In addition, in different scenarios, the label sets of the same target object vary greatly, which is not conducive to the calculation of the model's association degree. Therefore, the label set in a specific scenario and the label relationship between the label and the object can be obtained by manual labeling. Therefore, in order to speed up the training speed and the accuracy of the model, manually labeled samples can also be added to the sample data during the training process, and the manually labeled data can be used to guide the model to learn in a more precise direction to obtain a more accurate model.
[0158] Based on the above analysis, in an embodiment of the present application, a labeling prompt can be output for the text pair to mark the target object with a manually marked label; and the association prediction model can be updated based on the text pair consisting of the identification text of the target object and the text of the manually marked label.
[0159] In addition, in the initial stage of training the association prediction model, the association prediction effect is poor, and the prediction results can also be manually marked to correct the association score output by the model; until the accuracy of the output result of the association prediction model exceeds the preset threshold, the manual marking can be cancelled and the target object can be marked directly according to the association score output by the model.
[0160] S12-13. Use the trained correlation prediction model to obtain the correlation score.
[0161] In an embodiment of the present application, a text pair consisting of the identification text of the target object and the text of the label to be marked is input into a pre-trained association prediction model, and the association prediction model predicts and scores the association between the target object and the label to be marked to obtain a association score.
[0162] After the above steps and sufficient training of the correlation prediction model, the output correlation score will be close to the actual situation, and the correlation prediction model will have practical application value. The value of the correlation score can be within a preset range. For example, the correlation score can be any value from 0 to 1, where when the correlation score is 0, it means that there is absolutely no correlation between the current target object and the label to be marked; when the correlation score is 1, it means that there is an absolute correlation between the current target object and the label to be marked.
[0163] Step S13: When the relevance score is greater than a preset threshold, mark the target object with the tag to be marked.
[0164] This embodiment uses a preset threshold to convert the problem of calculating the degree of association between the tag to be marked and the target object into the problem of whether the two are associated, that is, associated or not associated, which is equivalent to modeling the association problem between the tag to be marked and the target object into a binary classification problem. For any given target object and a tag to be marked, a conclusion of whether they are associated can be directly drawn. Due to the good scalability of the binary classification model, for a newly added tag, this embodiment can directly input the text pair consisting of its text and the identification text of the target object into the association prediction model, without the need to retrain the model.
[0165] Referring to the numerical example of the association score in the above step S12, correspondingly, the preset threshold value can also be any value between 0 and 1, such as 0.5, 0.6, etc. When the association score is greater than the preset threshold value, it means that the association between the target object and the label to be marked can be judged as positive, that is, there is an association between the target object and the label to be marked, then the label to be marked is marked for the target object, and the label marking is completed.
[0166] The preset threshold is adaptively adjusted according to the training result of the association prediction model.
[0167] In this embodiment, the application of label marking to the target object can be completed by the server operation, and can be specifically used for searching and recalling the target retrieval object, obtaining similar objects of the target retrieval object, displaying object information to the user terminal, actively pushing objects to the user terminal, and adding description text of the object.
[0168] Application 1: Use the label marking method to search and recall the target retrieval object in the search scenario.
[0169] After the accuracy of the relevance score of the relevance prediction model reaches a certain level, the accuracy of the label marking can have a certain practical application value. At this time, the label marking method of this embodiment can be applied to the object search scenario. When the user searches for an object, the search text is matched with the object identification text, and the search text is matched with the label text to provide the user with search results containing one or more target retrieval objects.
[0170] Figure 3 is a flowchart of the steps of a method for generating a first search result provided by an embodiment of the present application. Figure 3 As shown, in the embodiment of the present application, after marking the target object with the to-be-marked tag, the following steps may also be included:
[0171] Step S13-11, obtaining the search text;
[0172] Step S13-12, matching the search text with texts of multiple tags marked for multiple objects respectively, wherein the multiple objects include a first target search object;
[0173] Step S13-13, when the search text matches the text of a tag marked for the object, determining all objects marked by the tag as the first target retrieval object;
[0174] Step S13 - 14 , adding the first target search object to a first target search object set, generating a first search result according to the first target search object set, wherein the first search result includes all the first target search objects.
[0175] The multiple objects may be all objects in the object database that meet the search screening conditions. The first target retrieval object is an object that matches the search text and is the first batch of objects that may meet the user's search requirements. In this embodiment, the search text is matched one by one with all the label texts marked for the multiple objects, and the first target retrieval object that matches the search text is obtained through the label relationship between each matching label and the object.
[0176] Example 1: On a shopping website, if a user sets search filter conditions such as "free shipping" and "place of shipment H", then "multiple objects" are all products under the filter conditions for the user to search; when the user searches for "milk candy", the search text "milk candy" is matched one by one with the text of all tags that have been marked on these products. For example, a tag with the text "milk candy" must match the search text "milk candy"; all products marked with the tag "milk candy" can be used as the first target retrieval object, and the first target retrieval object can be added to the first target retrieval object set; search results are generated based on the first target retrieval object set and returned to the user.
[0177] Example 2: On a group buying website, if a user sets search filter conditions such as "restaurant", "area J", "merchant", etc., then the "multiple objects" are all merchants under the filter conditions for the user to search; the user searches for "Ming Yue", and the search text "Ming Yue" is matched one by one with the text of all tags that have been marked with these merchants. For example, a tag with the text "Ming Yue" must match the search text "Ming Yue"; all merchants marked with the tag "Ming Yue" can be taken as the first target retrieval objects, and the first target retrieval objects can be added to the first target retrieval object set; search results are generated based on the first target retrieval object set and returned to the user.
[0178] This embodiment can also combine the two methods of matching the search text with the label text of the object and matching the search text with the object identification text. When the search text matches the label text of any object and the search text matches the identification text of the object, the object is used as the target retrieval object, and search results are generated and returned to the user.
[0179] Application 2: Use the label marking method to obtain similar objects to the target retrieval object.
[0180] Considering that there may be close or similar relationships between tags, in order to avoid the aforementioned steps, some target retrieval objects may still be missed during the tag marking and search recall process.
[0181] Figure 4 is a flowchart of the steps of the second search result generation method provided by an embodiment of the present application. Figure 3As shown, in this embodiment, after marking the to-be-marked tag for the target object, the following steps may also be implemented:
[0182] Step S13-21, obtaining all target search objects in the first target search object set;
[0183] Step S13-22, comparing the label of the first target search object with the labels respectively marked for a plurality of objects, the plurality of objects including the second target search object;
[0184] Step S13-23, when the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity, determining the object as a second target retrieval object;
[0185] Step S13-24, adding the second target search object to the second target search object set, generating a second search result according to the first target search object set and the second target search object set, wherein the second search result includes all the first target search objects and the second target search objects.
[0186] The above method can be simply understood as follows: the first target retrieval object is an object associated with the label text by directly matching the label text through the user's search text; and the second target retrieval object is an object similar to the first target retrieval object determined by calculating the similarity between labels, so that the search results are more comprehensive.
[0187] In an application example of the present embodiment, a user searches for "Coca-Cola". At this time, the label "Coca-Cola" is used as an independent recall source, and multiple first target retrieval objects associated with "Coca-Cola" are recalled; for a certain first target retrieval object K, its associated labels include "Coke", "Coca-Cola", and "Carbonated Beverages". If a label "carbonated beverages" associated with another object L that has not been recalled exceeds a preset threshold value in similarity with the label "carbonated beverages" of the first target retrieval object K, then object L will be determined as a second target retrieval object as a similar object and added to the search results.
[0188] Application three: using the label marking method to display object information to the user terminal.
[0189] After generating the search results, the server will return the target search object in the search results to the user terminal and control the user terminal to display the object information. Specifically, the priority ranking of the target search object can be determined according to the screening conditions set by the user or according to the information such as the current location and historical behavior of the user, and according to the priority ranking, the user terminal is controlled to display one or more objects to the user in the form of a chart column.
[0190] Figure 5 It is an example diagram of a user terminal displaying object information. Figure 5 As shown, the user searches for "Grandma" in the group buying software, and obtains multiple target retrieval objects by applying the label marking method in this application. According to the user's current location, the "Grandma's Braised Pork Rice (Hongpailou Store)" closest to the user is ranked first and displayed to the user.
[0191] Application 4: Use the tag marking method to actively push objects to user terminals.
[0192] After generating the search results, the server will be informed of the user's possible needs, and can also actively push the target search objects in the search results to the user, so that the user can quickly select the required objects.
[0193] Application 5: Use the label marking method to add descriptive text to the object.
[0194] The label text can reflect the relevant information of the tagged object to a certain extent, including brand, attribute, category, etc. If the label text of an object is added to the description text of the object and displayed to the user on the object display page, the user can understand the information of the object more intuitively.
[0195] Therefore, in an embodiment of the present application, after marking the target object with the tag to be marked, the text of the tag to be marked can also be added to the description text of the target object; and the display terminal is controlled to display the description text of the target object including the text of the tag to be marked. In this embodiment, the text of the marked tag is added to the description text of the object, and the terminal displays the description text, thereby enriching the description display of the object.
[0196] Referring to the application example described in step 12, in one application example of the present application, after marking "Chuandong Spicy" with the label "Hot Pot Restaurant", the text of "Hot Pot Restaurant" is added to the description text of "Chuandong Spicy". Then, when a user searches for "hot pot restaurant", a web page or application APP will be used to control the user terminal to display the description text containing "hot pot restaurant", such as "hot pot restaurant Chuandong Spicy". If the user only sees the description text of "Chuandong Spicy" of the merchant on the display page, it may not be certain what the merchant does, but if the description text of the merchant is "hot pot restaurant Chuandong Spicy", the user can clearly and intuitively understand that this is a hot pot restaurant.
[0197] Finally, considering that the label texts from multiple sources in the above method will be continuously updated, the amount of data in the label text library will gradually increase, and the amount of data calculation for label marking will also increase accordingly, in order to reduce the load pressure of the server database and achieve long-term and efficient operation of the database, in the embodiment of the present application, it is also possible to set a periodic timing according to the load capacity of the database, and regularly remove the label text data and label marking data on the target object that have not triggered the positive search behavior for a long time, so as to reduce the load pressure of the database and improve the computational efficiency of the object search. The positive search behavior includes the user's search behavior and the user's corresponding target behavior after obtaining the search results, including using the search text to search, clicking to enter the object details page in the search behavior result page, generating purchase behavior, etc.
[0198] Based on the same inventive concept, an embodiment of the present application provides a label marking device. Figure 6 The structure diagram of the label marking device provided by an embodiment of the present application is shown. Figure 6 As shown, the device specifically includes:
[0199] A text acquisition module 61 is used to obtain the text of the label to be marked and the identification text of the target object;
[0200] The relevance prediction module 62 is used to input the text pair consisting of the identification text of the target object and the text of the tag to be marked into a pre-trained relevance prediction model to obtain a relevance score;
[0201] The tag marking module 63 is used to mark the target object with the tag to be marked when the relevance score is greater than a preset threshold.
[0202] Optionally, the device further comprises:
[0203] A description text adding module, used for adding the text of the tag to be marked to the description text of the target object;
[0204] The description text display module is used to control the display terminal to display the description text of the target object including the text of the tag to be marked.
[0205] Optionally, the device further comprises:
[0206] A search text acquisition module, used to obtain the search text;
[0207] A text matching module, used to match the search text with texts of multiple tags marked for multiple objects, respectively, wherein the multiple objects include a first target retrieval object;
[0208] A first target retrieval object determination module, configured to determine all objects marked by the tag as first target retrieval objects when the search text matches the text of a tag marked for the object;
[0209] The first search result generating module is used to add the first target search object to a first target search object set, and generate a first search result according to the first target search object set, wherein the first search result includes all the first target search objects.
[0210] Optionally, the device further comprises:
[0211] A tag acquisition module, used to obtain all target search objects in the first target search object set;
[0212] A similarity comparison module, configured to compare the label of the first target search object with the labels respectively marked for a plurality of objects, the plurality of objects including the second target search object;
[0213] A second target retrieval object determination module, configured to determine the object as a second target retrieval object when the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity;
[0214] The second search result generating module is used to add the second target search object to the second target search object set, and generate a second search result according to the first target search object set and the second target search object set, wherein the second search result includes all the first target search objects and the second target search objects.
[0215] Optionally, the device further comprises:
[0216] A first tag text adding module, used to obtain a target search text, extract a tag text from the target search text and add the tag text to a tag text set, wherein the target search text is determined based on at least one of a corresponding search number, a conversion rate, and a search result;
[0217] A second label text adding module is used to perform entity recognition on the text of each object, obtain the label text and add it to the label text set, wherein the text of an object includes at least one of the following: description text of the object, identification text of the object, brand text of the object, and attribute text of the object;
[0218] The label text acquisition module is used to obtain the text of the label to be marked, including: determining the text of the label to be marked from the label text set.
[0219] Optionally, the device further comprises:
[0220] A search log acquisition module, used to obtain multiple search logs of sample objects, each search log includes a sample search request input by a user and a sample search result containing the sample object output by a search engine;
[0221] A first sample pair acquisition module, configured to generate a positive sample pair according to the search request and the sample object when a user has performed a preset operation on the sample object included in the sample search result; and / or
[0222] In the case that the user has not performed the preset operation on the sample object included in the sample search results, a negative sample pair is generated according to the search request and the sample object.
[0223] Optionally, the device further comprises:
[0224] An entity recognition module is used to perform entity recognition on the text of the sample object and obtain entity recognition results;
[0225] A second sample pair acquisition module is used to generate a positive sample pair according to the entity recognition result and the sample object; and / or
[0226] A negative sample pair is generated according to other texts except the entity recognition result and the sample object.
[0227] Optionally, the device further comprises:
[0228] A basic model acquisition module, used to obtain a basic model, wherein the basic model is used to identify the relevance between two texts, and the basic model is obtained by training the second preset model using sample text pairs marked with relevance labels as training samples;
[0229] The basic model training module is used to train the basic model using the multiple sample pairs as training samples to obtain the association prediction model.
[0230] Optionally, the device further comprises:
[0231] A manual labeling module, for outputting a labeling prompt for the text pair, so as to label the target object with a manually labeled label;
[0232] The prediction model updating module is used to update the relevance prediction model according to the text pair consisting of the identification text of the target object and the text of the manually marked label.
[0233] Based on the same inventive concept, another embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the method described in any of the above embodiments of the present application are implemented.
[0234] Based on the same inventive concept, another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executed, implements the steps of the method described in any of the above embodiments of the present application.
[0235] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0236] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0237] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0238] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0239] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0240] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0241] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0242] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0243] The above is a detailed introduction to a label marking method, device, electronic device and readable storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for label marking, characterized in that, it includes: obtaining the text of the label to be marked and the identification text of the target object; inputting the text pair composed of the identification text of the target object and the text of the label to be marked into a pre-trained correlation prediction model to obtain a correlation score; when the correlation score is greater than a preset threshold, marking the label to be marked for the target object; after marking the label to be marked for the target object, the method further includes: obtaining a search text; respectively matching the search text with the texts of multiple labels marked for multiple objects, where the multiple objects include a first target retrieval object; when the search text matches the text of a label marked for the object, determining all the objects marked with this label as the first target retrieval object; adding the first target retrieval object to the first target retrieval object set, and generating a first search result according to the first target retrieval object set, where the first search result includes all the first target retrieval objects; after marking the label to be marked for the target object, the method further includes: obtaining all the target retrieval objects in the first target retrieval object set; comparing the similarity between the labels of the first target retrieval object and each of the labels marked for multiple objects respectively, where the multiple objects include a second target retrieval object; when the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity, determining this object as the second target retrieval object; adding the second target retrieval object to the second target retrieval object set, and generating a second search result according to the first target retrieval object set and the second target retrieval object set, where the second search result includes all the first target retrieval objects and the second target retrieval objects.
2. The method according to claim 1, characterized in that, after marking the label to be marked for the target object, the method further includes: adding the text of the label to be marked to the description text of the target object; controlling the display terminal to display the description text of the target object including the text of the label to be marked.
3. The method according to claim 1, characterized in that, the method further includes: obtaining a target search text, extracting the text of the label from the target search text and adding it to the label text set, where the target search text is determined according to at least one of the corresponding search times, conversion rate, and search results; or performing entity recognition on the text of each object to obtain the text of the label and adding it to the label text set, where the text of an object includes at least one of the following: the description text of the object, the identification text of the object, the brand text of the object, and the attribute text of the object; obtaining the text of the label to be marked, including: determining the text of the label to be marked from the label text set.
4. The method according to claim 1, characterized in that, the method further includes: Obtaining multiple search logs of sample objects, each search log including a sample search request input by a user and a sample search result including the sample object output by a search engine; In a case where the user has performed a preset operation on the sample object included in the sample search results, a positive sample pair is generated based on the search request and the sample object; and / or in a case where the user has not performed the preset operation on the sample object included in the sample search results, a negative sample pair is generated based on the search request and the sample object, and the preset operation includes a click behavior.
5. The method according to claim 1, It is characterized in that The method further comprises: Perform entity recognition on the text of the sample object to obtain the entity recognition result; A positive sample pair is generated according to the entity recognition result and the sample object; and / or a negative sample pair is generated according to other texts other than the entity recognition result and the sample object.
6. The method according to any one of claims 1 to 5, It is characterized in that The method further comprises: Obtaining a basic model, the basic model is used to identify the relevance between two texts, the basic model is obtained by training the second preset model using sample text pairs marked with relevance labels as training samples; The basic model is trained using the multiple sample pairs as training samples to obtain the correlation prediction model.
7. The method according to any one of claims 1 to 5, It is characterized in that The method further comprises: For the text pair, output a labeling prompt to label the target object with a manually labeled label; The relevance prediction model is updated according to a text pair consisting of the identification text of the target object and the text of the manually marked label.
8. A label marking device, It is characterized in that The device comprises: The text acquisition module is used to obtain the text of the label to be marked and the identification text of the target object. After marking the label to be marked for the target object, the text acquisition module is also used to: Get the search text; Matching the search text with texts of multiple tags marked for multiple objects respectively, the multiple objects including a first target retrieval object; When the search text matches the text of a tag marked for the object, determining all objects marked by the tag as the first target retrieval object; Adding the first target search object to a first target search object set, generating a first search result according to the first target search object set, wherein the first search result includes all the first target search objects; After marking the target object with the to-be-marked label, the text acquisition module is further used to: Obtaining all target search objects in the first target search object set; Comparing the label of the first target retrieval object with the labels respectively marked for a plurality of objects for similarity, the plurality of objects including the second target retrieval object; When the similarity between the label of the first target retrieval object and a label marked for the object is greater than a preset similarity, determining the object as a second target retrieval object; Add the second target retrieval object to the second target retrieval object set, and generate a second search result according to the first target retrieval object set and the second target retrieval object set, where the second search result includes all the first target retrieval objects and the second target retrieval object; The relevance prediction module is configured to input a text pair composed of the identification text of the target object and the text of the to-be-labeled tag into a pre-trained relevance prediction model to obtain a relevance score; The tag marking module is configured to mark the to-be-labeled tag for the target object when the relevance score is greater than a preset threshold.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, the steps in the method according to any one of claims 1-7 are implemented.
10. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes, the steps of the method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Label determination method and device, computer equipment and storage medium
CN110674319A