Vector library-based text label automatic generation method and system

Through a vector library-based method, text labels are generated using word embedding model and vector similarity search, which solves the problems of inefficiency and unstable label quality in the prior art, and realizes efficient and accurate automatic generation of text labels.

CN119938929APending Publication Date: 2025-05-06INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510031882.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing text label generation methods are inefficient, rely solely on surface semantics, unstable labeling quality and difficult to adapt to large-scale text data.

Method used

Using a vector library-based method, an effective keyword collection is constructed through word embedding model and keyword extraction, and a label is generated using vector similarity search.

Benefits of technology

It realizes efficient and accurate automatic generation of text tags, adapts to text data of different fields and scales, and improves the efficiency and quality of tag generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938929A_ABST
    Figure CN119938929A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a text label automatic generation method and system based on a vector library. According to the text label automatic generation method based on the vector library, a word embedding model and keyword extraction are combined to construct an effective keyword set, and the effective keyword set is stored in a vector database; and extracting keywords of the to-be-labeled text data, determining keyword matching information corresponding to the to-be-labeled text data in combination with the word embedding model and the vector database, and generating a tag to which the to-be-labeled text data belongs based on the keyword matching information corresponding to the to-be-labeled text data. According to the text label automatic generation method and system based on the vector library, the topic information of the text can be accurately captured, the high-quality label can be generated, rapid matching and label generation of the text data are achieved, the label generation efficiency is greatly improved, the method and system can adapt to the text data of different fields and scales, and high universality and expandability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software development, and in particular to a method and system for automatically generating text labels based on a vector library. Background Art

[0002] In existing text tag generation methods, they usually rely on manual tagging or keyword matching hits. However, this method has the problems of low efficiency, reliance on surface semantics, unstable annotation quality, and difficulty in adapting to large-scale text data.

[0003] In order to solve the above problems, based on the above situation, the present invention proposes a method and system for automatically generating text labels based on a vector library. Summary of the invention

[0004] In order to overcome the defects of the prior art, the present invention provides a simple and efficient method and system for automatically generating text labels based on a vector library.

[0005] The present invention is achieved through the following technical solutions:

[0006] A method for automatically generating text labels based on a vector library, characterized in that it includes the following steps:

[0007] Step S1: Combine the word embedding model and keyword extraction to construct a valid keyword set and store it in a vector database;

[0008] Step S1.1: The management personnel customize and determine the typical keyword set corresponding to each tag based on industry experience and the characteristics of the current tagging business;

[0009] Step S1.2: For each typical keyword set corresponding to a tag, use a word embedding model to encode the keywords therein, and store the keywords, keyword encoding vectors and corresponding tags in a vector database;

[0010] Step S1.3, construct a text data set by randomly selecting a certain number of text data, extracting keywords from each text data in the text data set by combining word segmentation and keyword extraction methods, constructing an overall keyword set, and performing deduplication processing;

[0011] Step S1.4, using the word embedding model to encode the keywords in the overall keyword set, and storing the keywords, keyword encoding vectors and generated keyword unique identifiers in the vector database;

[0012] Step S1.5: for each tag, select keywords from the typical keyword set corresponding to the tag from the vector database in turn as the currently selected keyword; and use the encoding vector of the currently selected keyword as the current query vector, perform vector similarity retrieval based on the vector database, retrieve keywords whose similarity with the currently selected keyword is greater than the custom threshold, and record the similarity between each retrieved keyword and the currently selected keyword, and merge the retrieved keywords into the keyword pending set corresponding to the current tag; after no more keywords are added to the keyword pending set, perform deduplication processing and retain the maximum value of the similarity corresponding to each keyword;

[0013] Step S1.6: After determining the keyword pending set corresponding to the tag, if the keyword pending sets of each tag do not contain the same keyword, the keyword pending set corresponding to each tag is determined as the respective valid keyword set;

[0014] If the keyword pending sets of each tag contain the same keywords, the same keywords will be assigned to the keyword pending set corresponding to the tag with the highest similarity, and the keyword pending sets of tags that do not contain the same keywords will be used as the valid keyword sets of each tag;

[0015] Step S1.7: According to the effective keyword set of the tag, the corresponding tags are matched for the keywords in the overall keyword set in the vector database, and the keywords in the effective keyword set that do not belong to all tags are deleted, and the final effective keyword set is obtained after merging with the typical keyword set to remove duplicates;

[0016] In the step S1.3, when extracting keywords from each text data, the extraction quantity or keyword weight threshold is pre-defined to determine the extraction quantity of keywords;

[0017] When constructing the overall keyword set based on text data, the part of speech of the keyword is set to improve the accuracy of keyword extraction; at the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

[0018] Step S2: extract keywords from the text data to be annotated, and determine keyword matching information corresponding to the text data to be annotated by combining the word embedding model and the vector database;

[0019] Step S2.1, extract keywords from the text data to be annotated by combining word segmentation and keyword extraction technology, obtain a text keyword set, calculate keyword weights using the TF-IDF method, and record the weight corresponding to each keyword;

[0020] Step S2.2, customizing a keyword in the text keyword set to be selected as the currently selected text keyword;

[0021] Use the word embedding model to obtain the encoding vector of the currently selected text keyword as the query vector. Based on the vector database, use the vector similarity retrieval method to retrieve the similarity between the currently selected text keyword and the matching keyword, and sort the similarities from high to low. Select the first K matching keywords, where K is the return quantity threshold of the custom-set search keyword.

[0022] Obtain the associated label corresponding to each matching keyword, as well as the similarity between each matching keyword and the currently selected text keyword, and use the product of the similarity and the corresponding weight of the currently selected text keyword as the weight of the corresponding matching keyword; thereby constructing a triple, including the matching keyword, the associated label and the weight, as the matching information of each matching keyword;

[0023] Step S2.3, based on the text keyword set, select the keywords therein in turn according to the weights as the currently selected text keywords, repeat step S2.2, and obtain all matching information as the keyword matching information corresponding to the text data to be annotated;

[0024] In the step S2.1, when extracting keywords from the text data to be annotated, the extraction quantity or keyword weight threshold is pre-customized to determine the extraction quantity of keywords;

[0025] When constructing a text keyword set based on the text data to be annotated, set the part of speech of the keyword to improve the accuracy of keyword extraction;

[0026] At the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

[0027] Step S3: Generate a label for the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated;

[0028] Step S3.1, based on the keyword matching information corresponding to the text data to be annotated, group the data according to the associated tags therein, use the associated tags corresponding to each group as the potential tags of the current text data, and add the weights in the keyword matching information in each group as the matching degree between the current text data and the potential tags, which is used to measure the matching degree between the text data and the potential tags; the greater the matching degree, the greater the possibility that the text data belongs to the corresponding tag;

[0029] Step S3.2: for the single label generation task, based on the potential label matching information of the text data to be annotated, select the potential label with the greatest matching degree as the label of the current text data;

[0030] For the multi-label generation task, based on the matching information of the potential labels of the text data to be annotated, the matching ratio of each potential label is calculated; then, the primary and secondary factor analysis method is used to generate multiple labels for the text data to be annotated.

[0031] In step S3.2, the matching ratio of the potential tag is the ratio of the matching ratio of the potential tag to the total matching ratio of all potential tags, and the calculation formula is as follows:

[0032]

[0033] Among them, s j is the matching ratio of the jth potential label, p j is the matching degree of the jth potential label, and n is the total number of potential labels.

[0034] In step S3.2, the process of using the primary and secondary factor analysis method to generate multiple labels for the text data to be annotated is as follows:

[0035] Based on actual needs, customize the cumulative matching ratio threshold λ, 0<λ<1;

[0036] Arrange the potential belonging tags in descending order based on their matching ratio values, select the matching ratio values ​​corresponding to the potential belonging tags from the front to the back, add them up to obtain the cumulative matching ratio value, and add the matching ratio values ​​of all potential belonging tags until the ratio of the above two is not less than the cumulative matching ratio threshold, and use the selected potential belonging tags as multiple belonging tags of the text data to be annotated;

[0037] The calculation formula is as follows:

[0038]

[0039] Among them, n is the total number of potential labels, k is the total number of selected potential labels, and s i is the matching percentage value corresponding to the potential label i at the i-th position in descending order;

[0040] In descending order, select the first k potential labels that make f(k)≥λ and have the smallest number as the k labels of the text data to be annotated.

[0041] A text label automatic generation system based on a vector library, comprising a keyword set building module and a label automatic generation module;

[0042] The keyword set construction module is responsible for combining the word embedding model and keyword extraction to construct a valid keyword set and store it in the vector database;

[0043] The automatic label generation module is responsible for extracting keywords from the text data to be annotated, combining the word embedding model and the vector database to determine the keyword matching information corresponding to the text data to be annotated, and generating the label of the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated.

[0044] A device for automatically generating text labels based on a vector library, characterized in that it comprises a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the above method steps when executing the computer program.

[0045] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program implements the above method steps when executed by a processor.

[0046] The beneficial effects of the present invention are as follows: the method and system for automatically generating text tags based on a vector library can accurately capture the subject information of the text, generate high-quality tags, realize fast matching and tag generation of text data, greatly improve the efficiency of tag generation, can adapt to text data of different fields and scales, and has strong versatility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Attached Figure 1 It is a schematic diagram of the method for automatically generating text labels based on a vector library of the present invention. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0050] The method for automatically generating text labels based on a vector library comprises the following steps:

[0051] Step S1: Combine the word embedding model and keyword extraction to construct a valid keyword set and store it in a vector database;

[0052] Step S1.1: The management personnel customize and determine the typical keyword set corresponding to each tag based on industry experience and the characteristics of the current tagging business;

[0053] Step S1.2: For each typical keyword set corresponding to a tag, use a word embedding model to encode the keywords therein, and store the keywords, keyword encoding vectors and corresponding tags in a vector database;

[0054] Step S1.3, construct a text data set by randomly selecting a certain number of text data, extracting keywords from each text data in the text data set by combining word segmentation and keyword extraction methods, constructing an overall keyword set, and performing deduplication processing;

[0055] Step S1.4, using the word embedding model to encode the keywords in the overall keyword set, and storing the keywords, keyword encoding vectors and generated keyword unique identifiers in the vector database;

[0056] Step S1.5: for each tag, select keywords from the typical keyword set corresponding to the tag from the vector database in turn as the currently selected keyword; and use the encoding vector of the currently selected keyword as the current query vector, perform vector similarity retrieval based on the vector database, retrieve keywords whose similarity with the currently selected keyword is greater than the custom threshold, and record the similarity between each retrieved keyword and the currently selected keyword, and merge the retrieved keywords into the keyword pending set corresponding to the current tag; after no more keywords are added to the keyword pending set, perform deduplication processing and retain the maximum value of the similarity corresponding to each keyword;

[0057] Step S1.6: After determining the keyword pending set corresponding to the tag, if the keyword pending sets of each tag do not contain the same keyword, the keyword pending set corresponding to each tag is determined as the respective valid keyword set;

[0058] If the keyword pending sets of each tag contain the same keywords, the same keywords will be assigned to the keyword pending set corresponding to the tag with the highest similarity, and the keyword pending sets of tags that do not contain the same keywords will be used as the valid keyword sets of each tag;

[0059] Step S1.7: According to the effective keyword set of the tag, the corresponding tags are matched for the keywords in the overall keyword set in the vector database, and the keywords in the effective keyword set that do not belong to all tags are deleted, and the final effective keyword set is obtained after merging with the typical keyword set to remove duplicates;

[0060] In the step S1.3, when extracting keywords from each text data, the extraction quantity or keyword weight threshold is pre-defined to determine the extraction quantity of keywords;

[0061] When constructing the overall keyword set based on text data, the part of speech of the keyword is set to improve the accuracy of keyword extraction; at the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

[0062] Step S2: extract keywords from the text data to be annotated, and determine keyword matching information corresponding to the text data to be annotated by combining the word embedding model and the vector database;

[0063] Step S2.1, extract keywords from the text data to be annotated by combining word segmentation and keyword extraction technology, obtain a text keyword set, calculate keyword weights using the TF-IDF (the product of term frequency and inverse document frequency) method, and record the weight corresponding to each keyword;

[0064] Step S2.2, customizing a keyword in the text keyword set to be selected as the currently selected text keyword;

[0065] Use the word embedding model to obtain the encoding vector of the currently selected text keyword as the query vector. Based on the vector database, use the vector similarity retrieval method to retrieve the similarity between the currently selected text keyword and the matching keyword, and sort the similarities from high to low. Select the first K matching keywords, where K is the return quantity threshold of the custom-set search keyword.

[0066] Obtain the associated label corresponding to each matching keyword, as well as the similarity between each matching keyword and the currently selected text keyword, and use the product of the similarity and the corresponding weight of the currently selected text keyword as the weight of the corresponding matching keyword; thereby constructing a triple, including the matching keyword, the associated label and the weight, as the matching information of each matching keyword, specifically expressed as (matching keyword, associated label, weight);

[0067] An example of calculating the weight of matching keywords is as follows: if the current text keyword is "digitalization", its weight is 0.8, and the matching keyword matched in the vector database is "digital transportation", the similarity between the text keyword and the matching keyword is 0.9, and its label is information technology, then the generated matching information is (digital transportation, information technology, 0.72).

[0068] Step S2.3, based on the text keyword set, select the keywords therein in turn according to the weights as the currently selected text keywords, repeat step S2.2, and obtain all matching information as the keyword matching information corresponding to the text data to be annotated;

[0069] In the step S2.1, when extracting keywords from the text data to be annotated, the extraction quantity or keyword weight threshold is pre-customized to determine the extraction quantity of keywords;

[0070] When constructing a text keyword set based on the text data to be annotated, set the part of speech of the keyword to improve the accuracy of keyword extraction;

[0071] At the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

[0072] Step S3: Generate a label for the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated;

[0073] Step S3.1, based on the keyword matching information corresponding to the text data to be annotated, group the data according to the associated tags therein, use the associated tags corresponding to each group as the potential tags of the current text data, and add the weights in the keyword matching information in each group as the matching degree between the current text data and the potential tags, which is used to measure the matching degree between the text data and the potential tags; the greater the matching degree, the greater the possibility that the text data belongs to the corresponding tag;

[0074] Step S3.2: for the single label generation task, based on the potential label matching information of the text data to be annotated, select the potential label with the greatest matching degree as the label of the current text data;

[0075] For the multi-label generation task, based on the matching information of the potential labels of the text data to be annotated, the matching ratio of each potential label is calculated; then, the primary and secondary factor analysis method is used to generate multiple labels for the text data to be annotated.

[0076] In step S3.2, the matching ratio of the potential tag is the ratio of the matching ratio of the potential tag to the total matching ratio of all potential tags, and the calculation formula is as follows:

[0077]

[0078] Among them, s j is the matching ratio of the jth potential label, p j is the matching degree of the jth potential label, and n is the total number of potential labels.

[0079] In step S3.2, the process of using the primary and secondary factor analysis method to generate multiple labels for the text data to be annotated is as follows:

[0080] Based on actual needs, customize the cumulative matching ratio threshold λ, 0<λ<1;

[0081] Arrange the potential belonging tags in descending order based on their matching ratio values, select the matching ratio values ​​corresponding to the potential belonging tags from the front to the back, add them up to obtain the cumulative matching ratio value, and add the matching ratio values ​​of all potential belonging tags until the ratio of the above two is not less than the cumulative matching ratio threshold, and use the selected potential belonging tags as multiple belonging tags of the text data to be annotated;

[0082] The results of the descending sorting of the matching percentages of potential labels are as follows: (potential label 1, s1), (potential label 2, s2), ..., (potential label n, s n ), based on this, the calculation formula for the cumulative matching percentage of the first k potential labels is as follows:

[0083]

[0084] Among them, n is the total number of potential labels, k is the total number of selected potential labels, and s i is the matching percentage value corresponding to the potential label i at the i-th position in descending order;

[0085] In descending order, select the first k potential labels that make f(k)≥λ and have the smallest number as the k labels of the text data to be annotated.

[0086] An example of automatically generating labels based on the keyword matching information corresponding to the text data to be annotated is as follows:

[0087] If the keyword matching information corresponding to the text data to be annotated is: (server, information technology, 0.8), (informatization, information technology, 0.7), (printer, office supplies, 0.6), (projector, office supplies, 0.4), (rangefinder, engineering technology, 0.5);

[0088] The matching information contains two related tags: information technology and office supplies. After group calculation, the matching degree of information technology is 1.5, the matching degree of office supplies is 1, and the matching degree of engineering technology is 0.5.

[0089] For the single-label generation task, the label of the text data is information technology;

[0090] For multi-label generation tasks, calculate the matching ratio of each potential label:

[0091] The matching ratio of information technology is 1.5 / (1.5+1+0.5)=0.5;

[0092] The matching ratio of office supplies is 1 / (1.5+1+0.5)=0.33;

[0093] The matching ratio of engineering technology is 0.5 / (1.5+1+0.5)=0.17.

[0094] If the cumulative matching percentage threshold is set to 0.8, the primary and secondary factor analysis method is used, and the labels of the text data are information technology and office supplies.

[0095] The text label automatic generation system based on the vector library includes a keyword set construction module and a label automatic generation module;

[0096] The keyword set construction module is responsible for combining the word embedding model and keyword extraction to construct a valid keyword set and store it in the vector database;

[0097] The automatic label generation module is responsible for extracting keywords from the text data to be annotated, combining the word embedding model and the vector database to determine the keyword matching information corresponding to the text data to be annotated, and generating the label of the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated.

[0098] The text label automatic generation device based on the vector library comprises a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the above method steps when executing the computer program.

[0099] The readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method steps are implemented.

[0100] In summary, the method and system for automatically generating text labels based on the vector library have the following characteristics:

[0101] Efficiency: By building a vector library, fast matching and label generation of text data is achieved, greatly improving the efficiency of label generation.

[0102] Accuracy: Using pre-trained word embedding models and keyword extraction technology, we can effectively extract key information from the text, deeply understand the text semantics, and then automatically generate labels for text data efficiently and accurately.

[0103] Adaptability: The method of the present invention can adapt to text data of different fields and scales, and has strong versatility and scalability.

[0104] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for automatically generating text labels based on a vector library, characterized in that: The following steps are involved: Step S1: Combine the word embedding model and keyword extraction to construct a valid keyword set and store it in a vector database; Step S1.1: The management personnel customize and determine the typical keyword set corresponding to each tag based on industry experience and the characteristics of the current tagging business; Step S1.2: For each typical keyword set corresponding to a tag, use a word embedding model to encode the keywords therein, and store the keywords, keyword encoding vectors and corresponding tags in a vector database; Step S1.3, construct a text data set by randomly selecting text data, extract the keywords of each text data in the text data set by combining word segmentation and keyword extraction methods, construct an overall keyword set, and perform deduplication processing; Step S1.4, using the word embedding model to encode the keywords in the overall keyword set, and storing the keywords, keyword encoding vectors and generated keyword unique identifiers in the vector database; Step S1.5: for each tag, select keywords from the typical keyword set corresponding to the tag from the vector database in turn as the currently selected keyword; and use the encoding vector of the currently selected keyword as the current query vector, perform vector similarity retrieval based on the vector database, retrieve keywords whose similarity with the currently selected keyword is greater than the custom threshold, and record the similarity between each retrieved keyword and the currently selected keyword, and merge the retrieved keywords into the keyword pending set corresponding to the current tag; after no more keywords are added to the keyword pending set, perform deduplication processing and retain the maximum value of the similarity corresponding to each keyword; Step S1.6: After determining the keyword pending set corresponding to the tag, if the keyword pending sets of each tag do not contain the same keyword, the keyword pending set corresponding to each tag is determined as the respective valid keyword set; If the keyword pending sets of each tag contain the same keywords, the same keywords will be assigned to the keyword pending set corresponding to the tag with the highest similarity, and the keyword pending sets of tags that do not contain the same keywords will be used as the valid keyword sets of each tag; Step S1.7: According to the effective keyword set of the tag, the corresponding tags are matched for the keywords in the overall keyword set in the vector database, and the keywords in the effective keyword set that do not belong to all tags are deleted, and the final effective keyword set is obtained after merging with the typical keyword set to remove duplicates; Step S2: extract keywords from the text data to be annotated, and determine keyword matching information corresponding to the text data to be annotated by combining the word embedding model and the vector database; Step S2.1, extract keywords from the text data to be annotated by combining word segmentation and keyword extraction technology, obtain a text keyword set, calculate keyword weights using the TF-IDF method, and record the weight corresponding to each keyword; Step S2.2, customizing a keyword in the text keyword set to be selected as the currently selected text keyword; Use the word embedding model to obtain the encoding vector of the currently selected text keyword as the query vector. Based on the vector database, use the vector similarity retrieval method to retrieve the similarity between the currently selected text keyword and the matching keyword, and sort the similarities from high to low. Select the first K matching keywords, where K is the return quantity threshold of the custom-set search keyword. Obtain the associated label corresponding to each matching keyword, as well as the similarity between each matching keyword and the currently selected text keyword, and use the product of the similarity and the corresponding weight of the currently selected text keyword as the weight of the corresponding matching keyword; thereby constructing a triple, including the matching keyword, the associated label and the weight, as the matching information of each matching keyword; Step S2.3, based on the text keyword set, select the keywords therein in turn according to the weights as the currently selected text keywords, repeat step S2.2, and obtain all matching information as the keyword matching information corresponding to the text data to be annotated; Step S3: Generate a label for the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated; Step S3.1, based on the keyword matching information corresponding to the text data to be annotated, group the data according to the associated tags therein, use the associated tags corresponding to each group as the potential tags of the current text data, and add the weights in the keyword matching information in each group as the matching degree between the current text data and the potential tags, which is used to measure the matching degree between the text data and the potential tags; Step S3.2: for the single label generation task, based on the potential label matching information of the text data to be annotated, select the potential label with the greatest matching degree as the label of the current text data; For multi-label generation tasks, based on the potential label matching information of the text data to be annotated, calculate the matching percentage of each potential label; Then, the primary and secondary factor analysis method is used to generate multiple labels for the text data to be annotated.

2. The method for automatically generating text labels based on a vector library according to claim 1, characterized in that: In the step S1.3, when extracting keywords from each text data, the extraction quantity or keyword weight threshold is pre-defined to determine the extraction quantity of keywords; When constructing the overall keyword set based on text data, the part of speech of the keyword is set to improve the accuracy of keyword extraction; at the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

3. The method for automatically generating text labels based on a vector library according to claim 1, characterized in that: In the step S2.1, when extracting keywords from the text data to be annotated, the extraction quantity or keyword weight threshold is pre-customized to determine the extraction quantity of keywords; When constructing a text keyword set based on the text data to be annotated, set the part of speech of the keyword to improve the accuracy of keyword extraction; At the same time, by counting the number of times each keyword appears, the keywords whose number of appearances exceeds the custom threshold and are irrelevant to the annotation task are added to the stop word list.

4. The method for automatically generating text labels based on a vector library according to claim 1, characterized in that: In step S3.2, the matching ratio of the potential tag is the ratio of the matching ratio of the potential tag to the total matching ratio of all potential tags, and the calculation formula is as follows: Among them, s j is the matching ratio of the jth potential label, p j is the matching degree of the jth potential label, and n is the total number of potential labels.

5. The method for automatically generating text labels based on a vector library according to claim 1, characterized in that: In step S3.2, the process of using the primary and secondary factor analysis method to generate multiple labels for the text data to be annotated is as follows: Based on actual needs, customize the cumulative matching ratio threshold λ, 0<λ<1; Arrange the potential belonging tags in descending order based on their matching ratio values, select the matching ratio values ​​corresponding to the potential belonging tags from the front to the back, add them up to obtain the cumulative matching ratio value, and add the matching ratio values ​​of all potential belonging tags until the ratio of the above two is not less than the cumulative matching ratio threshold, and use the selected potential belonging tags as multiple belonging tags of the text data to be annotated; The calculation formula is as follows: Among them, n is the total number of potential labels, k is the total number of selected potential labels, and s i is the matching percentage value corresponding to the potential label i at the i-th position in descending order; In descending order, select the first k potential labels that make f(k)≥λ and have the smallest number as the k labels of the text data to be annotated.

6. A text label automatic generation system based on a vector library, characterized by: Includes keyword set building module and label automatic generation module; The keyword set construction module is responsible for combining the word embedding model and keyword extraction to construct a valid keyword set and store it in the vector database; The automatic label generation module is responsible for extracting keywords from the text data to be annotated, combining the word embedding model and the vector database to determine the keyword matching information corresponding to the text data to be annotated, and generating the label of the text data to be annotated based on the keyword matching information corresponding to the text data to be annotated.

7. A text label automatic generation device based on a vector library, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the method according to any one of claims 1 to 5 when executing the computer program.

8. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Method and system for identifying disability grade based on natural language

    CN120636480A