Method and device for automatically mining target associated words and electronic equipment
By calculating the word TF-IDF' value and correlation score in the document set, the target conjunction is automatically mined, which solves the inefficiency problem of relying on manual annotation and simple statistics in the existing technology, and achieves efficient target conjunction extraction.
Patent Information
- Application Number
- CN202510082560.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the target conjunction extraction method relies on manual annotation or simple statistical methods, is inefficient and difficult to adapt to the needs of large-scale data and complex scenarios.
By obtaining the target-related and unrelated document sets, word segmentation is performed, and the correlation score of each word is calculated using the TF-IDF' value, and the target conjunction is automatically mined.
It achieves accurate and automatic mining of target conjunctions, is efficient, can adapt to large-scale data and complex scenarios, and improves user experience.
Smart Images

Figure CN120012773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method, device and electronic device for automatically mining target associated words. Background Art
[0002] With the rapid development of the Internet, a large amount of text data is generated and stored. The growth of data volume increases the richness on the one hand, but also causes information redundancy on the other hand, which may lead to the generation of a large amount of irrelevant information, and then lead to inefficiency and inaccuracy in information search or processing. Therefore, in fields such as e-commerce and search engines, how to accurately mine target associated words is crucial to improving user experience and business efficiency. At present, the target associated word extraction method often relies on manual annotation or simple statistical methods. This method is not only inefficient, but also has a small amount of data that can be processed, and it is difficult to adapt to the needs of large-scale data and complex scenarios. Therefore, developing a technology that can automatically mine target associated words has important practical application value. Summary of the invention
[0003] In order to solve the problems existing in the prior art, the present invention provides the following technical solutions.
[0004] A first aspect of the present invention provides a method for automatically mining target associated words, comprising:
[0005] Obtain a target-related document set and a target-irrelevant document set;
[0006] Perform word segmentation on the document texts in all document sets;
[0007] The TF-IDF′ value of each word in each document is calculated using the following formula:
[0008] TF′(w,d)=log(number of times word w appears in document d+1);
[0009]
[0010] TF-IDF′(w,d)=TF′(w,d)×IDF′(w);
[0011] Calculate the weighted average of the TF-IDF′ value of each word in the target related document set and the target irrelevant document set, respectively, and record them as S(w1,D1) and S(w2,D2) respectively; where w1 is the word in the target related document set D1, and w2 is the word in the target irrelevant document set D2;
[0012] The relevance score R(w1) of each word in the target related document set is calculated using the following formula:
[0013]
[0014] Among them, R(w1) is the relevance score of word w1, and T is the threshold;
[0015] The words whose relevance scores meet the requirements are used as target associated words.
[0016] Preferably, the data of the target-related document set and the target-irrelevant document set are taken from description texts of electronic, clothing and / or food products on the e-commerce platform and / or patent texts in a patent database.
[0017] Preferably, the word segmentation process further includes the following steps:
[0018] Use a stop word list to remove words that have no real meaning;
[0019] Convert all characters to lowercase and unify the format;
[0020] Filter out words that are less than two characters long or too common;
[0021] Generate a document vocabulary database, including document code, document name, words, and the number of times the words appear in the document.
[0022] Preferably, the method further comprises the step of recording the target associated words and their relevance scores in the document word library database.
[0023] A second aspect of the present invention provides a device for automatically mining target associated words, comprising:
[0024] A data acquisition module, used for acquiring a target-related document set and a target-irrelevant document set;
[0025] The word segmentation module is used to perform word segmentation on the document texts in all document sets;
[0026] The first calculation module is used to calculate the TF-IDF′ value of each word in each document using the following formula:
[0027] TF′(w, d) = log(number of times word w appears in document d + 1);
[0028]
[0029] TF-IDF′(w,d)=TF′(w,d)×IDF′(w);
[0030] The second calculation module is used to calculate the weighted average of the TF-IDF′ value of each word in the target related document set and the target irrelevant document set, which are recorded as S(w1, D1) and S(w2, D2) respectively; wherein w1 is a word in the target related document set D1, and w2 is a word in the target irrelevant document set D2;
[0031] The third calculation module is used to calculate the relevance score R(w1) of each word in the target related document set using the following formula:
[0032]
[0033] Among them, R(w1) is the relevance score of word w1, and T is the threshold;
[0034] The target associated word determination module is used to take words whose relevance scores meet the requirements as target associated words.
[0035] Preferably, the data of the target-related document set and the target-irrelevant document set are taken from description texts of electronic, clothing and / or food products on the e-commerce platform and / or patent texts in a patent database.
[0036] Preferably, the device further comprises a preprocessing module, which is used for:
[0037] Use a stop word list to remove words that have no real meaning;
[0038] Convert all characters to lowercase and unify the format;
[0039] Filter out words that are less than two characters long or too common;
[0040] Generate a document vocabulary database, including document code, document name, words, and the number of times the words appear in the document.
[0041] Preferably, the device further comprises a recording module, which is used to record the target associated words and their relevance scores in the document word library database.
[0042] A third aspect of the present invention provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for automatically mining target associated words as described in the first aspect.
[0043] The fourth aspect of the present invention provides an electronic device, comprising a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, which can be loaded and executed by the processor so that the processor can execute the method for automatically mining target associated words as described in the first aspect.
[0044] The beneficial effects of the present invention are as follows: the method, device and electronic device provided by the present invention for automatically mining target associated words first calculate the TF-IDF' value of each word in each document in the target related document set and the target unrelated document set, then calculate the weighted average of the TF-IDF' value of each word in the target related document set and the target unrelated document set, and then use the weighted average of the TF-IDF' value to calculate the relevance score of each word in the target related document set, and finally determine the target associated words according to the relevance score. The technical solution provided by the present invention can realize accurate automatic mining of target associated words, which is not only efficient, but also can better adapt to the needs of large-scale data and complex scenarios, improve user experience, and has important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the process of the method for automatically mining target associated words according to the present invention;
[0046] Figure 2 The figure is a functional structure diagram of the device for automatically mining target associated words according to the present invention. DETAILED DESCRIPTION
[0047] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0048] The method provided by the present invention can be implemented in the following terminal environment, and the terminal may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method described in the following embodiment.
[0049] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the entire terminal, and executes various functions of the terminal and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory.
[0050] The memory may include random access memory (RAM) or read-only memory (ROM). The memory may be used to store instructions, programs, codes, code sets or instructions.
[0051] The display screen is used to display the user interface of each application.
[0052] In addition, those skilled in the art can understand that the structure of the above terminal does not constitute a limitation on the terminal, and the terminal may include more or fewer components, or combine certain components, or arrange the components differently. For example, the terminal also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, and a power supply, which will not be described in detail here.
[0053] Embodiment 1
[0054] like Figure 1 As shown, an embodiment of the present invention provides a method for automatically mining target associated words, including:
[0055] S101, obtaining a target-related document set and a target-irrelevant document set;
[0056] S102, performing word segmentation processing on the document texts in all document sets;
[0057] S103, calculate the TF-IDF′ value of each word in each document using the following formula:
[0058] TF′(w,d)=log(number of times word w appears in document d+1);
[0059]
[0060] TF-IDF′(w,d)=TF′(w,d)×IDF′(w);
[0061] S104, respectively calculating the weighted average of the TF-IDF′ value of each word in the target relevant document set and the target irrelevant document set, which are recorded as S(w1, D1) and S(w2, D2) respectively; wherein w1 is a word in the target relevant document set D1, and w2 is a word in the target irrelevant document set D2;
[0062] S105, using the following formula to calculate the relevance score R(w1) of each word in the target related document set:
[0063]
[0064] Among them, R(w1) is the relevance score of word w1, and T is the threshold;
[0065] S106, taking the words whose relevance scores meet the requirements as target associated words.
[0066] In the above method provided by the present invention, a smoothing mechanism is introduced when calculating TF-IDF′ to avoid the problem that the weight of some words w is underestimated due to the very few times they appear in the document d. The improved calculation formula is as follows:
[0067] TF′(w,d)=log(number of times word w appears in document d+1);
[0068]
[0069] By improving the formula, even if a word is relatively rare in the entire document set, it can still get a higher weight if it appears frequently in a specific document. Therefore, the calculation formula of the improved TF-IDF′ is:
[0070] TF-IDF′(w,d)=TF′(w,d)×IDF′(w);
[0071] According to the above improved TF-IDF' value calculation formula, the TF-IDF' value of each word in each document is calculated, and then the weighted average of the TF-IDF' value of each word in the target relevant document set D1 and the irrelevant document set D2 is summarized, which are recorded as S(w1,D1) respectively.
[0072] S(w2,D2). For the same word, if it exists in both D1 and D2, and the weighted average of the TF-IDF′ value is not much different, it is considered a neutral word. By setting a threshold T, such neutral words are filtered out. The formula for the relevance score R(w1) of all words in the final target related document set D1 is:
[0073]
[0074] The data of the target-related document set and the target-unrelated document set are taken from the description texts of electronic, clothing and / or food products on the e-commerce platform and / or the patent texts in the patent database.
[0075] In a preferred embodiment of the present invention, the word segmentation processing also includes the following steps: applying a stop word list to remove words that have no practical meaning; converting all characters to lowercase and unifying the format; filtering out words that are less than two characters in length or are too common; and generating a document vocabulary database, including document codes, document names, words, and the number of times the words appear in the document.
[0076] In another preferred embodiment of the present invention, the method further comprises the step of recording the target associated words and their relevance scores in the document word library database.
[0077] Embodiment 2
[0078] like Figure 2 As shown, another aspect of the present invention also includes a functional module architecture that is completely consistent with the aforementioned method flow, that is, the embodiment of the present invention also provides a device for automatically mining target associated words, including:
[0079] A data acquisition module 201 is used to acquire a target-related document set and a target-irrelevant document set;
[0080] The word segmentation module 202 is used to perform word segmentation processing on the document texts in all document sets;
[0081] The first calculation module 203 is used to calculate the TF-IDF′ value of each word in each document using the following formula:
[0082] TF′(w,d)=log(number of times word w appears in document d+1);
[0083]
[0084] TF-IDF′(w,d)=TF′(w,d)×IDF′(w);
[0085] The second calculation module 204 is used to calculate the weighted average of the TF-IDF′ value of each word in the target related document set and the target irrelevant document set, which are recorded as S(w1, D1) and S(w2, D2) respectively; wherein w1 is a word in the target related document set D1, and w2 is a word in the target irrelevant document set D2;
[0086] The third calculation module 205 is used to calculate the relevance score R(w1) of each word in the target related document set using the following formula:
[0087]
[0088] Among them, R(w1) is the relevance score of word w1, and T is the threshold;
[0089] The target associated word determination module 206 is used to take words whose relevance scores meet the requirements as target associated words.
[0090] The data of the target-related document set and the target-unrelated document set are taken from the description texts of electronic, clothing and / or food products on the e-commerce platform and / or the patent texts in the patent database.
[0091] Furthermore, the device also includes a preprocessing module, which is used to: apply a stop word list to remove words that have no practical meaning; convert all characters to lowercase and unify the format; filter out words that are less than two characters in length or are too common; and generate a document vocabulary database, which includes document codes, document names, words, and the number of times the words appear in the document.
[0092] Furthermore, the device also includes a recording module, which is used to record the target associated words and their relevance scores in the document word library database.
[0093] The device can be implemented by the method for automatically mining target associated words provided in the above-mentioned embodiment 1. The specific implementation method can be found in the description of embodiment 1, which will not be repeated here.
[0094] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for automatically mining target associated words as described in the first embodiment.
[0095] The present invention also provides an electronic device, comprising a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for automatically mining target associated words as described in Example 1.
[0096] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for automatically mining target related words, characterized in that: include: Obtain a target-related document set and a target-irrelevant document set; Perform word segmentation on the document texts in all document sets; The TF-IDF' value of each word in each document is calculated using the following formula: TF'(w,d) = log(number of times word w appears in document d + 1); TF-IDF'(w,d)=TF'(w,d)×IDF'(w); Calculate the weighted average of the TF-IDF' value of each word in the target related document set and the target irrelevant document set, respectively, and record them as S(w1,D1) and S(w2,D2) respectively; where w1 is the word in the target related document set D1, and w2 is the word in the target irrelevant document set D2; The relevance score R(w1) of each word in the target related document set is calculated using the following formula: Among them, R(w1) is the relevance score of word w1, and T is the threshold; The words whose relevance scores meet the requirements are used as target associated words.
2. The method for automatically mining target associated words according to claim 1, characterized in that: The data of the target-related document set and the target-irrelevant document set are taken from the description texts of electronic, clothing and / or food products on the e-commerce platform and / or the patent texts in the patent database.
3. The method for automatically mining target associated words according to claim 1, characterized in that: The word segmentation process further includes the following steps: Use a stop word list to remove words that have no real meaning; Convert all characters to lowercase and unify the format; Filter out words that are less than two characters long or too common; Generate a document vocabulary database, including document code, document name, words, and the number of times the words appear in the document.
4. The method for automatically mining target associated words according to claim 3, characterized in that: The method further comprises the step of recording the target associated words and their relevance scores in the document word library database.
5. A device for automatically mining target related words, characterized in that: include: A data acquisition module, used for acquiring a target-related document set and a target-irrelevant document set; The word segmentation module is used to perform word segmentation on the document texts in all document sets; The first calculation module is used to calculate the TF-IDF' value of each word in each document using the following formula: TF'(w,d) = log(number of times word w appears in document d + 1); TF-IDF'(w,d)=TF'(w,d)×IDF'(w); The second calculation module is used to calculate the weighted average of the TF-IDF' value of each word in the target related document set and the target irrelevant document set, which are recorded as S(w1,D1) and S(w2,D2) respectively; wherein w1 is a word in the target related document set D1, and w2 is a word in the target irrelevant document set D2; The third calculation module is used to calculate the relevance score R(w1) of each word in the target related document set using the following formula: Among them, R(w1) is the relevance score of word w1, and T is the threshold; The target associated word determination module is used to take words whose relevance scores meet the requirements as target associated words.
6. The device for automatically mining target associated words according to claim 5, characterized in that: The data of the target-related document set and the target-irrelevant document set are taken from the description texts of electronic, clothing and / or food products on the e-commerce platform and / or the patent texts in the patent database.
7. The device for automatically mining target associated words according to claim 5, characterized in that: The device also includes a pre-processing module, which is used to: Use a stop word list to remove words that have no real meaning; Convert all characters to lowercase and unify the format; Filter out words that are less than two characters long or too common; Generate a document vocabulary database, including document code, document name, words, and the number of times the words appear in the document.
8. The method for automatically mining target associated words according to claim 7, characterized in that: The device also includes a recording module, which is used to record the target associated words and their relevance scores in the document word library database.
9. A memory, characterized in that: A plurality of instructions are stored, and the instructions are used to implement the method for automatically mining target associated words as described in any one of claims 1-4.
10. An electronic device, characterized in that: It comprises a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for automatically mining target associated words as described in any one of claims 1-4.