Word extraction method, device, electronic device and storage medium
By comprehensively considering the corpus frequency of the first and second industries, candidate parameters are determined, thus improving the accuracy of word extraction, which is particularly suitable for hot word extraction unique to the first industry.
Patent Information
- Application Number
- CN202211439902.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The word extraction method in the prior art has low accuracy in hot word extraction scenarios.
By comprehensively considering the first corpus of the first industry and the second corpus of the second industry, the first frequency and second frequency of the candidate words are obtained, the candidate parameters are determined, and the target words are extracted from the multiple candidate words.
It improves the accuracy of word extraction and is especially suitable for hot word extraction scenarios unique to the first industry.
Smart Images

Figure CN115757681B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of energy technologies, and particularly to a method, apparatus, electronic device, and storage medium for word extraction. Background Art
[0002] Currently, with the continuous development of big data technologies, word extraction has been widely applied in the field of data processing. For example, hot words can be extracted from documents in a certain industry to understand the technical focus, technical hotspots, industry development trends, etc. in that industry. However, the word extraction methods in related technologies, especially in the scenario of hot word extraction, have the problem of low accuracy. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems in the above technologies to some extent.
[0004] To this end, one objective of the present invention is to propose a word extraction method, which can comprehensively consider the first corpus of the first industry and the second corpus of the second industry to determine the candidate parameters of the first candidate words, so as to extract target words from multiple first candidate words, improving the accuracy of word extraction, and particularly applicable to the hot word extraction scenario unique to the first industry.
[0005] The second objective of the present invention is to propose a word extraction apparatus.
[0006] The third objective of the present invention is to propose an electronic device.
[0007] The fourth objective of the present invention is to propose a computer-readable storage medium.
[0008] The first aspect embodiment of the present invention proposes a word extraction method, including: obtaining multiple first candidate words of the first industry; obtaining the first frequency of the first candidate words in the first corpus of the first industry and the second frequency of the first candidate words in the second corpus of the second industry; determining the candidate parameters of the first candidate words based on the first frequency and the second frequency; and extracting target words from multiple first candidate words based on the candidate parameters.
[0009] According to the word extraction method of an embodiment of the present invention, a plurality of first candidate words in a first industry are obtained, the first frequency of the first candidate words in the first corpus of the first industry is obtained, and the second frequency of the first candidate words in the second corpus of the second industry is obtained. Based on the first frequency and the second frequency, the candidate parameter of the first candidate word is determined, and based on the candidate parameter, the target word is extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameter of the first candidate word, so as to extract the target word from the plurality of first candidate words, improving the accuracy of word extraction, and is particularly applicable to the scenario of extracting hot words unique to the first industry.
[0010] In addition, the word extraction method proposed according to the above embodiment of the present invention may further have the following additional technical features:
[0011] In an embodiment of the present invention, the determining the candidate parameter of the first candidate word based on the first frequency and the second frequency includes: identifying whether there is a second candidate word nested in the first candidate word in the first corpus; in response to the non-existence of the second candidate word in the first corpus, determining the candidate parameter based on the first frequency and the second frequency.
[0012] In an embodiment of the present invention, the determining the candidate parameter based on the first frequency and the second frequency includes: obtaining the sum value between the first frequency and the second frequency; obtaining the first ratio between the first frequency and the sum value; determining the candidate parameter based on the first ratio.
[0013] In an embodiment of the present invention, it further includes: in response to the existence of the second candidate word in the first corpus, obtaining the third frequency of the second candidate word in the first corpus; determining the candidate parameter based on the first frequency, the second frequency and the third frequency.
[0014] In an embodiment of the present invention, the determining the candidate parameter based on the first frequency, the second frequency and the third frequency includes: obtaining the average value of the plurality of third frequencies; determining the candidate parameter based on the first frequency, the second frequency and the average value.
[0015] In an embodiment of the present invention, the determining the candidate parameter based on the first frequency, the second frequency and the average value includes: obtaining the sum value between the first frequency and the second frequency; obtaining the difference between the first frequency and the average value; obtaining the second ratio between the difference and the sum value; determining the candidate parameter based on the second ratio.
[0016] In one embodiment of the present invention, the obtaining of a plurality of first candidate words in the first industry includes: obtaining the original text of the first industry; performing word segmentation on the original text to obtain a plurality of initial words; and extracting a plurality of the first candidate words from the plurality of initial words.
[0017] In one embodiment of the present invention, the extracting of a target word from a plurality of the first candidate words based on the candidate parameter includes: sorting the plurality of candidate parameters in descending order; and determining the first candidate words corresponding to the top N sorted candidate parameters as the target word, where N is a positive integer.
[0018] In one embodiment of the present invention, the second industry includes at least one industry different from the first industry.
[0019] A second aspect embodiment of the present invention provides a word extraction device, including: a first obtaining module, configured to obtain a plurality of first candidate words in a first industry; a second obtaining module, configured to obtain a first frequency of the first candidate words in a first corpus of the first industry and a second frequency of the first candidate words in a second corpus of a second industry; a determining module, configured to determine a candidate parameter of the first candidate words based on the first frequency and the second frequency; and an extracting module, configured to extract a target word from a plurality of the first candidate words based on the candidate parameter.
[0020] The word extraction device according to the embodiments of the present invention obtains a plurality of first candidate words in a first industry, obtains a first frequency of the first candidate words in a first corpus of the first industry and a second frequency of the first candidate words in a second corpus of a second industry, determines a candidate parameter of the first candidate words based on the first frequency and the second frequency, and extracts a target word from a plurality of the first candidate words based on the candidate parameter. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameter of the first candidate words, so as to extract the target word from a plurality of the first candidate words, improving the accuracy of word extraction, and is particularly applicable to the scenario of extracting hot words unique to the first industry.
[0021] In addition, the word extraction device according to the above embodiments of the present invention may further have the following additional technical features:
[0022] In one embodiment of the present invention, the determining module is further configured to: identify whether there is a second candidate word nested in the first candidate word in the first corpus; and in response to the non-existence of the second candidate word in the first corpus, determine the candidate parameter based on the first frequency and the second frequency.
[0023] In an embodiment of the present invention, the determining module is further configured to: obtain a sum value between the first frequency and the second frequency; obtain a first ratio between the first frequency and the sum value; and determine the candidate parameter based on the first ratio.
[0024] In an embodiment of the present invention, the determining module is further configured to: in response to the existence of the second candidate word in the first corpus, obtain a third frequency of the second candidate word in the first corpus; and determine the candidate parameter based on the first frequency, the second frequency, and the third frequency.
[0025] In an embodiment of the present invention, the determining module is further configured to: obtain an average value of the multiple third frequencies; and determine the candidate parameter based on the first frequency, the second frequency, and the average value.
[0026] In an embodiment of the present invention, the determining module is further configured to: obtain a sum value between the first frequency and the second frequency; obtain a difference value between the first frequency and the average value; obtain a second ratio between the difference value and the sum value; and determine the candidate parameter based on the second ratio.
[0027] In an embodiment of the present invention, the first obtaining module is further configured to: obtain an original text of the first industry; perform word segmentation processing on the original text to obtain a plurality of initial words; and extract a plurality of the first candidate words from the plurality of initial words.
[0028] In an embodiment of the present invention, the extracting module is further configured to: sort the plurality of candidate parameters in descending order; and determine the first candidate words corresponding to the top N candidate parameters as the target words, where N is a positive integer.
[0029] In an embodiment of the present invention, the second industry includes at least one industry different from the first industry.
[0030] An embodiment of the third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the word extraction method as described in the embodiment of the first aspect of the present invention is implemented.
[0031] In the electronic device according to the embodiment of the present invention, a computer program stored in a memory is executed by a processor to obtain a plurality of first candidate words in a first industry, obtain a first frequency of the first candidate words in a first corpus of the first industry, and a second frequency of the first candidate words in a second corpus of a second industry. Based on the first frequency and the second frequency, candidate parameters of the first candidate words are determined, and based on the candidate parameters, target words are extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameters of the first candidate words, so as to extract target words from the plurality of first candidate words, improving the accuracy of word extraction, and especially applicable to the scenario of extracting hot words unique to the first industry.
[0032] In the fourth aspect of the present application, an embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the word extraction method described in the first aspect of the present invention is implemented.
[0033] In the computer-readable storage medium according to the embodiment of the present invention, by storing a computer program and executing it by a processor, a plurality of first candidate words in a first industry are obtained, a first frequency of the first candidate words in a first corpus of the first industry, and a second frequency of the first candidate words in a second corpus of a second industry are obtained. Based on the first frequency and the second frequency, candidate parameters of the first candidate words are determined, and based on the candidate parameters, target words are extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameters of the first candidate words, so as to extract target words from the plurality of first candidate words, improving the accuracy of word extraction, and especially applicable to the scenario of extracting hot words unique to the first industry.
[0034] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0035] The above-mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0036] Figure 1 It is a schematic flowchart of a word extraction method according to an embodiment of the present invention;
[0037] Figure 2 It is a schematic flowchart of determining candidate parameters in a word extraction method according to an embodiment of the present invention;
[0038] Figure 3 It is a schematic flowchart of determining candidate parameters in a word extraction method according to another embodiment of the present invention;
[0039] Figure 4 Flow chart of the word extraction method according to a specific example of the present invention;
[0040] Figure 5 Structural diagram of the word extraction device according to an embodiment of the present invention;
[0041] Figure 6 Structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0042] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0043] The word extraction method, device, electronic device, and storage medium according to the embodiments of the present invention will be described below with reference to the drawings.
[0044] Figure 1 Flow chart of the word extraction method according to an embodiment of the present invention.
[0045] As Figure 1 shown, the word extraction method according to the embodiment of the present invention includes:
[0046] S101, obtaining a plurality of first candidate words in a first industry.
[0047] It should be noted that there are no excessive limitations on the first industry. For example, the first industry includes energy, vehicles, home appliances, online sales, logistics, etc.
[0048] It should be noted that there are no excessive limitations on the first candidate words. For example, the first candidate words may include terms. For example, when the first industry is energy, the first candidate words may include wind power generation, solar energy, clean energy, renewable, etc. When the first industry is vehicles, the first candidate words may include autonomous driving, lane lines, maps, obstacles, etc.
[0049] It should be noted that there are no excessive limitations on the number of the first candidate words. For example, the number of the first candidate words may include 3, 10, etc.
[0050] In one implementation manner, obtaining a plurality of first candidate words in a first industry includes extracting a plurality of first candidate words from the database of the first industry. It can be understood that a database may be established in advance for the first industry, and the database is used to store data such as documents, pictures, and videos of the first industry.
[0051] In one implementation, obtaining a plurality of first candidate words in a first industry includes obtaining the original text of the first industry, performing word segmentation on the original text to obtain a plurality of initial words, and extracting a plurality of first candidate words from the plurality of initial words. It should be noted that there are no excessive restrictions on the original text. For example, the original text may include news, comments, papers, patents, etc.
[0052] In some examples, extracting a plurality of first candidate words from the plurality of initial words includes performing keyword extraction on the plurality of initial words to obtain a plurality of first candidate words.
[0053] It should be noted that there are no excessive restrictions on the specific methods of word segmentation and keyword extraction. For example, the NLPIR algorithm can be used for word segmentation and keyword extraction. Among them, the NLPIR algorithm refers to a word segmentation algorithm based on NLP (Natural Language Processing).
[0054] S102, obtaining the first frequency of the first candidate word in the first corpus of the first industry, and the second frequency of the first candidate word in the second corpus of the second industry.
[0055] It should be noted that the first industry is different from the second industry.
[0056] In one implementation, the second industry includes at least one industry different from the first industry. For example, when the first industry is energy, the second industry includes vehicles, home appliances, online sales, and logistics. Or, when the first industry is vehicles, the second industry includes energy, home appliances, online sales, and logistics.
[0057] In the embodiments of the present disclosure, a first corpus can be established in advance for the first industry, and a second corpus can be established in advance for the second industry.
[0058] For example, when the first industry is energy and the second industry includes vehicles, the first corpus includes corpus 1 of energy, and the second corpus includes corpus 2 of vehicles. If the first candidate word is "solar energy", the first frequency of "solar energy" in corpus 1 can be obtained, and the second frequency of "solar energy" in corpus 2 can be obtained.
[0059] In one implementation, obtaining the second frequency of the first candidate word in the second corpus of the second industry includes obtaining the second frequency of the first candidate word in a plurality of second corpora.
[0060] For example, the first industry is energy, the second industry includes vehicles, home appliances, online sales, and logistics. The first corpus includes Corpus 1 of energy, and the second corpus includes Corpus 2 of vehicles, Corpus 3 of home appliances, Corpus 4 of online sales, and Corpus 5 of logistics. If the first candidate word is "solar energy", the first frequency of "solar energy" in Corpus 1 can be obtained, and the second frequency of "solar energy" in Corpora 2 to 5 can be obtained.
[0061] It can be understood that different first candidate words can correspond to different first frequencies and second frequencies.
[0062] S103. Based on the first frequency and the second frequency, determine the candidate parameter of the first candidate word.
[0063] In one implementation, based on the first frequency and the second frequency, determining the candidate parameter of the first candidate word includes inputting the first frequency and the second frequency into a set algorithm or model for processing to obtain the candidate parameter. It should be noted that the set algorithm or model is not overly limited. For example, the set algorithm or model can be pre-set or generated in real time.
[0064] It can be understood that different first candidate words can correspond to different candidate parameters.
[0065] S104. Based on the candidate parameter, extract the target word from multiple first candidate words.
[0066] It should be noted that the number of target words is not overly limited. For example, the number of target words is at least one. The target word is part or all of the first candidate words.
[0067] In one implementation, based on the candidate parameter, extracting the target word from multiple first candidate words includes sorting the multiple candidate parameters in descending order, and determining the first candidate words corresponding to the top N candidate parameters as the target words, where N is a positive integer. It should be noted that N is not overly limited. For example, N can be 1, 3, etc. When N is 1, the first candidate word corresponding to the largest candidate parameter is determined as the target word.
[0068] In one implementation, based on the candidate parameter, extracting the target word from multiple first candidate words includes determining the first candidate words corresponding to the candidate parameters greater than a set threshold as the target words. It should be noted that the set threshold is not overly limited.
[0069] In one implementation, it further includes determining the target word as a hot word in the first industry.
[0070] In summary, according to the word extraction method of the embodiments of the present invention, a plurality of first candidate words in the first industry are obtained, the first frequency of the first candidate words in the first corpus of the first industry and the second frequency of the first candidate words in the second corpus of the second industry are obtained, based on the first frequency and the second frequency, the candidate parameters of the first candidate words are determined, and based on the candidate parameters, target words are extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameters of the first candidate words, so as to extract target words from the plurality of first candidate words, improving the accuracy of word extraction, especially suitable for the scenario of extracting hot words unique to the first industry.
[0071] Based on any of the above embodiments, as Figure 2 shown, in step S103, based on the first frequency and the second frequency, determining the candidate parameters of the first candidate words includes:
[0072] S201, identifying whether there is a second candidate word nested with the first candidate word in the first corpus.
[0073] S202, in response to the absence of the second candidate word in the first corpus, determining the candidate parameters based on the first frequency and the second frequency.
[0074] It should be noted that the second candidate word nested with the first candidate word includes the case where the second candidate word contains the first candidate word. For example, if the first candidate word is "renewable" and there is "renewable energy" in the first corpus, it can be identified that there is "renewable energy" nested with "renewable" in the first corpus, that is, "renewable energy" is the second candidate word.
[0075] In one implementation manner, determining the candidate parameters based on the first frequency and the second frequency includes obtaining the sum value between the first frequency and the second frequency, obtaining the first ratio between the first frequency and the sum value, and determining the candidate parameters based on the first ratio.
[0076] In some examples, determining the candidate parameters based on the first ratio includes obtaining the length of the first candidate word, and determining the candidate parameters based on the length, the first frequency and the first ratio. Thus, in this method, the length, the first frequency and the first ratio of the first candidate word can be comprehensively considered to determine the candidate parameters, improving the accuracy of the candidate parameters.
[0077] For example, determining the candidate parameters based on the first frequency and the second frequency can be achieved through the following formula:
[0078]
[0079] Among them, the SC-value is a candidate parameter, s is the length, sf(s) is the first frequency, bf(s) is the second frequency, and bf(s) + sf(s) is the sum value. is the first ratio.
[0080] Thus, when the second candidate word does not exist in the first corpus in this method, the candidate parameter can be determined based on the first frequency and the second frequency.
[0081] Based on any of the above embodiments, as Figure 3 shown, determining the candidate parameter of the first candidate word based on the first frequency and the second frequency in step S103 includes:
[0082] S301, identifying whether there is a second candidate word nested with the first candidate word in the first corpus.
[0083] For the relevant content of step S301, reference can be made to the above embodiments and will not be elaborated here.
[0084] S302, in response to the existence of the second candidate word in the first corpus, obtaining the third frequency of the second candidate word in the first corpus.
[0085] It should be noted that the number of the second candidate words is not overly limited. For example, the number of the second candidate words is at least one.
[0086] For example, if the first industry is energy and the first corpus includes corpus 1 of energy, if the first candidate word is "renewable", it can be identified that there are second candidate words "renewable energy" and "renewable resources" in corpus 1, obtaining the third frequency of "renewable energy" in corpus 1 and obtaining the third frequency of "renewable resources" in corpus 1.
[0087] S303, determining the candidate parameter based on the first frequency, the second frequency, and the third frequency.
[0088] In one implementation, determining the candidate parameter based on the first frequency, the second frequency, and the third frequency includes inputting the first frequency, the second frequency, and the third frequency into a set algorithm or model for processing to obtain the candidate parameter. It should be noted that the set algorithm or model is not overly limited. For example, the set algorithm or model can be pre-set or generated in real time.
[0089] In one implementation, determining the candidate parameter based on the first frequency, the second frequency, and the third frequency includes obtaining the average value of multiple third frequencies and determining the candidate parameter based on the first frequency, the second frequency, and the average value.
[0090] In some examples, based on a first frequency, a second frequency, and an average value, a candidate parameter is determined, including obtaining a sum value between the first frequency and the second frequency, obtaining a difference between the first frequency and the average value, obtaining a second ratio between the difference and the sum value, and determining the candidate parameter based on the second ratio.
[0091] In some examples, based on the second ratio, a candidate parameter is determined, including obtaining the length of a first candidate word, and determining the candidate parameter based on the length, the first frequency, and the second ratio. Thus, in this method, the length, the first frequency, and the second ratio of the first candidate word can be comprehensively considered to determine the candidate parameter, improving the accuracy of the candidate parameter.
[0092] For example, based on a first frequency, a second frequency, and a third frequency, a candidate parameter is determined, which can be implemented through the following formula:
[0093]
[0094] where SC-value is the candidate parameter, s is the length, sf(s) is the first frequency, bf(s) is the second frequency, sf(b i ) is the third frequency of the i-th second candidate word in the first corpus, sc(s) is the number of second candidate words, is the average value, bf(s)+sf(s) is the sum value, is the difference, is the second ratio, 1≤i≤sc(s), and i and sc(s) are positive integers.
[0095] Thus, when there are second candidate words in the first corpus in this method, the third frequency of the second candidate words in the first corpus can be obtained, and the candidate parameter is determined based on the first frequency, the second frequency, and the third frequency.
[0096] To make those skilled in the art understand the present invention more clearly, Figure 4 As shown in the flowchart of the word extraction method according to a specific example of the present invention, Figure 4 as shown, the method may include the following steps:
[0097] S401, obtaining a plurality of first candidate words in a first industry.
[0098] S402, obtaining the first frequency of the first candidate word in the first corpus of the first industry, and the second frequency of the first candidate word in the second corpus of the second industry.
[0099] S403, identifying whether there are second candidate words nested with the first candidate word in the first corpus.
[0100] If yes, then execute step S404; if no, then execute step S406.
[0101] S404. Obtain the third frequency of the second candidate word in the first corpus.
[0102] S405. Determine the candidate parameter based on the first frequency, the second frequency, and the third frequency.
[0103] S406. Determine the candidate parameter based on the first frequency and the second frequency.
[0104] S407. Sort the multiple candidate parameters in descending order.
[0105] S408. Determine the first candidate words corresponding to the top N sorted candidate parameters as the target words, where N is a positive integer.
[0106] For the relevant content of steps S401 - S408, reference can be made to the above - mentioned embodiments, which will not be elaborated here.
[0107] To implement the above - mentioned embodiments, the present invention also proposes a word extraction device.
[0108] Figure 5 It is a schematic structural diagram of a word extraction device according to an embodiment of the present invention.
[0109] As Figure 5 shown, the word extraction device 100 of the embodiment of the present invention includes: a first acquisition module 110, a second acquisition module 120, a determination module 130, and an extraction module 140.
[0110] The first acquisition module 110 is configured to acquire multiple first candidate words in the first industry.
[0111] The second acquisition module 120 is configured to acquire the first frequency of the first candidate word in the first corpus of the first industry and the second frequency of the first candidate word in the second corpus of the second industry.
[0112] The determination module 130 is configured to determine the candidate parameter of the first candidate word based on the first frequency and the second frequency.
[0113] The extraction module 140 is configured to extract target words from the multiple first candidate words based on the candidate parameter.
[0114] In an embodiment of the present invention, the determination module 130 is further configured to: identify whether there is a second candidate word nested with the first candidate word in the first corpus; in response to the non - existence of the second candidate word in the first corpus, determine the candidate parameter based on the first frequency and the second frequency.
[0115] In one embodiment of the present invention, the determining module 130 is further configured to: obtain the sum value between the first frequency and the second frequency; obtain the first ratio between the first frequency and the sum value; and determine the candidate parameter based on the first ratio.
[0116] In one embodiment of the present invention, the determining module 130 is further configured to: in response to the existence of the second candidate word in the first corpus, obtain the third frequency of the second candidate word in the first corpus; and determine the candidate parameter based on the first frequency, the second frequency, and the third frequency.
[0117] In one embodiment of the present invention, the determining module 130 is further configured to: obtain the average value of multiple third frequencies; and determine the candidate parameter based on the first frequency, the second frequency, and the average value.
[0118] In one embodiment of the present invention, the determining module 130 is further configured to: obtain the sum value between the first frequency and the second frequency; obtain the difference value between the first frequency and the average value; obtain the second ratio between the difference value and the sum value; and determine the candidate parameter based on the second ratio.
[0119] In one embodiment of the present invention, the first obtaining module 110 is further configured to: obtain the original text of the first industry; perform word segmentation on the original text to obtain a plurality of initial words; and extract a plurality of the first candidate words from the plurality of initial words.
[0120] In one embodiment of the present invention, the extracting module 140 is further configured to: sort the plurality of candidate parameters in descending order; and determine the first candidate words corresponding to the top N candidate parameters as the target words, where N is a positive integer.
[0121] In one embodiment of the present invention, the second industry includes at least one industry different from the first industry.
[0122] It should be noted that for the details not disclosed in the word extraction device of the embodiments of the present invention, please refer to the details disclosed in the word extraction method of the embodiments of the present invention, which will not be elaborated here.
[0123] In summary, the word extraction device according to the embodiment of the present invention obtains a plurality of first candidate words in the first industry, obtains the first frequency of the first candidate words in the first corpus of the first industry, and the second frequency of the first candidate words in the second corpus of the second industry. Based on the first frequency and the second frequency, the candidate parameter of the first candidate word is determined, and based on the candidate parameter, the target word is extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameter of the first candidate word, so as to extract the target word from the plurality of first candidate words, improving the accuracy of word extraction, and especially applicable to the hot word extraction scenario unique to the first industry.
[0124] To implement the above embodiment, as Figure 6 shown, an electronic device 200 according to an embodiment of the present invention includes: a memory 210, a processor 220, and a computer program stored on the memory 210 and executable on the processor 220. When the processor 220 executes the program, the above-mentioned word extraction method is implemented.
[0125] The electronic device according to the embodiment of the present invention obtains a plurality of first candidate words in the first industry by the processor executing the computer program stored on the memory, obtains the first frequency of the first candidate words in the first corpus of the first industry, and the second frequency of the first candidate words in the second corpus of the second industry. Based on the first frequency and the second frequency, the candidate parameter of the first candidate word is determined, and based on the candidate parameter, the target word is extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameter of the first candidate word, so as to extract the target word from the plurality of first candidate words, improving the accuracy of word extraction, and especially applicable to the hot word extraction scenario unique to the first industry.
[0126] To implement the above embodiment, an embodiment of the present invention proposes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned word extraction method is implemented.
[0127] The computer-readable storage medium according to the embodiment of the present invention obtains a plurality of first candidate words in the first industry by storing a computer program and executing it by a processor, obtains the first frequency of the first candidate words in the first corpus of the first industry, and the second frequency of the first candidate words in the second corpus of the second industry. Based on the first frequency and the second frequency, the candidate parameter of the first candidate word is determined, and based on the candidate parameter, the target word is extracted from the plurality of first candidate words. Thus, the first corpus of the first industry and the second corpus of the second industry can be comprehensively considered to determine the candidate parameter of the first candidate word, so as to extract the target word from the plurality of first candidate words, improving the accuracy of word extraction, and especially applicable to the hot word extraction scenario unique to the first industry.
[0128] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0129] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0130] In the present invention, unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0131] In the present invention, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0132] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0133] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for extracting words, characterized in that, Including: Obtain a plurality of first candidate words in the first industry; Obtain the first frequency of the first candidate word in the first corpus of the first industry and the second frequency of the first candidate word in the second corpus of the second industry, where the first industry is different from the second industry; Based on the first frequency and the second frequency, determine the candidate parameter of the first candidate word; Based on the candidate parameter, extract a target word from the plurality of first candidate words and determine the target word as a hot word unique to the first industry; where The calculation formula of the candidate parameter includes: is the candidate parameter, is the length of the first candidate word, is the first frequency, is the second frequency.
2. The method according to claim 1, wherein The determining the candidate parameter of the first candidate word based on the first frequency and the second frequency includes: Identify whether there is a second candidate word nested with the first candidate word in the first corpus; In response to the non-existence of the second candidate word in the first corpus, determine the candidate parameter based on the first frequency and the second frequency.
3. The method according to claim 2, wherein The determining the candidate parameter based on the first frequency and the second frequency includes: Obtain the sum value between the first frequency and the second frequency; Obtain the first ratio between the first frequency and the sum value; Based on the first ratio, determine the candidate parameter.
4. The method according to claim 2, wherein Also including: In response to the existence of the second candidate word in the first corpus, obtain the third frequency of the second candidate word in the first corpus; Based on the first frequency, the second frequency and the third frequency, determine the candidate parameter.
5. The method according to claim 4, wherein The determining the candidate parameter based on the first frequency, the second frequency and the third frequency includes: Obtain the average value of the plurality of third frequencies; Based on the first frequency, the second frequency and the average value, determine the candidate parameter.
6. The method according to claim 5, wherein The determining the candidate parameter based on the first frequency, the second frequency and the average value includes: Obtain the sum value between the first frequency and the second frequency; Obtain the difference value between the first frequency and the average value; Obtain the second ratio between the difference value and the sum value; Based on the second ratio, determine the candidate parameter.
7. The method according to claim 1, characterized in that, The obtaining a plurality of first candidate words in the first industry includes: Obtain the original text of the first industry; Perform word segmentation on the original text to obtain a plurality of initial words; Extract a plurality of the first candidate words from the plurality of initial words.
8. The method according to claim 1, characterized in that, The extracting a target word from the plurality of first candidate words based on the candidate parameter includes: Sort the plurality of candidate parameters in descending order; Determine the first candidate words corresponding to the top N sorted candidate parameters as the target words, where N is a positive integer.
9. The method according to any one of claims 1-8, characterized in that, The second industry includes at least one industry different from the first industry.
10. A word extraction device, characterized in that, Including: A first obtaining module for obtaining a plurality of first candidate words in the first industry; A second obtaining module for obtaining the first frequency of the first candidate word in the first corpus of the first industry and the second frequency of the first candidate word in the second corpus of the second industry, where the first industry is different from the second industry; A determination module, configured to determine candidate parameters of the first candidate word based on the first frequency and the second frequency; An extraction module, configured to extract a target word from multiple first candidate words based on the candidate parameters, and determine the target word as a hot word unique to the first industry; wherein, The calculation formula of the candidate parameters is: is the candidate parameter, is the length of the first candidate word, is the first frequency, is the second frequency.
11. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the word extraction method according to any one of claims 1-9 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the word extraction method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Hot word analysis and statistic system and method
CN105205048A
Method and system for automatically extracting hot words for agriculture public opinion
CN107967299A