Method and apparatus for determining characteristics of text

By marking and hash mapping operations on advertising text, the problem of difficulty in accurately determining the digital characteristics of advertising text in the prior art is solved, and more efficient and accurate industry recognition is achieved.

CN113868420BActive Publication Date: 2025-05-30GUANGZHOU YOUMI INFORMATION TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111153504.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-05-30
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to accurately determine the digital characteristics of advertising text, which affects the accuracy of industry identification.

Method used

By marking the text of the industry to be recognized, obtaining its hash value, and mapping the hash value, the feature vector of the text is obtained to determine the industry category that matches the text of the industry to be recognized.

Benefits of technology

The accuracy and efficiency of the hash value determination operation of text is improved, and the amount of word data in the text can be reduced without relying on a fixed vocabulary list, thereby improving the ability to quickly determine accurate text feature vectors, and enhancing the accuracy and efficiency of industry category recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113868420B_ABST
    Figure CN113868420B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for determining the characteristics of text. After determining the text of the industry to be recognized, by first performing a marking operation on the text of the industry to be recognized, it is beneficial to improve the accuracy and efficiency of performing the hash value determination operation on the text, and then automatically performing a mapping operation on the determined hash value of the text, and not relying on a fixed vocabulary, which can reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby being beneficial to improving the rapid determination of the accurate feature vector of the text and being beneficial to improving the accuracy and efficiency of recognizing the industry category matching the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular, to a method and device for determining the characteristics of text. Background Art

[0002] As an important channel for businesses and enterprises in different industries to promote and market their products, Internet advertisements often contain the brands, names of the corresponding promoted products, as well as relevant introductions, ingredients, and slogans. Effectively classifying them by industry helps to explore the advertising forms and the brands and categories contained in the advertisements in different industries. Among them, in order to accurately identify the industry of the advertisement text, it is very necessary to first determine the digital characteristics of the advertisement text. The traditional method is to use the one-hot encoding method to process traditional long strings of words to obtain the digital characteristics of the text. However, it is found in practice that the digital characteristics of the advertisement text cannot be accurately determined by the one-hot encoding method. Therefore, it is particularly important to propose a solution for accurately determining the digital characteristics of the advertisement text. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for determining the characteristics of text, which can accurately determine the digital characteristics of the advertisement text.

[0004] To solve the above technical problem, in a first aspect of the present invention, a method for determining the characteristics of text is disclosed, and the method includes:

[0005] Performing a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text, where the target text is the text of the industry to be recognized after marking, and the text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text;

[0006] Obtaining the hash value of the target text, and performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text, where the feature vector of the target text is used to determine the industry category matching the text of the industry to be recognized.

[0007] As an optional implementation manner, in the first aspect of the present invention, the obtaining the hash value of the target text includes:

[0008] Inputting each word in the target text into the determined hash function one by one for analysis, and obtaining the analysis result output by the hash function as the hash value of each word in the target text;

[0009] Determining the hash values of all the words in the target text as the hash value of the target text.

[0010] As an alternative implementation, in the first aspect of the present invention, the number of hash values of each word in the target text is greater than 1;

[0011] Among them, the performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text includes:

[0012] Mapping the hash value of each word in the target text to a pre-determined set to obtain the feature vector of each word;

[0013] Determining the feature vectors of all the words in the target text as the feature vector of the target text.

[0014] As an alternative implementation, in the first aspect of the present invention, the determining the feature vectors of all the words in the target text as the feature vector of the target text includes:

[0015] When the text of the industry to be recognized is the Chinese text, determining the characters included in each word of the text of the industry to be recognized, and determining the sum of the feature vectors of all the characters included in each word as the feature vector of the word, and determining the feature vectors of all the words in the target text as the feature vector of the target text;

[0016] When the text of the industry to be recognized is the English text, determining the feature vectors of all the words in the target text as the feature vector of the target text.

[0017] As an alternative implementation, in the first aspect of the present invention, after performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text, the method further includes:

[0018] Performing an industry classification learning operation on the feature vector of each word in the target text based on the determined bottleneck layer to obtain the bottleneck vector of each word in the target text;

[0019] Inputting the bottleneck vector of each word in the target text into the determined bidirectional encoder stack for analysis to obtain the target vector of each word in the target text, and the target vector of each word includes semantic information of adjacent words to the word, and the target vector of each word in the target text is used to determine the industry recognition model.

[0020] As an alternative implementation, in the first aspect of the present invention, the target vectors of each of the words in the target text are concatenated to obtain the concatenated target text, and a basic industry recognition model determined by training based on the concatenated target text is used to obtain a trained industry recognition model, where the industry recognition model is used to analyze the text of the industry to be recognized and obtain an industry category that matches the text of the industry to be recognized;

[0021] When the text of the industry to be recognized is the Chinese text, the trained industry recognition model is a Chinese text industry recognition model;

[0022] When the text of the industry to be recognized is the English text, the trained industry recognition model is an English text industry recognition model.

[0023] As an alternative implementation, in the first aspect of the present invention, after performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text, the method further includes:

[0024] After obtaining the feature vector of the Chinese text and the feature vector of the English text, determine the length of the feature vector of the Chinese text and the length of the feature vector of the English text;

[0025] Judge whether the length of the feature vector of the Chinese text and the length of the feature vector of the English text are both less than the corresponding determined length threshold to obtain a judgment result;

[0026] Match an industry recognition model corresponding to the judgment result according to the judgment result, and analyze the text of the industry to be recognized according to the industry recognition model corresponding to the judgment result to obtain an industry category that matches the text of the industry to be recognized, where the industry recognition model corresponding to the judgment result includes a Chinese text industry recognition model or an English text industry recognition model.

[0027] As an alternative implementation, in the first aspect of the present invention, the performing a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text includes:

[0028] If the number of lines of the text of the industry to be recognized is greater than or equal to 1, add corresponding marks at the beginning and end of each line of the text of the industry to be recognized to obtain a target text; or,

[0029] If the number of sentences of the text of the industry to be recognized is greater than or equal to 1, add corresponding marks at the beginning and end of each sentence of the text of the industry to be recognized to obtain a target text; or,

[0030] Add corresponding marks between every two adjacent words of the text of the industry to be recognized to obtain a target text.

[0031] The second aspect of the present invention discloses a device for determining the characteristics of a text, the device comprising:

[0032] A marking module, configured to perform a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text, where the target text is the text of the industry to be recognized after marking, and the text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text;

[0033] An obtaining module, configured to obtain the hash value of the target text;

[0034] A mapping module, configured to perform a mapping operation on the hash value of the target text to obtain a feature vector of the target text, where the feature vector of the target text is used to determine an industry category matching the text of the industry to be recognized.

[0035] As an optional implementation manner, in the second aspect of the present invention, the manner in which the obtaining module obtains the hash value of the target text is specifically:

[0036] Input each word in the target text into the determined hash function one by one for analysis, and obtain the analysis result output by the hash function as the hash value of each word in the target text;

[0037] Determine the hash values of all the words in the target text as the hash value of the target text.

[0038] As an optional implementation manner, in the second aspect of the present invention, the number of hash values of each word in the target text is greater than 1;

[0039] Among them, the manner in which the mapping module performs a mapping operation on the hash value of the target text to obtain a feature vector of the target text is specifically:

[0040] Map the hash value of each word in the target text to a pre-determined set to obtain a feature vector of each word;

[0041] Determine the feature vectors of all the words in the target text as the feature vector of the target text.

[0042] As an optional implementation manner, in the second aspect of the present invention, the manner in which the mapping module determines the feature vectors of all the words in the target text as the feature vector of the target text is specifically:

[0043] When the text of the industry to be recognized is the Chinese text, determine the characters included in each word of the text of the industry to be recognized, and determine the sum of the feature vectors of all the characters included in each word as the feature vector of the word, and determine the feature vectors of all the words in the target text as the feature vector of the target text;

[0044] When the text of the industry to be recognized is the English text, determine the feature vectors of all the words in the target text as the feature vector of the target text.

[0045] As an optional implementation manner, in the second aspect of the present invention, the device further includes:

[0046] A classification module, configured to, when the text of the industry to be recognized is a sample text, and after the mapping module performs a mapping operation on the hash value of the target text to obtain the feature vector of the target text, perform an industry classification learning operation on the feature vector of each word in the target text based on the determined bottleneck layer to obtain the bottleneck vector of each word in the target text;

[0047] A first analysis module, configured to input the bottleneck vector of each word in the target text into the determined bidirectional encoder stack for analysis to obtain the target vector of each word in the target text, where the target vector of each word includes semantic information of adjacent words to the word, and the target vector of each word in the target text is used to determine an industry recognition model.

[0048] As an optional implementation manner, in the second aspect of the present invention, the device further includes:

[0049] A determination module, configured to, after the mapping module performs a mapping operation on the hash value of the target text to obtain the feature vector of the target text, and after obtaining the feature vector of the Chinese text and the feature vector of the English text, determine the length of the feature vector of the Chinese text and the length of the feature vector of the English text;

[0050] A judgment module, configured to judge whether the length of the feature vector of the Chinese text and the length of the feature vector of the English text are both less than the corresponding determined length threshold to obtain a judgment result;

[0051] A second analysis module, configured to match an industry recognition model corresponding to the judgment result according to the judgment result, and analyze the text of the industry to be recognized according to the industry recognition model corresponding to the judgment result to obtain an industry category matching the text of the industry to be recognized, where the industry recognition model corresponding to the judgment result includes a Chinese text industry recognition model or an English text industry recognition model.

[0052] As an alternative implementation, in the second aspect of the present invention, the way that the marking module performs a marking operation on the text of the industry to be recognized according to the determined marking method to obtain the target text is specifically as follows:

[0053] If the number of lines of the text of the industry to be recognized is greater than or equal to 1, corresponding marks are added at the beginning and end of each line of the text of the industry to be recognized to obtain the target text; or,

[0054] If the number of sentences of the text of the industry to be recognized is greater than or equal to 1, corresponding marks are added at the beginning and end of each sentence of the text of the industry to be recognized to obtain the target text; or,

[0055] Corresponding marks are added between every two adjacent words of the text of the industry to be recognized to obtain the target text.

[0056] The third aspect of the present invention discloses another device for determining the characteristics of text, and the device includes:

[0057] A memory storing executable program code;

[0058] A processor coupled to the memory;

[0059] The processor calls the executable program code stored in the memory and executes some or all of the steps in the method for determining the characteristics of text disclosed in the first aspect of the present invention.

[0060] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute some or all of the steps in the method for determining the characteristics of text disclosed in the first aspect of the present invention when being called.

[0061] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0062] In an embodiment of the present invention, a marking operation is performed on the text of the industry to be recognized according to the determined marking method to obtain a target text. The target text is the text of the industry to be recognized after marking. The text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text. The hash value of the target text is obtained, and a mapping operation is performed on the hash value of the target text to obtain a feature vector of the target text. The feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized. It can be seen that after determining the text of the industry to be recognized, by first performing a marking operation on the text of the industry to be recognized, it is beneficial to improve the accuracy and efficiency of the operation for determining the hash value of the text. Then, the mapping operation is automatically performed on the determined hash value of the text, and it does not depend on a fixed vocabulary. It can reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby facilitating the rapid determination of the accurate feature vector of the text and improving the accuracy and efficiency of recognizing the industry category that matches the text. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 is a flowchart showing a method for determining the features of a text disclosed in an embodiment of the present invention;

[0065] Figure 2 is a flowchart showing another method for determining the features of a text disclosed in an embodiment of the present invention;

[0066] Figure 3 is a structural diagram showing a device for determining the features of a text disclosed in an embodiment of the present invention;

[0067] Figure 4 is a structural diagram showing another device for determining the features of a text disclosed in an embodiment of the present invention;

[0068] Figure 5 is a structural diagram showing yet another device for determining the features of a text disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0070] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0071] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0072] The present invention discloses a method and device for determining the characteristics of text. After determining the text of the industry to be recognized, by first performing a marking operation on the text of the industry to be recognized, it is beneficial to improve the accuracy and efficiency of performing the hash value determination operation on the text, and then automatically performing a mapping operation on the determined hash value of the text, and not relying on a fixed vocabulary, it is possible to reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby facilitating the rapid determination of accurate feature vectors of the text and improving the accuracy and efficiency of identifying the industry category matching the text. The following will be described in detail respectively.

[0073] Embodiment 1

[0074] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a method for determining the characteristics of text disclosed in an embodiment of the present invention. Among them, Figure 1 The described method can be applied to a device for determining the characteristics of text. The device for determining the characteristics of text includes one of a text processing server, a text processing system, a text processing platform, and a text processing device, which is not limited in the embodiments of the present invention. As Figure 1As shown, the method for determining the features of the text may include the following operations:

[0075] 101. Perform a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text, where the target text is the text of the industry to be recognized after marking, and the text of the industry to be recognized includes Chinese text or English text, and the Chinese text and English text are texts extracted from the same original text.

[0076] In the embodiments of the present invention, the industry includes, but is not limited to, at least one of the game industry, financial industry, culture and entertainment industry, integrated e-commerce industry, education and training industry, medical and health industry, catering and food industry, real estate industry, life service industry, wedding service industry, social and dating industry, automotive industry, digital home appliance industry, beauty and personal care industry, clothing and footwear industry, mother and baby and children industry, food and beverage industry, smart home industry, building materials industry, etc. Further, the industry includes multiple sub-industries. For example, the game industry includes one or more of the role-playing industry, action adventure industry, strategy game industry, simulation operation industry, board game industry, sports racing industry, flight shooting industry, casual puzzle industry, and other game industries.

[0077] In the embodiments of the present invention, the target text of the industry to be recognized is the text of one of the above-mentioned all industries. Optionally, the text of the industry to be recognized includes any advertising text or non-advertising text that needs to identify its industry. Further, the text of the industry to be recognized includes one or more of the text obtained by identifying the packaging of goods or non-goods, the text of the industry to be recognized obtained from the storage unit, and the text of the industry to be recognized input by the user. Still further, the text of the industry to be recognized also includes one or more of the title, product description (such as: usage instructions), and store information (such as: store name).

[0078] 102. Obtain the hash value of the target text.

[0079] 103. Perform a mapping operation on the hash value of the target text to obtain a feature vector of the target text, and the feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized.

[0080] In the embodiments of the present invention, the dimension of the feature vector of the target text may include three dimensions, four dimensions, etc.

[0081] It can be seen that the implementation Figure 1The described method can, after determining the text of the industry to be recognized, first perform a tagging operation on the text of the industry to be recognized, which is beneficial to improving the accuracy and efficiency of performing the hash value determination operation on the text, and then automatically perform a mapping operation on the hash value of the determined text, and does not rely on a fixed vocabulary, and can reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby being beneficial to improving the rapid determination of the feature vector of the accurate text and being beneficial to improving the accuracy and efficiency of recognizing the industry category matching the text.

[0082] In an alternative embodiment, obtaining the hash value of the target text includes:

[0083] Input each word in the target text into the determined hash function one by one for analysis, and obtain the analysis result output by the hash function as the hash value of each word in the target text;

[0084] Determine the hash values of all words in the target text as the hash value of the target text.

[0085] In this alternative embodiment, optionally, the hash function includes any function that can analyze the hash value of each word in the target text, such as: the murmurhash hash function.

[0086] In this alternative embodiment, optionally, each word in the target text is represented by a pre-determined dimensional vector, such as: 3-dimensional.

[0087] In this alternative embodiment, for Chinese text, the word input into the hash function can be understood as a single character, such as: red, and for English text, the word input into the hash function can be understood as a single word, such as: red.

[0088] It can be seen that this alternative embodiment can quickly and accurately analyze the words of the text by inputting each word of the text into the hash function one by one, improving the accuracy and efficiency of obtaining the hash value of the text; and the processing objects are single characters of Chinese text and single words of English text, which is beneficial to improving the analysis accuracy of the words of the text, thereby improving the accuracy of determining the hash value of the text words.

[0089] In another alternative embodiment, the number of hash values of each word in the target text is greater than 1; wherein, performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text includes:

[0090] Map the hash value of each word in the target text to a pre-determined set to obtain the feature vector of each word;

[0091] Determine the feature vectors of all words in the target text as the feature vector of the target text.

[0092] In this alternative embodiment, different hash functions result in different numbers of hash values.

[0093] In this alternative embodiment, the pre-determined set can be any set that meets the conditions, such as: a three-dimensional (ternary) set: {-1, 0, 1}.

[0094] In this alternative embodiment, determining the feature vectors of all words in the target text as the feature vectors of the target text includes:

[0095] When the text of the industry to be recognized is in Chinese, determine the characters included in each word of the text of the industry to be recognized, and determine the sum of the feature vectors of all characters included in each word as the feature vector of the word, and determine the feature vectors of all words in the target text as the feature vectors of the target text;

[0096] When the text of the industry to be recognized is in English, determine the feature vectors of all words in the target text as the feature vectors of the target text.

[0097] In this alternative embodiment, the character is a single Chinese character.

[0098] It can be seen that in this alternative embodiment, by mapping the hash value of each word of the text to the corresponding set, it is beneficial to represent relevant words (such as words with the same or similar semantics) with the same feature vector, and can reduce the word data volume of the text while ensuring the retention of the key words of the text, improving the accuracy and reliability of industry recognition for text matching.

[0099] In another alternative embodiment, when the text of the industry to be recognized is a sample text, after performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text, the method may further include the following operations:

[0100] Perform industry classification learning operations on the feature vectors of each word in the target text based on the determined bottleneck layer to obtain the bottleneck vectors of each word in the target text;

[0101] Input the bottleneck vectors of each word in the target text into the determined bidirectional encoder stack for analysis to obtain the target vectors of each word in the target text. The target vector of each word contains semantic information of adjacent words to the word, and the target vectors of each word in the target text are used to determine the industry recognition model.

[0102] In this alternative embodiment, the dimension of the target vector of each word is smaller than the dimension of the bottleneck vector of each word. For example, the dimension of the feature vector of each word is 512 dimensions, and the dimension of the target vector of each word is 64 dimensions.

[0103] In this alternative embodiment, the bidirectional encoder includes, but is not limited to, an encoder determined based on one or more of a bidirectional QRNN encoder, a bidirectional LSTM encoder, a bidirectional GRU encoder, a bidirectional PQRNN encoder, and a transformer encoder. It should be noted that the bidirectional QRNN encoder is preferably selected. By using the bidirectional QRNN encoder, the ability to process multiple words in the sample text in parallel can be improved, the training efficiency of the industry recognition model can be enhanced, and it is beneficial to increase the update and iteration speed of the industry recognition model.

[0104] It can be seen that in this alternative embodiment, by performing an industry classification operation on the feature vectors of each word in the target text through the bottleneck layer, the dimension of the feature vectors of the words can be reduced while retaining the key features, which is beneficial to improving the extraction accuracy and efficiency of the context semantic information corresponding to each word in the target text, and thus is beneficial to improving the determination accuracy and efficiency of the vectors of the words containing the semantic information of adjacent words.

[0105] In yet another alternative embodiment, performing a marking operation on the text of the industry to be recognized according to the determined marking method to obtain the target text includes:

[0106] If the number of lines of the text of the industry to be recognized is greater than or equal to 1, adding corresponding marks at the beginning and end of each line of the text of the industry to be recognized to obtain the target text; or,

[0107] If the number of sentences of the text of the industry to be recognized is greater than or equal to 1, adding corresponding marks at the beginning and end of each sentence of the text of the industry to be recognized to obtain the target text; or,

[0108] Adding corresponding marks between every two adjacent words in the text of the industry to be recognized to obtain the target text.

[0109] In this alternative embodiment, it is preferably to mark the sentences of the text of the industry to be recognized.

[0110] It can be seen that in this alternative embodiment, by providing multiple marking methods for marking the lines or sentences or words of the text of the recognized industry, it is beneficial to improve the marking flexibility of the text of the industry to be recognized and the intelligent function of the device for determining the rich features of the text.

[0111] Embodiment 2

[0112] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for determining the features of text disclosed in the embodiments of the present invention. Among them, Figure 2The described method can be applied to a device for determining the characteristics of text. The device for determining the characteristics of text includes one of a text processing server, a text processing system, a text processing platform, and a text processing device, which is not limited in the embodiments of the present invention. As Figure 2 shown, the method for determining the characteristics of the text may include the following operations:

[0113] 201. Perform a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text. The target text is the text of the industry to be recognized after marking. The text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text.

[0114] 202. Obtain the hash value of the target text.

[0115] 203. Perform a mapping operation on the hash value of the target text to obtain a feature vector of the target text. The feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized.

[0116] 204. After obtaining the feature vector of the Chinese text and the feature vector of the English text, determine the length of the feature vector of the Chinese text and the length of the feature vector of the English text.

[0117] 205. Determine whether the lengths of the feature vector of the Chinese text and the feature vector of the English text are both less than the corresponding determined length threshold to obtain a judgment result.

[0118] 206. Match the industry recognition model corresponding to the judgment result according to the judgment result. According to the industry recognition model corresponding to the judgment result, analyze the text of the industry to be recognized to obtain the industry category that matches the text of the industry to be recognized. The industry recognition model corresponding to the judgment result includes a Chinese text industry recognition model or an English text industry recognition model.

[0119] In the embodiments of the present invention, it should be noted that for the relevant descriptions of steps 201-step 203, please refer to the detailed descriptions of steps 101-step 103 in Embodiment 1, and the embodiments of the present invention will not be repeated.

[0120] It can be seen that the implementation Figure 2The described method can, after determining the text of the industry to be recognized, first perform a marking operation on the text of the industry to be recognized, which is beneficial to improving the accuracy and efficiency of performing the hash value determination operation on the text, and then automatically perform a mapping operation on the determined hash value of the text. Without relying on a fixed vocabulary, it can reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby facilitating the rapid determination of the feature vector of the accurate text and improving the accuracy and efficiency of recognizing the industry category matching the text. In addition, after performing the mapping operation on the hash value of the text of the industry to be recognized, further determine the industry category matching the text of the industry to be recognized according to the length of the feature vector of the English text and the length of the feature vector of the Chinese text, which can improve the accuracy of recognizing the industry category matching the text and is beneficial to improving the accuracy and reliability of exploring the brands and categories contained in the texts of different industries (such as advertising texts).

[0121] In an alternative embodiment, determining the industry category matching the text of the industry to be recognized according to the length of the feature vector of the Chinese text and the length of the feature vector of the English text includes:

[0122] Judge whether the lengths of the feature vectors of the Chinese text and the English text are both less than the corresponding determined length thresholds to obtain a judgment result;

[0123] When the judgment result is used to indicate that the length of the feature vector of the Chinese text is less than the corresponding length threshold and the length of the feature vector of the English text is greater than or equal to the corresponding length threshold, determine the English text industry recognition model corresponding to the feature vector of the English text as the industry recognition model corresponding to the judgment result, input the feature vector of the English text into the English text industry recognition model for analysis, and obtain the industry analysis result output by the English text industry recognition model as the industry category matching the text of the industry to be recognized; when the judgment result is used to indicate that the length of the feature vector of the Chinese text is greater than or equal to the corresponding length threshold and the length of the feature vector of the English text is less than the corresponding length threshold, determine the Chinese text industry recognition model corresponding to the feature vector of the Chinese text as the industry recognition model corresponding to the judgment result, input the feature vector of the Chinese text into the Chinese text industry recognition model for analysis, and obtain the industry analysis result output by the Chinese text industry recognition model as the industry category matching the text of the industry to be recognized; when the judgment result is used to indicate that the length of the feature vector of the Chinese text is greater than or equal to the corresponding length threshold and the length of the feature vector of the English text is greater than or equal to the corresponding length threshold, determine the Chinese text industry recognition model corresponding to the feature vector of the Chinese text and the English text industry recognition model corresponding to the feature vector of the English text as the industry recognition model corresponding to the judgment result, input the feature vector of the Chinese text into the Chinese text industry recognition model for analysis, and obtain the industry analysis result output by the Chinese text industry recognition model as the industry category matching the text of the industry to be recognized, input the feature vector of the English text into the English text industry recognition model for analysis, and obtain the industry analysis result output by the English text industry recognition model as the industry category matching the text of the industry to be recognized. If the industry labels included in the industry analysis result output by the English text industry recognition model are the same as the industry labels included in the industry analysis result output by the Chinese text industry recognition model, determine the industry labels of both as the industry category matching the text of the industry to be recognized; if they are different, calculate the confidence level of the industry label included in the industry analysis result output by the Chinese text industry recognition model and the weight factor corresponding to the Chinese text industry recognition model to obtain the first industry label score, calculate the confidence level of the industry label included in the industry analysis result output by the English text industry recognition model and the weight factor corresponding to the English text industry recognition model to obtain the second industry label score; if the first industry label score is greater than the second industry label score, determine the industry label included in the industry analysis result output by the Chinese text industry recognition model as the industry category matching the text of the industry to be recognized; if the first industry label score is less than the second industry label score, determine the industry label included in the industry analysis result output by the English text industry recognition model as the industry category matching the text of the industry to be recognized; if the first industry label score is equal to the second industry label score, determine that the industry matching the text is empty.

[0124] It can be seen that, in this alternative embodiment, by comparing the lengths of the feature vectors of the Chinese text and the English text with the corresponding length thresholds, the Chinese text industry recognition model and / or the English text industry recognition model are determined, which can improve the determination accuracy and reliability of the required industry recognition model; and by inputting the word set into the corresponding industry recognition model for analysis, the analysis accuracy and reliability of the text can be improved; and when it is determined that the similarity of the industry labels of the industry analysis results output by the Chinese text industry recognition model and the industry analysis results output by the English text industry recognition model is small, further by calculating the confidence level of the industry label and the industry label score of the weight factor of the corresponding industry recognition model, and taking the industry label with the larger industry label score as the industry matching the text, the probability of accurately identifying the industry category matching the text can be further improved.

[0125] In another alternative embodiment, the method may further include the following steps:

[0126] Determine the position where each word in the text of the industry to be recognized appears in the text of the industry to be recognized and / or the glyph complexity;

[0127] According to the position where each word appears in the text of the industry to be recognized and / or the glyph complexity, correct the lengths of the feature vectors of the Chinese text and the English text, and trigger the execution of step 205.

[0128] It should be noted that the more complex the glyph, the longer the length of the corresponding word set; the closer the appearance position is to the beginning or end position of the target text, the longer the length of the corresponding word set.

[0129] It can be seen that, in this alternative embodiment, by combining the position where the word in the text of the industry to be recognized appears in the text and the glyph complexity to correct the lengths of the feature vectors of the Chinese text and the English text, the accuracy and reliability of determining the length of the feature vector of the text can be improved, which is beneficial to further improving the matching accuracy and reliability of the industry recognition model, and further beneficial to improving the accuracy and reliability of identifying the industry category matching the text.

[0130] Embodiment III

[0131] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a device for determining the features of a text disclosed in an embodiment of the present invention. Among them, the device for determining the features of the text includes any one of a text processing server, a text processing system, a text processing platform, and a text processing device. As Figure 3 shown, the device for determining the features of the text may include:

[0132] A tagging module 301 is configured to perform a tagging operation on the text of the industry to be recognized according to the determined tagging method, so as to obtain a target text, where the target text is the tagged text of the industry to be recognized, and the text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text.

[0133] An obtaining module 302 is configured to obtain the hash value of the target text.

[0134] A mapping module 303 is configured to perform a mapping operation on the hash value of the target text to obtain a feature vector of the target text, and the feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized.

[0135] It can be seen that the apparatus for determining the features of the described text can, after determining the text of the industry to be recognized, first perform a tagging operation on the text of the industry to be recognized, which is beneficial to improving the accuracy and efficiency of performing the operation of determining the hash value of the text, and then automatically perform a mapping operation on the determined hash value of the text, and does not depend on a fixed vocabulary, and can reduce the word data volume of the text while ensuring the retention of the required words of the text, thereby being beneficial to quickly determining the accurate feature vector of the text and being beneficial to improving the accuracy and efficiency of recognizing the industry category that matches the text. Figure 3 In an optional embodiment, as

[0136] shown, the specific manner in which the obtaining module 302 obtains the hash value of the target text is as follows: Figure 3

[0137] Each word in the target text is input into the determined hash function one by one for analysis, and the analysis result output by the hash function is obtained as the hash value of each word in the target text;

[0138] The hash values of all words in the target text are determined as the hash value of the target text.

[0139] Figure 3 It can be seen that the apparatus for determining the features of the described text can also analyze each word of the text quickly and accurately by inputting each word of the text into the hash function one by one, improving the accuracy and efficiency of obtaining the hash value of the text; and the objects to be processed are single Chinese characters of Chinese text and single words of English text, which is beneficial to improving the analysis accuracy of the words of the text, thereby improving the accuracy of determining the hash value of the words of the text. Figure 3 In another optional embodiment, the number of hash values of each word in the target text is greater than 1; and as

[0140] shown Figure 3As shown in the figure, the mapping module 303 performs a mapping operation on the hash value of the target text to obtain the feature vector of the target text in the following specific manner:

[0141] Map the hash value of each word in the target text to a pre-determined set to obtain the feature vector of each word;

[0142] Determine the feature vectors of all words in the target text as the feature vector of the target text.

[0143] In this optional embodiment, the manner in which the mapping module 303 determines the feature vectors of all words in the target text as the feature vector of the target text is as follows:

[0144] When the text of the industry to be recognized is a Chinese text, determine the characters included in each word of the text of the industry to be recognized, and determine the sum of the feature vectors of all characters included in each word as the feature vector of the word, and determine the feature vectors of all words in the target text as the feature vector of the target text;

[0145] When the text of the industry to be recognized is an English text, determine the feature vectors of all words in the target text as the feature vector of the target text.

[0146] It can be seen that the Figure 3 The device for determining the features of the described text can also map the hash value of each word of the text to the corresponding set, which is beneficial to representing related words (such as words with the same or similar semantics) with the same feature vector, and can reduce the word data volume of the text while ensuring the retention of the key words of the text, improving the accuracy and reliability of industry recognition for text matching.

[0147] In yet another optional embodiment, as Figure 4 shown in the figure, the device further includes:

[0148] A classification module 304, configured to, when the text of the industry to be recognized is a sample text and after the mapping module 303 performs a mapping operation on the hash value of the target text to obtain the feature vector of the target text, perform industry classification learning operations on the feature vectors of each word in the target text based on the determined bottleneck layer to obtain the bottleneck vector of each word in the target text.

[0149] A first analysis module 305, configured to input the bottleneck vector of each word in the target text into the determined bidirectional encoder stack for analysis to obtain the target vector of each word in the target text. The target vector of each word includes semantic information of adjacent words to the word, and the target vector of each word in the target text is used to determine the industry recognition model.

[0150] It can be seen that the Figure 4The device for determining the characteristics of the described text can also perform an industry classification operation on the feature vectors of each word in the target text through a bottleneck layer, which can reduce the dimension of the feature vectors of the words while retaining the key features, facilitating the improvement of the extraction accuracy and efficiency of the context semantic information corresponding to each word in the target text, and thus facilitating the improvement of the determination accuracy and efficiency of the vectors of the words containing the semantic information of adjacent words.

[0151] In yet another alternative embodiment, as Figure 4 shown, the device may further include:

[0152] A determination module 306, configured to determine the length of the feature vector of the Chinese text and the length of the feature vector of the English text after the mapping module 303 performs a mapping operation on the hash value of the target text to obtain the feature vector of the target text, and after obtaining the feature vector of the Chinese text and the feature vector of the English text.

[0153] A judgment module 307, configured to judge whether the length of the feature vector of the Chinese text and the length of the feature vector of the English text are both less than the corresponding determined length threshold to obtain a judgment result.

[0154] A second analysis module 308, which matches the industry recognition model corresponding to the judgment result according to the judgment result, and analyzes the text of the industry to be recognized according to the industry recognition model corresponding to the judgment result to obtain the industry category matching the text of the industry to be recognized. The industry recognition model corresponding to the judgment result includes a Chinese text industry recognition model or an English text industry recognition model.

[0155] It can be seen that implementing Figure 4 The device for determining the characteristics of the described text can also, after performing a mapping operation on the hash value of the text of the industry to be recognized, further determine the industry category matching the text of the industry to be recognized according to the length of the feature vector of the English text and the length of the feature vector of the Chinese text, which can improve the accuracy of identifying the industry category matching the text, and is conducive to improving the accuracy and reliability of exploring the brands and categories contained in the texts of different industries (such as advertising texts).

[0156] In yet another alternative embodiment, as Figure 3 or 4 shown, the specific manner in which the marking module 301 performs a marking operation on the text of the industry to be recognized according to the determined marking method to obtain the target text is:

[0157] If the number of lines of the text of the industry to be recognized is greater than or equal to 1, add corresponding marks at the beginning and end of each line of the text of the industry to be recognized to obtain the target text; or,

[0158] The number of sentences in the text of the industry to be recognized is greater than or equal to 1. Add corresponding tags at the beginning and end of each sentence in the text of the industry to be recognized to obtain the target text; or,

[0159] Add corresponding tags between every two adjacent words in the text of the industry to be recognized to obtain the target text. In this alternative embodiment, it is preferred to tag the sentences in the text of the industry to be recognized.

[0160] It can be seen that the device for determining the characteristics of the text described in Figure 3 or Embodiment 4 can also provide multiple tagging methods for tagging lines, sentences, or words in the text of the recognized industry, which is beneficial to improving the tagging flexibility of the text of the industry to be recognized and enriching the intelligent functions of the device for determining the characteristics of the text.

[0161] Embodiment Four

[0162] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another device for determining the characteristics of the text disclosed in the embodiments of the present invention. Among them, the device for determining the characteristics of the text includes any one of a text processing server, a text processing system, a text processing platform, and a text processing device. As Figure 5 shown, the device may include:

[0163] A memory 501 storing executable program code;

[0164] A processor 502 coupled to the memory 501;

[0165] Further, it may further include an input interface 503 and an output interface 504 coupled to the processor 502;

[0166] Among them, the processor 502 calls the executable program code stored in the memory 501 and executes the steps in the method for determining the characteristics of the text disclosed in Embodiment 1 or Embodiment 2 of the present invention.

[0167] Embodiment Five

[0168] The embodiments of the present invention disclose a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the steps in the method for determining the characteristics of the text disclosed in Embodiment 1 or Embodiment 2 of the present invention when being called.

[0169] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0170] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0171] Finally, it should be noted that: the determination method and device for the features of a text disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for determining the characteristics of a text, characterized in that, the method includes: performing a marking operation on the text of the industry to be recognized according to the determined marking method to obtain a target text, where the target text is the text of the industry to be recognized after marking, and the text of the industry to be recognized includes Chinese text or English text, and the Chinese text and the English text are texts extracted from the same original text; obtaining the hash value of the target text and performing a mapping operation on the hash value of the target text to obtain the feature vector of the target text, where the feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized; after performing the mapping operation on the hash value of the target text to obtain the feature vector of the target text, the method further includes: after obtaining the feature vector of the Chinese text and the feature vector of the English text, determining the length of the feature vector of the Chinese text and the length of the feature vector of the English text; judging whether the length of the feature vector of the Chinese text and the length of the feature vector of the English text are both less than the corresponding determined length threshold to obtain a judgment result; matching the industry recognition model corresponding to the judgment result according to the judgment result, and analyzing the text of the industry to be recognized according to the industry recognition model corresponding to the judgment result to obtain the industry category that matches the text of the industry to be recognized; wherein, the analyzing the text of the industry to be recognized according to the industry recognition model corresponding to the judgment result to obtain the industry category that matches the text of the industry to be recognized includes: when the industry recognition model corresponding to the judgment result is the Chinese text industry recognition model, inputting the feature vector of the Chinese text into the Chinese text industry recognition model for analysis, and obtaining the industry analysis result output by the Chinese text industry recognition model as the industry category that matches the text of the industry to be recognized; when the industry recognition model corresponding to the judgment result is the English text industry recognition model, inputting the feature vector of the English text into the English text industry recognition model for analysis, and obtaining the industry analysis result output by the English text industry recognition model as the industry category that matches the text of the industry to be recognized; when the industry recognition model corresponding to the judgment result is the Chinese text industry recognition model and the English text industry recognition model, if the industry labels included in the industry analysis result output by the English text industry recognition model are the same as the industry labels included in the industry analysis result output by the Chinese text industry recognition model, then determining their industry labels as the industry category that matches the text of the industry to be recognized; if they are not the same, then determining the first industry label score corresponding to the industry analysis result output by the Chinese text industry recognition model and the second industry label score corresponding to the industry analysis result output by the English text industry recognition model, and screening the industry label corresponding to the higher score from the first industry label score and the second industry label score as the industry category that matches the text of the industry to be recognized.

2. The method for determining the characteristics of the text according to claim 1, wherein, the obtaining of the hash value of the target text includes: inputting each word in the target text into the determined hash function one by one for analysis, and obtaining the analysis result output by the hash function as the hash value of each word in the target text; determining the hash values of all the words in the target text as the hash value of the target text.

3. The method for determining the characteristics of the text according to claim 2, wherein, the number of hash values of each word in the target text is greater than 1; wherein, the performing of the mapping operation on the hash value of the target text to obtain the feature vector of the target text includes: mapping the hash values of each word in the target text to a pre-determined set to obtain the feature vector of each word; determining the feature vectors of all the words in the target text as the feature vector of the target text.

4. The method for determining the characteristics of the text according to claim 3, wherein, the determining of the feature vectors of all the words in the target text as the feature vector of the target text includes: when the text of the industry to be recognized is the Chinese text, determining the characters included in each word of the text of the industry to be recognized, and determining the sum of the feature vectors of all the characters included in each word as the feature vector of the word, and determining the feature vectors of all the words in the target text as the feature vector of the target text; when the text of the industry to be recognized is the English text, determining the feature vectors of all the words in the target text as the feature vector of the target text.

5. The method for determining the characteristics of the text according to any one of claims 1-3, wherein, when the text of the industry to be recognized is the sample text, after performing the mapping operation on the hash value of the target text to obtain the feature vector of the target text, the method further includes: performing an industry classification learning operation on the feature vectors of each word in the target text based on the determined bottleneck layer to obtain the bottleneck vector of each word in the target text; inputting the bottleneck vectors of each word in the target text into the determined bidirectional encoder stack for analysis to obtain the target vector of each word in the target text, and the target vector of each word includes the semantic information of the adjacent words of the word, and the target vectors of each word in the target text are used to determine the industry recognition model.

6. The method for determining the characteristics of the text according to any one of claims 1-4, wherein, the performing of the marking operation on the text of the industry to be recognized according to the determined marking method to obtain the target text includes: the number of lines of the text of the industry to be recognized is greater than or equal to 1, adding corresponding marks at the beginning and end of each line of the text of the industry to be recognized to obtain the target text; or, The number of sentences in the text of the industry to be recognized is greater than or equal to 1. Add corresponding markers at the beginning and end of each sentence in the text of the industry to be recognized to obtain the target text; or, Add corresponding markers between every two adjacent words in the text of the industry to be recognized to obtain the target text.

7. An apparatus for determining the characteristics of a text, characterized in that, the apparatus is configured to execute the method for determining the characteristics of a text according to any one of claims 1-6, and the apparatus includes: a marking module, configured to perform a marking operation on the text of the industry to be recognized according to the determined marking method to obtain the target text, where the target text is the text of the industry to be recognized after marking, and the text of the industry to be recognized includes a Chinese text or an English text, and the Chinese text and the English text are texts extracted from the same original text; an obtaining module, configured to obtain the hash value of the target text; a mapping module, configured to perform a mapping operation on the hash value of the target text to obtain the feature vector of the target text, and the feature vector of the target text is used to determine the industry category that matches the text of the industry to be recognized.

8. An apparatus for determining the characteristics of a text, characterized in that, the apparatus includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method for determining the characteristics of a text according to any one of claims 1-6.

9. A computer storage medium, characterized in that, the computer storage medium stores computer instructions, and when the computer instructions are called, the method for determining the characteristics of a text according to any one of claims 1-6 is executed.

Citation Information

Patent Citations

  • An industry classification method and terminal device based on machine learning

    CN109388712A

  • Entity recognition model training method, and threat intelligence entity extraction method and device

    CN112149420A

  • Text classification method and device, storage medium and electronic device

    CN112597764A