Text concept recognition method, device, computer equipment and storage medium

By building a target noun database and concept database and using a pre-trained language model to recognize text concepts, the problems of low accuracy and high cost of text concept recognition in the existing technology are solved, and more efficient and accurate recognition effects are achieved.

CN119397378BActive Publication Date: 2025-05-02HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510007019.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-02
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

In the prior art, in text concept recognition, the recognition accuracy is low and the cost is high. It is mainly due to the need to conceptually annotate a large number of texts, which leads to the accuracy of the labeling information affecting the recognition accuracy.

Method used

By constructing a target noun database and concept database, the text concept of the text to be recognized is determined using a pre-trained language model, which avoids concept labeling of a large number of texts, reduces the recognition cost, and improves the recognition accuracy.

Benefits of technology

This method effectively reduces the cost of text concept labeling, improves the accuracy of text concept recognition, and solves the problems of low recognition accuracy and high cost in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397378B_ABST
    Figure CN119397378B_ABST
Patent Text Reader

Abstract

The present application relates to a text concept recognition method, device, computer equipment and storage medium. It includes: constructing a target noun database based on sample text information; determining the concept word vector corresponding to the sample text information through a pre-trained language model; constructing a concept database based on the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; obtaining the text to be recognized, and determining the text concept of the text to be recognized based on the target noun database and the concept database through a pre-trained language model. The above scheme avoids the concept annotation of a large amount of text, reduces the cost of text concept annotation, and improves the accuracy of text concept recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and in particular to a text concept recognition method, apparatus, computer equipment and storage medium. Background Art

[0002] Currently, when extracting conceptual information from text, the conceptual information in the text is often extracted and annotated according to the application scenario of the text, and the algorithm model is trained through the annotated text, so that the trained algorithm model can perform concept recognition on the text to be recognized, thereby determining the text concept of the text to be recognized. The recognition accuracy of the algorithm model for text concepts depends on the number of annotated texts, and the recognition accuracy is affected by the accuracy of the annotated information. If the annotation is wrong, it may lead to low prediction accuracy of the algorithm model. Therefore, how to improve the recognition accuracy of text concepts is a problem that needs to be solved. Summary of the invention

[0003] Based on this, it is necessary to provide a text concept recognition method, device, computer equipment and storage medium that can improve the recognition accuracy of text concepts in response to the above technical problems.

[0004] In a first aspect, the present application provides a text concept recognition method, the method comprising:

[0005] Construct a target noun database based on sample text information;

[0006] Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0007] A concept database is constructed according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0008] A text to be recognized is obtained, and a text concept of the text to be recognized is determined based on the target noun database and the concept database through the pre-trained language model.

[0009] In one embodiment, constructing a target noun database based on sample text information includes:

[0010] Construct a corpus noun database based on sample text information;

[0011] Determining high-frequency subsequences in the sample text information;

[0012] A target noun database is constructed according to the high-frequency subsequences and the corpus noun database.

[0013] In one embodiment, constructing a target noun database according to the high-frequency subsequences and the corpus noun database includes:

[0014] Determine noun identifiers and noun features of high-frequency subsequences based on the corpus noun database;

[0015] Determining the confidence of the high-frequency subsequence according to the noun identifier and the noun feature through a regression model;

[0016] A target subsequence is determined from the high-frequency subsequence according to the confidence of the high-frequency subsequence, and a target noun database is constructed according to the target subsequence and the corpus noun database.

[0017] In one embodiment, before determining the concept word vector corresponding to the sample text information through the pre-trained language model, the method further includes:

[0018] Determining a target short sentence according to the sample text information;

[0019] Masking the concept nouns in the target short sentence, and using a language representation model to predict concept words for the target short sentence after the masking process, and determining a cross entropy loss function of the language representation model according to the concept word prediction result and the concept nouns in the target short sentence;

[0020] Using the language representation model to predict the sentence type of the target sentence, determine the predicted sentence type, and determine the classification loss function of the language representation model according to the predicted sentence type and the actual sentence type of the target sentence;

[0021] Performing entity modification on the target sentence to determine a modified sentence, and using the language representation model to determine a model training loss function according to a sentence vector of the target sentence and a sentence vector of the modified sentence;

[0022] Determining whether the language representation model is trained according to the cross entropy loss function, the classification loss function and the model training loss function;

[0023] If so, the trained language representation model is used as a pre-trained language model.

[0024] In one embodiment, determining a target phrase according to the sample text information includes:

[0025] Segment the sample text information based on the preset text segmentation length to determine the corpus short sentences;

[0026] The corpus short sentences are deduplicated to determine the target short sentences.

[0027] In one embodiment, determining whether the language representation model is trained according to the cross entropy loss function, the classification loss function, and the model training loss function includes:

[0028] Performing a weighted summation on the cross entropy loss function, the classification loss function, and the model training loss function to determine a target loss function;

[0029] If the target loss function is greater than or equal to a preset loss function threshold, it is determined that the language representation model is not trained;

[0030] If the target loss function is less than the loss function threshold, it is determined that the language representation model training is completed.

[0031] In one embodiment, obtaining a text to be recognized, and determining a text concept of the text to be recognized based on the target noun database and the concept database through the pre-trained language model includes:

[0032] Acquire a text to be recognized, and determine, according to the target noun database, a noun subsequence corresponding to the text to be recognized as a noun to be recognized;

[0033] Determining target text information corresponding to the text to be recognized according to the noun to be recognized and the target noun database;

[0034] Determining the concept word vector corresponding to the target text information according to the target text information through the pre-trained language model;

[0035] The text concept of the text to be recognized is determined according to the concept word vector corresponding to the concept database and the target text information.

[0036] In a second aspect, the present application further provides a text concept recognition device, the device comprising:

[0037] A noun database construction module is used to construct a target noun database based on sample text information;

[0038] A concept word vector determination module is used to determine the concept word vector corresponding to the sample text information through a pre-trained language model;

[0039] A concept database construction module is used to construct a concept database according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0040] The text concept determination module is used to obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and the concept database through the pre-trained language model.

[0041] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:

[0042] Construct a target noun database based on sample text information;

[0043] Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0044] A concept database is constructed according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0045] A text to be recognized is obtained, and a text concept of the text to be recognized is determined based on the target noun database and the concept database through the pre-trained language model.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0047] Construct a target noun database based on sample text information;

[0048] Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0049] A concept database is constructed according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0050] A text to be recognized is obtained, and a text concept of the text to be recognized is determined based on the target noun database and the concept database through the pre-trained language model.

[0051] The above-mentioned text concept recognition method, device, computer equipment and storage medium construct a target noun database based on sample text information; determine the concept word vector corresponding to the sample text information through a pre-trained language model; construct a concept database based on the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and the concept database through a pre-trained language model. The problem of high recognition cost of text concepts and low recognition accuracy of text concepts when extracting concept information from text through an algorithm model is solved. The above scheme first constructs a target noun database based on the nouns in the sample corpus information, and then constructs a concept database based on the expected concepts and concept word vectors corresponding to the concept words in the sample corpus information. Through the pre-trained language model, the text concept of the text to be recognized is determined based on the target noun database and the concept database, avoiding concept annotation of a large amount of text, reducing the cost of text concept annotation, and improving the accuracy of text concept recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 An application environment diagram of a text concept recognition method in an embodiment;

[0053] Figure 2 A flowchart of a text concept recognition method in one embodiment;

[0054] Figure 3 A flowchart of a text concept recognition method in another embodiment;

[0055] Figure 4 A flowchart of a text concept recognition method in another embodiment;

[0056] Figure 5 is a structural block diagram of a text concept recognition device in one embodiment;

[0057] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] The text concept recognition method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The server 104 builds a target noun database based on the sample text information; determines the concept word vector corresponding to the sample text information through the pre-trained language model; builds a concept database based on the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the concept database table contains concept identifiers, concept nouns of corpus concepts, and the mapping relationship between corpus concepts and concept identifiers; obtains the text to be recognized, determines the text concept of the text to be recognized based on the target noun database and the concept database through the pre-trained language model, and sends the text concept of the text to be recognized to the terminal 102 through the communication network. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0060] In one embodiment, Figure 2 As shown, a text concept recognition method is provided. This embodiment takes the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0061] S210: Construct a target noun database based on the sample text information.

[0062] The sample text information may be corpus content obtained from the Internet, for example, it may be entry information obtained from Internet sources such as Wikipedia. The target noun database includes: sample text information, noun names corresponding to the sample text information, noun identifiers, noun types, noun subsequences, noun occurrence frequencies, noun features, mapping relationships between noun identifiers and noun names, and noun position information.

[0063] Specifically, sample text information is obtained, and nouns contained in the sample text information are extracted, the nouns extracted from the sample text information are marked, and a target noun database is constructed according to the sample text information, the nouns contained in the sample text information, and the nouns extracted from the marked sample text information. The target noun database may include: a table phrase database table_phrases, a pattern database Patterns, an identification database pattern2id, a frequency database id2ends, a noun segmentation database wordToken, a text information database Sentences, and a feature database Features.

[0064] Among them, table_phrases records the noun names of nouns extracted from the sample text information, the noun tags of nouns extracted from the sample text information, and the noun types of nouns extracted from the sample text information; the pattern database Patterns records the nouns extracted from the sample text information, the noun subsequences of nouns extracted from the sample text information, and the frequency of occurrence of nouns extracted from the sample text information; pattern2id records the noun names and noun identifiers; id2ends records the noun names and noun occurrence frequencies; wordTokens records noun segmentation information; Sentences records sample text information; Features records noun names, noun features, and feature values.

[0065] S220. Determine the concept word vector corresponding to the sample text information through the pre-trained language model.

[0066] The pre-trained language model may be a trained BERT (language representation model). The concept word vector refers to the word vector of a noun with a noun concept in the sample corpus information. The noun concept refers to information that can represent the attributes of a noun. A noun with a noun concept can be used as a concept word.

[0067] Specifically, seed concepts are determined in advance, which are predefined concept information, including objects, scenes, food, buildings, plants, etc. Through the pre-trained language model, concept words in the sample corpus information are determined based on the seed concepts, and the concept word vectors of the concept words are determined.

[0068] S230: construct a concept database according to the sample text information, the concept word vector corresponding to the sample text information, and the corpus concepts of the sample text information.

[0069] The concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers.

[0070] Exemplarily, after constructing the concept database, the concept database can also be updated, specifically by obtaining the concepts to be stored, determining the vector similarity between the concept vector of the concepts to be stored and the concept vector of the seed concept, determining the concepts to be stored corresponding to the concept vectors of the K concepts to be stored with the highest similarity as extended concepts, where K is a positive integer, and storing the extended concepts as new seed concepts in the concept database.

[0071] S240, obtaining a text to be recognized, and determining a text concept of the text to be recognized based on a target noun database and a concept database through a pre-trained language model.

[0072] Specifically, obtain the text to be recognized, determine the noun subsequence corresponding to the text to be recognized based on the target noun database, determine the text information corresponding to the text to be recognized from the target noun database according to the noun subsequence corresponding to the text to be recognized, determine the concept word vector of the text information corresponding to the text to be recognized through the pre-trained language model, and determine the text concept of the text to be recognized according to the concept word vector of the text information corresponding to the text to be recognized and the concept database.

[0073] In the above text concept recognition method, a target noun database is constructed according to sample text information; the concept word vector corresponding to the sample text information is determined through a pre-trained language model; a concept database is constructed according to the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the text to be recognized is obtained, and the text concept of the text to be recognized is determined based on the target noun database and the concept database through a pre-trained language model. The problem of high recognition cost of text concepts and low recognition accuracy of text concepts when extracting concept information from text through an algorithm model is solved. The above scheme first constructs a target noun database according to the nouns in the sample corpus information, and then constructs a concept database according to the expected concepts and concept word vectors corresponding to the concept words in the sample corpus information. The text concept of the text to be recognized is determined based on the target noun database and the concept database through a pre-trained language model, avoiding the problem of excessively high text annotation costs caused by concept annotation of a large amount of text, reducing the text concept annotation cost, and improving the text concept recognition accuracy.

[0074] In one embodiment, Figure 3 As shown, a target noun database is constructed based on sample text information, including:

[0075] S310: Construct a corpus noun database based on the sample text information.

[0076] The corpus noun database includes sample text information, noun names corresponding to all noun subsequences in the sample text information, noun identifiers, noun types, noun subsequences, noun occurrence frequencies, noun features, noun position information, and mapping relationships between noun identifiers and noun names.

[0077] Specifically, all noun subsequences in the sample text information are determined, one noun subsequence corresponds to one noun information, and a corpus noun database is constructed according to the noun subsequences.

[0078] S320: Determine high-frequency subsequences in the sample text information.

[0079] Among them, high-frequency subsequences refer to noun subsequences that appear frequently in sample text information.

[0080] Specifically, the high-frequency subsequences in the sample text information are determined according to the corpus noun database and a preset frequency threshold, that is, the noun occurrence frequency of the noun subsequence in the sample text information is determined according to the corpus noun database, and the noun subsequences whose noun occurrence frequency is greater than the preset frequency threshold are taken as high-frequency subsequences.

[0081] S330, constructing a target noun database according to the high-frequency subsequences and the corpus noun database.

[0082] Specifically, database information corresponding to high-frequency subsequences is extracted from the corpus noun database, and a target noun database is constructed according to the database information corresponding to the high-frequency subsequences.

[0083] Optionally, the noun subsequence corresponding to the sample corpus information may be determined, and an invalid subsequence may be determined from the noun subsequence, and the invalid subsequence may be stored in an invalid noun database. The invalid subsequence refers to a noun subsequence corresponding to an invalid noun.

[0084] The above scheme constructs a target noun database based on high-frequency subsequences in sample text information, which can improve the data quality in the target noun database.

[0085] In one embodiment, a target noun database is constructed based on high-frequency subsequences and a corpus noun database, including:

[0086] Based on the corpus noun database, the noun identifier and noun feature of the high-frequency subsequence are determined; through the regression model, the confidence of the high-frequency subsequence is determined according to the noun identifier and noun feature; according to the confidence of the high-frequency subsequence, the target subsequence is determined from the high-frequency subsequence, and the target noun database is constructed according to the target subsequence and the corpus noun database.

[0087] Optionally, after the target noun database is constructed, it is possible to continue to obtain new text information, and update the target noun database according to the new text information and the noun subsequences in the new text information.

[0088] The above scheme determines the confidence of the high-frequency subsequence according to the noun identifier and noun feature of the high-frequency subsequence, and can select the high-frequency subsequence with higher confidence as the target subsequence, so as to construct the target noun database according to the target subsequence and the corpus noun database, thereby further improving the data quality in the target noun database.

[0089] In one embodiment, Figure 4 As shown, before determining the concept word vector corresponding to the sample text information through the pre-trained language model, it also includes:

[0090] S410: Determine a target short sentence according to the sample text information.

[0091] The sample text information refers to the corpus information used for model training.

[0092] Specifically, the sample text information is segmented into a plurality of short sentences, and the short sentences determined after the sample text information is segmented can be used as target short sentences.

[0093] S420, performing masking on the concept nouns in the target short sentence, and using a language representation model to predict concept words for the target short sentence after the masking, and determining a cross entropy loss function of the language representation model based on the concept word prediction result and the concept nouns in the target short sentence.

[0094] Among them, ‌Mask processing‌ is a technique for encoding and decoding data using masks, which is mainly used to protect the privacy of data and prevent unauthorized access. Masking hides certain parts of the data while retaining the readability of other parts by converting the original data into another format. ‌The cross entropy loss function‌ is a loss function commonly used in machine learning and deep learning, especially for classification problems. It is used to evaluate the difference between the predicted probability distribution and the true distribution of the model, and can measure the accuracy of the model's prediction. Generally speaking, the smaller the cross entropy loss, the closer the model's prediction result is to the true label.

[0095] Specifically, the concept nouns in the target short sentence are masked, and then the language representation model is used to predict the concept nouns in the target short sentence based on the masked target short sentence to determine the concept word prediction result. The cross entropy loss function of the language representation model is determined based on the concept word prediction result and the masked concept nouns.

[0096] S430, using the language representation model to predict the sentence type of the target sentence, determining the predicted sentence type, and determining a classification loss function of the language representation model according to the predicted sentence type and the actual sentence type of the target sentence.

[0097] The actual sentence type may be the sentence type of the pre-labeled target sentence. The sentence type may be a judgment sentence, an imperative sentence, a declarative sentence, an exclamatory sentence, an elliptical sentence, an affirmative sentence, or a negative sentence. The classification loss function is a function used to measure the difference between the class label distribution predicted by the model and the true label distribution.

[0098] S440, perform entity modification on the target short sentence, determine the modified short sentence, and use the language representation model to determine the model training loss function according to the short sentence vector of the target short sentence and the short sentence vector of the modified short sentence.

[0099] Among them, entity modification of the target short sentence means modifying part of the content of the target short sentence. The ‌model training loss function‌ is a function used in machine learning to evaluate the difference between the model prediction value and the true value. Its main function is to guide model optimization and improve the prediction accuracy of the model by minimizing the loss function.

[0100] S450: Determine whether the language representation model is trained according to the cross entropy loss function, the classification loss function, and the model training loss function.

[0101] Exemplarily, if the cross entropy loss function is less than the first preset loss function, the classification loss function is less than the second preset loss function, and the model training loss function is less than the third preset loss function, it is determined that the language representation model training is completed. If the cross entropy loss function is greater than or equal to the first preset loss function, the classification loss function is greater than or equal to the second preset loss function, or the model training loss function is greater than or equal to the third preset loss function, it is determined that the language representation model training is not completed. Among them, the first preset loss function, the second preset loss function and the third preset loss function can all be set according to actual needs.

[0102] S460: If yes, the trained language representation model is used as a pre-trained language model.

[0103] Specifically, if the language representation model is trained, the trained language representation model is used as the pre-trained language model. It is understandable that if the language representation model is not trained, the language representation model is continuously iterated until the language representation model is trained.

[0104] The above scheme performs mask processing on the target short sentence corresponding to the sample text information, and determines the cross entropy loss function of the language representation model through the language representation model and the target short sentence after mask processing; predicts the short sentence type of the target short sentence through the language representation model, and determines the classification loss function of the language representation model according to the prediction result of the short sentence type; performs entity modification on the target short sentence to determine the modified short sentence, determines the short sentence vector of the modified short sentence according to the language representation model, and determines the model training loss function according to the short sentence vector of the modified short sentence and the short sentence vector of the target short sentence; determines whether the language representation model is trained based on the cross entropy loss function, the classification loss function and the model training loss function, which can improve the reliability of the pre-trained language model.

[0105] In one embodiment, determining a target short sentence according to sample text information includes:

[0106] The sample text information is divided into sentences based on the preset text segmentation length to determine the corpus short sentences; the corpus short sentences are deduplicated to determine the target short sentences.

[0107] The text segmentation length can be set according to actual needs. The length of the short sentences in the corpus meets the requirements of the text segmentation length.

[0108] The above scheme divides the sample text information into sentences based on the preset text segmentation length, which can better ensure that the target short sentence contains the conceptual words in the sample text information. Deduplication of the corpus short sentences can reduce the model training overhead of the language representation model.

[0109] In one embodiment, determining whether the language representation model is trained according to the cross entropy loss function, the classification loss function, and the model training loss function includes:

[0110] The cross entropy loss function, classification loss function and model training loss function are weightedly summed to determine the target loss function; if the target loss function is greater than or equal to the preset loss function threshold, it is determined that the language representation model has not been trained; if the target loss function is less than the loss function threshold, it is determined that the language representation model has been trained.

[0111] Specifically, a weighted sum is performed on the cross entropy loss function, the classification loss function, and the model training loss function, and the result of the weighted sum is used as the target loss function. If the target loss function is greater than or equal to a preset loss function threshold, it is determined that the language representation model has not been trained; if the target loss function is less than the loss function threshold, it is determined that the language representation model has been trained. If the language representation model has not been trained, the language representation model is continuously iterated.

[0112] The above scheme determines the target loss function by weighted summation of the cross entropy loss function, the classification loss function and the model training loss function, and determines whether the language representation model is trained based on the target loss function. This can improve the reliability of the language representation model while improving the training efficiency of the language representation model.

[0113] In one embodiment, a text to be recognized is obtained, and a text concept of the text to be recognized is determined based on a target noun database and a concept database through a pre-trained language model, including:

[0114] Obtain the text to be recognized, and determine the noun subsequence corresponding to the text to be recognized as the noun to be recognized according to the target noun database; determine the target text information corresponding to the text to be recognized according to the noun to be recognized and the target noun database; determine the concept word vector corresponding to the target text information according to the target text information through the pre-trained language model; determine the text concept of the text to be recognized according to the concept database and the concept word vector corresponding to the target text information.

[0115] The text to be recognized is the corpus information whose text concept needs to be determined.

[0116] Specifically, the text to be recognized is obtained, and a noun subsequence of the text to be recognized is extracted from the text to be recognized as the subsequence to be recognized. The subsequence to be recognized is matched with the noun subsequence in the target noun database for similarity, and the noun subsequence with the highest similarity to the subsequence to be recognized is determined as the noun to be recognized. The sample text information corresponding to the noun to be recognized is determined according to the target noun database, and the sample text information corresponding to the recognized noun is integrated to determine the target text information. The concept word vector in the target text information is extracted through the pre-trained language model, and the concept database is searched based on the concept word vector corresponding to the target text information to determine the text concept of the text to be recognized.

[0117] The above scheme can determine the text concepts of the text to be recognized based on the target noun database and the concept database through the pre-trained language model, thereby improving the recognition accuracy of the text concepts of the text to be recognized.

[0118] Exemplarily, based on the above embodiment, the text concept recognition method includes:

[0119] Determine all noun subsequences in the sample text information, one noun subsequence corresponds to one noun information, and build a corpus noun database based on the noun subsequences. Determine the high-frequency subsequences in the sample text information based on the corpus noun database and the preset frequency threshold, that is, determine the noun occurrence frequency of the noun subsequence in the sample text information based on the corpus noun database, and take the noun subsequence whose noun occurrence frequency is greater than the preset frequency threshold as the high-frequency subsequence. Determine the noun identifier and noun feature of the high-frequency subsequence based on the corpus noun database; determine the confidence of the high-frequency subsequence based on the noun identifier and noun feature through the regression model; determine the target subsequence from the high-frequency subsequence based on the confidence of the high-frequency subsequence, and build the target noun database based on the target subsequence and the corpus noun database.

[0120] Based on the preset text segmentation length, the sample text information is divided into sentences to determine the corpus sentences; the corpus sentences are deduplicated to determine the target sentences, the concept nouns in the target sentences are masked, and then the language representation model is used to predict the concept nouns in the target sentences based on the masked target sentences to determine the concept word prediction results. The cross entropy loss function of the language representation model is determined based on the concept word prediction results and the masked concept nouns. The language representation model is used to predict the sentence type of the target sentence to determine the predicted sentence type, and the classification loss function of the language representation model is determined based on the predicted sentence type and the actual sentence type of the target sentence. The target sentence is modified to determine the modified sentence, and the language representation model is used to determine the model training loss function based on the sentence vector of the target sentence and the sentence vector of the modified sentence. The cross entropy loss function, classification loss function and model training loss function are weighted and summed, and the result of the weighted sum is used as the target loss function. If the target loss function is greater than or equal to the preset loss function threshold, it is determined that the language representation model has not been trained; if the target loss function is less than the loss function threshold, it is determined that the language representation model has been trained. If the language representation model has not been trained, continue to iterate the language representation model.

[0121] Predetermine the seed concepts, which are predefined concept information, including objects, scenes, food, buildings, plants, etc. Through the pre-trained language model, determine the concept words in the sample corpus information based on the seed concepts, and determine the concept word vectors of the concept words.

[0122] A concept database is constructed based on sample text information, concept word vectors corresponding to the sample text information, and corpus concepts of the sample text information. The concept database table contains concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers. After the concept database is constructed, the concept database can also be updated, specifically: obtaining the concept to be stored, determining the vector similarity between the concept vector of the concept to be stored and the concept vector of the seed concept, determining the concept to be stored corresponding to the concept vectors of the K concepts to be stored with the highest similarity as the extended concept, K is a positive integer, and storing the extended concept as a new seed concept in the concept database.

[0123] Obtain the text to be recognized, extract the noun subsequence of the text to be recognized as the subsequence to be recognized, perform similarity matching between the subsequence to be recognized and the noun subsequence in the target noun database, determine the noun subsequence with the highest similarity to the subsequence to be recognized as the noun to be recognized, determine the sample text information corresponding to the noun to be recognized based on the target noun database, integrate the sample text information corresponding to the recognized noun, and determine the target text information. Extract the concept word vector in the target text information through the pre-trained language model, search the concept database based on the concept word vector corresponding to the target text information, and determine the text concept of the text to be recognized.

[0124] In the above text concept recognition method, a target noun database is constructed based on sample text information; the concept word vector corresponding to the sample text information is determined through a pre-trained language model; a concept database is constructed based on the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the text to be recognized is obtained, and the text concept of the text to be recognized is determined based on the target noun database and the concept database through a pre-trained language model. The problem of high recognition cost of text concepts and low recognition accuracy of text concepts when extracting concept information from text through an algorithm model is solved. The above scheme first constructs a target noun database based on the nouns in the sample corpus information, and then constructs a concept database based on the expected concepts and concept word vectors corresponding to the concept words in the sample corpus information. Through the pre-trained language model, the text concept of the text to be recognized is determined based on the target noun database and the concept database, avoiding concept annotation of a large amount of text, reducing the cost of text concept annotation, and improving the accuracy of text concept recognition.

[0125] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0126] Based on the same inventive concept, the embodiment of the present application also provides a text concept recognition device for implementing the text concept recognition method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more text concept recognition device embodiments provided below can refer to the limitations of the text concept recognition method above, and will not be repeated here.

[0127] In one embodiment, Figure 5 As shown, a text concept recognition device is provided, including: a noun database construction module 501, a concept word vector determination module 502, a concept database construction module 503 and a text concept determination module 504, wherein:

[0128] A noun database construction module 501 is used to construct a target noun database according to sample text information;

[0129] A concept word vector determination module 502 is used to determine the concept word vector corresponding to the sample text information through a pre-trained language model;

[0130] The concept database construction module 503 is used to construct a concept database according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores the concept identifier, the concept noun of the corpus concept, and the mapping relationship between the corpus concept and the concept identifier;

[0131] The text concept determination module 504 is used to obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and the concept database through the pre-trained language model.

[0132] Exemplarily, the noun database construction module 501 is specifically used for:

[0133] Construct a corpus noun database based on sample text information;

[0134] Determine high frequency subsequences in sample text information;

[0135] The target noun database is constructed based on the high-frequency subsequences and the corpus noun database.

[0136] Exemplarily, the noun database construction module 501 is also specifically used for:

[0137] Determine the noun identifiers and noun features of high-frequency subsequences based on the corpus noun database;

[0138] Through the regression model, the confidence of high-frequency subsequences is determined according to noun identification and noun features;

[0139] A target subsequence is determined from the high-frequency subsequences according to the confidence of the high-frequency subsequences, and a target noun database is constructed according to the target subsequences and the corpus noun database.

[0140] Exemplarily, the concept word vector determination module 502 is specifically used for:

[0141] Determine the target phrase based on the sample text information;

[0142] Masking the concept nouns in the target short sentence, and using the language representation model to predict the concept words of the target short sentence after masking, and determining the cross entropy loss function of the language representation model according to the concept word prediction results and the concept nouns in the target short sentence;

[0143] The language representation model is used to predict the sentence type of the target sentence, and the predicted sentence type is determined, and the classification loss function of the language representation model is determined according to the predicted sentence type and the actual sentence type of the target sentence;

[0144] Perform entity modification on the target sentence, determine the modified sentence, and use the language representation model to determine the model training loss function according to the sentence vector of the target sentence and the sentence vector of the modified sentence;

[0145] Determine whether the language representation model is trained based on the cross entropy loss function, classification loss function and model training loss function;

[0146] If so, the trained language representation model is used as a pre-trained language model.

[0147] Exemplarily, the concept word vector determination module 502 is further specifically used for:

[0148] Segment the sample text information based on the preset text segmentation length to determine the corpus short sentences;

[0149] Remove duplicate sentences from the corpus and determine the target sentences.

[0150] Exemplarily, the concept word vector determination module 502 is further specifically used for:

[0151] Perform weighted summation on the cross entropy loss function, classification loss function, and model training loss function to determine the target loss function;

[0152] If the target loss function is greater than or equal to the preset loss function threshold, it is determined that the language representation model is not trained;

[0153] If the target loss function is less than the loss function threshold, it is determined that the language representation model training is completed.

[0154] Exemplarily, the text concept determination module 504 is further specifically configured to:

[0155] Obtain the text to be recognized, and determine the noun subsequence corresponding to the text to be recognized as the noun to be recognized according to the target noun database;

[0156] Determine the target text information corresponding to the text to be recognized according to the noun to be recognized and the target noun database;

[0157] By pre-training the language model, the concept word vector corresponding to the target text information is determined according to the target text information;

[0158] The text concept of the text to be recognized is determined based on the concept word vector corresponding to the concept database and the target text information.

[0159] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be realized through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a text concept recognition method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0160] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0161] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0162] Step 1: Build a target noun database based on sample text information;

[0163] Step 2: Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0164] Step 3: construct a concept database based on the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0165] Step 4: Obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and concept database through the pre-trained language model.

[0166] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0167] Step 1: Build a target noun database based on sample text information;

[0168] Step 2: Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0169] Step 3: construct a concept database based on the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0170] Step 4: Obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and concept database through the pre-trained language model.

[0171] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0172] Step 1: Build a target noun database based on sample text information;

[0173] Step 2: Determine the concept word vector corresponding to the sample text information through the pre-trained language model;

[0174] Step 3: construct a concept database based on the sample text information, the concept word vector corresponding to the sample text information, and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers;

[0175] Step 4: Obtain the text to be recognized, and determine the text concept of the text to be recognized based on the target noun database and concept database through the pre-trained language model.

[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0177] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0178] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0179] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A text concept recognition method, characterized in that: include: Construct a corpus noun database based on sample text information; Determining high-frequency subsequences in the sample text information; Determine noun identifiers and noun features of high-frequency subsequences based on the corpus noun database; Determining the confidence of the high-frequency subsequence according to the noun identifier and the noun feature through a regression model; Determining a target subsequence from the high-frequency subsequence according to the confidence of the high-frequency subsequence, and constructing a target noun database according to the target subsequence and the corpus noun database; Determine the concept word vector corresponding to the sample text information through the pre-trained language model; A concept database is constructed according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers; Acquire a text to be recognized, and determine, according to the target noun database, a noun subsequence corresponding to the text to be recognized as a noun to be recognized; Determining target text information corresponding to the text to be recognized according to the noun to be recognized and the target noun database; Determining the concept word vector corresponding to the target text information according to the target text information through the pre-trained language model; The text concept of the text to be recognized is determined according to the concept word vector corresponding to the concept database and the target text information.

2. The method according to claim 1, characterized in that Before determining the concept word vector corresponding to the sample text information through the pre-trained language model, it also includes: Determining a target short sentence according to the sample text information; Masking the concept nouns in the target short sentence, and using a language representation model to predict concept words for the target short sentence after the masking process, and determining a cross entropy loss function of the language representation model according to the concept word prediction result and the concept nouns in the target short sentence; Using the language representation model to predict the sentence type of the target sentence, determine the predicted sentence type, and determine the classification loss function of the language representation model according to the predicted sentence type and the actual sentence type of the target sentence; Performing entity modification on the target sentence to determine a modified sentence, and using the language representation model to determine a model training loss function according to a sentence vector of the target sentence and a sentence vector of the modified sentence; Determining whether the language representation model is trained according to the cross entropy loss function, the classification loss function and the model training loss function; If so, the trained language representation model is used as a pre-trained language model.

3. The method according to claim 2, characterized in that Determining a target short sentence according to the sample text information includes: Segment the sample text information based on the preset text segmentation length to determine the corpus short sentences; The corpus short sentences are deduplicated to determine the target short sentences.

4. The method according to claim 2, characterized in that: Determining whether the language representation model is trained according to the cross entropy loss function, the classification loss function, and the model training loss function includes: Performing a weighted summation on the cross entropy loss function, the classification loss function, and the model training loss function to determine a target loss function; If the target loss function is greater than or equal to a preset loss function threshold, it is determined that the language representation model is not trained; If the target loss function is less than the loss function threshold, it is determined that the language representation model training is completed.

5. A text concept recognition device, characterized in that: The text concept recognition device comprises: A noun database construction module is used to construct a corpus noun database according to sample text information; determine a high-frequency subsequence in the sample text information; determine a noun identifier and a noun feature of the high-frequency subsequence based on the corpus noun database; determine the confidence of the high-frequency subsequence according to the noun identifier and the noun feature through a regression model; determine a target subsequence from the high-frequency subsequence according to the confidence of the high-frequency subsequence, and construct a target noun database according to the target subsequence and the corpus noun database; A concept word vector determination module is used to determine the concept word vector corresponding to the sample text information through a pre-trained language model; A concept database construction module is used to construct a concept database according to the sample text information, the concept word vector corresponding to the sample text information and the corpus concept of the sample text information; the concept database table stores concept identifiers, concept nouns of corpus concepts, and mapping relationships between corpus concepts and concept identifiers; A text concept determination module is used to obtain the text to be recognized, determine the noun subsequence corresponding to the text to be recognized as the noun to be recognized according to the target noun database; determine the target text information corresponding to the text to be recognized according to the noun to be recognized and the target noun database; determine the concept word vector corresponding to the target text information according to the target text information through the pre-trained language model; and determine the text concept of the text to be recognized according to the concept database and the concept word vector corresponding to the target text information.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Text classification method and obtained text classifier

    CN106951565A

  • Text enhancement method, text classification method and related devices

    CN112906392A