Text tracing method and device, equipment and medium
The pretrained model is generated by deep learning models and combined with time extraction algorithms, which solves the problem of low text traceability efficiency in the existing technology, and achieves efficient and accurate text traceability, reducing resource consumption.
Patent Information
- Application Number
- CN202411287140.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-13
AI Technical Summary
When processing large-scale text data, the prior art requires a lot of resources and time to trace text through document marking method, and is relatively inefficient.
A deep learning model is adopted to generate a pre-trained model by training the text data set, which is used to generate entry definitions, and a time extraction algorithm is used to sort and filter the descriptive feature representation and candidate text feature representation to realize text traceability.
It improves the efficiency and accuracy of text traceability, reduces resource consumption, eliminates the need to mark and maintain text data in the library, and saves computing resources and maintenance time.
Smart Images

Figure CN120011480A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of text processing, and in particular to a text tracing method, apparatus, device, and medium. Background Art
[0002] In the era of big data, the demand for tracing the source of text information is growing. Accurately tracing the source of text can ensure the reliability and authenticity of the information, help resolve disputes over information sources, and prevent the spread of false information.
[0003] In the related art, the term traceability is achieved by specifically marking documents. When processing large-scale data sets, each document is assigned a unique mark (such as a mark in the form of a code, symbol, etc.). When the term traceability is required, the document containing the term can be quickly located by searching for the mark related to the term.
[0004] However, the document tagging method requires preprocessing and continuous maintenance of the document, which consumes resources and time. The traceability process also consumes a lot of computing resources, and the efficiency is low when tracing massive texts. Summary of the invention
[0005] The embodiments of the present application provide a text tracing method, apparatus, device, and medium, which can improve the efficiency of text tracing and reduce resource consumption. The technical solution is as follows:
[0006] In one aspect, a text tracing method is provided, the method comprising:
[0007] Acquire a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries;
[0008] Acquire search term content and perform feature extraction on the search term content to obtain term feature representation, wherein the search term content is a term input by a user for text tracing;
[0009] Based on the matching degree between the term feature representation and a plurality of candidate text feature representations in the feature representation library, determining at least one candidate text feature representation from the feature representation library as a first feature representation set;
[0010] Generate a term interpretation text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the term interpretation text to obtain a feature representation of the interpretation;
[0011] Based on a time extraction algorithm, the interpretation feature representation and the first feature representation set are sorted and filtered to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
[0012] In another aspect, a text tracing device is provided, the device comprising:
[0013] A training module, used to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries;
[0014] A feature extraction module is used to obtain search term content and perform feature extraction on the search term content to obtain a term feature representation, wherein the search term content is a term input by a user for text tracing;
[0015] A matching module, configured to determine at least one candidate text feature representation from the feature representation library as a first feature representation set based on a degree of matching between the feature representation of the term and a plurality of candidate text feature representations in the feature representation library;
[0016] A definition generation module, used to generate a definition text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the definition text to obtain a definition feature representation;
[0017] A tracing module is used to sort and filter the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
[0018] In an optional embodiment, the training module is also used to obtain the training text data set; train the deep learning model based on the multiple sentences in the training text data set to obtain a first-stage model; adjust the model architecture and model parameters of the first-stage model to obtain a second-stage model; train the second-stage model based on the multiple entries annotated with explanations in the training text data set to obtain the pre-trained model.
[0019] In an optional embodiment, the training module is also used to obtain multiple corpus sentences and multiple entries marked with interpretations, wherein the multiple corpus sentences are sentence contents obtained after splitting multiple text contents; synonym replacement is performed on the multiple corpus sentences to obtain a first sentence set, wherein the i-th sentence in the first sentence set is obtained after synonym replacement for each term in the i-th corpus sentence in the multiple corpus sentences; the multiple corpus sentences are screened based on the word frequency and inverse document frequency method to obtain a second sentence set, wherein the word frequency refers to the frequency of occurrence of each term in the multiple corpus sentences, and the inverse document frequency is used to indicate the importance of each term in the multiple corpus sentences; the first sentence set and the second sentence set are determined as the multiple sentences.
[0020] In an optional embodiment, the training module is also used to add a custom feature extraction layer to the first-stage model to obtain the first-stage model after adjusting the model architecture; based on preset parameter standards, the learning rate, number of training rounds and batch size of the first-stage model after adjusting the model architecture are adjusted to obtain the second-stage model, wherein the learning rate is used to determine the step size of the model parameter update during the training process, the number of training rounds refers to the number of times the training text data set is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.
[0021] In an optional embodiment, the matching module is also used to obtain a candidate text library, which contains multiple candidate text contents; after performing feature extraction on the multiple candidate text contents respectively, the candidate text feature representations in the feature representation library are obtained; based on the similarity between the term feature representation and the multiple candidate text feature representations, the multiple candidate text feature representations are sorted to obtain serial numbers corresponding to the multiple candidate text feature representations respectively; and the at least one candidate text feature representation whose serial number meets the preset requirements is used as the first feature representation set.
[0022] In an optional embodiment, the interpretation generation module is also used to input the term feature representation into the pre-trained model, and output a hidden state corresponding to the term feature representation, wherein the hidden state refers to the encoded representation of the search term content by the pre-trained model, and is used to reflect the pre-trained model's understanding of the search term content; and generate a term interpretation text corresponding to the search term content based on the hidden state through a decoder.
[0023] In an optional embodiment, the tracing module is used to parse the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain timestamp information corresponding to the interpretation feature representation and all candidate text feature representations in the first feature representation set, and the timestamp information is used to indicate the release time of the search term content and the candidate text feature representation; sort the interpretation feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; filter at least one feature representation from the first feature representation set based on the similarity between the interpretation feature representation and the second feature representation set to obtain a second feature representation set; determine the text content corresponding to the second feature representation set as the text content set containing the search term; and obtain the text tracing record based on the sorting result and the text content set.
[0024] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a text tracing method as described in any of the above-mentioned embodiments of the present application.
[0025] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a text tracing method as described in any of the above-mentioned embodiments of the present application.
[0026] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text tracing method described in any of the above embodiments.
[0027] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:
[0028] By adjusting the architecture of the deep learning model and using targeted training text data sets to train the deep learning model, a pre-trained model capable of generating text interpretation content is obtained. The pre-trained model generates interpretations for the search terms input by the user. Based on the candidate texts that meet the matching requirements and the interpretation content with the semantic similarity of the search content, the search terms can be understood from a natural language perspective, the accuracy of text tracing can be improved, the source of the search terms can be effectively tracked and traced, and the efficiency of text tracing can be improved. Compared with the related art, which uses document tagging to annotate and continuously maintain massive text data, this solution does not require tagging and maintenance operations on the text data in the library, which can save computing resources and maintenance time. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0030] Figure 1 is a schematic diagram of a text tracing system provided by an exemplary embodiment of the present application;
[0031] Figure 2 is a flowchart of a text tracing method provided by an exemplary embodiment of the present application;
[0032] Figure 3 is a structural block diagram of a text tracing device provided by an exemplary embodiment of the present application;
[0033] Figure 4 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0035] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0036] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0037] It should be noted that the information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0038] It should be understood that, although the terms first, second, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0039] First, a brief introduction is given to the terms involved in the embodiments of this application:
[0040] Word2Vec model (Word to Vector): The main purpose is to generate word vectors, that is, to map words into vector space. Each word is represented as a vector of fixed dimension, and these vectors can capture the semantic relationship between words.
[0041] In this application, feature representation refers to a quantity, and the process of extracting features from text to obtain feature representation can be regarded as the process of converting text into a vector.
[0042] BERT model (Bidirectional Encoder Representations from Transformers): A pre-trained language model based on the Transformer architecture that can also convert text into vector form. Compared with the Word2Vector model, the BERT model generates context-related word vectors. It can simultaneously consider the contextual information on the left and right sides of a word through a multi-layer self-attention mechanism to generate deep bidirectional representations.
[0043] Among them, the deep learning model used to generate term interpretations in this application refers to the BERT model. A pre-trained model is obtained by training and fine-tuning the BERT model. The pre-trained model has the functions of the BERT model and can be applied to the task of generating text interpretations.
[0044] Time extraction algorithm: An algorithm used to extract time information from natural language text, combining natural language processing technology and specific rule patterns to identify time expressions in text. Use vocabulary features, grammatical structure, and contextual information to accurately locate and extract time information. For example, the algorithm can identify clear date expressions: "August 29, 2024", relative time expressions: "yesterday", "next week", "next month", etc., and fuzzy time range expressions: "recently", "for some time", etc. In text data such as news reports and social media, time extraction algorithms can help analyze the time of occurrence of events and construct a timeline of events, so as to better understand the development context and evolution process of events.
[0045] TF-IDF (Term Frequency-Inverse Document Frequency): In text processing, TF-IDF can be used to extract important keywords in a document. Term frequency (TF) indicates the frequency of a word appearing in a document. Words with high frequency of appearance may be the key content of the document. Inverse document frequency (IDF) measures the rarity of a word in the entire document collection. If a word appears in many documents, its IDF value is relatively low; if a word appears in only a few documents, its IDF value is relatively high. The TF-IDF method can highlight words that appear frequently in specific documents and are relatively rare in the entire document collection. These words are regarded as key information of the document.
[0046] In the era of big data, with the development of information technology, data is increasing exponentially, and the demand for tracing the source of text information is also growing. Accurately tracing the source of text can ensure the reliability and authenticity of information, effectively resolve disputes about the source of information, and prevent the widespread dissemination of false information.
[0047] In the related art, the tracing operation of a term is achieved by specifically marking the document. When processing a large-scale data set, a unique tag is assigned to each document in the data set. The tag can be in the form of a code or other forms such as a symbol. When it is necessary to trace the source of a certain term, the document containing the term can be quickly located by searching for the tag related to the term.
[0048] However, using document markup to trace the source of text requires preprocessing of the document, and in the subsequent use process, it is also necessary to continue to maintain the document, which consumes resources and time. The tracing process consumes a lot of computing resources and is relatively inefficient.
[0049] Figure 1 It is a schematic diagram of a text tracing system provided by an exemplary embodiment of the present application, which can accurately and efficiently realize text tracing.
[0050] The text tracing system 100 involves a user terminal 110 and a server 120 , and the user terminal 110 and the server terminal 120 are connected via a communication network 130 .
[0051] The user terminal 110 inputs the search term content and sends it to the server terminal 120, and the server terminal 120 analyzes the search term content.
[0052] The server side 120 generates an interpretation of the search term content and selects candidate text content with a high semantic similarity to the search term content in the candidate text library, and determines the release time of each content by time extraction of the interpretation and candidate text content, and sorts them to obtain a complete text traceability record, and feeds the text traceability record back to the user side 110.
[0053] Among them, a pre-trained model is deployed in the server side 120. The pre-trained model is a model obtained by adjusting the architecture of the BERT model and training it using targeted sentences and term contents in the training text data set. The pre-trained model can perform semantic analysis on the search term content input by the user side 110 and generate a term interpretation of the search content term.
[0054] Before the user terminal 110 inputs the search term content to the server terminal 120, a candidate text library is prepared in advance, which contains multiple candidate text contents. The candidate text contents include text contents in the form of news, articles, reports, etc., which belong to the same field as the search term content.
[0055] Feature extraction is performed on multiple candidate content texts to obtain candidate text feature representations corresponding to each candidate text content.
[0056] After feature extraction of the search term content input by the user terminal 110, a term feature representation is obtained, and similarity matching is performed between the term feature representation and multiple candidate text feature representations, and the candidate text feature representation with a higher similarity to the term feature representation is determined as the first feature representation set.
[0057] After obtaining the interpretation of the search term content, feature extraction is performed on the interpretation to obtain the interpretation feature representation. Based on the time extraction algorithm, the text content corresponding to the interpretation feature representation and the first feature representation set is analyzed to determine the release time of these text contents, and sort them to obtain the sorting results of the interpretation feature representation and the first feature representation set.
[0058] A similarity analysis is performed between the interpretation feature representation and all feature representations in the first feature representation set, the semantic association between the above feature representations is determined, and the feature representations with higher similarity are screened as the second feature representation set, wherein the semantic association between the feature representations in the second feature representation set is high.
[0059] The text contents corresponding to the feature representations in the second feature representation set are reordered according to the sorting result to obtain a text tracing record, which can indicate the text content containing the search term content and the development context of the text content.
[0060] The above-mentioned user terminal 110 can be a terminal device in various forms such as a mobile phone, a tablet computer, a desktop computer, a portable laptop computer, a smart TV, a car terminal, a smart home device, etc., and the embodiments of the present application are not limited to this.
[0061] It is worth noting that the above-mentioned server end 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.
[0062] Among them, cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to realize data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which is used on demand and flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0063] In some embodiments, the above server can also be implemented as a node in a blockchain system.
[0064] In combination with the above-mentioned noun introduction and application scenarios, the text tracing method provided by the present application is described. The method can be executed by a server or a terminal, or can be executed by a server and a terminal together. In the embodiment of the present application, the method is described by taking the execution of the server as an example. Figure 2 As shown, Figure 2 1 is a flowchart of a text tracing method provided by an exemplary embodiment of the present application. The method includes the following steps.
[0065] Step 210: Obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model.
[0066] The training text dataset contains multiple sentences and multiple entries annotated with definitions, and the pre-trained model is used to generate entry definitions.
[0067] The deep learning model uses the BERT model, and the pre-trained model is obtained after fine-tuning and training the BERT model.
[0068] The content in the training text dataset may belong to the same field as the search term content entered by the user. For example, if the search term content entered by the user belongs to the terminology in the field of artificial intelligence, the content in the training text dataset may also be the sentence content obtained by splitting articles in the field, or the content obtained by annotating and defining multiple terms in the field.
[0069] Optionally, a training text dataset is obtained, and a deep learning model is trained based on multiple sentences in the training text dataset to obtain a first-stage model.
[0070] Exemplarily, a plurality of corpus sentences and a plurality of entries marked with definitions are obtained, and the plurality of corpus sentences are sentence contents obtained by splitting a plurality of text contents.
[0071] The field to which the search term content belongs is determined as the target field, and multiple text contents in the target field are obtained, including but not limited to professional literature, historical documents, academic papers, etc.
[0072] Remove redundant information such as reference lists, irrelevant links, annotations, and format errors from multiple text contents, and perform at least one of the following operations on them: remove stop words, remove content irrelevant to the target field, remove incorrect annotations, duplicate entries, etc.
[0073] The sentence contents obtained after splitting multiple text contents are as follows: [u1, u2, ..., um], where m is a positive integer and um refers to a corpus sentence.
[0074] Synonym replacement is performed on multiple corpus sentences to obtain a first sentence set, wherein the i-th sentence in the first sentence set is obtained by performing synonym replacement on each word item of the i-th corpus sentence in the multiple corpus sentences.
[0075] That is, for each word / term wi in each corpus sentence ui (i is a positive integer), the Word2Vec model is used to calculate its feature representation: vi=Word2Vec(word=wi).
[0076] For the synonym wj (j is a positive integer) of word wi, use the Word2Vec model to calculate its feature representation vj: vj=Word2Vec(word=wj).
[0077] Calculate the cosine similarity between the word feature representation and the synonym feature representation, select vj with the highest similarity to vi for replacement, and obtain the first sentence set.
[0078] Based on the word frequency and inverse document frequency method, multiple corpus sentences are screened to obtain a second sentence set, wherein the word frequency refers to the frequency of each word appearing in the multiple corpus sentences, and the inverse document frequency is used to indicate the importance of each word in the multiple corpus sentences.
[0079] Exemplarily, multiple corpus sentences are input: [u1, u2, ..., um], and for each word wi, its vector representation vi is calculated, vi = Word2Vec (word = wi).
[0080] For word wi, calculate its TF value in the corpus sentence: TF(wi) = number of occurrences / total number of words in the sentence, calculate the IDF value of word wi in multiple corpus sentences: IDF(wi) = log(total number of corpus sentences / number of corpus sentences containing the word), and construct the TF-IDF matrix: TF-IDF(wi) = TF(wi)*IDF(wi).
[0081] For each corpus sentence, multiply the TF-IDF value of each word in it by the corresponding word embedding vector: TF-IDF(wi, d)*vi. Add or average the weighted word embedding vectors of all words in the corpus sentence to get the vector representation of the corpus sentence: doc_vector = sum(TF-IDF(wi, d)*vi) /
[0082] (The total number of terms in corpus sentence d).
[0083] According to the TF-IDF weighted score of the corpus sentence, the probability of the corpus sentence being sampled is determined: sampling_probability(d)=sum(TF-IDF(wi,d)).
[0084] Using the weighted random sampling method, the second sentence set [t1, t2, ..., tm] is sampled from the training corpus to obtain a vector representation.
[0085] The first sentence set and the second sentence set are determined as a plurality of sentences.
[0086] Optionally, the model architecture and model parameters of the first-stage model are adjusted to obtain a second-stage model.
[0087] Exemplarily, a custom feature extraction layer is added to the first-stage model to obtain a first-stage model after adjusting the model architecture.
[0088] Among them, the custom feature extraction layer can be a fully connected layer.
[0089] The input vector for each Transformer (feature transformation) layer is [X1, X2, ..., Xn], and the output vector is [Y1, Y2, ..., Yn].
[0090] After adding the custom feature extraction layer, the new input vector is [X1, X2, …, Xn, Z1] and the new output vector is [Y1, Y2, …, Yn, Z2].
[0091] The weight of the custom feature extraction layer is W_custom, the input vector is x, and the output vector y is calculated as: y = σ(Wcustom·x), where σ represents an activation function, such as a ReLU function (Rectified Linear Unit, ReLU, rectified linear unit) or a Sigmoid function (logistic function).
[0092] The output vector of the last layer of the original BERT model is h0, h1, ..., hn. By adding a fully connected layer, the output vector can be changed to f(h0), f(h1), ..., f(hn), where f represents the transformation function of the fully connected layer.
[0093] Based on the preset parameter standards, the learning rate, number of training rounds and batch size of the first-stage model after adjusting the model architecture are adjusted to obtain the second-stage model.
[0094] Among them, the learning rate is used to determine the step size of the model parameter update during the training process, the number of training rounds refers to the number of times the training text dataset is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.
[0095] For example, during the update process of gradient descent, the learning rate adjustment is updated based on the following formula: Where α is the learning rate.
[0096] When adjusting the batch size, the number of samples used to update the weights each time during model training is the batch size.
[0097] Adjusting the number of training epochs, i.e. the number of times the model is iterated over the entire training dataset, usually involves multiple passes over the entire dataset for parameter updates.
[0098] Optionally, the second-stage model is trained based on a plurality of entries annotated with definitions in the training text dataset to obtain a pre-trained model.
[0099] A small number of key words in the target field are selected (for example, words whose frequency reaches a preset frequency threshold as key words) and their definitions are compiled. These words and definitions are used as multiple words annotated with definitions in the training text dataset, hereinafter referred to as fine-tuning corpus.
[0100] Exemplarily, the fine-tuning corpus is used to fine-tune the second-stage model. By modifying the model's attention weights and strategies, the model is enabled to understand the contextual meaning of terms and generate accurate and concise explanations. At the same time, the model is given a certain generalization ability on unseen data, so that term interpretations can be automatically generated for new terms.
[0101] For each sentence or sentence fragment of the fine-tuning corpus, receive the sentence or sentence fragment [t1, t2, ..., tm] as input. Tokenize each ti (word) and convert it into a sequence of tokens that the model can understand. Add a special token [CLS] at the beginning of the sequence to aggregate sentence-level semantic information. Add a [SEP] token at the end of the sequence to indicate the end of the sentence.
[0102] Assume that the tokenization result of the input sequence is [101, t1, t2, ..., tm, 102], where 101 and 102 are the tag IDs of [CLS] and [SEP], respectively.
[0103] Load the trained model parameters. The model's word embedding matrix W maps each tag ID to the corresponding word feature representation, resulting in an embedding matrix E of size (m+2)×d, where d is the dimension of the word feature representation.
[0104] The feature representation corresponding to the [CLS] tag is obtained from the output of the second-stage model as the sentence feature representation, or the sentence feature representation is obtained by aggregating all word feature representations (such as average pooling).
[0105] The sentence feature representation S is calculated in the following way: using the embedding of the [CLS] tag: S = E[1,:], using average pooling: S = mean(E[1:m+1,:], 1), and ignoring the [SEP] tag.
[0106] The attention weight matrix of the second-stage model is modified by the attention matrix weighted correction method, the output weight of the self-attention layer is extracted from each Transformer layer of the second-stage model, and for each term ti, the weighted sum of the self-attention weights of all layers is calculated.
[0107] Assume that the self-attention weight matrix of the l-th layer Transformer is A^l, with a size of (m+2)×
[0108] (m+2). For each term ti, the weighted sum of its self-attention weights in all layers can be expressed as: importance(ti)=sum(l=1toL,sum(A^l[:,i])), where L is the number of Transformer layers.
[0109] For all j, Softmax normalization is applied: importance_normalized(ti)=exp(importance(ti)) / sum(exp(importance(tj))).
[0110] Adjust the weight strategy and use Softmax normalization to enhance the dependencies between words. By modifying the attention weights and attention strategy of the second-stage model, the dependencies between words related to the paraphrase generation task are strengthened so that the model can more accurately focus on the words or phrases related to the paraphrase generation task, thereby better understanding the contextual meaning of the terms.
[0111] After training the second-stage model using the above method, the obtained pre-trained model can understand the semantics of the input text in combination with the context and generate the interpretation content of the input text.
[0112] Step 220, obtain the search term content and perform feature extraction on the search term content to obtain a term feature representation.
[0113] The search term content is the term input by the user for text tracing.
[0114] The search term content is input into the Word2Vec model for feature extraction. The Word2Vec model converts the search term content into a high-dimensional term feature representation user_input_ids, which captures the semantic information of the search term content. For example, if the search term content is "artificial intelligence", the Word2Vec model will output a vector representation, such as [0.1, 0.3, 0.7, ...].
[0115] Step 230: Based on the matching degree between the term feature representation and a plurality of candidate text feature representations in the feature representation library, determine at least one candidate text feature representation from the feature representation library as a first feature representation set.
[0116] Optionally, a candidate text library is obtained, where the candidate text library contains a plurality of candidate text contents.
[0117] After extracting features from multiple candidate text contents respectively, the candidate text feature representations in the feature representation library are obtained.
[0118] Among them, a candidate text content contains a term and its interpretation in the target domain. For the candidate text content in the candidate text library, after word segmentation and encoding, the numerical ID sequence input_ids is added with a special tag [CLS] at the beginning of the sequence to aggregate the semantic information at the sentence level, and the [SEP] tag is added at the end of the sequence to indicate the end of the sentence.
[0119] Pass the processed input data input_ids into the BERT model for forward propagation, and get the embedding representation of each token (word) from the model output embeddings = BERT (input_ids). Use the last layer output of the BERT model as the feature representation of the term.
[0120] First, perform average pooling: text_vector = mean(layer_output, axis = 1), and then do weighted summation on the results: text_vector = sum(weight*layer_output, axis = 1), where weight is a weight customized according to the task.
[0121] [CLS] tag: text_vector = embeddings[0,:], directly using the embedding representation of the [CLS] tag initialized during BERT model training.
[0122] For each candidate text content, you can choose to average and weight the embedding representations of all tokens to obtain the final candidate text feature representation.
[0123] Optionally, based on the similarity between the term feature representation and multiple candidate text feature representations, multiple candidate text feature representations are sorted to obtain serial numbers corresponding to the multiple candidate text feature representations, and at least one candidate text feature representation whose serial number meets preset requirements is used as the first feature representation set.
[0124] Exemplarily, a pre-trained HNSW model (Hierarchical Navigable Small World, high-speed channel model) is used to calculate multiple candidate text contents to determine the first feature representation.
[0125] The training process is as follows: obtain a corpus containing multiple text contents, including but not limited to articles, news, poems, etc., perform text preprocessing on the corpus, including removing special contents in the text content (such as punctuation, numbers, special characters, etc.), converting uppercase characters to lowercase, deleting stop words, etc., split the text content into single sentence text [u1, u2, ..., um], use the BERT model to vectorize each sentence, and obtain the sentence vector representation [s1, s2, ..., sm]. Input the vector representation [s1, s2, ..., sm] into the HNSW model for training.
[0126] Use the trained HNSW model to calculate the cosine similarity between each feature representation vj and the term feature representation si from multiple candidate text features, sort the multiple candidate text feature representations according to the cosine similarity scores, and return the highest ranked K most similar candidate text feature representations as the first feature representation set, where K is a positive integer.
[0127] Step 240, generating a term interpretation text corresponding to the search term content based on the pre-trained model, and performing feature extraction on the term interpretation text to obtain a feature representation of the interpretation.
[0128] Optionally, the term feature representation is input into a pre-trained model, and the hidden state corresponding to the term feature representation is output. The hidden state refers to the encoded representation of the search term content by the pre-trained model, which is used to reflect the pre-trained model's understanding of the search term content.
[0129] The decoder generates a term interpretation text corresponding to the search term content based on the hidden state.
[0130] Exemplarily, a pre-trained model is loaded, which can capture the deep characteristics of language.
[0131] The pre-trained model receives user_input_ids and generates a sequence of hidden states h_1, h_2, ..., h_n.
[0132] The decoder generates the next token y_t of the paraphrase at each time step t using the following formula: yt = decoder(ht-1, Attention(ht-1, H, H)). Where H is the encoder's hidden state set and ht-1 is the decoder's hidden state at the previous time step. In this way, the pre-trained model can generate paraphrases that capture the semantics of the term, thus achieving the term paraphrase generation task.
[0133] The generated definitions are processed, including removing stop words, correcting grammatical errors, etc., to obtain the final output definition text of the entry, for example, "Artificial intelligence refers to the intelligence exhibited by systems created by humans, usually achieved through machine learning and data mining."
[0134] The Word2Vec model is used to extract features from the word interpretation text to obtain the interpretation feature representation.
[0135] For example, the process splits the entry definition text into multiple words, uses the Word2Vec model to extract features of each word, and obtains a set of word feature vectors corresponding to the entry definition text.
[0136] Step 250, sorting and screening the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain the text tracing record.
[0137] The text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
[0138] Exemplarily, the interpretation feature representation is combined with the first feature representation set and then deduplication is performed to remove text content corresponding to the repeated feature representations.
[0139] The first feature representation set is S, with a dimension of d. The set contains n word feature representations, and the dimension of each word feature representation is also d. The sentence feature representation S is obtained through the output feature representation of the [CLS] tag of the last layer of the BERT model.
[0140] The word feature representation set corresponding to the term interpretation / interpretation feature representation is V = {v1, v2, ..., vn}, where each vi is the d-dimensional feature representation of word i. The combined feature representation set is C = {S∪
[0141] V}={s1,s2,...,sn,v1,v2,...,vn}, where si is the feature representation from S and vi is the feature representation from V.
[0142] For any two word feature representations vi and vj in set C, calculate the cosine similarity between the two. If the cosine similarity reaches a preset threshold, vi and vj are considered to be duplicates, and one of the feature representations is retained, and the other is removed from set C. The retained feature representation can be arbitrary.
[0143] Optionally, the interpretation feature representation and the first feature representation set are parsed based on a time extraction algorithm to obtain timestamp information corresponding to the interpretation feature representation and all candidate text feature representations in the first feature representation set, and the timestamp information is used to indicate the release time of the search term content and the candidate text feature representation.
[0144] The interpretation feature representation and the first feature representation set are sorted based on the timestamp information to obtain a sorting result.
[0145] Exemplarily, a time extraction algorithm is applied to the text content corresponding to each feature representation in set C, the extracted string time expression is parsed into a timestamp, and converted into a unified time format, such as a Unix timestamp.
[0146] The Unix timestamp is the total number of seconds from 00:00:00 (Coordinated Universal Time, UTC) on January 1, 1970 to the current time. For example, the Unix timestamp of January 1, 2020 is 1577644800, which means that 1577644800 seconds have passed between January 1, 1970 and January 1, 2020.
[0147] That is, each feature represents a text content corresponding to a Unix timestamp. The larger the value, the later the text content was published.
[0148] Optionally, based on the similarity between the interpretation feature representation and the second feature representation set, at least one feature representation is selected from the first feature representation set to obtain the second feature representation set; the text content corresponding to the second feature representation set is determined as the text content set containing the search terms; and the text tracing record is obtained based on the sorting result and the text content set.
[0149] Exemplarily, the degree of matching between the feature representation of the first sentence and the feature representations of the plurality of first contents is determined based on the cosine distance, the Pearson distance, and the BM25 correlation value.
[0150] Among them, cosine distance is a method to measure the difference between two vectors. The distance is calculated based on the cosine similarity of the vectors. Cosine similarity measures the degree of similarity between two vectors in direction. Pearson distance refers to the complement of the Pearson correlation coefficient, which is used to measure the degree of correlation between two variables. The BM25 (Best Matching 25) algorithm is a relevance scoring algorithm based on the TF-IDF (Term Frequency-Inverse Document Frequency) model, and improves it to consider factors such as document length. The BM25 algorithm evaluates the relevance between a document and a query by calculating the frequency (TF) and inverse document frequency (IDF) of the query terms in the document, and adjusting it based on the document length.
[0151] In order to obtain the historical development context and semantic relevance of at least one text content corresponding to the first feature representation set, first obtain the set S and the set V, and calculate the conversion distances such as the cosine distance, Pearson distance, and BM25 correlation value between each feature representation in the set V and each feature representation in the set S.
[0152] The obtained cosine distance, Pearson distance and BM25 correlation value are used as sorting features, and the sorting results are obtained through the refined sorting model. The sorting results are integrated into a coherent historical record to form a complete text traceability record.
[0153] For example, the search term input by the user is "artificial intelligence" and the target field is field A. The text content corresponding to the second feature representation set is 10 text contents in field A (serial numbers 0 to 9 respectively). All 10 text contents contain the search term content, and the release order is as follows: serial number 4, serial number 5, serial number 1, serial number 2, serial number 3, serial number 6, serial number 7, serial number 8, serial number 9, serial number 0.
[0154] The text traceability record is integrated in the order of publication, which can show the development process of the search term content "artificial intelligence".
[0155] In summary, the method provided by the present application adjusts the architecture of the deep learning model, uses a targeted training text data set to train the deep learning model, and obtains a pre-trained model that can generate text interpretation content. The pre-trained model generates interpretations for the search term content input by the user, and based on the candidate texts that meet the matching requirements of the semantic similarity with the search content and the interpretation content, it can understand the search term content from a natural language perspective, improve the accuracy of text tracing, effectively track and trace the source of the search term content, and improve the efficiency of text tracing. Compared with the method of using document tagging to annotate and continuously maintain massive text data in the related art, this solution does not need to mark and maintain the text data in the library, which can save computing resources and maintenance time.
[0156] Figure 3 is a structural block diagram of a text tracing device provided by an exemplary embodiment of the present application, such as Figure 3 As shown, the device includes the following parts.
[0157] A training module 310 is used to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries;
[0158] A feature extraction module 320 is used to obtain search term content and perform feature extraction on the search term content to obtain a term feature representation, wherein the search term content is a term input by a user for text tracing;
[0159] A matching module 330, configured to determine at least one candidate text feature representation from the feature representation library as a first feature representation set based on a degree of matching between the term feature representation and a plurality of candidate text feature representations in the feature representation library;
[0160] The interpretation generation module 340 is used to generate an interpretation text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the interpretation text to obtain an interpretation feature representation;
[0161] The tracing module 350 is used to sort and filter the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
[0162] In an optional embodiment, the training module 310 is also used to obtain the training text data set; train the deep learning model based on the multiple sentences in the training text data set to obtain a first-stage model; adjust the model architecture and model parameters of the first-stage model to obtain a second-stage model; train the second-stage model based on the multiple entries annotated with explanations in the training text data set to obtain the pre-trained model.
[0163] In an optional embodiment, the training module 310 is further used to obtain multiple corpus sentences and multiple entries marked with interpretations, wherein the multiple corpus sentences are sentence contents obtained by splitting multiple text contents; synonym replacement is performed on the multiple corpus sentences to obtain a first sentence set, wherein the i-th sentence in the first sentence set is obtained by synonym replacement of each term in the i-th corpus sentence in the multiple corpus sentences; the multiple corpus sentences are screened based on the word frequency and inverse document frequency method to obtain a second sentence set, wherein the word frequency refers to the frequency of each term appearing in the multiple corpus sentences, and the inverse document frequency is used to indicate the importance of each term in the multiple corpus sentences; the first sentence set and the second sentence set are determined as the multiple sentences.
[0164] In an optional embodiment, the training module 310 is also used to add a custom feature extraction layer to the first-stage model to obtain a first-stage model after adjusting the model architecture; based on preset parameter standards, the learning rate, number of training rounds and batch size of the first-stage model after adjusting the model architecture are adjusted to obtain the second-stage model, wherein the learning rate is used to determine the step size of the parameter update of the model during the training process, the number of training rounds refers to the number of times the training text data set is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.
[0165] In an optional embodiment, the matching module 330 is also used to obtain a candidate text library, which contains multiple candidate text contents; after performing feature extraction on the multiple candidate text contents respectively, the candidate text feature representations in the feature representation library are obtained; based on the similarity between the term feature representation and the multiple candidate text feature representations, the multiple candidate text feature representations are sorted to obtain the serial numbers corresponding to the multiple candidate text feature representations respectively; and the at least one candidate text feature representation whose serial number meets the preset requirements is used as the first feature representation set.
[0166] In an optional embodiment, the interpretation generation module 340 is further used to input the term feature representation into the pre-trained model, and output a hidden state corresponding to the term feature representation, wherein the hidden state refers to the encoded representation of the search term content by the pre-trained model, and is used to reflect the pre-trained model's understanding of the search term content; and generate a term interpretation text corresponding to the search term content based on the hidden state through a decoder.
[0167] In an optional embodiment, the tracing module 350 is used to parse the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain timestamp information corresponding to the interpretation feature representation and all candidate text feature representations in the first feature representation set, and the timestamp information is used to indicate the release time of the search term content and the candidate text feature representation; sort the interpretation feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; filter at least one feature representation from the first feature representation set based on the similarity between the interpretation feature representation and the second feature representation set to obtain a second feature representation set; determine the text content corresponding to the second feature representation set as the text content set containing the search term; and obtain the text tracing record based on the sorting result and the text content set.
[0168] In summary, the text tracing device provided by the present application adjusts the architecture of the deep learning model, uses a targeted training text data set to train the deep learning model, and obtains a pre-trained model that can generate text interpretation content. The pre-trained model generates interpretations for the search term content input by the user, and based on the candidate texts that meet the matching requirements of the semantic similarity with the search content and the interpretation content, it can understand the search term content from a natural language perspective, improve the accuracy of text tracing, effectively track and trace the source of the search term content, and improve the efficiency of text tracing. Compared with the method of using document tagging to annotate and continuously maintain massive text data in the related art, this solution does not need to mark and maintain the text data in the library, which can save computing resources and maintenance time.
[0169] It should be noted that the text tracing device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the text tracing device provided in the above embodiment belongs to the same concept as the text tracing method embodiment. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0170] Figure 4 The block diagram of the structure of a computer device 400 provided by an exemplary embodiment of the present application is shown. The computer device 400 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Compression Standard Audio Layer 3), an MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) player, a laptop computer or a desktop computer. The computer device 400 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0171] Typically, the computer device 400 includes a processor 401 and a memory 402 .
[0172] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0173] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, which is used to be executed by the processor 401 to implement the text tracing method provided in the method embodiment of the present application.
[0174] In some embodiments, the computer device 400 further includes some other components 403, and the type and quantity of the other components 403 can be selected based on the functional requirements of the computer device 400. Those skilled in the art will appreciate that Figure 4 The structure shown in the figure does not constitute a limitation on the computer device 400, and the computer device 400 may include more or less components than those shown in the figure, or combine some components, or adopt a different arrangement of components.
[0175] Optionally, the computer readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid state drive (SSD), or an optical disk. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0176] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a text tracing method as described in any of the above embodiments of the present application.
[0177] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a text tracing method as described in any of the above embodiments of the present application.
[0178] The present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the text tracing method described in any of the above embodiments.
[0179] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0180] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A text tracing method, characterized in that: The method comprises: Acquire a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries; Acquire search term content and perform feature extraction on the search term content to obtain term feature representation, wherein the search term content is a term input by a user for text tracing; Based on the matching degree between the term feature representation and a plurality of candidate text feature representations in the feature representation library, determining at least one candidate text feature representation from the feature representation library as a first feature representation set; Generate a term interpretation text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the term interpretation text to obtain a feature representation of the interpretation; Based on a time extraction algorithm, the interpretation feature representation and the first feature representation set are sorted and filtered to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
2. The method according to claim 1, characterized in that The step of obtaining a training text data set and training a deep learning model based on the training text data set to obtain a pre-trained model includes: Obtaining the training text dataset; Training the deep learning model based on the multiple sentences in the training text dataset to obtain a first-stage model; Adjusting the model architecture and model parameters of the first-stage model to obtain a second-stage model; The second-stage model is trained based on the multiple entries annotated with definitions in the training text data set to obtain the pre-trained model.
3. The method according to claim 2, characterized in that The step of obtaining the training text data set includes: Acquire a plurality of corpus sentences and the plurality of entries marked with definitions, wherein the plurality of corpus sentences are sentence contents obtained by splitting a plurality of text contents; Performing synonym replacement on the multiple corpus sentences to obtain a first sentence set, wherein the i-th sentence in the first sentence set is obtained by performing synonym replacement on each word item of the i-th corpus sentence in the multiple corpus sentences; The plurality of corpus sentences are screened based on a word frequency and an inverse document frequency method to obtain a second sentence set, wherein the word frequency refers to the frequency at which each word item appears in the plurality of corpus sentences, and the inverse document frequency is used to indicate the importance of each word item in the plurality of corpus sentences; The first sentence set and the second sentence set are determined as the plurality of sentences.
4. The method according to claim 2, characterized in that: The adjusting of the model architecture and model parameters of the first-stage model to obtain the second-stage model includes: Adding a custom feature extraction layer to the first-stage model to obtain a first-stage model after adjusting the model architecture; The learning rate, number of training rounds and batch size of the first-stage model after adjusting the model architecture are adjusted based on preset parameter standards to obtain the second-stage model, wherein the learning rate is used to determine the step size of parameter update of the model during training, the number of training rounds refers to the number of times the training text data set is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.
5. The method according to any one of claims 1 to 4, characterized in that: The step of determining at least one candidate text feature representation from the feature representation library as a first feature representation set based on the matching degree between the feature representation of the term and a plurality of candidate text feature representations in the feature representation library comprises: Acquire a candidate text library, wherein the candidate text library contains a plurality of candidate text contents; After extracting features from the plurality of candidate text contents respectively, obtaining the candidate text feature representations in the feature representation library; Based on the similarity between the term feature representation and the plurality of candidate text feature representations, the plurality of candidate text feature representations are sorted to obtain sequence numbers corresponding to the plurality of candidate text feature representations respectively; The at least one candidate text feature representation whose sequence number meets the preset requirement is used as the first feature representation set.
6. The method according to any one of claims 1 to 4, characterized in that: The generating of the term interpretation text corresponding to the search term content based on the pre-trained model includes: Inputting the term feature representation into the pre-trained model, and outputting a hidden state corresponding to the term feature representation, wherein the hidden state refers to the encoding representation of the search term content by the pre-trained model, and is used to reflect the understanding of the search term content by the pre-trained model; A decoder is used to generate a term interpretation text corresponding to the search term content based on the hidden state.
7. The method according to any one of claims 1 to 4, characterized in that: The method of sorting and filtering the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain a text tracing record includes: Parsing the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain timestamp information corresponding to the interpretation feature representation and all candidate text feature representations in the first feature representation set, wherein the timestamp information is used to indicate the release time of the search term content and the candidate text feature representation; sorting the interpretation feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; Filtering at least one feature representation from the first feature representation set based on the similarity between the interpretation feature representation and the second feature representation set to obtain a second feature representation set; Determining the text content corresponding to the second feature representation set as the text content set containing the search term; The text tracing record is obtained based on the sorting result and the text content set.
8. A text tracing device, characterized in that: The device comprises: A training module, used to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries; A feature extraction module is used to obtain search term content and perform feature extraction on the search term content to obtain a term feature representation, wherein the search term content is a term input by a user for text tracing; A matching module, configured to determine at least one candidate text feature representation from the feature representation library as a first feature representation set based on a degree of matching between the feature representation of the term and a plurality of candidate text feature representations in the feature representation library; A definition generation module, used to generate a definition text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the definition text to obtain a definition feature representation; A tracing module is used to sort and filter the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.
9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the text tracing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the text tracing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Text enhancement method and system based on word interpretation
CN113591469A
Text sorting method and device
CN113987161A
Text model training method, device and equipment
CN116150621A
Cloud-based payroll management system
KR1020250014209A