Text Tracing Method, Device, Equipment, Medium

Through the deep learning model generation and the method of combining time extraction algorithms, the problem of low text traceability efficiency in the existing technology is solved, efficient and accurate text traceability is achieved, and resource consumption is reduced.

CN120011480BActive Publication Date: 2025-06-24RICHFIT INFORMATION TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411287140.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-06-24
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

When processing large-scale text data, the prior art requires a lot of resources and time to trace text through document marking method, and is relatively inefficient.

Method used

A deep learning model is adopted to generate a pre-trained model by training the text data set, which is used to generate entry definitions, and a time extraction algorithm is used to sort and filter the descriptive feature representation and candidate text feature representation to realize text traceability.

Benefits of technology

It improves the efficiency and accuracy of text traceability, reduces resource consumption, and avoids the tagging and maintenance operations of massive text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011480B_ABST
    Figure CN120011480B_ABST
Patent Text Reader

Abstract

The present application discloses a text traceability method, device, equipment, and medium. The present application relates to the field of text processing. The method includes the following steps: obtaining a training text data set and training a deep learning model based on the training text data set to obtain a pre-trained model; obtaining the content of a search term and extracting features from the content of the search term to obtain a term feature representation; determining at least one candidate text feature representation as a first feature representation set from the feature representation library based on the matching degree between the term feature representation and multiple candidate text feature representations in the feature representation library; generating a term interpretation text corresponding to the content of the search term based on the pre-trained model, and extracting features from the term interpretation text to obtain an interpretation feature representation; sorting and screening the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text traceability record. It can improve the efficiency during text traceability and reduce resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of text processing, and particularly to a text traceability method, apparatus, device, and medium. Background Art

[0002] In the big data era, the demand for tracing text information is increasing day by day. Accurately tracing the source of text can ensure the reliability and authenticity of information, and help to resolve disputes over information sources, prevent the spread of false information, etc.

[0003] In related technologies, term traceability is achieved by making specific marks on documents. When processing a large-scale data set, a unique mark (such as a mark in the form of coding, symbol, etc.) is assigned to each document. When term traceability is required, the document containing the term can be quickly located by searching for the mark related to the term.

[0004] However, the document marking method requires preprocessing and continuous maintenance of documents, consuming resources and time, and the traceability process also requires a large amount of computing resources, resulting in low efficiency when tracing a large amount of text. Summary of the Invention

[0005] Embodiments of the present application provide a text traceability method, apparatus, device, and medium, which can improve the efficiency of text traceability and reduce resource consumption. The technical solutions are as follows:

[0006] On the one hand, a text traceability method is provided, and the method includes:

[0007] Obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model. The training text data set includes multiple sentences and multiple terms annotated with interpretations, and the pre-trained model is used to generate term interpretation texts;

[0008] Obtain the content of a search term and perform feature extraction on the content of the search term to obtain a term feature representation. The content of the search term is the term input by the user for text traceability;

[0009] Based on the matching degree between the term feature representation and multiple candidate text feature representations in the feature representation library, determine at least one candidate text feature representation from the feature representation library as the first feature representation set;

[0010] Generate a term interpretation text corresponding to the content of the search term based on the pre-trained model, and perform feature extraction on the term interpretation text to obtain an interpretation feature representation;

[0011] Based on a time extraction algorithm, sort and filter the paraphrase feature representation and the first feature representation set to obtain a text traceability record, which is used to indicate a set of text contents containing the content of the search term, and the text traceability record contains the text contents corresponding to the filtered feature representations.

[0012] On the other hand, a text traceability device is provided, and the device includes:

[0013] A training module, configured to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model. The training text data set includes multiple sentences and multiple terms annotated with paraphrases, and the pre-trained model is used to generate term paraphrases;

[0014] A feature extraction module, configured to obtain the content of a search term and perform feature extraction on the content of the search term to obtain a term feature representation, where the content of the search term is a term input by a user for text traceability;

[0015] A matching module, configured to determine at least one candidate text feature representation as a first feature representation set from the feature representation library based on the matching degree between the term feature representation and multiple candidate text feature representations in the feature representation library;

[0016] A paraphrase generation module, configured to generate a term paraphrase text corresponding to the content of the search term based on the pre-trained model, and perform feature extraction on the term paraphrase text to obtain a paraphrase feature representation;

[0017] A traceability module, configured to sort and filter the paraphrase feature representation and the first feature representation set based on a time extraction algorithm to obtain a text traceability record, which is used to indicate a set of text contents containing the content of the search term, and the text traceability record contains the text contents corresponding to the filtered feature representations.

[0018] In an optional embodiment, the training module is further configured to obtain the training text data set; train the deep learning model based on the multiple sentences in the training text data set to obtain a first-stage model; adjust the model architecture and model parameters of the first-stage model to obtain a second-stage model; and train the second-stage model based on the multiple terms annotated with paraphrases in the training text data set to obtain the pre-trained model.

[0019] In an optional embodiment, the training module is further configured to obtain a plurality of corpus sentences and the plurality of entries annotated with paraphrases, where the plurality of corpus sentences are sentence contents obtained by splitting a plurality of text contents; perform synonym replacement on the plurality of corpus sentences to obtain a first set of sentences, where the i-th sentence in the first set of sentences is obtained by performing synonym replacement on each term in the i-th corpus sentence among the plurality of corpus sentences; screen the plurality of corpus sentences based on the term frequency-inverse document frequency method to obtain a second set of sentences, where the term frequency refers to the frequency of occurrence of each term in the plurality of corpus sentences, and the inverse document frequency is used to indicate the importance of each term in the plurality of corpus sentences; and determine the first set of sentences and the second set of sentences as the plurality of sentences.

[0020] In an optional embodiment, the training module is further configured to add a custom feature extraction layer to the first-stage model to obtain the first-stage model with an adjusted model architecture; adjust the learning rate, the number of training epochs, and the batch size of the first-stage model with the adjusted model architecture based on a preset parameter standard to obtain the second-stage model, where the learning rate is used to determine the step size of parameter update during model training, the number of training epochs refers to the number of times the training text dataset is completely trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.

[0021] In an optional embodiment, the matching module is further configured to obtain a candidate text library, where the candidate text library contains a plurality of candidate text contents; perform feature extraction on the plurality of candidate text contents respectively to obtain the candidate text feature representations in the feature representation library; sort the plurality of candidate text feature representations based on the similarity between the entry feature representation and the plurality of candidate text feature representations to obtain the serial numbers corresponding to the plurality of candidate text feature representations respectively; and use at least one candidate text feature representation whose serial number meets the preset requirements as the first feature representation set.

[0022] In an optional embodiment, the paraphrase generation module is further configured to input the entry feature representation into the pre-trained model to output the hidden state corresponding to the entry feature representation, where the hidden state is the encoded representation of the search entry content by the pre-trained model and is used to reflect the understanding of the search entry content by the pre-trained model; and generate an entry paraphrase text corresponding to the search entry content through a decoder based on the hidden state.

[0023] In an optional embodiment, the traceability module is configured to parse the paraphrase feature representation and the first feature representation set based on the time extraction algorithm to obtain timestamp information corresponding to all candidate text feature representations in the paraphrase feature representation and the first feature representation set, where the timestamp information is used to indicate the release times of the search term content and the candidate text feature representations; sort the paraphrase feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; screen at least one feature representation from the first feature representation set based on the similarity between the paraphrase feature representation and the first feature representation set to obtain a second feature representation set; determine the text content corresponding to the second feature representation set as the text content set including the search term; and obtain the text traceability record based on the sorting result and the text content set.

[0024] On the other hand, a computer device is provided, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text traceability method as described in any one of the embodiments of the present application above.

[0025] On the other hand, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the text traceability method as described in any one of the embodiments of the present application above.

[0026] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text traceability method as described in any one of the above embodiments.

[0027] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0028] By adjusting the architecture of the deep learning model and training the deep learning model using a targeted training text dataset, a pre-trained model capable of generating text paraphrase content is obtained. The pre-trained model generates a paraphrase for the search term content input by the user. Based on the candidate text that meets the matching requirements for semantic similarity with the search content and the paraphrase content, it can understand the search term content from the perspective of natural language, improve the accuracy during text traceability, effectively track and trace the source of the search term content, and improve the efficiency during text traceability. Compared with the method in the related technology of using document marking method to annotate and continuously maintain a large amount of text data, this solution does not require marking and maintenance operations on the text data in the library, and can save computing resources and maintenance time. Brief Description of the Drawings

[0029] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0030] Figure 1 is a schematic diagram of a text traceability system provided by an exemplary embodiment of the present application;

[0031] Figure 2 is a flowchart of a text traceability method provided by an exemplary embodiment of the present application;

[0032] Figure 3 is a structural block diagram of a text traceability device provided by an exemplary embodiment of the present application;

[0033] Figure 4 is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed Description of the Embodiments

[0034] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the drawings.

[0035] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0036] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0037] It should be noted that the information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0038] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0039] First, a brief introduction to the nouns involved in the embodiments of this application:

[0040] Word2Vec model (Word to Vector): The main purpose is to generate word vectors, that is, to map words into a vector space. Each word is represented as a vector of a fixed dimension, and these vectors can capture the semantic relationships between words.

[0041] In this application, the feature representation refers to a vector, and the process of feature extraction from text to obtain the feature representation can be regarded as a process of converting text into a vector.

[0042] BERT model (Bidirectional Encoder Representations from Transformers): A pre-trained language model based on the Transformer architecture that can also convert text into vector form. Compared with the Word2Vector model, the BERT model generates context-related word vectors. It can consider the context information on both the left and right sides of the word simultaneously through multiple layers of self-attention mechanisms to generate deep bidirectional representations.

[0043] Among them, the deep learning model used in this application to generate the definition of a term refers to the BERT model. A pre-trained model is obtained by training and fine-tuning the BERT model. This pre-trained model has the functions of the BERT model and can be applied to the task of generating text definitions.

[0044] Time extraction algorithm: An algorithm used to extract time information from natural language text, which combines natural language processing techniques and specific rule patterns to identify time expressions in the text. It uses lexical features, syntactic structures, and context information to accurately locate and extract time information. For example, this algorithm can identify explicit date expressions such as "August 29, 2024", relative time expressions such as "yesterday", "next week", "next month", etc., and vague time range expressions such as "recently", "for some time", etc. In text data such as news reports and social media, the time extraction algorithm can help analyze the occurrence time of events, construct the timeline of events, and thus better understand the development context and evolution process of events.

[0045] TF-IDF (Term Frequency-Inverse Document Frequency): In text processing, TF-IDF can be used to extract important keywords in a document. Term Frequency (TF) represents the frequency of a word in a document. Words with a high frequency of occurrence may be the key content of the document. Inverse Document Frequency (IDF) measures the rarity of a word in the entire document collection. If a word appears in many documents, its IDF value is low; if a word appears in only a few documents, its IDF value is high. Through the TF-IDF method, words with a high frequency of occurrence in a specific document and relatively rare in the entire document collection can be highlighted, and these words are regarded as the key information of the document.

[0046] In the big data era, with the development of information technology, data has increased exponentially, and the demand for tracing text information has also grown day by day. Accurately tracing the source of text can ensure the reliability and authenticity of information, effectively solve the controversy over information sources, and prevent the widespread spread of false information, etc.

[0047] In related technologies, the tracing operation of terms is achieved by making specific marks on documents. When dealing with a large-scale dataset, a unique mark is assigned to each document in the dataset. The mark can be in the form of a code or other forms such as symbols. When it is necessary to trace a certain term, the relevant mark associated with this term can be searched, and thus the document containing this term can be quickly located.

[0048] However, when using the document tagging method to trace the source of text, it is necessary to preprocess the document, and during subsequent use, it is also necessary to continuously maintain the document, consuming resources and time. The tracing process consumes a large amount of computing resources and has relatively low efficiency.

[0049] Figure 1 FIG. 4 is a schematic diagram of a text tracing system provided by an exemplary embodiment of the present application, which can accurately and efficiently implement text tracing.

[0050] The text tracing system 100 involves a user terminal 110 and a server 120, and the user terminal 110 and the server 120 are connected through a communication network 130.

[0051] Among them, the user terminal 110 inputs the search term content and sends it to the server 120, and the server 120 analyzes the search term content.

[0052] The server 120 generates an interpretation of the search term content and filters out candidate text content with a relatively high semantic similarity to the search term content in the candidate text library. By performing time extraction on the interpretation and the candidate text content, the release time of each content is determined and sorted to obtain a complete text tracing record, and the text tracing record is fed back to the user terminal 110.

[0053] Among them, a pre-trained model is deployed inside the server 120. The pre-trained model is a model obtained by adjusting the architecture of the BERT model and training it with the sentences and term content specifically set in the training text dataset. The pre-trained model can perform semantic analysis on the search term content input by the user terminal 110 and generate the term interpretation of the search content term.

[0054] Before the user terminal 110 inputs the search term content to the server 120, a candidate text library is prepared in advance. The library contains multiple candidate text contents, and the candidate text contents include text contents in the forms of news, articles, reports, etc., and belong to the same field as the search term content.

[0055] Feature extraction is performed on multiple candidate content texts to obtain candidate text feature representations corresponding to each candidate text content respectively.

[0056] After feature extraction is performed on the search term content input by the user terminal 110, a term feature representation is obtained. Based on the term feature representation and multiple candidate text feature representations, similarity matching is performed, and the candidate text feature representation with a relatively high similarity to the term feature representation is determined as the first feature representation set.

[0057] After obtaining the entry definition of the search entry content, feature extraction is performed on the entry definition to obtain a definition feature representation. Based on a time extraction algorithm, the text contents corresponding to the definition feature representation and the first feature representation set are analyzed respectively to determine the release times of these text contents, and they are sorted to obtain the sorting result of the definition feature representation and the first feature representation set.

[0058] Perform a similarity analysis between all the feature representations in the definition feature representation and the first feature representation set to determine the semantic correlation degree between the above-mentioned feature representations, and filter the feature representations with higher similarity as the second feature representation set. Among them, the semantic correlation degree between the feature representations in the second feature representation set is high.

[0059] Re-sort the text contents corresponding to each feature representation in the second feature representation set according to the sorting result to obtain a text traceability record, which can indicate the text content containing the search entry content and the development context of the text content.

[0060] The above-mentioned client 110 can be various forms of terminal devices such as mobile phones, tablet computers, desktop computers, portable laptops, smart TVs, vehicle-mounted terminals, and smart home devices, and the embodiments of the present application do not limit this.

[0061] It should be noted that the above-mentioned server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0062] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system support, which can only be achieved through cloud computing.

[0063] In some embodiments, the above server can also be implemented as a node in a blockchain system.

[0064] Combined with the above noun introduction and application scenarios, the text traceability method provided in this application is described. This method can be executed by a server or a terminal, or jointly executed by a server and a terminal. In the embodiments of this application, taking the method being executed by the server as an example for description, as Figure 2 shown, Figure 2 is a flowchart of the text traceability method provided by an exemplary embodiment of this application. The method includes the following steps.

[0065] Step 210, obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model.

[0066] Among them, the training text data set contains multiple sentences and multiple entries annotated with paraphrases, and the pre-trained model is used to generate entry paraphrases.

[0067] The deep learning model uses a BERT model, and a pre-trained model is obtained after fine-tuning and training the BERT model.

[0068] The content in the training text data set can be in the same field as the content of the search entry input by the user. For example, if the content of the search entry input by the user belongs to the terms in the field of artificial intelligence, the content in the training text data set can also be the sentence content obtained by splitting the articles in this field, as well as the content obtained by annotating and paraphrasing multiple entries in this field.

[0069] Optionally, obtain a training text data set, and train a deep learning model based on multiple sentences in the training text data set to obtain a first-stage model.

[0070] Exemplarily, obtain multiple corpus sentences and multiple entries annotated with paraphrases, and the multiple corpus sentences are the sentence content obtained by splitting multiple text contents.

[0071] Determine the field to which the search entry content belongs as the target field, and obtain multiple text contents in the target field, including but not limited to professional literature, historical documents, academic papers, etc.

[0072] Remove redundant information such as reference lists, irrelevant links, annotations, and format errors in the multiple text contents, and perform at least one of the following operations on them: remove stop words, remove content irrelevant to the target field, mislabeling, duplicate entries, etc.

[0073] The sentence content obtained by splitting the multiple text contents is as follows: [u1, u2,..., um], where m is a positive integer, and um refers to a corpus sentence.

[0074] Perform synonym replacement on multiple corpus sentences to obtain a first set of sentences. Among them, the i-th sentence in the first set of sentences is obtained by performing synonym replacement on each term in the i-th corpus sentence among the multiple corpus sentences.

[0075] That is, for each word / term wi in each corpus sentence ui (i is a positive integer), use the Word2Vec model to calculate its feature representation: vi = Word2Vec(word = wi).

[0076] For a synonym wj (j is a positive integer) of the word wi, use the Word2Vec model to calculate its feature representation vj: vj = Word2Vec(word = wj).

[0077] Calculate the cosine similarity between the word feature representation and the synonym feature representation, and select the vj with the highest similarity to vi for replacement to obtain the first set of sentences.

[0078] Based on the term frequency and inverse document frequency method, screen multiple corpus sentences to obtain a second set of sentences. Among them, the term frequency refers to the frequency of each term in multiple corpus sentences, and the inverse document frequency is used to indicate the importance of each term in multiple corpus sentences.

[0079] Exemplarily, input multiple corpus sentences: [u1, u2,..., um]. For each word wi, calculate its vector representation vi, vi = Word2Vec(word = wi).

[0080] For the word wi, calculate its TF value in the corpus sentence: TF(wi) = number of occurrences / total number of words in the sentence. Calculate the IDF value of the word wi in multiple corpus sentences: IDF(wi) = log(total number of corpus sentences / number of corpus sentences containing this word), and construct a TF-IDF matrix: TF-IDF(wi) = TF(wi) * IDF(wi).

[0081] For each corpus sentence, multiply the TF-IDF value of each word in it by the corresponding word embedding vector: TF-IDF(wi, d) * vi. Add or average the weighted word embedding vectors of all words in the corpus sentence to obtain the vector representation of the corpus sentence: doc_vector = sum(TF-IDF(wi, d) * vi) / (total number of terms in corpus sentence d).

[0082] According to the TF-IDF weighted score of the corpus sentence, determine the probability that the corpus sentence is sampled: sampling_probability(d) = sum(TF-IDF(wi, d)).

[0083] Using the method of weighted random sampling, sample from the training corpus to obtain the vector representations of the second set of sentences [t1, t2, ..., tm].

[0084] Determine the first set of sentences and the second set of sentences as multiple sentences.

[0085] Optionally, adjust the model architecture and model parameters of the first-stage model to obtain the second-stage model.

[0086] Exemplarily, add a custom feature extraction layer to the first-stage model to obtain the first-stage model with the adjusted model architecture.

[0087] Among them, the custom feature extraction layer can be a fully connected layer.

[0088] The input vectors for each Transformer (feature transformation) layer are [X1, X2, …, Xn], and the output vectors are [Y1, Y2, …, Yn].

[0089] After adding the custom feature extraction layer, the new input vectors are [X1, X2, …, Xn, Z1], and the new output vectors are [Y1, Y2, …, Yn, Z2].

[0090] The weight of the custom feature extraction layer is W_custom, and the input vector is x. Then the output vector y is calculated as: y = σ (Wcustom · x), where σ represents the activation function, such as the ReLU function (Rectified Linear Unit, ReLU, rectified linear unit) or the Sigmoid function (logical function).

[0091] The output vectors of the last layer of the original BERT model are h0, h1, ..., hn. By adding a layer of fully connected layers, the output vectors can be changed to f(h0), f(h1), ..., f(hn), where f represents the transformation function of the fully connected layer.

[0092] Based on the preset parameter criteria, adjust the learning rate, number of training epochs, and batch size of the first-stage model with the adjusted model architecture to obtain the second-stage model.

[0093] Among them, the learning rate is used to determine the step size of parameter updates during model training, the number of training epochs refers to the number of times the training text dataset is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.

[0094] For example, during the update process of gradient descent, adjust the learning rate based on the following formula: Wnew = Wold - α · ▽L (Wold), where α is the learning rate.

[0095] When adjusting the batch size, during the model training process, the number of samples used for each weight update is the batch size.

[0096] Adjust the number of training epochs, which is the number of iterations of the model over the entire training dataset. It usually involves multiple passes through the entire dataset for parameter updates.

[0097] Optionally, train the second-stage model based on multiple terms with paraphrases annotated in the training text dataset to obtain a pre-trained model.

[0098] Select a small number of keyword terms in the target domain (e.g., terms with an occurrence frequency reaching a preset frequency threshold as keyword terms) and write paraphrases for them. Use these terms and paraphrases as multiple terms with paraphrases annotated in the training text dataset, hereinafter simply referred to as the fine-tuning corpus.

[0099] Exemplarily, use the fine-tuning corpus to fine-tune the second-stage model. By modifying the attention weights and strategies of the model, enable the model to understand the context meaning of the terms and generate accurate and concise explanations. At the same time, make the model have a certain generalization ability on unseen data, so that it can automatically generate term paraphrases for new terms.

[0100] For each sentence or sentence fragment in the fine-tuning corpus, receive the sentence or sentence fragment [t1, t2,..., tm] as input. Tokenize each item ti (word) and convert it into a token sequence that the model can understand. Add a special token [CLS] at the beginning of the sequence to aggregate sentence-level semantic information. Add the [SEP] token at the end of the sequence to indicate the end of the sentence.

[0101] Assume that the tokenization result of the input sequence is [101, t1, t2,..., tm, 102], where 101 and 102 are the token IDs of [CLS] and [SEP] respectively.

[0102] Load the trained model parameters. The word embedding matrix W of the model maps each token ID to the corresponding word feature representation to obtain the embedding matrix E, with a size of (m + 2) × d, where d is the dimension of the word feature representation.

[0103] Obtain the feature representation corresponding to the [CLS] token from the output of the second-stage model as the sentence feature representation, or obtain the sentence feature representation by aggregating all word feature representations (such as average pooling).

[0104] Calculate the sentence feature representation S in the following way: Use the embedding of the [CLS] token: S = E[1, :]; use average pooling: S = mean(E[1:m + 1, :], 1), ignoring the [SEP] token.

[0105] Modify the attention weight matrix of the second-stage model through a correction method weighted by the attention matrix. Extract the output weights of the self-attention layer from each Transformer layer of the second-stage model. For each term ti, calculate the weighted sum of the self-attention weights of all layers.

[0106] Assume that the self-attention weight matrix of the l-th layer Transformer is A^l, with a size of (m + 2) × (m + 2). For each term ti, the weighted sum of its self-attention weights across all layers can be expressed as: importance(ti) = sum(l = 1 to L, sum(A^l[:, i])), where L is the number of Transformer layers.

[0107] For all j, apply Softmax normalization: importance_normalized(ti) = exp(importance(ti)) / sum(exp(importance(tj))).

[0108] Adjust the weight strategy and use Softmax normalization to enhance the dependencies between words. By modifying the attention weights and attention strategy of the second-stage model, strengthen the dependencies between words related to the paraphrase generation task, so that the model can more accurately focus on words or phrases related to the paraphrase generation task, and thus better understand the context meaning of the entry.

[0109] After training the second-stage model in the above manner, the obtained pre-trained model can understand the semantics of the input text in combination with the context and generate the paraphrase content of the input text.

[0110] Step 220, obtain the search entry content and perform feature extraction on the search entry content to obtain the entry feature representation.

[0111] Among them, the search entry content is the entry input by the user for text tracing.

[0112] Input the search entry content into the Word2Vec model for feature extraction. The Word2Vec model converts the search entry content into a high-dimensional entry feature representation user_input_ids, which captures the semantic information of the search entry content. For example, if the search entry content is "artificial intelligence", the Word2Vec model will output a vector representation, such as [0.1, 0.3, 0.7,...].

[0113] Step 230, based on the matching degree between the entry feature representation and multiple candidate text feature representations in the feature representation library, determine at least one candidate text feature representation from the feature representation library as the first feature representation set.

[0114] Optionally, obtain a candidate text library, which contains multiple candidate text contents.

[0115] After performing feature extraction on multiple candidate text contents respectively, obtain the candidate text feature representations in the feature representation library.

[0116] Among them, one candidate text content contains an entry in the target domain and its paraphrase. For the candidate text contents in the candidate text library, the numerical ID sequence input_ids after word segmentation and encoding, add a special marker [CLS] at the beginning of the sequence to aggregate sentence-level semantic information, and add the [SEP] marker at the end of the sequence to indicate the end of the sentence.

[0117] Pass the processed input data input_ids into the BERT model for forward propagation, and obtain the embedding representation embeddings = BERT(input_ids) of each token (word) from the model output. Use the output of the last layer of the BERT model as the feature representation of the entry.

[0118] First, perform average pooling: text_vector = mean(layer_output, axis = 1), and then perform weighted summation on the result: text_vector = sum(weight * layer_output, axis = 1), where weight is a weight customized according to the task.

[0119] [CLS] marker: text_vector = embeddings[0, :], directly use the embedding representation of the [CLS] marker initialized during the training of the BERT model.

[0120] For each candidate text content, you can choose to average and perform weighted summation on the embedding representations of all tokens to obtain the final candidate text feature representation.

[0121] Optionally, based on the similarity between the entry feature representation and multiple candidate text feature representations, sort the multiple candidate text feature representations to obtain the serial numbers corresponding to the multiple candidate text feature representations respectively, and use at least one candidate text feature representation that meets the preset requirements as the first feature representation set.

[0122] Exemplarily, use the pre-trained HNSW model (Hierarchical Navigable Small World model) to calculate multiple candidate text contents to determine the first feature representation.

[0123] The training process is as follows: Obtain a corpus, which contains multiple segments of text content, including but not limited to articles, news, poems, etc. Perform text preprocessing on the corpus, including removing special content (such as punctuation marks, numbers, special characters, etc.) from the text content, converting uppercase characters to lowercase, deleting stop words, etc. Split the text content into single-sentence texts [u1, u2, …, um]. Use the BERT model to vectorize each sentence to obtain the vector representation of the sentence [s1, s2, ..., sm]. Input the vector representation [s1, s2, ..., sm] into the HNSW model for training.

[0124] Use the trained HNSW model to calculate the cosine similarity between each feature representation vj of multiple candidate text features and the entry feature representation si. Sort the multiple candidate text feature representations according to the scores of the cosine similarity, and return the top K most similar candidate text feature representations with the highest ranking as the first feature representation set, where K is a positive integer.

[0125] Step 240: Generate an entry interpretation text corresponding to the search entry content based on a pre-trained model, and perform feature extraction on the entry interpretation text to obtain an interpretation feature representation.

[0126] Optionally, input the entry feature representation into the pre-trained model, and output the hidden state corresponding to the entry feature representation. The hidden state refers to the encoded representation of the search entry content by the pre-trained model, which is used to reflect the understanding of the search entry content by the pre-trained model.

[0127] Generate an entry interpretation text corresponding to the search entry content through a decoder based on the hidden state.

[0128] Exemplarily, load a pre-trained model, which can capture deep features of language.

[0129] The pre-trained model receives user_input_ids and generates a series of hidden states h_1, h_2, ..., h_n.

[0130] The decoder uses the following formula to generate the next token y_t of the interpretation at each time step t: yt = decoder(ht-1, Attention(ht-1, H, H)). Where H is the set of hidden states of the encoder, and ht-1 is the hidden state of the decoder at the previous time step. In this way, the pre-trained model can generate an interpretation that captures the semantics of the entry, achieving the task of entry interpretation generation.

[0131] Process the generated paraphrases, including removing stop words, correcting grammar errors, etc., to obtain the final output of the glossary paraphrase text. For example, "Artificial intelligence refers to the intelligence demonstrated by systems created by humans, usually achieved through machine learning and data mining."

[0132] Use the Word2Vec model to extract features from the glossary paraphrase text to obtain the paraphrase feature representation.

[0133] For example, this process splits the glossary paraphrase text into multiple words, and for each word, uses the Word2Vec model to extract features of the word to obtain a set of word feature vectors corresponding to the glossary paraphrase text.

[0134] Step 250, based on the time extraction algorithm, sort and filter the paraphrase feature representation and the first feature representation set to obtain the text traceability record.

[0135] The text traceability record is used to indicate the set of text contents containing the content of the search term, and the text traceability record contains the text contents corresponding to the filtered feature representations.

[0136] Exemplarily, after merging the paraphrase feature representation and the first feature representation set, perform deduplication processing to remove the text contents corresponding to the duplicate feature representations.

[0137] The first feature representation set is S, with a dimension of d. The set contains n word feature representations, and the dimension of each word feature representation is also d. The sentence feature representation S is obtained through the output feature representation of the [CLS] token in the last layer of the BERT model.

[0138] The set of word feature representations corresponding to the glossary paraphrase / paraphrase feature representation is V={v1, v2,..., vn}, where each vi is the d-dimensional feature representation of word i. The merged feature representation set is C={S∪V}={s1, s2,..., sn, v1, v2,..., vn}, where si is the feature representation from S and vi is the feature representation from V.

[0139] For any two word feature representations vi and vj in the set C, calculate the cosine similarity between them. If the cosine similarity reaches the preset threshold, then vi and vj are considered duplicates. Keep one of the feature representations and remove the other from the set C. The retained feature representation can be arbitrary.

[0140] Optionally, based on the time extraction algorithm, parse the paraphrase feature representation and the first feature representation set to obtain the timestamp information corresponding to all candidate text feature representations in the paraphrase feature representation and the first feature representation set. The timestamp information is used to indicate the publication time of the search term content and the candidate text feature representations.

[0141] Sort the paraphrase feature representation and the first feature representation set based on the timestamp information to obtain a sorting result.

[0142] Exemplarily, apply a time extraction algorithm to the text content corresponding to each feature representation in set C, parse the extracted string time expression into a timestamp, and convert it into a unified time format, such as, Unix timestamp.

[0143] The Unix timestamp is the total number of seconds from January 1, 1970, 00:00:00 (Coordinated Universal Time, UTC) to the current time. For example, the Unix timestamp of January 1, 2020 is 1577644800, indicating that 1577644800 seconds have passed between January 1, 1970 and January 1, 2020.

[0144] That is, there is a Unix timestamp for the text content corresponding to each feature representation, and the larger the value, the later the publication time of the text content.

[0145] Optionally, screen at least one feature representation from the first feature representation set based on the similarity between the paraphrase feature representation and the first feature representation set to obtain a second feature representation set; determine the text content corresponding to the second feature representation set as the text content set containing the search term; obtain a text traceability record based on the sorting result and the text content set.

[0146] Exemplarily, determine the matching degree between the first sentence feature representation and the feature representations of multiple first contents based on the cosine distance, Pearson distance, and BM25 correlation value.

[0147] Among them, the cosine distance is a method to measure the difference between two vectors, which calculates the distance based on the cosine similarity of the vectors, and the cosine similarity measures the similarity degree of the two vectors in the direction. The Pearson distance refers to the complement of the Pearson correlation coefficient, which is used to measure the correlation degree between two variables. The BM25 (Best Matching 25) algorithm is a relevance scoring algorithm, which is based on the TF-IDF (Term Frequency-Inverse Document Frequency) model and is improved to consider factors such as document length. The BM25 algorithm evaluates the relevance between a document and a query by calculating the frequency (TF) of the query term appearing in the document and the inverse document frequency (IDF), and adjusting it in combination with the document length.

[0148] To obtain the historical development context and semantic relevance of at least one text content corresponding to the first feature representation set, first obtain set S and set V, and calculate conversion distances such as the cosine distance, Pearson distance, and BM25 correlation value between each feature representation in set V and each feature representation in set S.

[0149] Take the obtained cosine distance, Pearson distance, and BM25 correlation value as sorting features, obtain a sorting result through a fine-rank model, and form a complete text traceability record by integrating them into a coherent historical record according to the sorting result.

[0150] Exemplarily, the content of the search term input by the user is "artificial intelligence", and the target field is field A. Then the text content corresponding to the second feature representation set is 10 text contents in field A (the serial numbers are 0 to 9 respectively). All 10 text contents contain the search term content, and the release order is as follows: serial number 4, serial number 5, serial number 1, serial number 2, serial number 3, serial number 6, serial number 7, serial number 8, serial number 9, serial number 0.

[0151] Integrate the text traceability record according to the release order, and this text traceability record can show the development process of the search term content "artificial intelligence".

[0152] In summary, the method provided by this application adjusts the architecture of the deep learning model, trains the deep learning model using a targeted training text data set, obtains a pre-trained model capable of generating text paraphrase content. This pre-trained model generates a paraphrase for the content of the search term input by the user. Based on the candidate text that meets the matching requirements of the semantic similarity with the search content text and this paraphrase content, it can understand the content of the search term from the perspective of natural language, improve the accuracy during text traceability, effectively track and trace the source of the search term content, and improve the efficiency during text traceability. Compared with the method of using the document marking method to label and continuously maintain a large amount of text data in the related technology, this solution does not require marking and maintenance operations on the text data in the library, and can save computing resources and maintenance time.

[0153] Figure 3 It is the structural block diagram of a text traceability device provided by an exemplary embodiment of this application. As Figure 3 shown, this device includes the following parts.

[0154] A training module 310, configured to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model. The training text data set includes multiple sentences and multiple entries marked with paraphrases. The pre-trained model is used to generate entry paraphrases;

[0155] A feature extraction module 320, configured to obtain the content of a search term and perform feature extraction on the content of the search term to obtain a feature representation of the term, where the content of the search term is a term input by a user for text traceability;

[0156] A matching module 330, configured to determine at least one candidate text feature representation as a first feature representation set from the feature representation library based on the matching degree between the feature representation of the term and multiple candidate text feature representations in the feature representation library;

[0157] An interpretation generation module 340, configured to generate an interpretation text corresponding to the content of the search term based on the pre-trained model, and perform feature extraction on the interpretation text to obtain an interpretation feature representation;

[0158] A traceability module 350, configured to sort and filter the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text traceability record, where the text traceability record is used to indicate a set of text contents including the content of the search term, and the text traceability record includes the text contents corresponding to the filtered feature representations.

[0159] In an optional embodiment, the training module 310 is further configured to obtain the training text data set; train the deep learning model based on the multiple sentences in the training text data set to obtain a first-stage model; adjust the model architecture and model parameters of the first-stage model to obtain a second-stage model; train the second-stage model based on the multiple terms with interpretations annotated in the training text data set to obtain the pre-trained model.

[0160] In an optional embodiment, the training module 310 is further configured to obtain multiple corpus sentences and the multiple terms with interpretations annotated, where the multiple corpus sentences are sentence contents obtained by splitting multiple text contents; perform synonym replacement on the multiple corpus sentences to obtain a first sentence set, where the i-th sentence in the first sentence set is obtained by performing synonym replacement on each term in the i-th corpus sentence among the multiple corpus sentences; filter the multiple corpus sentences based on the term frequency and inverse document frequency method to obtain a second sentence set, where the term frequency refers to the frequency of each term in the multiple corpus sentences, and the inverse document frequency is used to indicate the importance of each term in the multiple corpus sentences; determine the first sentence set and the second sentence set as the multiple sentences.

[0161] In an optional embodiment, the training module 310 is further configured to add a custom feature extraction layer to the first-stage model to obtain the first-stage model with an adjusted model architecture; adjust the learning rate, the number of training epochs, and the batch size of the first-stage model with the adjusted model architecture based on a preset parameter standard to obtain the second-stage model, where the learning rate is used to determine the step size of parameter update during model training, the number of training epochs refers to the number of times the training text dataset is completely trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.

[0162] In an optional embodiment, the matching module 330 is further configured to obtain a candidate text library, where the candidate text library contains multiple candidate text contents; perform feature extraction on the multiple candidate text contents respectively to obtain the candidate text feature representations in the feature representation library; sort the multiple candidate text feature representations based on the similarity between the entry feature representation and the multiple candidate text feature representations to obtain the serial numbers corresponding to the multiple candidate text feature representations respectively; use at least one candidate text feature representation whose serial number meets the preset requirements as the first feature representation set.

[0163] In an optional embodiment, the paraphrase generation module 340 is further configured to input the entry feature representation into the pre-trained model, and output the hidden state corresponding to the entry feature representation, where the hidden state refers to the encoded representation of the search entry content by the pre-trained model and is used to reflect the understanding of the search entry content by the pre-trained model; generate a paraphrase text corresponding to the search entry content through a decoder based on the hidden state.

[0164] In an optional embodiment, the traceability module 350 is configured to parse the paraphrase feature representation and the first feature representation set based on the time extraction algorithm to obtain the timestamp information corresponding to all candidate text feature representations in the paraphrase feature representation and the first feature representation set, where the timestamp information is used to indicate the release time of the search entry content and the candidate text feature representations; sort the paraphrase feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; screen at least one feature representation from the first feature representation set based on the similarity between the paraphrase feature representation and the first feature representation set to obtain a second feature representation set; determine the text content corresponding to the second feature representation set as the text content set containing the search entry; obtain the text traceability record based on the sorting result and the text content set.

[0165] In summary, the text traceability device provided by the present application adjusts the architecture of the deep learning model, trains the deep learning model using a targeted training text data set, and obtains a pre-trained model capable of generating text paraphrase content. The pre-trained model generates a paraphrase for the search term content input by the user. Based on the candidate text that meets the matching requirements of the semantic similarity with the search content text and the paraphrase content, it can understand the search term content from the perspective of natural language, improve the accuracy during text traceability, effectively track and trace the source of the search term content, and improve the efficiency during text traceability. Compared with the method of using document marking to annotate and continuously maintain a large amount of text data in the related art, this solution does not require marking and maintenance operations on the text data in the library, and can save computing resources and maintenance time.

[0166] It should be noted that: for the text traceability device provided in the above embodiment, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the text traceability device provided in the above embodiment and the text traceability method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0167] Figure 4 The structural block diagram of a computer device 400 provided by an exemplary embodiment of the present application is shown. The computer device 400 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The computer device 400 may also be referred to by other names such as a user device, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0168] Generally, the computer device 400 includes: a processor 401 and a memory 402.

[0169] The processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 401 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0170] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the text traceability method provided in the method embodiments of this application.

[0171] In some embodiments, the computer device 400 further includes some other components 403, and the type and quantity of the other components 403 may be selected based on the functional requirements of the computer device 400. Those skilled in the art can understand that Figure 4 the structure shown in does not constitute a limitation on the computer device 400, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.

[0172] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid state drives (SSD), or optical discs, etc. Among them, the random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0173] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the text traceability method as described in any one of the embodiments of the present application above.

[0174] An embodiment of the present application further provides a computer-readable storage medium. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the text traceability method as described in any one of the embodiments of the present application above.

[0175] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text traceability method as described in any one of the above embodiments.

[0176] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be read-only memory, a magnetic disk, or an optical disc, etc.

[0177] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A text tracing method, characterized in that: The method comprises: Acquire a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries; Acquire search term content and perform feature extraction on the search term content to obtain term feature representation, wherein the search term content is a term input by a user for text tracing; Based on the matching degree between the term feature representation and a plurality of candidate text feature representations in the feature representation library, determining at least one candidate text feature representation from the feature representation library as a first feature representation set; Generate a term interpretation text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the term interpretation text to obtain a feature representation of the interpretation; Based on a time extraction algorithm, the interpretation feature representation and the first feature representation set are sorted and filtered to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.

2. The method according to claim 1, characterized in that The step of obtaining a training text data set and training a deep learning model based on the training text data set to obtain a pre-trained model includes: Obtaining the training text dataset; Training the deep learning model based on the multiple sentences in the training text dataset to obtain a first-stage model; Adjusting the model architecture and model parameters of the first-stage model to obtain a second-stage model; The second-stage model is trained based on the multiple entries annotated with definitions in the training text data set to obtain the pre-trained model.

3. The method according to claim 2, characterized in that The step of obtaining the training text data set includes: Acquire a plurality of corpus sentences and the plurality of entries marked with definitions, wherein the plurality of corpus sentences are sentence contents obtained by splitting a plurality of text contents; Performing synonym replacement on the multiple corpus sentences to obtain a first sentence set, wherein the i-th sentence in the first sentence set is obtained by performing synonym replacement on each word item of the i-th corpus sentence in the multiple corpus sentences; The plurality of corpus sentences are screened based on a word frequency and an inverse document frequency method to obtain a second sentence set, wherein the word frequency refers to the frequency at which each word item appears in the plurality of corpus sentences, and the inverse document frequency is used to indicate the importance of each word item in the plurality of corpus sentences; The first sentence set and the second sentence set are determined as the plurality of sentences.

4. The method according to claim 2, characterized in that: The adjusting of the model architecture and model parameters of the first-stage model to obtain the second-stage model includes: Adding a custom feature extraction layer to the first-stage model to obtain a first-stage model after adjusting the model architecture; The learning rate, number of training rounds and batch size of the first-stage model after adjusting the model architecture are adjusted based on preset parameter standards to obtain the second-stage model, wherein the learning rate is used to determine the step size of parameter update of the model during training, the number of training rounds refers to the number of times the training text data set is fully trained in the model, and the batch size is used to indicate the number of samples used by the model in one parameter update.

5. The method according to any one of claims 1 to 4, characterized in that: The step of determining at least one candidate text feature representation from the feature representation library as a first feature representation set based on the matching degree between the feature representation of the term and a plurality of candidate text feature representations in the feature representation library comprises: Acquire a candidate text library, wherein the candidate text library contains a plurality of candidate text contents; After extracting features from the plurality of candidate text contents respectively, obtaining the candidate text feature representations in the feature representation library; Based on the similarity between the term feature representation and the plurality of candidate text feature representations, the plurality of candidate text feature representations are sorted to obtain sequence numbers corresponding to the plurality of candidate text feature representations respectively; The at least one candidate text feature representation whose sequence number meets the preset requirement is used as the first feature representation set.

6. The method according to any one of claims 1 to 4, characterized in that: The generating of the term interpretation text corresponding to the search term content based on the pre-trained model includes: Inputting the term feature representation into the pre-trained model, and outputting a hidden state corresponding to the term feature representation, wherein the hidden state refers to the encoding representation of the search term content by the pre-trained model, and is used to reflect the understanding of the search term content by the pre-trained model; A decoder is used to generate a term interpretation text corresponding to the search term content based on the hidden state.

7. The method according to any one of claims 1 to 4, characterized in that: The method of sorting and filtering the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain a text tracing record includes: Parsing the interpretation feature representation and the first feature representation set based on the time extraction algorithm to obtain timestamp information corresponding to the interpretation feature representation and all candidate text feature representations in the first feature representation set, wherein the timestamp information is used to indicate the release time of the search term content and the candidate text feature representation; sorting the interpretation feature representation and the first feature representation set based on the timestamp information to obtain a sorting result; Filtering at least one feature representation from the first feature representation set based on the similarity between the interpretation feature representation and the first feature representation set to obtain a second feature representation set; Determining the text content corresponding to the second feature representation set as the text content set containing the search term; The text tracing record is obtained based on the sorting result and the text content set.

8. A text tracing device, characterized in that: The device comprises: A training module, used to obtain a training text data set and train a deep learning model based on the training text data set to obtain a pre-trained model, wherein the training text data set includes a plurality of sentences and a plurality of entries marked with definitions, and the pre-trained model is used to generate definition of the entries; A feature extraction module is used to obtain search term content and perform feature extraction on the search term content to obtain a term feature representation, wherein the search term content is a term input by a user for text tracing; A matching module, configured to determine at least one candidate text feature representation from the feature representation library as a first feature representation set based on a degree of matching between the feature representation of the term and a plurality of candidate text feature representations in the feature representation library; A definition generation module, used to generate a definition text corresponding to the search term content based on the pre-trained model, and perform feature extraction on the definition text to obtain a definition feature representation; A tracing module is used to sort and filter the interpretation feature representation and the first feature representation set based on a time extraction algorithm to obtain a text tracing record, wherein the text tracing record is used to indicate a text content set containing the search term content, and the text tracing record contains the text content corresponding to the filtered feature representation.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the text tracing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the text tracing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text enhancement method and system based on word interpretation

    CN113591469A

  • Text model training method, device and equipment

    CN116150621A