Text data processing method and system
Through the combination of multimodal pre-trained neural network and information flow propagation algorithm, the inefficiency problem of traditional data processing methods in complex multimodal data analysis is solved, and efficient automated analysis and report generation of text and tabular data is realized.
Patent Information
- Application Number
- CN202510491402.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional data processing methods are inefficient when facing complex multimodal data, cannot meet the needs of large-scale data real-time processing and analysis, and rely on more manual intervention, and have poor adaptability and scalability.
Multimodal pre-trained neural networks, contrast learning and attention mechanisms are used to fusion of cross-modal features of text and tabular data, and semantic relationship diagrams are constructed in combination with information flow propagation algorithms, and key information is extracted automatically and reports are generated.
It realizes efficient automated analysis of text and tabular data, improves the processing efficiency and accuracy of large-scale data sets, and simplifies complex data processing processes.
Smart Images

Figure CN120493934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a text data processing method and system. Background Art
[0002] With the rapid development of various data platforms and the continuous growth of data volumes, many data systems need to continuously update and analyze data to ensure its timeliness and accuracy. This repetitive data analysis process consumes significant human and material resources. Traditional data processing methods are particularly inefficient when faced with highly complex and large datasets. To meet the growing demand for data analysis, many industries are seeking more efficient solutions that reduce manual intervention and increase processing speed.
[0003] Traditional data analysis techniques often require manual data cleaning, feature extraction, model training, and validation. This process is not only time-consuming but also susceptible to human interference. Furthermore, because traditional techniques often rely on explicit rules and manual adjustments, they lack adaptability and scalability when faced with complex multimodal and heterogeneous data, and cannot meet the demands of large-scale real-time data processing and analysis. Consequently, traditional methods have significant limitations when processing high-dimensional and multi-source data.
[0004] In recent years, machine learning technology has matured and become a crucial tool for data analysis. Machine learning can automatically learn patterns from data to make predictions and decisions, reducing manual intervention and improving the efficiency of data analysis. However, existing language-related machine learning models mostly focus on areas such as text generation, machine translation, sentiment analysis, and dialogue systems. They tend to process a single modality and are often unable to effectively integrate and analyze cross-modal data.
[0005] Therefore, the present invention proposes a text data processing method and system. Summary of the Invention
[0006] The embodiments of the present invention provide a text data processing method and system to solve the above-mentioned technical problems in the prior art.
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, the following is a brief summary. This summary is not intended to be a comprehensive review, identify key or essential elements, or delineate the scope of these embodiments. Its sole purpose is to present some concepts in a simplified form as a prelude to the detailed description that follows.
[0008] According to a first aspect of an embodiment of the present invention, a method for processing text data is provided.
[0009] In one embodiment, the text data processing method includes:
[0010] Acquire data resources and preprocess the text data and table data in the data resources separately. Based on feature fusion technology, perform cross-modal feature fusion on the text data and table data.
[0011] Based on the word frequency-inverse document frequency method, keywords in the fused data are extracted as feature texts, and the semantic relationship between keywords is analyzed using the information flow propagation algorithm. An information flow propagation graph is constructed based on co-occurrence frequency and semantic similarity.
[0012] Determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram. Based on the text recognition model and text frequency recognition algorithm, extract the data fields corresponding to the characteristic text data in the resource data, generate tables and corresponding charts, and import the report template to generate the report.
[0013] In one embodiment, the acquiring of data resources, and preprocessing of text data and table data in the data resources respectively, and cross-modal feature fusion of the text data and table data based on feature fusion technology include:
[0014] Obtain various resource data to be processed, convert text data in different formats into a standard format, parse table files in the resource data, identify fields, and mark data types;
[0015] Based on the multimodal pre-trained neural network model, the features of text data and tabular data are extracted, and the features of text data and tabular data are fused using contrastive learning and attention mechanism.
[0016] In one embodiment, extracting features of text data and tabular data based on a multimodal pre-trained neural network model and fusing the features of text data and tabular data using contrastive learning and an attention mechanism includes:
[0017] Based on the multimodal pre-trained neural network model, the text data and table data are encoded separately, and the fields in the text data and table are converted into feature vector representations;
[0018] In the cross-modal contrastive learning framework, the contrastive loss function is used to align text features and table features, and the attention mechanism is introduced to dynamically adjust the weights of text features and table features.
[0019] The text features and table features processed by contrastive learning and attention mechanism are fused using the splicing method to obtain a joint feature vector, where the joint feature vector includes the key information in the text and table data.
[0020] In one embodiment, the method of aligning text features and table features using a contrastive loss function under a cross-modal contrastive learning framework and introducing an attention mechanism to dynamically adjust the weights of text features and table features includes:
[0021] Map text features and table features to the same feature space, calculate the distance between them using a similarity metric, and optimize the alignment between them based on a contrastive loss function.
[0022] Among them, the expression of the contrast loss function is:
[0023]
[0024] Where L represents the contrast loss, N represents the number of sample pairs, and x i 、x i ′ represents the text features and table features in the i-th pair of samples, yi represents the label, D(x i ,x′ i ) represents the text feature x i and the form feature x i ′, m represents the minimum distance between positive and negative sample pairs;
[0025] Based on the attention module, the attention weights of text features and table features are dynamically adjusted based on the similarity between them.
[0026] In one embodiment, the method of extracting keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyzing the semantic relationship between keywords using the information flow propagation algorithm, and constructing an information flow propagation graph based on co-occurrence frequency and semantic similarity includes:
[0027] The word frequency-inverse document frequency method is used to calculate the word frequency and inverse document frequency of each word in the fused data. The word frequency-inverse document frequency weight of each word is determined based on the word frequency and inverse document frequency. Keywords are selected as feature texts based on the weight ranking results.
[0028] Construct a co-occurrence matrix between keywords, count the frequency of each pair of keywords appearing in the same context, and construct a similarity measure between keywords based on the co-occurrence frequency; use the word vector model to calculate the semantic similarity between keywords, and generate a similarity matrix based on the semantic similarity between keywords;
[0029] An information flow propagation graph is constructed by taking keywords as nodes and the co-occurrence frequency and semantic similarity between keywords as edge weights. An information flow propagation algorithm is used to sort and cluster keywords in the information flow propagation graph, and the propagation path and intensity of information flow in the information flow propagation graph are analyzed to determine the importance and influence of keywords in the entire semantic network.
[0030] In one embodiment, constructing a co-occurrence matrix between keywords, counting the frequency of each pair of keywords appearing in the same context, and constructing a similarity measure between keywords based on the co-occurrence frequency; calculating the semantic similarity between keywords using a word vector model, and generating a similarity matrix based on the semantic similarity between keywords includes:
[0031] Determine the context window, traverse all feature texts, record the frequency of each pair of keywords appearing in the same context window, and obtain the co-occurrence frequency; based on the point mutual information, construct a similarity measure between keywords according to the co-occurrence frequency;
[0032] Among them, the expression of point mutual information is:
[0033]
[0034] Where, PMI(t i ,t j ) represents word t i With the word t j The correlation between i ,t j ) represents word t i With the word t j The joint probability of appearing in the corpus at the same time, P(t i ) represents word t i The probability of occurrence, P(t j ) represents word t j Probability of occurrence;
[0035] Use the word vector model to map each keyword into a vector of fixed dimension, calculate the cosine similarity between keywords, and analyze the semantic similarity based on the cosine similarity between the cosine similarities;
[0036] A similarity matrix is constructed based on the semantic similarity values, and each element in the similarity matrix represents the semantic similarity between keyword pairs. The co-occurrence frequency and semantic similarity are weighted and fused to generate the final keyword similarity matrix.
[0037] In one embodiment, the information flow propagation graph is constructed by using keywords as nodes and the co-occurrence frequency and semantic similarity between keywords as edge weights; the keywords are sorted and clustered in the information flow propagation graph using an information flow propagation algorithm, and the propagation path and intensity of information flow in the information flow propagation graph are analyzed to determine the importance and influence of keywords in the entire semantic network. The steps include:
[0038] Using keywords as nodes, we construct edge weights based on the co-occurrence frequency and semantic similarity between keywords, and generate a weighted graph using the co-occurrence frequency and semantic similarity to obtain an information flow propagation graph.
[0039] The information flow propagation algorithm is used to calculate the importance of each node. Based on the importance value of the node, the keywords are clustered to identify the key groups in the semantic network.
[0040] Analyze the propagation path from one node to another in the information flow propagation graph, calculate the propagation intensity, and evaluate the role and contribution of different nodes in information propagation based on the propagation intensity;
[0041] According to the results of the information flow propagation algorithm, the importance and influence of each keyword in the semantic network are determined, and the ranking and clustering results of the keywords and their importance and influence in the information flow propagation graph are obtained;
[0042] The calculation formula for the importance of a node is:
[0043]
[0044] The formula for calculating the propagation intensity is:
[0045] I(v)=∑ u∈N(v) ω vu I(u)
[0046] Where PR(v) represents the importance of node v, d represents the damping factor, M represents the total number of nodes in the graph, In(v) represents all nodes pointing to node v, C(u) represents the number of edges of node u, PR(u) represents the importance of node u, I(v) represents the propagation strength of node v, N(v) represents all neighboring nodes of node v, and ω vu represents the edge weight between node v and its neighbor node u, and I(u) represents the propagation strength of node u.
[0047] In one embodiment, the method of determining corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extracting data fields corresponding to the characteristic text data in the resource data based on the text recognition model and the text frequency recognition algorithm, and generating a table and corresponding chart, and importing the report template to generate the report includes:
[0048] Mapping feature text data to corresponding nodes in the information flow propagation graph, analyzing the propagation path in the information flow propagation graph, and identifying keyword nodes and their adjacent nodes related to the feature text data;
[0049] Based on the length, weight, and propagation strength of the propagation path in the information flow propagation diagram, the nodes that are highly correlated with the feature text data are screened out, and the priority of the data field is determined based on the importance or propagation strength of the node;
[0050] Analyze resource data using a text recognition model to locate potential data fields corresponding to identified keywords. Analyze the frequency of occurrence of potential data fields in resource data based on a text frequency recognition algorithm, and determine data fields that are highly correlated with feature text based on data priority.
[0051] Organize the extracted data fields in a structured form, build standardized tables, and select corresponding chart forms for visualization based on the type and characteristics of the data in the table;
[0052] Import the generated tables and charts into the preset report template, adjust the report structure according to the propagation path and node importance in the information flow diagram, highlight the core information, and generate a complete report.
[0053] In one embodiment, the method of analyzing the occurrence frequency of potential data fields in resource data based on a text frequency recognition algorithm and determining data fields that are highly relevant to the feature text in combination with the data priority includes:
[0054] Based on the word frequency-inverse document frequency method, the frequency, inverse document frequency and word frequency-inverse document frequency of each potential data field in the resource data are calculated;
[0055] A preset number of data fields are selected based on the ranking results of word frequency-inverse document frequency, and data fields that are highly relevant to the feature text are screened out based on the priority of the data fields.
[0056] According to a second aspect of an embodiment of the present invention, a text data processing system is provided.
[0057] In one embodiment, the text data processing system includes a data resource acquisition module, a keyword extraction and semantic analysis module, and a data field extraction and report generation module;
[0058] The data resource acquisition module is used to acquire data resources and pre-process the text data and table data in the data resources respectively, and perform cross-modal feature fusion on the text data and table data based on feature fusion technology;
[0059] The keyword extraction and semantic analysis module is used to extract keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyze the semantic relationship between keywords using the information flow propagation algorithm, and construct an information flow propagation graph based on co-occurrence frequency and semantic similarity;
[0060] The data field extraction and report generation module is used to determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extract the data fields corresponding to the characteristic text data in the resource data based on the text recognition model and text frequency recognition algorithm, generate tables and corresponding charts, and import the report template to generate a report.
[0061] According to a third aspect of an embodiment of the present invention, a computer device is provided.
[0062] In some embodiments, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0063] According to a fourth aspect of embodiments of the present invention, a computer-readable storage medium is provided.
[0064] In one embodiment, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0065] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0066] 1) This invention combines advanced technologies such as multimodal pre-trained neural networks, contrastive learning, and attention mechanisms to achieve cross-modal feature fusion of text and tabular data. It also constructs a semantic relationship graph between keywords through an information flow propagation algorithm, effectively extracting and analyzing key information from the data. This technology not only improves the automation level of data processing but also provides more efficient and accurate analysis capabilities for large-scale datasets, addressing the shortcomings of traditional data processing methods in complex data analysis.
[0067] 2) The present invention is based on the existing machine learning text data processing model, modifies its output layer, and integrates it with the data analysis system to achieve automated search, reporting and data processing for specific texts and tables, simplifying and automating various highly repetitive data processing processes, thereby effectively solving the problems of low efficiency, small scale and lack of scalability in the existing technology.
[0068] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0070] Figure 1 is a flowchart of a method for processing text data according to an exemplary embodiment;
[0071] Figure 2 is a structural block diagram of a text data processing system according to an exemplary embodiment;
[0072] Figure 3 The figure is a schematic diagram showing the structure of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION
[0073] The following description and accompanying drawings sufficiently illustrate the specific embodiments herein to enable those skilled in the art to practice them. Portions and features of some embodiments may be included in or substituted for portions and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims, including all available equivalents thereof. Herein, the terms "first," "second," and the like are used solely to distinguish one element from another and do not require or imply any actual relationship or order between these elements. In practice, the first element can also be referred to as the second element, and vice versa. Furthermore, the terms "comprise," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a structure, device, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such structure, device, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the structure, device, or apparatus comprising the element. The various embodiments herein are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Similar or identical parts between the various embodiments can be referenced to each other.
[0074] The terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like used herein to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are intended only to facilitate the description of this document and simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention. In the description herein, unless otherwise specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, they can be mechanical or electrical connections, or they can be internal connections between two elements, they can be directly connected, or they can be indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to the specific circumstances.
[0075] As used herein, unless otherwise specified, the term "plurality" means two or more.
[0076] In this document, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0077] In this article, the term "and / or" is used to describe the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.
[0078] It should be understood that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0079] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software so that the processor can call and execute the operations corresponding to the above modules.
[0080] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0081] Figure 1An embodiment of a text data processing method of the present invention is shown.
[0082] In this optional embodiment, the text data processing method includes:
[0083] Step S101: Acquire data resources, pre-process text data and table data in the data resources, and perform cross-modal feature fusion on the text data and table data based on feature fusion technology;
[0084] The acquiring of data resources, preprocessing of text data and table data in the data resources, and cross-modal feature fusion of the text data and table data based on feature fusion technology include:
[0085] 1) Obtain various resource data to be processed (supports Markdown, TXT, PDF, DOCX, HTML, XLS, XLSX, and CSV file types), convert text data in different formats in the resource data into a standard format, and parse the table files in the resource data to identify fields and mark data types;
[0086] Data resource acquisition: Collect and import different types of data resources, including text files (such as Markdown, TXT, PDF, DOCX, HTML, etc.) and table data files (such as XLS, XLSX, CSV, etc.). Data can be acquired and imported in the following ways:
[0087] Text data obtains document content from files, web pages, or databases. Text information in documents (such as plain text in PDF and HTML) can be extracted using OCR technology or parsing tools.
[0088] Extract tabular data from Excel, CSV, database and other table formats, ensuring that each field in the table has a clear definition;
[0089] Text data preprocessing: Remove irrelevant characters from the text, such as extra spaces, punctuation, and line breaks, and convert different file formats to standard text formats to ensure that their content can be further processed (such as UTF-8 encoded text). For files in formats such as PDF, HTML, and DOCX, extract the plain text content and remove useless symbols and formatting.
[0090] Preprocessing of table data: Remove blank rows, invalid columns, and redundant null values in tables to ensure data validity; fill missing values in tables using appropriate methods (such as mean filling, interpolation, and deleting missing rows) based on business needs; parse table files (such as XLS, XLSX, and CSV), identify fields (column names, cell values, etc.), and annotate data types to ensure that table data can be integrated with text data and analyzed together in subsequent steps;
[0091] 2) Based on a multimodal pre-trained neural network model, features of text and tabular data are extracted and then fused using contrastive learning and attention mechanisms;
[0092] Specifically, based on the multimodal pre-trained neural network model, the features of text data and tabular data are extracted, and the features of text data and tabular data are fused using contrastive learning and attention mechanisms, including:
[0093] 21) Based on a multimodal pre-trained neural network model (such as the improved CLIP), encode text data and tabular data separately, and convert the text data and table fields (such as numerical data and column names) into feature vector representations;
[0094] 22) In the cross-modal contrastive learning framework, the contrastive loss function is used to align text features and table features, and the attention mechanism is introduced to dynamically adjust the weights of text features and table features. Specifically,
[0095] Map text features and table features to the same feature space, calculate the distance between them using a similarity metric, and optimize the alignment between them based on a contrastive loss function.
[0096] Among them, the expression of the contrast loss function is:
[0097]
[0098] Where L represents the contrast loss, N represents the number of sample pairs, and x i 、x i ′ represents the text features and table features in the i-th pair of samples, yi represents the label, D(x i ,x i ′) represents the text feature x i and the form feature x i ′, m represents the minimum distance between positive and negative sample pairs;
[0099] Based on the attention module, the attention weights of text features and table features are dynamically adjusted based on the similarity between them.
[0100] Specifically, an attention mechanism is designed to assign dynamic weights to text and table features. This mechanism can adjust the contribution of text and table features at different training stages based on task requirements. Through the attention mechanism, the weights of text and table features are adjusted according to their importance to the task. This step helps to flexibly adjust feature weights in different situations (for example, when text information is more abundant or table information is more important).
[0101] 23) Using the concatenation method, the text features and table features processed by contrastive learning and attention mechanism are fused to obtain a joint feature vector, where the joint feature vector includes key information in the text and table data;
[0102] Step S102: extract keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyze the semantic relationships between keywords using the information flow propagation algorithm, and construct an information flow propagation graph based on co-occurrence frequency and semantic similarity;
[0103] The method of extracting keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyzing the semantic relationship between keywords using the information flow propagation algorithm, and constructing an information flow propagation graph based on co-occurrence frequency and semantic similarity includes:
[0104] 1) Use the word frequency-inverse document frequency method to calculate the word frequency and inverse document frequency of each word in the fused data, determine the word frequency-inverse document frequency weight of each word based on the word frequency and inverse document frequency, and select keywords as feature texts based on the weight ranking results (in addition, you can also check the feature texts, manually add, delete or modify the automatically generated feature texts, and select the required feature texts);
[0105] Specifically, TF-IDF (term frequency–inverse document frequency) is a statistical method used to assess the importance of a word to a document within a collection or corpus. The importance of a word increases with the number of times it appears in a document, but decreases inversely with its frequency in the corpus. The key idea is that if a word appears with a high TF frequency in one article and rarely appears in other articles, it is considered to have good classificatory power and is suitable for classification.
[0106] Term frequency (TF) indicates the frequency with which a term (keyword) appears in a text:
[0107]
[0108] Where TFi,j Indicates word frequency, n i,j Indicates the number of times word i appears in file j, n k,j represents the number of all words in file j, and k represents the total number of words;
[0109] Inverse Document Frequency (IDF): The IDF of a particular word can be calculated by dividing the total number of documents by the number of documents containing the word and taking the logarithm of the quotient:
[0110]
[0111] Where, IDF j represents the inverse document frequency, D represents the total number of documents in the data, and d i Represents the number of documents containing word i;
[0112] If the word is not in the corpus, the denominator will be zero, so |d is generally used. i |+1 is used as the numerator. If the number of documents containing the term t is smaller, the larger the IDF is, which means that the term has good category discrimination ability.
[0113] Term Frequency-Inverse Document Frequency (TF-IDF) combines the frequency of a field in resource data with its importance. The calculation formula is as follows:
[0114] TF-IDF(i,j)=TF i,j IDF j
[0115] Where TF-IDF(i,j) represents the weight of word i in document j;
[0116] 2) Construct a co-occurrence matrix between keywords, count the frequency of each pair of keywords appearing in the same context, and construct a similarity measure between keywords based on the co-occurrence frequency; use the word vector model to calculate the semantic similarity between keywords, and generate a similarity matrix based on the semantic similarity between keywords;
[0117] Specifically, the method of constructing a co-occurrence matrix between keywords, counting the frequency of each pair of keywords appearing in the same context, and constructing a similarity measure between keywords based on the co-occurrence frequency; calculating the semantic similarity between keywords using a word vector model, and generating a similarity matrix based on the semantic similarity between keywords includes:
[0118] 21) Determine the context window, traverse all feature texts, record the frequency of each pair of keywords appearing in the same context window, and obtain the co-occurrence frequency; construct a similarity measure between keywords based on the co-occurrence frequency based on the point mutual information;
[0119] Define the context window: Set a window size to determine the range within which word pairs are considered co-occurring. Common context window sizes are 3-5 words.
[0120] Counting co-occurrence frequencies: traverse all texts and record the frequency of each pair of keywords appearing in the same context window to form a co-occurrence matrix. Each matrix element represents the number of times two keywords co-occur.
[0121] Pointwise Mutual Information (PMI): measures the correlation between word pairs. The expression of pointwise mutual information is:
[0122]
[0123] Where, PMI(t i ,t j ) represents word t i With the word t j The correlation between i ,t j ) represents word t i With the word t j The joint probability of appearing in the corpus at the same time, P(t i ) represents word t i The probability of occurrence, P(t j ) represents word t j Probability of occurrence;
[0124] 22) Use the word vector model to map each keyword into a vector of fixed dimension, calculate the cosine similarity between the keywords, and analyze the semantic similarity based on the cosine similarity between the cosine similarities;
[0125] Specifically, the word vector model in this embodiment is the word2vec-CBOW word vector model. In order to enable computers to intuitively process text data, the text must first be converted into space vectors through the word vector model. The word vector model is a model that maps words in a vocabulary to real-number vector space. These vectors capture the semantic relationship between words, so that words with similar meanings are close to each other in the vector space.
[0126] set up Represents T words w1, w2......, w T The probability of a sentence composed of is:
[0127]
[0128] Where, Represents the probability of the entire sentence W, p(w1,w2......,w T) represents the probability of all words appearing together in the sentence, p(w1) represents the probability of the first word w1, and p(w2|w1) represents the probability of the second word w2 given the first word w1. It represents the probability of the third word w3 given the first two words w1 and w2. Indicates that given the first T-1 words, the Tth word w T probability;
[0129] After converting the text into probability numbers, we introduce a neural network model to construct an objective function for it, then optimize this objective function to find an optimal set of parameters. Finally, we use the model corresponding to this set of optimal parameters to make predictions. For statistical language models, using maximum likelihood, the objective function can be set as:
[0130] Πp(w|context(w))
[0131] In the formula, ∏ represents the multiplication symbol, that is, the conditional probabilities of all words are multiplied, Context(w) represents the context of word w, that is, the set of words around w;
[0132] For the convenience of calculation, the maximum log-likelihood is taken:
[0133] F = ∑ log p ( w | context ( w ) )
[0134] Where F represents the objective function, that is, the log-likelihood function;
[0135] The process of solving the parameters of the maximum log-likelihood of the objective function is the training process of the model;
[0136] 23) constructing a similarity matrix based on the semantic similarity values, wherein each element in the similarity matrix represents the semantic similarity between the keyword pairs, and weightedly fusing the co-occurrence frequency and the semantic similarity to generate a final keyword similarity matrix. The two can be combined by weighted averaging, multiplication, or other fusion methods to comprehensively consider the semantic relationship and co-occurrence frequency of the keywords;
[0137] 3) Using keywords as nodes and the co-occurrence frequency and semantic similarity between keywords as edge weights, we construct an information flow propagation graph. We use an information flow propagation algorithm to sort and cluster keywords in the graph, analyze the propagation path and intensity of information flow in the graph, and determine the importance and influence of keywords in the entire semantic network. This includes:
[0138] 31) Using keywords as nodes, construct edge weights based on the co-occurrence frequency and semantic similarity between keywords, and generate a weighted graph using the co-occurrence frequency and semantic similarity to obtain an information flow propagation graph;
[0139] Define nodes: Treat each keyword as a node in the graph. Each node represents a specific keyword;
[0140] Define edges and weights: Construct edge weights based on the co-occurrence frequency and semantic similarity between keywords. You can use weighted sum, weighted product, etc. to combine the co-occurrence frequency and semantic similarity as the edge weight.
[0141] Construct a weighted graph: A weighted graph is generated using co-occurrence frequency and semantic similarity, where the weight of each edge represents the strength of the relationship between two keywords. The edge weights of the graph can be calculated using a weighted formula that combines co-occurrence frequency and semantic similarity to ensure the combined influence of the two. This results in a complete information flow propagation graph, where each node represents a keyword and each edge represents the strength of the relationship between nodes.
[0142] 32) Use the information flow propagation algorithm to calculate the importance of each node, cluster the keywords based on the importance value of the node, and identify the key groups in the semantic network;
[0143] In this embodiment, the information flow propagation algorithm uses the PageRank method (PageRank algorithm is an algorithm used to evaluate the importance of web pages), and the importance value of the node is calculated by the PageRank value of the node;
[0144] 33) Analyze the propagation path from one node to another in the information flow propagation graph, calculate the propagation intensity, and evaluate the role and contribution of different nodes in information propagation based on the propagation intensity;
[0145] 34) Based on the results of the information flow propagation algorithm, determine the importance and influence of each keyword in the semantic network, and obtain the keyword ranking results, clustering results, and their importance and influence in the information flow propagation graph;
[0146] Specifically, the influence of keywords can be evaluated by the node's propagation strength, clustering results, and its connection with other nodes; the importance of keywords can be measured by its PageRank value in the graph, centrality and other graph indicators;
[0147] The calculation formula for the importance of a node is:
[0148]
[0149] The formula for calculating the propagation intensity is:
[0150] I(v)=∑ u∈N(v) ω vu I(u)
[0151] Where PR(v) represents the importance of node v, d represents the damping factor, M represents the total number of nodes in the graph, In(v) represents all nodes pointing to node v, C(u) represents the number of edges of node u, PR(u) represents the importance of node u, I(v) represents the propagation strength of node v, N(v) represents all neighboring nodes of node v, and ω vu represents the edge weight between node v and neighbor node u, and I(u) represents the propagation intensity of node u;
[0152] Step S103: Determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation graph, extract the data fields corresponding to the characteristic text data in the resource data based on the text recognition model and the text frequency recognition algorithm, generate a table and corresponding chart, and import them into the report template to generate a report;
[0153] The steps of determining the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extracting the data fields corresponding to the characteristic text data in the resource data based on the text recognition model and the text frequency recognition algorithm, and generating tables and corresponding charts, and importing the report template to generate the report include:
[0154] 1) Mapping the feature text data to the corresponding nodes in the information flow propagation graph, analyzing the propagation path in the information flow propagation graph, and identifying the keyword nodes related to the feature text data and their adjacent nodes;
[0155] 2) Based on the length, weight, and propagation strength of the propagation path in the information flow propagation graph, the nodes that are highly correlated with the feature text data are screened out, and the priority of the data field is determined based on the importance or propagation strength of the node;
[0156] 3) Analyze resource data using the text recognition model to locate potential data fields corresponding to the identified keywords. Analyze the frequency of occurrence of potential data fields in the resource data based on the text frequency recognition algorithm, and determine the data fields that are highly correlated with the feature text based on the data priority.
[0157] Specifically, the method of analyzing resource data using a text recognition model to locate potential data fields corresponding to the determined keywords, analyzing the frequency of occurrence of potential data fields in the resource data based on a text frequency recognition algorithm, and determining data fields that are highly relevant to the feature text in combination with the data priority includes:
[0158] 31) Analyze resource data using text recognition models
[0159] Choose an appropriate text recognition model:
[0160] Named Entity Recognition (NER) model: If the resource data contains named entities (such as names of people, places, and organizations), the NER model can be used to locate these entities.
[0161] BERT or other Transformer models: For more complex semantic analysis, you can use pre-trained language models such as BERT to identify potential data fields related to keywords.
[0162] Analyze text content in resource data: Analyze the text content in resource data to locate text fields that may contain keywords or be related to keywords; use text recognition models to identify field names, values, and the relationship between field values and identified keywords;
[0163] 32) Analyze potential data fields based on text frequency recognition algorithm
[0164] Extract potential data fields: Through text recognition model analysis, obtain the fields in the resource data that may correspond to keywords and record the frequency of occurrence of these fields;
[0165] Calculate text frequency: Based on the word frequency-inverse document frequency method, calculate the frequency of each potential data field in the resource data, the inverse document frequency, and the word frequency-inverse document frequency;
[0166] Calculating TF-IDF values: TF-IDF (Term Frequency-Inverse Document Frequency) combines the frequency of a field in the resource data with its importance. TF-IDF values are calculated for all potential fields, ranking and selecting the most relevant ones.
[0167] 33) Filter fields based on data priority: Filter out the fields most relevant to the feature text data and the identified keywords according to priority, and give priority to extracting those fields with high TF-IDF values and high correlation with the keywords;
[0168] 4) Organize the extracted data fields in a structured form, construct a standardized table, and select the corresponding chart form for visualization based on the type and characteristics of the data in the table;
[0169] Specifically, after the table is generated, it is necessary to check the generated table to ensure that the data is correct; various results will be calculated based on the data in the table, various charts will be generated, and the table will be stored in the database;
[0170] 5) Import the generated tables and charts into the preset report template, adjust the report structure according to the propagation path and node importance in the information flow propagation diagram, highlight the core information, and generate a complete report.
[0171] Figure 2An embodiment of a text data processing system of the present invention is shown.
[0172] In this optional embodiment, the text data processing system includes:
[0173] The data resource acquisition module 201 is used to acquire data resources, pre-process the text data and table data in the data resources, and perform cross-modal feature fusion on the text data and table data based on feature fusion technology;
[0174] The keyword extraction and semantic analysis module 202 is used to extract keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyze the semantic relationship between keywords using the information flow propagation algorithm, and construct an information flow propagation graph based on co-occurrence frequency and semantic similarity;
[0175] The data field extraction and report generation module 203 is used to determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extract the data fields corresponding to the characteristic text data in the resource data based on the text recognition model and text frequency recognition algorithm, generate tables and corresponding charts, and import the report template to generate a report.
[0176] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps of the above-mentioned method embodiment are implemented.
[0177] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0178] In addition, the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiment when executing the computer program.
[0179] In addition, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0180] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0181] The present invention is not limited to the structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A text data processing method, characterized in that: include: Acquire data resources and preprocess the text data and table data in the data resources separately. Based on feature fusion technology, perform cross-modal feature fusion on the text data and table data. Based on the word frequency-inverse document frequency method, keywords in the fused data are extracted as feature texts, and the semantic relationship between keywords is analyzed using the information flow propagation algorithm. An information flow propagation graph is constructed based on co-occurrence frequency and semantic similarity. Determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram. Based on the text recognition model and text frequency recognition algorithm, extract the data fields corresponding to the characteristic text data in the resource data, generate tables and corresponding charts, and import the report template to generate the report.
2. The text data processing method according to claim 1, characterized in that: The step of acquiring data resources, preprocessing text data and table data in the data resources, and fusing the text data and table data with cross-modal features based on feature fusion technology includes: Obtain various resource data to be processed, convert text data in different formats into a standard format, parse table files in the resource data, identify fields, and mark data types; Based on the multimodal pre-trained neural network model, the features of text data and tabular data are extracted, and the features of text data and tabular data are fused using contrastive learning and attention mechanism.
3. The text data processing method according to claim 2, characterized in that: The method of extracting features of text data and tabular data based on a multimodal pre-trained neural network model and fusing the features of text data and tabular data using contrastive learning and an attention mechanism includes: Based on the multimodal pre-trained neural network model, the text data and table data are encoded separately, and the fields in the text data and table are converted into feature vector representations; In the cross-modal contrastive learning framework, the contrastive loss function is used to align text features and table features, and the attention mechanism is introduced to dynamically adjust the weights of text features and table features. The text features and table features processed by contrastive learning and attention mechanism are fused using the splicing method to obtain a joint feature vector, where the joint feature vector includes the key information in the text and table data.
4. The text data processing method according to claim 3, characterized in that: In the cross-modal contrastive learning framework, the contrastive loss function is used to align text features and table features, and the attention mechanism is introduced to dynamically adjust the weights of text features and table features. The following steps are involved: Map text features and table features to the same feature space, calculate the distance between them using a similarity metric, and optimize the alignment between them based on a contrastive loss function. Among them, the expression of the contrast loss function is: Where L represents the contrast loss, N represents the number of sample pairs, and x i 、x i ′ respectively represent the text features and table features in the i-th pair of samples, y i represents the label, D(x i ,x′ i ) represents the text feature x i and the form feature x i ′, m represents the minimum distance between positive and negative sample pairs; Based on the attention module, the attention weights of text features and table features are dynamically adjusted based on the similarity between them.
5. The text data processing method according to claim 1, characterized in that: The method based on word frequency-inverse document frequency extracts keywords from the fused data as feature texts, analyzes the semantic relationship between keywords using the information flow propagation algorithm, and constructs an information flow propagation graph based on co-occurrence frequency and semantic similarity, including: The word frequency-inverse document frequency method is used to calculate the word frequency and inverse document frequency of each word in the fused data. The word frequency-inverse document frequency weight of each word is determined based on the word frequency and inverse document frequency. Keywords are selected as feature texts based on the weight ranking results. Construct a co-occurrence matrix between keywords, count the frequency of each pair of keywords appearing in the same context, and construct a similarity measure between keywords based on the co-occurrence frequency; use the word vector model to calculate the semantic similarity between keywords, and generate a similarity matrix based on the semantic similarity between keywords; An information flow propagation graph is constructed by taking keywords as nodes and the co-occurrence frequency and semantic similarity between keywords as edge weights. An information flow propagation algorithm is used to sort and cluster keywords in the information flow propagation graph, and the propagation path and intensity of information flow in the information flow propagation graph are analyzed to determine the importance and influence of keywords in the entire semantic network.
6. The text data processing method according to claim 5, characterized in that: The co-occurrence matrix between keywords is constructed, the frequency of each pair of keywords appearing in the same context is counted, and a similarity measure between keywords is constructed based on the co-occurrence frequency; The word vector model is used to calculate the semantic similarity between keywords. The similarity matrix generated based on the semantic similarity between keywords includes: Determine the context window, traverse all feature texts, record the frequency of each pair of keywords appearing in the same context window, and obtain the co-occurrence frequency; based on the point mutual information, construct a similarity measure between keywords according to the co-occurrence frequency; Among them, the expression of point mutual information is: Where, PMI(t i ,t j ) represents word t i With the word t j The correlation between i ,t j ) represents word t i With the word t j The joint probability of appearing in the corpus at the same time, P(t i ) represents word t i The probability of occurrence, P(t j ) represents word t j Probability of occurrence; Use the word vector model to map each keyword into a vector of fixed dimension, calculate the cosine similarity between keywords, and analyze the semantic similarity based on the cosine similarity between the cosine similarities; A similarity matrix is constructed based on the semantic similarity values, and each element in the similarity matrix represents the semantic similarity between keyword pairs. The co-occurrence frequency and semantic similarity are weighted and fused to generate the final keyword similarity matrix.
7. The text data processing method according to claim 5, characterized in that: The information flow propagation graph is constructed by using keywords as nodes and the co-occurrence frequency and semantic similarity between keywords as edge weights; the information flow propagation algorithm is used to sort and cluster keywords in the information flow propagation graph, and the propagation path and intensity of information flow in the information flow propagation graph are analyzed to determine the importance and influence of keywords in the entire semantic network. Using keywords as nodes, we construct edge weights based on the co-occurrence frequency and semantic similarity between keywords, and generate a weighted graph using the co-occurrence frequency and semantic similarity to obtain an information flow propagation graph. The information flow propagation algorithm is used to calculate the importance of each node. Based on the importance value of the node, the keywords are clustered to identify the key groups in the semantic network. Analyze the propagation path from one node to another in the information flow propagation graph, calculate the propagation intensity, and evaluate the role and contribution of different nodes in information propagation based on the propagation intensity; According to the results of the information flow propagation algorithm, the importance and influence of each keyword in the semantic network are determined, and the ranking and clustering results of the keywords and their importance and influence in the information flow propagation graph are obtained; The calculation formula for the importance of a node is: The formula for calculating the propagation intensity is: Where PR(v) represents the importance of node v, d represents the damping factor, M represents the total number of nodes in the graph, In(v) represents all nodes pointing to node v, C(u) represents the number of edges of node u, PR(u) represents the importance of node u, I(v) represents the propagation strength of node v, N(v) represents all neighboring nodes of node v, and ω vu represents the edge weight between node v and its neighbor node u, and I(u) represents the propagation strength of node u.
8. The text data processing method according to claim 1, characterized in that: The method of determining corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extracting data fields corresponding to the characteristic text data in the resource data based on the text recognition model and the text frequency recognition algorithm, and generating tables and corresponding charts, and importing the report template to generate the report includes: Mapping feature text data to corresponding nodes in the information flow propagation graph, analyzing the propagation path in the information flow propagation graph, and identifying keyword nodes and their adjacent nodes related to the feature text data; Based on the length, weight, and propagation strength of the propagation path in the information flow propagation diagram, the nodes that are highly correlated with the feature text data are screened out, and the priority of the data field is determined based on the importance or propagation strength of the node; Analyze resource data using a text recognition model to locate potential data fields corresponding to identified keywords. Analyze the frequency of occurrence of potential data fields in resource data based on a text frequency recognition algorithm, and determine data fields that are highly correlated with feature text based on data priority. Organize the extracted data fields in a structured form, build standardized tables, and select corresponding chart forms for visualization based on the type and characteristics of the data in the table; Import the generated tables and charts into the preset report template, adjust the report structure according to the propagation path and node importance in the information flow diagram, highlight the core information, and generate a complete report.
9. The text data processing method according to claim 8, characterized in that: The method of analyzing the occurrence frequency of potential data fields in resource data based on a text frequency recognition algorithm and determining data fields that are highly relevant to the feature text in combination with the priority of the data includes: Based on the word frequency-inverse document frequency method, the frequency, inverse document frequency and word frequency-inverse document frequency of each potential data field in the resource data are calculated; A preset number of data fields are selected based on the ranking results of word frequency-inverse document frequency, and data fields that are highly relevant to the feature text are screened out based on the priority of the data fields.
10. A text data processing system, characterized in that: Including data resource acquisition module, keyword extraction and semantic analysis module, data field extraction and report generation module; The data resource acquisition module is used to acquire data resources and pre-process the text data and table data in the data resources respectively, and perform cross-modal feature fusion on the text data and table data based on feature fusion technology; The keyword extraction and semantic analysis module is used to extract keywords from the fused data as feature texts based on the word frequency-inverse document frequency method, analyze the semantic relationship between keywords using the information flow propagation algorithm, and construct an information flow propagation graph based on co-occurrence frequency and semantic similarity; The data field extraction and report generation module is used to determine the corresponding data fields based on the characteristic text data and the propagation path of the information flow propagation diagram, extract the data fields corresponding to the characteristic text data in the resource data based on the text recognition model and text frequency recognition algorithm, generate tables and corresponding charts, and import the report template to generate a report.