Text data processing method based on big data and deep learning

Through text data processing methods based on big data and deep learning, the problem of traditional methods performing poorly when processing super-large data sets is solved, efficient and real-time text data processing is achieved, and processing performance and system stability are improved.

CN120067320APending Publication Date: 2025-05-30CHENGDU WANGDING SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510151256.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional text data processing methods perform poorly when processing super-large data sets, and it is difficult to meet the business needs of high real-time requirements.

Method used

Text data processing methods based on big data and deep learning are adopted, including data preprocessing, feature extraction, classification judgment and query screening, and complex queries are supported through distributed databases and powerful query engines.

Benefits of technology

It improves the timeliness and accuracy of text data processing, especially when facing massive data, and improves the stability and flexibility of the system through modular architecture design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067320A_ABST
    Figure CN120067320A_ABST
Patent Text Reader

Abstract

The invention discloses a text data processing method based on big data and deep learning, and the method comprises the steps: constructing a text big data set, collecting and storing text data, and setting parameters of a data collection layer, the data collection layer is responsible for obtaining text information of different formats from various sources, and using a data preprocessor to process the text information of different formats; an original input text is sorted and standardized conversion is completed, a feature extractor is selected, the preprocessed text is extracted to generate high-quality feature representation, a label is generated, a classification judgment room is set, a classifier cluster is set in the classification judgment room, and each sub-classifier focuses on fine-grained recognition of a certain category. Identifying and comparing labels generated by the feature extractor, setting a query engine, allowing a user to customize complex screening conditions and sorting logic, summarizing and counting, and returning to a front-end interface by means of a high-speed channel to display an achievement report. The invention belongs to the field of data processing, and particularly relates to a text data processing method based on big data and deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and specifically refers to a method for processing text data based on big data and deep learning. Background Art

[0002] With the explosive growth of Internet information volume, how to quickly and effectively classify, store, and retrieve massive unstructured text data has become an important research topic. Traditional text data processing methods usually rely on artificial rules or simple machine learning models, and are often unable to cope when faced with large-scale, complex, and ever-changing data sets, with low efficiency and poor accuracy.

[0003] Specific solutions of the prior art: The technical means commonly adopted in the industry at present mainly include, but are not limited to: Traditional indexing technology based on keyword matching: By establishing an inverted index through a series of pre-set keywords to achieve the function of quickly locating the document position. However, this method has poor support for emerging topics or long-tail words, and is prone to missed detection or over-detection; Application of a single neural network model: In recent years, with the development of deep learning, many researchers have begun to try to use a single type of deep neural network (such as convolutional neural network CNN, recurrent neural network RNN, etc.) to perform end-to-end learning tasks. Although breakthrough progress has been made in certain specific tasks, in actual engineering practice, problems such as long training time and unstable generalization performance are still faced; Hybrid models composed of integrating multiple machine learning algorithms: This strategy attempts to comprehensively utilize the advantages of different types of models to complement each other's weaknesses, such as constructing a new prediction framework by combining the characteristics of decision trees and support vector machines. However, such methods also face challenges such as high integration difficulty and high complexity of parameter tuning. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that traditional text data processing methods have limitations in data processing, especially when dealing with ultra-large data sets, and it is difficult to meet the business requirements with high real-time requirements.

[0005] To solve the above problems, the technical solution adopted by the present invention is as follows: The method for processing text data based on big data and deep learning proposed by the present invention includes the following steps:

[0006] S1: Construct a text big data set, collect and store text data, and set the parameters of the data collection layer. The data collection layer is responsible for obtaining different formats of text information from multiple sources;

[0007] S2: Use a data preprocessor for data preprocessing to organize the original input text and complete the standardization conversion;

[0008] S3: Select a feature extractor to extract features, generate high-quality feature representations from the preprocessed text, and generate labels.

[0009] S4: Set up a classification determination chamber for comparison and recognition, and set up a classifier cluster in the classification determination chamber. Each sub-classifier focuses on the fine-grained recognition of a certain category, and conducts recognition and comparison on the labels generated by the feature extractor.

[0010] S5: Set up a query engine to allow users to customize complex filtering conditions and sorting logics.

[0011] S6: Summarize and statistically analyze, and return the results report to the front-end interface through a high-speed channel for display.

[0012] Furthermore, the data preprocessor is responsible for cleaning the original input file, removing noise interference factors such as invalid characters and stop words, and completing the standardization conversion.

[0013] Furthermore, the feature extractor uses NLP technology to generate high-quality feature representations, including the TF-IDF weight matrix and the WordEmbedding embedding layer.

[0014] Furthermore, the classifier cluster is composed of several specialized sub-classifiers obtained through specialized training.

[0015] Furthermore, the text big data set is constructed on the storage management layer. The storage management layer uses distributed database technology to build a high-performance cache pool and a persistent warehouse to ensure that various intermediate results can be quickly read and written back.

[0016] Furthermore, the query engine supports an advanced search interface with SQL-like syntax, allowing users to customize complex filtering conditions and sorting logics.

[0017] Furthermore, the data preprocessor and the feature extractor use a high-computing power GPU server as the main node when selecting the hardware configuration; at the software level, the Hadoop HDFS component in the Apache Spark ecosystem is introduced to achieve cross-node resource sharing.

[0018] The beneficial effects obtained by the present invention with the above structure are as follows:

[0019] 1. The text data processing method based on big data and deep learning proposed in this solution improves the timeliness and accuracy of text data processing, especially showing excellent performance advantages when facing massive data.

[0020] 2. The text data processing method based on big data and deep learning proposed in this solution innovatively puts forward the design concept of a modular architecture, greatly enhancing the stability and reliability of the system as well as the space for future upgrades and improvements, with high flexibility.

[0021] 3. The text data processing method based on big data and deep learning proposed in this solution uses a distributed database for data management and also develops a powerful query engine, which can support complex query commands and provides a good user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the overall step flow of the present invention.

[0023] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present invention.

[0025] As Figure 1 shown, the text data processing method based on big data and deep learning proposed by the present invention includes a data pre-processor responsible for cleaning the original input file, removing noise interference factors such as invalid characters and stop words, and completing standardization conversion. The feature extractor uses NLP technology to generate high-quality feature representations, including TF-IDF weight matrices and WordEmbedding embedding layers. The classifier cluster is composed of several specialized sub-classifiers obtained through special training. The text big data set is constructed in the storage management layer, and the storage management layer uses distributed database technology to build a high-performance cache pool and a persistent warehouse to ensure that various intermediate results can be quickly read and written back. The query engine supports an advanced search interface with SQL-like syntax, allowing users to customize complex filtering conditions and sorting logics. The data pre-processor and the feature extractor use a high-computing-power GPU server as the main node in terms of hardware configuration; at the software level, the Hadoop HDFS component in the Apache Spark ecosystem is introduced to achieve cross-node resource sharing. The entire workflow starts with the client submitting a task instruction to be processed to the data pre-processor of the scheduling center, and then triggers a chain reaction mechanism to activate each functional module to cooperate in sequence.

[0026] Among them, in the data preprocessing stage, through preliminary parsing of the source code format, redundant information is removed while valid content is retained; then it is passed to the feature extraction link for further processing to form a standard format for subsequent steps.

[0027] Then enter the classification determination room in the core judgment and comparison recognition stage. After determining the category label to which each record belongs according to the pre-prepared mapping table, they are respectively sent to the corresponding partitions for temporary storage and wait for the next action instruction.

[0028] Users use the query engine to customize complex filtering conditions and sorting logics for data filtering. After finally summarizing and counting, it is returned to the front-end interface through the high-speed channel to display the final result report.

[0029] In the face of sudden traffic impacts in extreme situations, it is also possible to consider temporarily increasing the number of additional working threads to cope with the peak load pressure.

[0030] Specific steps:

[0031] S1: Build a text big data set, collect and store text data, and set the parameters of the data collection layer. The data collection layer is responsible for obtaining different format text information from multiple sources;

[0032] S2: Use the data preprocessor for data preprocessing to organize the original input text and complete the standardization conversion;

[0033] S3: Select the feature extractor for feature extraction. In the feature extraction stage, we use natural language processing (NLP) technology to generate high-quality feature representations. Specifically, the feature extractor uses the following two methods to generate feature representations:

[0034] ① TF-IDF weight matrix

[0035] TF-IDF (Term Frequency-Inverse Document Frequency) is a commonly used text feature extraction method for measuring the importance of a word in a document. Its calculation formula is as follows:

[0036]

[0037] Among them:

[0038] represents the word The term frequency (Term Frequency) of the word in document d, and its calculation formula is:

[0039]

[0040] represents the word The Inverse Document Frequency, and its calculation formula is:

[0041]

[0042] where N is the total number of documents in the document collection, is the number of documents containing the word ;

[0043] By calculating the TF-IDF value of each word, a weight matrix is generated to represent the features of the document;

[0044] ②WordEmbedding Embedding Layer

[0045] WordEmbedding is a technique that maps words to a low-dimensional vector space. We use a pre-trained BERT model to generate word vectors. The output of the BERT model can be expressed as:

[0046] E(ω) = BERT(ω)

[0047] where E(ω) is the embedding vector of the word ω, and the BERT model generates context-related word vectors through a multi-layer Transformer structure;

[0048] S4: Set up a classification judgment chamber for comparison. In the classification judgment chamber, we use multiple sub-classifiers to perform fine-grained recognition on the extracted features. Each sub-classifier is a specially trained neural network model, and its output can be expressed as:

[0049]

[0050] where:

[0051] is the probability that the input feature x belongs to the category ;

[0052] h is the feature vector generated by the feature extractor;

[0053] W and b are the weight matrix and bias vector of the classifier respectively;

[0054] The Softmax function is used to convert the output into a probability distribution, and its formula is:

[0055]

[0056] where, is the score of the i-th category, and K is the total number of categories;

[0057] S5: Query and filter settings for the query engine. The query engine supports users to customize complex filtering conditions and sorting logics. To achieve efficient query filtering, we introduce the following scoring functions:

[0058]

[0059] Where:

[0060] Score(d,q) represents the score of document d for query q;

[0061] is the weight of the term in the query, which can be adjusted according to user needs;

[0062] By calculating the scores of each document, the query engine can return the most relevant documents according to the sorting logic defined by the user (such as sorting in descending order of scores);

[0063] S6: Aggregate statistics. In the aggregate statistics stage, we use the following formula to calculate the statistical results of each category:

[0064]

[0065] Where:

[0066] CategoryScore(c) represents the average score of category c;

[0067] ∣c∣ is the number of documents in category c.

[0068] By calculating the average scores of each category, the system can generate a final result report and return it to the front-end interface for display through a high-speed channel.

[0069] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the spirit of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A text data processing method based on big data and deep learning, characterized in that: The following steps are involved: S1: Build a large text dataset, collect text data storage, and set the parameters of the data collection layer, which is responsible for obtaining text information in different formats from multiple sources; S2: Data preprocessing uses a data preprocessor to organize the original input text and complete the standardized transformation; S3: Extract features: Select a feature extractor to extract the preprocessed text to generate high-quality feature representation and generate labels; S4: Comparative identification sets up a classification judgment room, and sets up a classifier cluster in the classification judgment room. Each sub-classifier focuses on the fine-grained recognition of a certain category and performs recognition comparison on the labels generated by the feature extractor; S5: Query filter settings query engine, allowing users to customize complex filter conditions and sorting logic; S6: Summarize the statistics and return them to the front-end interface through the high-speed channel to display the results report.

2. The text data processing method based on big data and deep learning according to claim 1, characterized in that: The data preprocessor is responsible for cleaning the original input file, removing noise interference factors such as invalid characters and stop words, and completing the standardization conversion.

3. The text data processing method based on big data and deep learning according to claim 2 is characterized in that: The feature extractor uses NLP technology to generate high-quality feature representations, including TF-IDF weight matrix and Word Embedding layer. TF-IDF is a commonly used text feature extraction method used to measure the importance of a word in a document. Its calculation formula is as follows: By calculating the TF-IDF value of each word, a weight matrix is ​​generated to represent the characteristics of the document; WordEmbedding is a technique for mapping words to a low-dimensional vector space. We use the pre-trained BERT model to generate word vectors. The output of the BERT model can be expressed as: E(ω)=BERT(ω) Among them, E(ω) is the embedding vector of word ω, and the BERT model generates context-related word vectors through a multi-layer Transformer structure.

4. The text data processing method based on big data and deep learning according to claim 3 is characterized in that: The classifier cluster is composed of several specially trained specialized sub-classifiers, each of which is a specially trained neural network model, and its output can be expressed as:

5. The text data processing method based on big data and deep learning according to claim 4 is characterized in that: The text big data set is built on the storage management layer, and the storage management layer uses distributed database technology to build a high-performance cache pool and a persistent warehouse to ensure that various intermediate results can be quickly read and written back.

6. The text data processing method based on big data and deep learning according to claim 5 is characterized in that: The query engine supports an advanced search interface with SQL-like syntax, allowing users to customize complex filtering conditions and sorting logic.

7. The text data processing method based on big data and deep learning according to claim 5 is characterized in that: The data preprocessor and feature extractor use powerful GPU servers as master nodes in selecting hardware configurations; at the software level, Hadoop HDFS components in the Apache Spark ecosystem are introduced to achieve cross-node resource sharing.

Citation Information

Cited By

  • Text key feature extraction system and method based on deep learning

    CN120913219A

  • Table recognition method and device based on thinking chain, electronic equipment and storage medium

    CN121259854A