Emergency resource classification method and system based on natural language processing
By using natural language processing technology, combined with data cleaning, word segmentation, word frequency features, and convolutional neural network models, the problems of low efficiency and poor accuracy in emergency resource classification have been solved, achieving efficient and accurate automatic classification of emergency resources and improving classification accuracy and management efficiency.
Patent Information
- Application Number
- CN202511016181.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-11
AI Technical Summary
Existing emergency resource classification methods are inefficient, inaccurate, and lack generalization ability, especially when faced with resources that are multilingual and semantically complex.
By employing a natural language processing approach, including data cleaning, word segmentation, part-of-speech tagging, word frequency feature extraction, and convolutional neural network models, combined with an attention mechanism, an emergency resource classification model is constructed to achieve automatic classification of emergency resource texts.
It improves the accuracy and efficiency of emergency resource classification, enabling more accurate classification decisions in complex texts. The classification accuracy can be improved by 10%-20%, and it has continuous optimization capabilities, significantly improving resource management and retrieval efficiency.
Smart Images

Figure CN120929598A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data classification technology, specifically to an emergency resource classification method and system based on natural language processing. Background Technology
[0002] With the rapid development of the times, emergency resources have become more diverse and sophisticated. Whether it's emergency document management, emergency personnel, emergency supplies within government departments, or other types of emergency resources, efficient and accurate classification methods are needed so that users can quickly retrieve and utilize the emergency resources they need.
[0003] Traditional emergency resource classification methods mainly rely on manual labeling or classification systems based on simple rules. Manual labeling is inefficient, costly, and easily influenced by subjective factors, making it difficult to meet the classification needs of massive amounts of emergency resources. While rule-based classification systems can achieve a certain degree of automation, rule formulation requires significant manpower and time, and lacks flexibility, making it difficult to accurately classify newly emerging emergency resource types or semantically complex content.
[0004] The development of natural language processing technology has provided a new approach to solving the problem of emergency resource classification. By analyzing and processing the textual information in the resource content and extracting valuable features, it is possible to automatically classify emergency resources.
[0005] However, existing emergency resource classification methods based on natural language processing still have shortcomings in terms of accuracy, generalization ability, and processing efficiency. For example, they perform poorly when dealing with resources that are multilingual, semantically ambiguous, or highly domain-specific. Therefore, developing a more efficient, accurate, and generalizable emergency resource classification technology based on natural language processing is of significant practical importance. Summary of the Invention
[0006] This invention addresses the problems of low efficiency, poor accuracy, and insufficient generalization ability of existing emergency resource classification methods by providing an emergency resource classification method and system based on natural language processing, which enables accurate matching and automatic classification management of emergency resource data and emergency resource types.
[0007] Firstly, the present invention provides an emergency resource classification method based on natural language processing, and the technical solution adopted to solve the above-mentioned technical problems is as follows:
[0008] An emergency resource classification method based on natural language processing includes the following steps:
[0009] S1. For the input emergency resource text data, perform data cleaning, word segmentation and part-of-speech tagging preprocessing operations in sequence;
[0010] S2. Extract the word frequency features of the preprocessed emergency resource text, weight the words in the text, use the pre-trained word vector model to obtain the distributed representation of words and combine them into text vectors to capture the semantic features of the emergency resource text.
[0011] S3. Construct an emergency resource classification model based on a convolutional neural network architecture, and introduce an attention mechanism in the model's input layer or intermediate layer; use labeled emergency resource text data to train and optimize the emergency resource classification model so that the model outputs the probability distribution of emergency resource text belonging to different categories;
[0012] S4. When acquiring new emergency resource text, first execute steps S1 and S2 to obtain a text vector with the same format as the training data of the emergency resource classification model. Then, input the text vector into the emergency resource classification model optimized in step S3. The emergency resource classification model calculates the output results through forward propagation to obtain the probability distribution of the text belonging to different categories. Select the category with the highest probability as the classification result.
[0013] Optionally, step S1 specifically includes:
[0014] S1.1. For the input emergency resource text data, perform data cleaning to remove noise information;
[0015] S1.2. For the cleaned emergency resource text data, select a word segmentation tool based on its language type and perform word segmentation processing;
[0016] S1.3. Using a natural language processing toolkit, perform part-of-speech tagging on each word after word segmentation.
[0017] Alternatively, step S1.2 can be performed:
[0018] For Chinese emergency resource text data, the Jieba word segmentation tool is used to segment continuous text into individual words;
[0019] For emergency resource text data in English, words are segmented using spaces and punctuation marks.
[0020] Optionally, step S2 specifically includes:
[0021] S2.1. Use the bag-of-words model to extract word frequency features from emergency resource texts, including: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary, and the value on the dimension represents the frequency of the corresponding word.
[0022] S2.2. Based on the frequency of each word in the emergency resource text, the TF-IDF algorithm is used to weight the words to highlight the uniqueness of the text;
[0023] S2.3. Using a pre-trained word vector model, words are mapped to a low-dimensional vector space, so that words with similar meanings are closer in the vector space. Then, by using an averaging or weighted averaging method, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
[0024] Optionally, step S3 specifically includes:
[0025] S3.1 Construct an emergency resource classification model based on a convolutional neural network architecture. Utilize the powerful feature extraction capabilities of convolutional neural networks to automatically learn the local features of emergency resource texts. Set up multiple convolutional layers in the model structure, with each convolutional layer using a different size convolutional kernel to capture features of text fragments of different lengths, adapting to information of different granularities in emergency resource texts.
[0026] S3.2 Introduce an attention mechanism into the input layer or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model can focus on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy.
[0027] S3.3. Prepare labeled emergency resource text data and divide it into training, validation, and test sets according to a preset ratio. Using the cross-entropy loss function as the optimization objective, employ stochastic gradient descent or its variant optimization algorithm to iteratively train the model using the training set. Adjust the model parameters through the validation set and evaluate the performance on the test set. Finally, enable the model to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold range, thus completing model optimization.
[0028] Secondly, this invention provides an emergency resource classification system based on natural language processing, and the technical solution adopted to solve the above-mentioned technical problems is as follows:
[0029] An emergency resource classification system based on natural language processing, comprising:
[0030] The data preprocessing module is used to perform data cleaning, word segmentation, and part-of-speech tagging preprocessing operations on the input emergency resource text data in sequence.
[0031] The data feature extraction module is used to extract word frequency features from the preprocessed emergency resource text, weight the words in the text, use a pre-trained word vector model to obtain distributed word representations and combine them into text vectors, thereby capturing the semantic features of the emergency resource text.
[0032] The model building and training module is used to build an emergency resource classification model based on a convolutional neural network architecture, and to introduce an attention mechanism in the model input layer or intermediate layer; it is also used to train and optimize the emergency resource classification model using labeled emergency resource text data, so that the model outputs the probability distribution of emergency resource text belonging to different categories;
[0033] The classification prediction module is used to sequentially call the data preprocessing module and the data feature extraction module when acquiring new emergency resource text to obtain text vectors with the same training data format as the emergency resource classification model. Then, the emergency resource classification model is called to obtain the probability distribution of the text belonging to different categories, and the category with the highest probability is selected as the classification result.
[0034] Optionally, the data preprocessing modules involved specifically include:
[0035] The data cleaning unit is used to clean the input emergency resource text data to remove noise information;
[0036] The word segmentation processing unit is used to perform word segmentation processing on the cleaned emergency resource text data by selecting a word segmentation tool based on its language type.
[0037] Part-of-speech tagging (POS) units are used to tag each word after word segmentation with the help of natural language processing toolkits.
[0038] Preferably, for Chinese emergency resource text data, the word segmentation processing unit uses the Jieba word segmentation tool to segment the continuous text into individual words;
[0039] For emergency resource text data in English, the word segmentation unit uses spaces and punctuation marks to segment words.
[0040] Optionally, the data feature extraction module involved specifically includes:
[0041] The feature extraction unit is used to extract word frequency features of emergency resource text using the bag-of-words model. It includes: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary and the value of the dimension represents the frequency of the corresponding word.
[0042] The weighted processing unit is used to weight words based on their frequency of occurrence in the emergency resource text using the TF-IDF algorithm to highlight the uniqueness of the text.
[0043] The vector generation unit is used to map words to a low-dimensional vector space using a pre-trained word vector model, so that words with similar meanings are closer in the vector space. Then, by averaging or weighted averaging, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
[0044] Optionally, the model building and training modules involved may specifically include:
[0045] The model building unit is used to build an emergency resource classification model based on a convolutional neural network architecture, which automatically learns the local features of emergency resource text by utilizing the powerful feature extraction capabilities of convolutional neural networks.
[0046] The model setup unit is used to set up multiple convolutional layers in the model structure. Each convolutional layer uses a different size convolutional kernel to capture features of text fragments of different lengths, adapting to information of different granularities in emergency resource texts.
[0047] The attention introduction unit is used to introduce an attention mechanism into the input layer or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model focuses on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy.
[0048] The data preparation unit is used to prepare labeled emergency resource text data as training data.
[0049] The data partitioning unit is used to divide the training data into training set, validation set and test set according to a preset ratio;
[0050] The training optimization unit is used to optimize the model using the cross-entropy loss function, employing stochastic gradient descent or its variants. It iteratively trains the model using the training set, adjusts the model parameters using the validation set, and evaluates the performance on the test set. Ultimately, the model is able to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold, thus completing model optimization.
[0051] The emergency resource classification method and system based on natural language processing of the present invention have the following advantages compared with the prior art:
[0052] 1. This invention can automatically classify emergency resource data, saving time spent on manual data comparison and classification, and improving user work efficiency; it has the ability to continuously optimize and learn based on user feedback, further improving the accuracy of emergency resource classification; and it solves the problems of low efficiency, poor accuracy and insufficient generalization ability of existing emergency resource classification methods.
[0053] 2. This invention combines multiple feature extraction methods, especially the introduction of pre-trained word vectors and attention mechanisms, to capture the semantics and key information of emergency resource texts more comprehensively and accurately. This enables the emergency resource classification model to make more accurate classification decisions when faced with complex texts. Compared with traditional classification methods, the classification accuracy can be improved by 10%-20%.
[0054] 3. When dealing with the classification task of millions of emergency resource documents, the present invention can shorten the processing time by several times or even tens of times compared with manual annotation and classification systems based on complex rules, which greatly improves the efficiency of resource management and retrieval. Attached Figure Description
[0055] Appendix Figure 1 This is a flowchart of the method according to Embodiment 1 of the present invention;
[0056] Appendix Figure 2 This is a module connection block diagram of Embodiment 2 of the present invention. Detailed Implementation
[0057] To make the technical solution, the technical problem solved, and the technical effect of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments.
[0058] Example 1:
[0059] Reference Appendix Figure 1 This embodiment proposes an emergency resource classification method based on natural language processing, which includes the following steps:
[0060] S1. For the input emergency resource text data, perform preprocessing operations in sequence, including data cleaning, word segmentation, and part-of-speech tagging, specifically including:
[0061] S1.1 For the input emergency resource text data, perform data cleaning to remove noise information such as special characters, HTML tags, and redundant spaces; for example, for web page text containing HTML tags, use regular expressions to match and remove tags such as <html>, <body>, etc.
[0062] S1.2. For the cleaned emergency resource text data, select a word segmentation tool based on its language type and perform word segmentation processing; for example, for Chinese emergency resource text data, use the Jieba word segmentation tool to split the continuous text into individual words, and for English emergency resource text data, use spaces and punctuation marks to segment words;
[0063] S1.3. Using a natural language processing toolkit (such as NLTK), perform part-of-speech tagging on each word after word segmentation, such as noun, verb, adjective, etc.
[0064] The above steps help to extract features with greater classification value in the subsequent steps.
[0065] S2. Extract word frequency features from the preprocessed emergency resource text, weight the words in the text, and use a pre-trained word vector model (such as Word2Vec, GloVe) to obtain distributed word representations and combine them into text vectors to capture the semantic features of the emergency resource text. Specifically, this includes:
[0066] S2.1. Use the Bag-of-Words model to extract word frequency features from emergency resource texts, including: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary, and the value of the dimension represents the frequency of the corresponding word.
[0067] S2.2. Based on the frequency of each word in the emergency resource text, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used to weight the words to highlight the uniqueness of the text. The higher the TF-IDF value, the more important the word is in the current emergency resource text and the lower its frequency of occurrence in other texts, which can better highlight the unique characteristics of the emergency resource text. For example, a word "firefighting" that appears frequently in fire emergency handling documents but rarely appears in the entire document set will have a higher TF-IDF value.
[0068] S2.3. Using pre-trained word vector models (such as Word2Vec and GloVe), words are mapped to a low-dimensional vector space, so that words with similar meanings are closer in the vector space. Then, by using averaging or weighted averaging methods, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
[0069] S3. Construct an emergency resource classification model based on a convolutional neural network (CNN) architecture, introducing an attention mechanism into the model's input layer or intermediate layers; train and optimize the emergency resource classification model using labeled emergency resource text data, so that the model outputs the probability distribution of emergency resource text belonging to different categories, specifically including:
[0070] S3.1 Construct an emergency resource classification model based on a convolutional neural network (CNN) architecture. Utilize the powerful feature extraction capabilities of CNN to automatically learn the local features of emergency resource texts. Set up multiple convolutional layers in the model structure, with each convolutional layer using convolutional kernels of different sizes, such as 3-gram and 5-gram kernels, to capture features of text fragments of different lengths and adapt to information of different granularities in emergency resource texts.
[0071] S3.2 Introduce an attention mechanism into the input or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model can focus on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy. For example, for a document about "forest fire prevention emergency plan", the model can automatically assign higher attention weights to core words such as "forest fire prevention" and "plan".
[0072] S3.3. Prepare labeled emergency resource text data and divide it into training, validation, and test sets according to a preset ratio (usually 70%, 15%, and 15%). Using the cross-entropy loss function as the optimization objective, employ stochastic gradient descent (SGD) or its variants (such as Adagrad, Adadelta, etc.) to iteratively train the model using the training set. Adjust the model parameters using the validation set and evaluate the performance on the test set. Ultimately, the model should be able to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold, thus completing model optimization.
[0073] S4. When acquiring new emergency resource text, first execute steps S1 and S2 to obtain a text vector with the same format as the training data of the emergency resource classification model. Then, input the text vector into the emergency resource classification model optimized in step S3. The emergency resource classification model calculates the output results through forward propagation to obtain the probability distribution of the text belonging to different categories. Select the category with the highest probability as the classification result.
[0074] For example, if the model outputs that the probability of the text belonging to the "firefighting" category is 0.7, the probability of it belonging to the "medical" category is 0.2, and the probability of it belonging to the "flood control" category is 0.1, then the resource is classified as the "firefighting" category.
[0075] Example 2:
[0076] Reference Appendix Figure 2 This embodiment proposes an emergency resource classification system based on natural language processing, which includes:
[0077] The data preprocessing module is used to perform data cleaning, word segmentation, and part-of-speech tagging preprocessing operations on the input emergency resource text data in sequence.
[0078] The data feature extraction module is used to extract word frequency features from the preprocessed emergency resource text, weight the words in the text, and use pre-trained word vector models (such as Word2Vec and GloVe) to obtain distributed word representations and combine them into text vectors, thereby capturing the semantic features of the emergency resource text.
[0079] The model building and training module is used to build an emergency resource classification model based on a convolutional neural network (CNN) architecture, and introduces an attention mechanism in the model input layer or intermediate layer; it is also used to train and optimize the emergency resource classification model using labeled emergency resource text data, so that the model outputs the probability distribution of emergency resource text belonging to different categories;
[0080] The classification prediction module is used to sequentially call the data preprocessing module and the data feature extraction module when acquiring new emergency resource text to obtain text vectors with the same training data format as the emergency resource classification model. Then, the emergency resource classification model is called to obtain the probability distribution of the text belonging to different categories, and the category with the highest probability is selected as the classification result.
[0081] In this embodiment, the data preprocessing module specifically includes:
[0082] The data cleaning unit is used to clean the input emergency resource text data to remove noise information;
[0083] The word segmentation processing unit is used to perform word segmentation processing on the cleaned emergency resource text data by selecting a word segmentation tool based on its language type.
[0084] Part-of-speech tagging (POS) units are used to tag each word after word segmentation with the help of natural language processing toolkits.
[0085] It should be added that, for Chinese emergency resource text data, the word segmentation processing unit uses the Jieba word segmentation tool to divide the continuous text into individual words;
[0086] For emergency resource text data in English, the word segmentation unit uses spaces and punctuation marks to segment words.
[0087] In this embodiment, the data feature extraction module specifically includes:
[0088] The feature extraction unit is used to extract word frequency features of emergency resource text using the bag-of-words model. This includes: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary, and the value of the dimension represents the frequency of the corresponding word.
[0089] The weighted processing unit is used to weight words based on their frequency of occurrence in the emergency resource text using the TF-IDF (term frequency-inverse document frequency) algorithm to highlight the uniqueness of the text;
[0090] The vector generation unit is used to map words to a low-dimensional vector space using a pre-trained word vector model, so that words with similar meanings are closer in the vector space. Then, by averaging or weighted averaging, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
[0091] In this embodiment, the model building and training module specifically includes:
[0092] The model building unit is used to build an emergency resource classification model based on a convolutional neural network (CNN) architecture, which automatically learns the local features of emergency resource text by utilizing the powerful feature extraction capabilities of the convolutional neural network.
[0093] The model setup unit is used to set up multiple convolutional layers in the model structure. Each convolutional layer uses a different size convolutional kernel to capture features of text fragments of different lengths, adapting to information of different granularities in emergency resource texts.
[0094] The attention introduction unit is used to introduce an attention mechanism into the input layer or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model focuses on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy.
[0095] The data preparation unit is used to prepare labeled emergency resource text data as training data.
[0096] The data partitioning unit is used to divide the training data into training set, validation set and test set according to a preset ratio;
[0097] The training optimization unit is used to optimize the model using the cross-entropy loss function, employing stochastic gradient descent or its variants. It iteratively trains the model using the training set, adjusts the model parameters using the validation set, and evaluates the performance on the test set. Ultimately, the model is able to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold, thus completing model optimization.
[0098] In summary, the emergency resource classification method and system based on natural language processing of this invention can automatically classify emergency resource data, saving time spent on manual data comparison and classification, and improving user work efficiency; it has the ability to continuously optimize and learn based on user feedback, further improving the accuracy of emergency resource classification; and it solves the problems of low efficiency, poor accuracy, and insufficient generalization ability of existing emergency resource classification methods.
[0099] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An emergency resource classification method based on natural language processing, characterized in that, Includes the following steps: S1. For the input emergency resource text data, perform data cleaning, word segmentation and part-of-speech tagging preprocessing operations in sequence; S2. Extract the word frequency features of the preprocessed emergency resource text, weight the words in the text, use the pre-trained word vector model to obtain the distributed representation of words and combine them into text vectors to capture the semantic features of the emergency resource text. S3. Construct an emergency resource classification model based on a convolutional neural network architecture, and introduce an attention mechanism in the model's input layer or intermediate layer; use labeled emergency resource text data to train and optimize the emergency resource classification model so that the model outputs the probability distribution of emergency resource text belonging to different categories; S4. When acquiring new emergency resource text, first execute steps S1 and S2 to obtain a text vector with the same format as the training data of the emergency resource classification model. Then, input the text vector into the emergency resource classification model optimized in step S3. The emergency resource classification model calculates the output results through forward propagation to obtain the probability distribution of the text belonging to different categories. Select the category with the highest probability as the classification result.
2. The emergency resource classification method based on natural language processing according to claim 1, characterized in that, Step S1 specifically includes: S1.
1. For the input emergency resource text data, perform data cleaning to remove noise information; S1.
2. For the cleaned emergency resource text data, select a word segmentation tool based on its language type and perform word segmentation processing; S1.
3. Using a natural language processing toolkit, perform part-of-speech tagging on each word after word segmentation.
3. The emergency resource classification method based on natural language processing according to claim 2, characterized in that, Perform step S1.2: For Chinese emergency resource text data, the Jieba word segmentation tool is used to segment continuous text into individual words; For emergency resource text data in English, words are segmented using spaces and punctuation marks.
4. The emergency resource classification method based on natural language processing according to claim 2, characterized in that, Step S2 specifically includes: S2.
1. Use the bag-of-words model to extract word frequency features from emergency resource texts, including: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary, and the value on the dimension represents the frequency of the corresponding word. S2.
2. Based on the frequency of each word in the emergency resource text, the TF-IDF algorithm is used to weight the words to highlight the uniqueness of the text; S2.
3. Using a pre-trained word vector model, words are mapped to a low-dimensional vector space, so that words with similar meanings are closer in the vector space. Then, by using an averaging or weighted averaging method, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
5. The emergency resource classification method based on natural language processing according to claim 4, characterized in that, Step S3 specifically includes: S3.1 Construct an emergency resource classification model based on a convolutional neural network architecture. Utilize the powerful feature extraction capabilities of convolutional neural networks to automatically learn the local features of emergency resource texts. Set up multiple convolutional layers in the model structure, with each convolutional layer using a different size convolutional kernel to capture features of text fragments of different lengths, adapting to information of different granularities in emergency resource texts. S3.2 Introduce an attention mechanism into the input layer or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model can focus on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy. S3.
3. Prepare labeled emergency resource text data and divide it into training, validation, and test sets according to a preset ratio. Using the cross-entropy loss function as the optimization objective, employ stochastic gradient descent or its variant optimization algorithm to iteratively train the model using the training set. Adjust the model parameters through the validation set and evaluate the performance on the test set. Finally, enable the model to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold range, thus completing model optimization.
6. An emergency resource classification system based on natural language processing, characterized in that, It includes: The data preprocessing module is used to perform data cleaning, word segmentation, and part-of-speech tagging preprocessing operations on the input emergency resource text data in sequence. The data feature extraction module is used to extract word frequency features from the preprocessed emergency resource text, weight the words in the text, use a pre-trained word vector model to obtain distributed word representations and combine them into text vectors, thereby capturing the semantic features of the emergency resource text. The model building and training module is used to build an emergency resource classification model based on a convolutional neural network architecture, and to introduce an attention mechanism in the model input layer or intermediate layer; it is also used to train and optimize the emergency resource classification model using labeled emergency resource text data, so that the model outputs the probability distribution of emergency resource text belonging to different categories; The classification prediction module is used to sequentially call the data preprocessing module and the data feature extraction module when acquiring new emergency resource text to obtain text vectors with the same training data format as the emergency resource classification model. Then, the emergency resource classification model is called to obtain the probability distribution of the text belonging to different categories, and the category with the highest probability is selected as the classification result.
7. An emergency resource classification system based on natural language processing according to claim 6, characterized in that, The data preprocessing module specifically includes: The data cleaning unit is used to clean the input emergency resource text data to remove noise information; The word segmentation processing unit is used to perform word segmentation processing on the cleaned emergency resource text data by selecting a word segmentation tool based on its language type. Part-of-speech tagging (POS) units are used to tag each word after word segmentation with the help of natural language processing toolkits.
8. An emergency resource classification system based on natural language processing according to claim 7, characterized in that, For Chinese emergency resource text data, the word segmentation processing unit uses the Jieba word segmentation tool to segment continuous text into individual words; For emergency resource text data in English, the word segmentation unit uses spaces and punctuation marks to segment words.
9. An emergency resource classification system based on natural language processing according to claim 7, characterized in that, The data feature extraction module specifically includes: The feature extraction unit is used to extract word frequency features of emergency resource text using the bag-of-words model. It includes: constructing a vocabulary containing all words, counting the frequency of each word in the emergency resource text, and representing the emergency resource text as a vector, where each dimension of the vector corresponds to a word in the vocabulary and the value of the dimension represents the frequency of the corresponding word. The weighted processing unit is used to weight words based on their frequency of occurrence in the emergency resource text using the TF-IDF algorithm to highlight the uniqueness of the text. The vector generation unit is used to map words to a low-dimensional vector space using a pre-trained word vector model, so that words with similar meanings are closer in the vector space. Then, by averaging or weighted averaging, the word vectors of all words in the emergency resource text are combined into a text vector to capture the semantic features of the emergency resource text.
10. An emergency resource classification system based on natural language processing according to claim 9, characterized in that, The model building and training module specifically includes: The model building unit is used to build an emergency resource classification model based on a convolutional neural network architecture, which automatically learns the local features of emergency resource text by utilizing the powerful feature extraction capabilities of convolutional neural networks. The model setup unit is used to set up multiple convolutional layers in the model structure. Each convolutional layer uses a different size convolutional kernel to capture features of text fragments of different lengths, adapting to information of different granularities in emergency resource texts. The attention introduction unit is used to introduce an attention mechanism into the input layer or intermediate layer of the emergency resource classification model. By calculating the attention weight of each word in the text, the model focuses on the core words that play a key role in classification during the processing, thereby enhancing the model's ability to capture important information and improving the classification accuracy. The data preparation unit is used to prepare labeled emergency resource text data as training data. The data partitioning unit is used to divide the training data into training set, validation set and test set according to a preset ratio; The training optimization unit is used to optimize the model using the cross-entropy loss function, employing stochastic gradient descent or its variants. It iteratively trains the model using the training set, adjusts the model parameters using the validation set, and evaluates the performance on the test set. Ultimately, the model is able to repeatedly output the probability distribution of emergency resource text belonging to different categories with errors within a preset threshold, thus completing model optimization.
Citation Information
Patent Citations
A text classification method based on a bidirectional cyclic attention neural network
CN109472024A
Text classification method and device, computer device and storage medium
CN109857860A
Medical text classification method and device based on ATT-CN
CN113449106A
Text classification system based on natural language processing
CN114547305A
Text classification method and device based on natural language
CN116821801A