A structured data generation model based on natural language processing
Through the keyword extraction, similarity comparison and scoring module based on the BERT language model, combined with the RNN and XLNet models, the data bias and insufficient interpretability in the structured data generation model are solved, and efficient and accurate data analysis and report generation are achieved.
Patent Information
- Application Number
- CN202410854943.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-28
AI Technical Summary
The existing structured data generation model of natural language processing has problems of data bias, unfairness and insufficient interpretability.
The keyword extraction, similarity comparison and scoring module based on the BERT language model is adopted, combined with the RNN recurrent neural network and the XLNet natural regression model, the documents are classified, identified, preprocessed and generated structured data, and high-quality reports are generated using emotional tendency analysis and credibility judgment.
It improves the real-time and dynamic nature of semantic understanding, reduces data bias and unfairness, reduces the time cost and legal risks of manual repetitive labor, and improves the accuracy and efficiency of data analysis.
Smart Images

Figure CN118940719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and specifically to a structured data generation model based on natural language processing. Background Art
[0002] Natural language processing (NLP) is an interdisciplinary field of artificial intelligence and linguistics, which studies various theories and methods that can enable effective communication between humans and computers in natural language; the applications of natural language processing include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR, etc.; the applications of natural language processing are very extensive, including intelligent customer service, search engine optimization, social media analysis, intelligent writing assistants, etc.
[0003] The application of natural language processing in management and production is extensive and in-depth, which provides enterprises with more efficient and accurate data processing and analysis tools, thereby optimizing management and production processes; for example, natural language processing technology can perform sentiment analysis on text data such as corporate brand reputation, product reviews, and customer feedback. This analysis helps enterprises to timely understand market feedback, improve and optimize the company, thereby improving customer satisfaction and brand loyalty; at the same time, natural language processing can automatically extract specific types of information from text, such as human resources information, marketing information, economic data, etc. This enables enterprises to obtain the latest and most comprehensive information in the first time, greatly improving the decision-making efficiency.
[0004] Although NLP has extensive applications in management and production, there are still some challenges and limitations, such as limitations in semantic understanding, language diversity issues, data bias and unfairness, lack of interpretability, etc. Therefore, when enterprises apply NLP technology, they need to fully consider these factors and make selections and adjustments in combination with their actual situations. Summary of the Invention
[0005] The purpose of the present invention is to propose a structured data generation model based on natural language processing in order to solve problems such as data bias and unfairness, lack of interpretability, etc. existing in the existing structured data generation model of natural language processing.
[0006] The purpose of the present invention can be achieved by the following technical solutions: a structured data generation model based on natural language processing, including a business subsystem, a customer service platform subsystem, a demand retrieval module, and a database; wherein the business subsystem includes a keyword extraction module, a similarity comparison module, and a scoring module, and the customer service platform subsystem includes a text preprocessing module, an identification processing module, a credibility determination module, and a report generation module.
[0007] The keyword extraction module classifies the documents, uses the BERT language model to identify the entire document, captures type judgment elements in the document. The type judgment elements include keywords such as "rights", "fees", "terms", "confidentiality", "validity", etc. and the change images in the change order. Record the occurrence frequency N of various type judgment elements in the entire document. Each type judgment element corresponds to a preset judgment weight λ. Multiply the occurrence times of various type judgment elements by the judgment weight and then sum to obtain the type judgment parameter α. Each document type corresponds to a specific type judgment parameter range. The document types include tender documents, response documents, business contracts, change orders, and administrative contracts.
[0008] Each document type uses the corresponding regular expression to extract keywords from the document, forms the keywords into a data vector, and generates structured data.
[0009] For example, extract keywords about "time", "project name", "construction scale", and "contract estimated price" from the tender documents, extract keywords about certificate requirements, professional title requirements, qualification requirements, and reputation requirements in the bidder qualification section, extract keywords about "the highest bid price compiled by the tenderer", and form all keywords into a data vector. The keyword extraction module scans several tender documents, generates several data vectors, and forms a structured data table.
[0010] The similarity comparison module compares the documents with the same document type with each other in full text, marks the words with semantic differences in the relatively similar paragraphs as different characters, uses the BERT language model to analyze the meaning of the different characters and match a reference character, generates the meaning vector X=(x1, x2, x3,..., xn) of the different characters and the meaning vector Y=(y1, y2, y3,..., yn) of the reference character, where n represents the number of dimensions included in the coordinate system where the meaning vector is located, and the value of n depends on the learning depth of the BERT language model; analyze the part-of-speech of the different characters, and the part-of-speech includes proper nouns, common nouns, verbs, adjectives, modal particles, function words, etc. Each part-of-speech corresponds to a preset part-of-speech factor e1; analyze the location where the different words are located, and the location includes document directory, announcement chapter, appendix chapter, instructions chapter, general rules, main text, etc. The BERT language model uses the RNN recurrent neural network to learn to obtain the accurate location category, and each location corresponds to a location factor e2; analyze the context and context where the different words are located, and obtain a context factor e3 based on the analysis of the BERT language model. Through the formula obtain the difference score E, where λ1, λ2, λ3, λ4 are preset weight factors, highlight the different characters with the difference score E greater than 100, and add the difference score E to the data vector itself.
[0011] The scoring module analyzes based on the data vectors generated by the keyword extraction module. For each element in the data vectors, preset scoring factors are set. For example, for the project name element, preset scoring factor one. When the element contains city A, the value of scoring factor one is a1; when the element contains city B, the value of scoring factor one is a2. For the contract estimated price element, preset scoring factor two. When the element is greater than 1 million, the value of preset scoring factor two is b2; when the element is less than 1 million, the value of preset scoring factor two is b2, and so on. Sum all the scoring factors to obtain the scoring parameter T, and add the scoring parameter of each data vector to the data vector itself to become a new element, completing the first step of scoring. The scoring module, based on the difference scores of each data vector obtained by the similarity comparison module, sums the difference scores to obtain the difference parameter K, and adds the difference parameter of each data vector to the data vector itself to become a new element, completing the second step of scoring.
[0012] After completing the first and second steps of scoring, the scoring module inputs all the data vectors into the database.
[0013] The text preprocessing module preprocesses the customer service feedback information such as petitions, complaints, etc. received by the customer service platform subsystem. First, it obtains the information source of the customer service feedback information, obtains the name of the sender and the sending time, and encrypts the name of the sender using the AES encryption algorithm. Subsequently, each Chinese character and English character in the customer service feedback information is split to form separate strings, and the XLNet natural regression pre-trained language model is used to obtain the word formation probability P of each individual character and the words formed by the two adjacent characters before and after it. When the sum of the word formation probabilities of two adjacent characters is greater than the preset probability threshold , the two characters are combined into a word, and the meaning vector of the word is obtained in the language model. If the meaning vector of the word belongs to the category of slanderous and extremely negative adjectives, the word is marked as useless information. The number of characters contained in the useless information is recorded as the information shielding parameter k, and the characters of the useless information are replaced with *. Then, the preprocessed customer service feedback information is sent to the recognition and processing module in the form of a string.
[0014] The recognition processing module takes the string sent by the text preprocessing module, obtains forum replies and comments on the Internet as training data for training, analyzes the sentiment tendency contained in each sentence based on natural language processing (NLP) technology to form a refined sentiment dictionary, divides the sentiment tendency into three first-level sentiment tendencies: positive, negative, and neutral. Each first-level sentiment tendency contains refined second-level sentiment tendencies. For example, the positive sentiment tendency can be subdivided into gratitude, excitement, hope, etc. Each sentiment tendency corresponds to a preset sentiment degree parameter. The sentiment degree parameter of the positive sentiment tendency is a positive number, the sentiment degree parameter of the negative sentiment tendency is a negative number, and the sentiment degree parameter of the neutral sentiment tendency is 0. The recognition processing module extracts the meaning vectors of each word in the string, understands the part of speech, meaning, and sentiment tendency of the words based on the BERT language model, obtains the sentiment degree parameter of each word, and sums up the sentiment degree parameters of each customer service feedback message to obtain the sentiment index k1.
[0015] The recognition module obtains the key degree parameter of each word in each customer service feedback message based on the BERT language model. The key degree parameters of proper nouns and words such as "suggestion", "best", "optimize", etc. are the highest. The recognition module sums up all the key degree parameters of each customer service feedback message to obtain the key degree index k2. Subsequently, the recognition module outputs k, k1, and k2 to the credibility determination module.
[0016] The credibility determination module uses the formula to obtain the credibility index F of the customer service feedback message, and comprehensively combines the encrypted sender's name Na, the sending time Ti, generates the original text Te of the customer service feedback message and the credibility index F to generate the data vector (Na, Ti, Te, F) of the customer service feedback message and sends it to the report generation module.
[0017] After receiving the data vector of the customer service feedback message, the report generation module extracts the credibility index F of each data vector and conducts targeted analysis on the data vectors whose credibility index F is greater than the threshold. Subsequently, the report generation module understands the original text of the data vector based on the BERT language model, extracts the problems raised in the original text, records them as value problem data and adds them to the data vector itself.
[0018] The report generation module obtains the information report of the customer service platform on the Internet as a report sample, annotates and trains a large number of report samples based on the BERT language model, and learns how to generate high-quality reports according to the extracted information. After the model training is completed, the generation of data reports can begin. Based on the NLP method, it can be integrated and output according to the predefined templates and rules, combined with the elements in the data vector, and automatically generate a report that meets the specifications. The generated report can cover various data indicators, analysis results, suggestions, and other contents.
[0019] The report generation module evaluates the metrics of the generated report by comparing the generated report with the standard report template, including the accuracy metric h1, the integrity metric h2, and the readability metric h3 of the report. The comprehensive evaluation metric H is obtained by performing a weighted sum on the metrics h1, h2, and h3. When the comprehensive evaluation metric is less than the preset threshold, the report is regenerated; otherwise, the report is added as an element to the data vector itself. Finally, the report generation module sends all the data vectors to the database.
[0020] After receiving the data vector sent by the business subsystem, the database adds a new element, the number 1, in front of the data vector; after receiving the data vector sent by the customer service platform subsystem, the database adds a new element, the number 2, in front of the data vector, and archives all the data vectors for retrieval and extraction by the demand retrieval module at any time.
[0021] After receiving the query statement, the demand retrieval module understands the meaning of each word in the query instruction, extracts the keywords and performs word substitution processing according to the similarity. It retrieves in the database and then extracts the elements in the data vector for display at the display end. For example, when inputting the query instruction "What are the projects that required a registered constructor certificate at or above the second level in the water conservancy and hydropower engineering major in the tender documents yesterday" to the demand retrieval module, based on the BERT language model, it understands that the requirement of this instruction is a query, and then extracts the keywords "yesterday", "tender documents", and "registered constructor certificate at or above the second level in the water conservancy and hydropower engineering major". First, it obtains the date of the query, x month x day, 20xx, through the Internet, and calculates one day back to get the first retrieval element: the retrieval word "y year y month y day" in the time element. Then, it identifies the second retrieval element from the keyword "registered constructor certificate at or above the second level in the water conservancy and hydropower engineering major": the retrieval word "at or above the second level in the water conservancy and hydropower engineering major" in the positive number requirement. Then it retrieves in the database and outputs the element "project name" in the data vector that meets the first and second retrieval elements to the display end. The above-disclosed preferred embodiments of the present invention are only used to help illustrate the present invention. The preferred embodiments do not elaborate on all the details, nor do they limit the invention to only the specific implementation manners.
[0022] Compared with the prior art, the beneficial effects of the present invention are:
[0023] 1. A structured data generation model based on natural language processing trains the language model based on a neural network, and at the same time captures the context information of words in semantic understanding, improving the real-time and dynamic nature of the system's data analysis method and reducing the limitations of semantic understanding;
[0024] 2. A structured data generation model based on natural language processing captures the data, fuses and analyzes the data, and analyzes and quantifies the credibility in special contexts, reducing data deviation and the unfairness of data analysis.
[0025] 3. A structured data generation model based on natural language processing identifies similar paragraphs in a document through semantic analysis, highlights the tampered clauses and fields, reducing the time cost wasted by manual labor in mechanical repetitive work and also reducing the legal risks caused by human errors, and solving the inefficiency problem of repetitive labor. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings.
[0027] Figure 1 It is a principle block diagram of the present invention.
[0028] Figure 2 It is a schematic diagram of an embodiment of the programming language of the regular expression in the example of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0029] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0030] The usage scenario of this system is preset for the analysis of a large amount of information in business and customer service work in production and life.
[0031] Please refer to Figure 1 As shown, a structured data generation model based on natural language processing includes a business subsystem, a customer service platform subsystem, a requirement retrieval module, and a database; wherein the business subsystem includes a keyword extraction module, a similarity comparison module, and a scoring module, and the customer service platform subsystem includes a text preprocessing module, an identification processing module, a credibility determination module, and a report generation module.
[0032] The business subsystem is used to process various types of documents such as tender documents, response documents, business contracts, change orders, and administrative reports in daily business work. The business subsystem identifies, checks, and extracts keywords from the whole document; interprets the meaning of the document, obtains key fields, and enters them into the structured database through natural language processing. The business subsystem includes a keyword extraction module, a similarity comparison module, and a scoring module.
[0033] The keyword extraction module classifies the documents, uses the BERT language model to identify the whole document, captures the type judgment elements in the document. The type judgment elements include keywords such as "rights", "fees", "terms", "confidentiality", "validity", etc. and the change images in the change order. Record the occurrence frequency N of various type judgment elements in the whole document. Each type judgment element corresponds to a preset judgment weight λ. Multiply the occurrence times of various type judgment elements by the judgment weight and then sum to obtain the type judgment parameter α. Each document type corresponds to a specific type judgment parameter range. The document types include tender documents, response documents, business contracts, change orders and administrative contracts.
[0034] Please refer to Figure 2 As shown, each document type uses the corresponding regular expression to extract the keywords in the document, forms the keywords into a data vector, and generates structured data.
[0035] For example, in the tender documents, extract the keywords about "time", "project name", "construction scale" and "contract estimated price", extract the keywords about certificate requirements, professional title requirements, qualification requirements and reputation requirements in the bidder qualification section, extract the keywords about "the highest bid price prepared by the tenderer", and form all the keywords into a data vector. The keyword extraction module scans several tender documents, generates several data vectors, and forms a structured data table.
[0036] The similarity comparison module compares the documents with the same document type with each other in full text, marks the words with semantic differences in the relatively similar paragraphs as different characters, uses the BERT language model to analyze the meaning of the different characters and match a reference character, generates the meaning vector X=(x1, x2, x3,..., xn) of the different characters and the meaning vector Y=(y1, y2, y3,..., yn) of the reference character, where n represents the number of dimensions included in the coordinate system where the meaning vector is located, and the value of n depends on the learning depth of the BERT language model; analyze the part-of-speech of the different characters, and the part-of-speech includes proper nouns, common nouns, verbs, adjectives, modal particles, function words, etc. Each part-of-speech corresponds to a preset part-of-speech factor e1; analyze the location where the different words are located. The location includes document directory, announcement chapter, appendix chapter, instructions chapter, general rules, main text, etc. The BERT language model uses the RNN recurrent neural network to learn to obtain the accurate location category, and each location corresponds to a location factor e2; analyze the context and context of the different words, and obtain a context factor e3 based on the analysis of the BERT language model. Through the formula Get the difference score E, where λ1, λ2, λ3, λ4 are preset weight factors, highlight the different characters with the difference score E greater than 100, and add the difference score E to the data vector itself.
[0037] The scoring module analyzes based on the data vectors generated by the keyword extraction module. For each element in the data vectors, preset scoring factors are set. For example, for the project name element, preset scoring factor one. When the element contains city A, the value of scoring factor one is a1, and when the element contains city B, the value of scoring factor one is a2; for the contract estimated price element, preset scoring factor two. When the element is greater than 1 million, the preset value of scoring factor two is b2, and when the element is less than 1 million, the preset value of scoring factor two is b2, and so on; sum all the scoring factors to obtain the scoring parameter T, and add the scoring parameter of each data vector to the data vector itself to become a new element, completing the first step of scoring; based on the difference scores of each data vector obtained by the similarity comparison module, the scoring module sums the difference scores to obtain the difference parameter K, and adds the difference parameter of each data vector to the data vector itself to become a new element, completing the second step of scoring.
[0038] After completing the first and second steps of scoring, the scoring module inputs all the data vectors into the database.
[0039] The text preprocessing module preprocesses the customer service feedback information such as petitions, complaints, and grievances received by the customer service platform subsystem. First, it obtains the information source of the customer service feedback information, obtains the name of the sender and the sending time, and encrypts the name of the sender using the AES encryption algorithm. Subsequently, each Chinese character and English character in the customer service feedback information is split to form separate strings, and the XLNet natural regression pre-trained language model is used to obtain the word formation probability P of each individual character and the two adjacent characters before and after it forming a word. When the sum of the word formation probabilities of two adjacent characters is greater than the preset probability threshold , the two characters are combined into a word, and the meaning vector of the word is obtained in the language model. If the meaning vector of the word belongs to the category of slanderous and extremely negative adjectives, the word is marked as useless information. The number of characters contained in the useless information is recorded as the information shielding parameter k, and the characters of the useless information are replaced with *. Then, the preprocessed customer service feedback information is sent to the recognition and processing module in the form of a string.
[0040] The recognition processing module takes the string sent by the text preprocessing module, obtains forum replies and comments on the Internet as training data for training, analyzes the emotional tendencies contained in each sentence based on natural language processing (NLP) technology to form a refined emotional dictionary, divides the emotional tendencies into three primary emotional tendencies: positive, negative, and neutral. Each primary emotional tendency contains refined secondary emotional tendencies. For example, the positive emotional tendency can be subdivided into gratitude, excitement, hope, etc. Each emotional tendency corresponds to a preset emotional degree parameter. The emotional degree parameter of the positive emotional tendency is a positive number, the emotional degree parameter of the negative emotional tendency is a negative number, and the emotional degree parameter of the neutral emotional tendency is 0. The recognition processing module extracts the meaning vectors of each word in the string, understands the part of speech, meaning, and emotional tendency of the words based on the BERT language model, obtains the emotional degree parameter of each word, and sums up the emotional degree parameters of each customer service feedback message to obtain the emotional index k1.
[0041] The recognition module obtains the key degree parameter of each word in each customer service feedback message based on the BERT language model. The key degree parameters of proper nouns and words such as "suggestion", "best", "optimize", etc. are the highest. The recognition module sums up all the key degree parameters of each customer service feedback message to obtain the key degree index k2. Subsequently, the recognition module outputs k, k1, and k2 to the credibility determination module.
[0042] The credibility determination module uses the formula to obtain the credibility index F of the customer service feedback message, and comprehensively combines the encrypted sender's name Na, the sending time Ti, generates the original text Te of the customer service feedback message and the credibility index F to generate the data vector (Na, Ti, Te, F) of the customer service feedback message and sends it to the report generation module.
[0043] After receiving the data vector of the customer service feedback message, the report generation module extracts the credibility index F of each data vector and conducts targeted analysis on the data vectors whose credibility index F is greater than the threshold. Subsequently, the report generation module understands the original text of the data vector based on the BERT language model, extracts the problems raised in the original text, records them as valuable problem data and adds them to the data vector itself.
[0044] The report generation module obtains the information report of the customer service platform on the Internet as a report sample, annotates and trains a large number of report samples based on the BERT language model, and learns how to generate high-quality reports according to the extracted information. After the model training is completed, the generation of data reports can begin. Based on the NLP method, it can be integrated and output according to the predefined templates and rules, combined with the elements in the data vector, and automatically generate a report that meets the specifications. The generated report can cover various data indicators, analysis results, suggestions, and other contents.
[0045] The report generation module evaluates the metrics of the generated report by comparing the generated report with the standard report template, including the accuracy metric h1, the integrity metric h2, and the readability metric h3 of the report. The comprehensive evaluation metric H is obtained by weighted summation of the metrics h1, h2, and h3. When the comprehensive evaluation metric is less than the preset threshold, the report is regenerated. Otherwise, the report is added as an element to the data vector itself. Finally, the report generation module sends all the data vectors to the database.
[0046] After receiving the data vector sent by the business subsystem, the database adds a new element, the number 1, in front of the data vector; after receiving the data vector sent by the customer service platform subsystem, the database adds a new element, the number 2, in front of the data vector, and archives all the data vectors for retrieval and extraction by the demand retrieval module at any time.
[0047] After receiving the query statement, the demand retrieval module understands the meanings of the words in the query instruction, extracts the keywords and performs word substitution processing according to the similarity, retrieves in the database, and then extracts the elements in the data vector for display on the display end. For example, when inputting the query instruction "Which projects required a registered constructor certificate of level 2 or above in the water conservancy and hydropower engineering major in the tender documents yesterday" to the demand retrieval module, based on the BERT language model, it understands that the requirement of this instruction is a query, and then extracts the keywords "yesterday", "tender documents", and "registered constructor certificate of level 2 or above in the water conservancy and hydropower engineering major". First, it obtains the date of the query, x month x day, 20xx, through the Internet, and calculates one day back to get the first retrieval element: the retrieval word "20y year y month y day" in the time element. Then, it identifies the second retrieval element from the keyword "registered constructor certificate of level 2 or above in the water conservancy and hydropower engineering major": the retrieval word "level 2 or above in the water conservancy and hydropower engineering major" in the positive requirement. Then it retrieves in the database and outputs the element "project name" in the data vector that meets the first and second retrieval elements to the display end. The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not elaborate on all the details, nor do they limit the present invention to the specific embodiments only.
[0048] Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the relevant technical fields can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A structured data generation model based on natural language processing, comprising a keyword extraction module, a similarity comparison module, a scoring module, a text preprocessing module, an identification processing module, a credibility determination module, and a report generation module, characterized in that: The keyword extraction module classifies the documents received by the business subsystem, uses the BERT language model to identify the entire document, and captures type judgment elements in the document to determine the document type; Use the BERT language model to identify the entire document, capture type judgment elements in the document, record the occurrence frequencies of various type judgment elements in the entire document, each type judgment element corresponds to a preset judgment weight, calculate the value of the type judgment parameter and determine the document type; for each document type, use the corresponding regular expression to extract the keywords in the document, form the keywords into a data vector, and generate structured data; The similarity comparison module compares the full texts of documents with the same document type to mark the different characters, obtains a difference score based on the characteristic analysis of the different characters, and highlights the characters with a difference score greater than the threshold; Compare the full texts of documents with the same document type, mark the words with different semantics in the relatively similar paragraphs as different characters, use the BERT language model to analyze the meaning of the different characters and match a reference character, generate the meaning vector of the different characters and the meaning vector of the reference character; analyze the part-of-speech of the different characters to obtain the preset part-of-speech factor; Analyze the position where the different words are located to obtain the corresponding position factor; analyze the context and the context where the different words are located to obtain a context factor, calculate the difference score through operations, highlight the different characters with a difference score greater than the preset threshold, and add the difference score to the data vector itself; The scoring module scores the document quality based on the keywords and the difference score, and then generates structured data from all the data of the business subsystem and sends it to the database; The text preprocessing module splits the Chinese characters, encrypts sensitive information, and shields useless information in the customer service feedback information received by the customer service platform subsystem; The identification processing module quantifies the credibility index in the customer service feedback information based on the language model; The credibility determination module calculates the credibility index of the customer service feedback information through the credibility index, combines the encrypted sender's name, sending time, the original text of the customer service feedback information, and the credibility index, generates a data vector of the customer service feedback information and sends it to the report generation module; The report generation module extracts the keywords in the original text of the customer service feedback information and generates a data report for the original text, and then generates structured data from all the data of the customer service platform subsystem and sends it to the database.
2. The structured data generation model based on natural language processing according to claim 1, wherein: It further includes a requirement retrieval module and a database; The database is used to store the structured data sent by the business subsystem and the customer service platform subsystem; The requirement retrieval module uses natural language processing technology to analyze the proposed retrieval requirements, obtains the structured data from the database and displays it.
3. A structured data generation model based on natural language processing according to claim 1, characterized in that, The specific process of the scoring module scoring the document quality based on the keywords and the difference score is as follows: Analyze the data vectors generated by the keyword extraction module. Preset a scoring factor for each element in the data vector, match the elements to obtain the scoring factors, sum all the scoring factors to obtain a scoring parameter, and add the scoring parameter of each data vector to the data vector itself to become a new element, completing the first step of scoring; based on the difference scores of each data vector obtained by the similarity comparison module, sum the difference scores to obtain a difference parameter, and add the difference parameter of each data vector to the data vector itself to become a new element, completing the second step of scoring; after the first step and the second step of scoring are completed, input all the data vectors into the database.
4. A structured data generation model based on natural language processing according to claim 1, characterized in that, The specific process of the text preprocessing module for Chinese character splitting, sensitive information encryption, and useless information masking is as follows: First, obtain the information source of the customer service feedback information, obtain the name of the sender and the delivery time, encrypt the name of the sender using the AES encryption algorithm, and then split each Chinese character and English character in the customer service feedback information into separate strings. Use the XLNet natural regression pre-trained language model to obtain the word formation probability of each individual character and the two adjacent characters before and after it forming a word. When the sum of the word formation probabilities of two adjacent characters is greater than the preset probability threshold, form the two characters into a word and obtain the meaning vector of the word in the language model. If the meaning vector of the word belongs to the category of slanderous or extremely negative adjectives, mark the word as useless information. Record the number of characters contained in the useless information as the information masking parameter, and replace the characters of the useless information with *.
5. A structured data generation model based on natural language processing according to claim 1, characterized in that, The process of the recognition processing module for quantifying the credibility index in the customer service feedback information based on the language model is as follows: The recognition processing module analyzes the sentiment tendency parameters contained in each sentence of the string sent by the text preprocessing module based on the natural processing NLP technology, and sums the sentiment degree parameters of each customer service feedback information to obtain a sentiment index; based on the BERT language model, obtain the key degree parameters of each word in each customer service feedback information, and sum all the key degree parameters of each customer service feedback information to obtain a key degree index; Subsequently, the recognition module outputs the sentiment tendency parameters, key degree index, and information masking parameter to the credibility determination module.
6. A structured data generation model based on natural language processing according to claim 1, wherein The specific process of the report generation module for extracting keywords from the original text of the customer service feedback information and generating a data report for the original text is as follows: Understand the original text of the data vector based on the BERT language model, extract the questions raised in the original text, record them as value question data and add them to the data vector itself. Automatically generate a report that meets the specifications based on the NLP method, add the report that meets the specifications as an element to the data vector itself. Finally, the report generation module sends all the data vectors to the database.
7. A structured data generation model based on natural language processing according to claim 2, characterized in that, The specific process of the database for saving the structured data sent by the business subsystem and the customer service platform subsystem and the requirement retrieval module for obtaining and displaying the structured data from the database is as follows: After receiving the data vector sent by the business subsystem, the database adds a marker element one in front of the data vector; after receiving the data vector sent by the customer service platform subsystem, it adds a marker element two in front of the data vector, and archives all data vectors for retrieval and access by the demand retrieval module at any time; After receiving the query statement, the demand retrieval module understands the meaning of each word in the query instruction, extracts the keywords and performs word substitution processing according to the similarity, retrieves in the database, and then extracts elements in the data vector for display on the display terminal.
Citation Information
Patent Citations
Method and system for identifying text difference content in combination with semantic recognition
CN113051869A
Method and system for processing investment and research report in futures field
CN115358201A
Customer marketing scene data analysis system and method based on NLP algorithm
CN116663664A
Short video content label knowledge base quick retrieval method based on natural language processing
CN117009461A