Data processing methods, devices, and electronic equipment based on natural language models
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]现有的工单处理系统大多依赖于人工操作和简单的自动化流程,缺乏智能化的支持
[0015]本申请提出的基于自然语言模型的数据处理方法、装置、电子设备及存储介质,首先,获取金融数据信息以及待解答问题信息,对金融数据信息进行知识表示,从而实现对金融数据信息的精准分析,以创建与金融数据信息中问题信息对应的第一数据结构和第二数据结构,实现对常见问题和关联问题的数据结构的创建,再根据第一数据结构和第二数据结构创建保险数据集,实现对专业数据库的构建,使得保险数据集内能够融合多种问题信息,便于后续对问题的全面检索,之后,通过保险数据集训练预设的自然语言模型,使得自然语言模型能够自动化处理大量文本数据,得到自动问答语言模型,从而能够减少人工干预,响应于用户检索指令,对待解答问题进行文本分析,从而分析出待解答问题信息中需要进行检索的关键词,得到分析结果,实现对待解答问题信息的精准检索,提高检索准确性,再基于自动问答语言模型以及保险数据集对分析结果进行答案检索,实现对待解答问题信息的全面检索,得到检索集合,能够提高问题处理的效率和服务质量,当检索集合为空或者检索集合中不存在预设的答案语句,说明用户输入的待解答问题信息没有被解决,则需要将用户检索指令转变为转人工服务指令,从而能够人工回复待解答问题信息,避免出现解答与客户预期不一致的情况,实现对待解答问题信息的精准回答。本申请实施例通过对金融数据信息进行知识表示来创建多个数据结构,再通过多个数据结构创建保险数据集,以及训练自然语言模型,从而能够实现对待解答问题信息的自动问答,减少人工投入,能够提高问题处理的效率,并且在检索集合为空或者检索集合中不存在预设的答案语句的情况下,本申请实施例会进行转人工操作,从而避免出现解答与客户预期不一致的情况,实现对待解答问题信息的精准回答。
Smart Images

Figure CN119149702B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular to a data processing method, apparatus, electronic device and medium based on a natural language model. Background Technology
[0002] In the daily operations of the insurance industry, the stable operation of business systems is crucial for ensuring customer service quality and improving business processing efficiency. However, with the increasing complexity of insurance business and the continuous upgrading of technology systems, business system problems occur frequently. These problems are usually submitted to the system operations team in the form of work orders for processing.
[0003] Existing work order processing systems mostly rely on manual operation and simple automated processes, lacking intelligent support. This makes it difficult for the system to respond quickly to complex or unexpected problems, affecting the efficiency and quality of problem resolution. Furthermore, due to the low efficiency of work order processing and the frequent occurrence of recurring issues, customers often have to wait a long time for answers or solutions, and the answers may not meet their expectations. This not only reduces customer satisfaction but may also negatively impact the insurance company's brand image. Summary of the Invention
[0004] The main objective of this application is to propose a data processing method, apparatus, electronic device, and medium based on a natural language model, which can improve the efficiency of problem handling and service quality, and avoid situations where the answer is inconsistent with the customer's expectations.
[0005] To achieve the above objectives, a first aspect of this application proposes a data processing method based on a natural language model, the method comprising: Acquire financial data and unanswered questions, wherein the financial data includes multiple questions. The financial data information is represented by knowledge to create a first data structure and a second data structure corresponding to the question information, wherein the first data structure is a data structure corresponding to common questions, and the second data structure is a data structure corresponding to related questions; Create an insurance dataset based on the first data structure and the second data structure; An automatic question-answering language model is obtained by training a preset natural language model using the insurance dataset. In response to a user's search command, text analysis is performed on the information of the question to be answered to obtain the analysis results; Based on the automatic question-answering language model and the insurance dataset, answers are retrieved from the analysis results to obtain a retrieval set; When the search set is empty or the preset answer statement does not exist in the search set, the user's search instruction is converted into a manual service instruction.
[0006] In some embodiments, the step of performing knowledge representation on the financial data information to create a first data structure and a second data structure corresponding to the question information includes: The financial data information is cleaned. Synonym expansion is performed on the cleaned financial data to obtain a financial database. Determine frequently asked questions and their answers from the financial database; A first data structure is established based on the information of the frequently asked questions and the answers to the frequently asked questions. Part-of-speech analysis is performed on the financial database to determine related question information and related question answers; A second data structure is established based on the associated question information and the answers to the associated questions.
[0007] In some embodiments, the text analysis of the unanswered question information to obtain the analysis result includes: The unanswered question information is segmented into words to obtain multiple question texts; Part-of-speech tagging was performed on all the question texts to obtain tagging information; Based on the annotation information, keywords are extracted from the question text to determine the keyword information in the question text; Analysis results are generated based on the question text, the annotation information, and the keyword information.
[0008] In some embodiments, the user search instruction includes a first user search instruction and a second user search instruction; the step of retrieving answers based on the analysis results using the automatic question-answering language model and the insurance dataset to obtain a search set includes: Extract key entities from the analysis results based on the keyword information; In response to the first user's search instruction, the key entity is retrieved using the insurance dataset; If the key entity does not exist in the insurance dataset, the key entity is replaced with a synonym to obtain replacement information; The replacement information is retrieved using the insurance dataset; If the replacement information is not found in the insurance dataset, the analysis results are retrieved using the automatic question-answering language model to obtain a retrieval set.
[0009] In some embodiments, retrieving the analysis results using the automatic question-answering language model to obtain a retrieval set includes: In response to a second user's retrieval command, the analysis results are input into the automatic question-answering language model, so that the automatic question-answering language model performs word vectorization on the labeled information, thereby converting the labeled information into vectors of a preset length to obtain word vectors; The automatic question-answering language model is used to perform semantic analysis on the word vectors, and the analysis answer is output. The analyzed answers are organized to obtain the retrieval set.
[0010] In some embodiments, after converting the user search instruction into a call to human agent service, the method further includes: Collect human responses; Sensitive information detection is performed on the manually replied information; When the manual response information passes the sensitive information detection, a target response information corresponding to the question to be answered is generated based on the manual response information. The information on the unanswered question and the information on the target response are added to the insurance dataset.
[0011] In some embodiments, after retrieving answers from the analysis results based on the automatic question-answering language model and the insurance dataset to obtain a retrieval set, the method further includes: When the search set is not empty and a preset answer statement exists in the search set, the target answer statement is determined in the search set.
[0012] To achieve the above objectives, a second aspect of this application provides a data processing apparatus based on a natural language model, the apparatus comprising: The information acquisition module is used to acquire financial data information and unanswered questions information, wherein the financial data information includes multiple questions. The knowledge representation module is used to represent the financial data information to create a first data structure and a second data structure corresponding to the question information, wherein the first data structure is a data structure corresponding to common questions, and the second data structure is a data structure corresponding to related questions. A dataset creation module is used to create an insurance dataset based on the first data structure and the second data structure; The model training module is used to train a preset natural language model using the insurance dataset to obtain an automatic question-answering language model. The text analysis module is used to perform text analysis on the information of the question to be answered in response to the user's search command, and obtain the analysis results; The answer retrieval module is used to retrieve answers based on the automatic question-answering language model and the insurance dataset from the analysis results, and obtain a retrieval set; The module for transferring to human assistance is used to convert the user's search instruction into a human assistance service instruction when the search set is empty or the search set does not contain a preset answer statement.
[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the data processing method based on a natural language model as described in the first aspect.
[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data processing method based on a natural language model as described in the first aspect.
[0015] This application proposes a data processing method, apparatus, electronic device, and storage medium based on a natural language model. First, it acquires financial data and unanswered question information, performs knowledge representation on the financial data to achieve accurate analysis, and creates a first and second data structure corresponding to the question information in the financial data. This creates data structures for common and related questions. Then, based on the first and second data structures, it creates an insurance dataset, constructing a professional database. This allows the insurance dataset to integrate multiple question information, facilitating comprehensive question retrieval. Finally, it trains a pre-defined natural language model using the insurance dataset, enabling the model to automatically process large amounts of text data, resulting in an automatic question-answering language model, thereby reducing... In response to user search commands, manual intervention is employed to analyze the text of the question to be answered, thereby identifying the keywords to be searched within the question information. This analysis yields precise retrieval of the question information, improving search accuracy. Then, based on an automatic question-answering language model and an insurance dataset, the analysis results are used to retrieve answers, achieving a comprehensive retrieval of the question information and obtaining a search set. This improves the efficiency and quality of problem handling. If the search set is empty or does not contain a preset answer statement, it indicates that the user's input question information has not been resolved. In this case, the user's search command is converted into a request for human assistance, allowing for a human response to avoid discrepancies between the answer and the customer's expectations, thus ensuring accurate answers. This embodiment of the application creates multiple data structures by representing financial data information using knowledge representation, then uses these data structures to create an insurance dataset and train a natural language model. This enables automatic question-answering of the question information, reducing manual input and improving problem handling efficiency. Furthermore, in cases where the search set is empty or does not contain a preset answer statement, this embodiment of the application will initiate a human intervention to avoid discrepancies between the answer and the customer's expectations, ensuring accurate answers to the question information. Attached Figure Description
[0016] Figure 1 This is a flowchart of a data processing method based on a natural language model provided in an embodiment of this application; Figure 2 yes Figure 1 The flowchart of step S102 in the document; Figure 3 A flowchart illustrating a method for text analysis of information about a problem to be solved, provided in an embodiment of this application; Figure 4 yes Figure 1 The flowchart of step S106 in the process; Figure 5A flowchart illustrating a method for retrieving analysis results using an automatic question-answering language model, as provided in this application embodiment; Figure 6 This is a flowchart of a data processing method based on a natural language model provided in another embodiment of this application; Figure 7 This is a flowchart of a data processing method based on a natural language model provided in another embodiment of this application; Figure 8 This is a schematic diagram of the structure of the data processing device based on a natural language model provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] First, let's analyze some of the terms used in this application: Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0021] Frequently Asked Questions (FAQs): FAQs have become a feature of the Internet. They appear to have originated in many user groups as a way to familiarize new users with rules. There are tens of thousands of FAQs on the World Wide Web.
[0022] JavaScript Object Notation (JSON): JSON is an open standard file and data exchange format designed based on a subset of ECMAScript. It is easy for humans to read and write, and also easy for machines to parse and generate. JSON is a commonly used data format with various applications in electronic data interchange, including data exchange between web applications and servers. Its concise and clear hierarchical structure effectively improves network transmission efficiency, making it an ideal data exchange language.
[0023] Extensible Markup Language (XML) is a subset of Standard Generalized Markup Language (SGML). It can be used to mark up data and define data types, and is a source language that allows users to define their own markup languages. XML boasts advantages such as good extensibility, separation of content and form, adherence to strict syntax requirements, and good value preservation.
[0024] The data processing method, apparatus, electronic device, and storage medium based on natural language models provided in this application can avoid situations where the answer is inconsistent with the customer's expectations, thereby improving the efficiency of problem handling and service quality.
[0025] The following embodiments will be used to illustrate the specific implementation of the data processing method based on the natural language model in this application.
[0026] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0027] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, module management for online meeting systems, natural language processing, and machine learning / deep learning.
[0028] The data processing method based on a natural language model provided in this application relates to the field of data processing technology. This data processing method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the data processing method based on a natural language model, but is not limited to the above forms.
[0029] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0030] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of the embodiments of this application obtained.
[0031] In the daily operations of the insurance industry, the stable operation of business systems is crucial for ensuring customer service quality and improving business processing efficiency. However, with the increasing complexity of insurance business and the continuous upgrading of technology systems, business system problems occur frequently. These problems are usually submitted to the system operations team in the form of work orders for processing.
[0032] Existing work order processing systems mostly rely on manual operation and simple automated processes, lacking intelligent support. This makes it difficult for the system to respond quickly to complex or unexpected problems, affecting the efficiency and quality of problem resolution. Furthermore, due to the low efficiency of work order processing and the frequent occurrence of recurring issues, customers often have to wait a long time for answers or solutions, and the answers may not meet their expectations. This not only reduces customer satisfaction but may also negatively impact the insurance company's brand image.
[0033] To address the aforementioned issues, this embodiment provides a data processing method, apparatus, electronic device, and storage medium based on a natural language model. First, financial data and unanswered question information are acquired. Knowledge representation is then performed on the financial data to achieve accurate analysis. A first and second data structure corresponding to the question information within the financial data are created, establishing data structures for common and related questions. Next, an insurance dataset is created based on the first and second data structures, constructing a professional database. This allows the insurance dataset to integrate various question information, facilitating comprehensive question retrieval. Finally, a pre-defined natural language model is trained using the insurance dataset, enabling the model to automatically process large amounts of text data, resulting in an automatic question-answering language model. This approach reduces manual intervention, responds to user search commands, performs text analysis on the questions to be answered, identifies keywords for retrieval, and obtains analysis results. This enables precise retrieval of the questions, improving search accuracy. Then, based on an automatic question-answering language model and an insurance dataset, the analysis results are used to retrieve answers, achieving a comprehensive retrieval of the questions and obtaining a search set. This improves the efficiency and quality of problem handling. If the search set is empty or does not contain a preset answer, it indicates that the user's input question has not been resolved. In this case, the user's search command is converted into a request for human assistance, allowing for a human response and preventing discrepancies between the answer and the customer's expectations, thus ensuring accurate answers. This embodiment creates multiple data structures by representing financial data, then uses these data structures to create an insurance dataset and train a natural language model. This enables automatic question-answering, reducing manual input and improving efficiency. Furthermore, if the search set is empty or does not contain a preset answer, this embodiment will initiate a human intervention to prevent discrepancies between the answer and the customer's expectations, ensuring accurate answers.
[0034] The following is a detailed explanation with reference to the accompanying drawings.
[0035] Figure 1 This is an optional flowchart of the data processing method based on a natural language model provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S107.
[0036] Step S101: Obtain financial data information and unanswered questions information. The financial data information includes multiple question information.
[0037] In step S101 of some embodiments, financial data information and unanswered question information are obtained. The financial data information includes multiple question information to facilitate subsequent retrieval of unanswered question information.
[0038] It should be noted that the financial data information in this application embodiment can be common information collected from insurance companies' daily work order processing records, FAQs, technical documents, and expert experience, etc. This application embodiment does not impose specific limitations.
[0039] Step S102: Perform knowledge representation on the financial data information to create a first data structure and a second data structure corresponding to the problem information.
[0040] It should be noted that the first data structure corresponds to the common problems, and the second data structure corresponds to the related problems.
[0041] In step S102 of some embodiments, financial data information is represented by knowledge to improve the overall quality of the data, ensure data quality, and reduce the risk of making wrong decisions based on erroneous data. This is done to create a first data structure and a second data structure corresponding to the problem information, which can avoid processing delays or errors caused by human factors, thereby improving the overall quality of the data, ensuring data quality, and reducing the risk of making wrong decisions based on erroneous data.
[0042] It is understood that the first data structure in the embodiments of this application is structured data, and the second data structure is semi-structured data.
[0043] Step S103: Create an insurance dataset based on the first data structure and the second data structure.
[0044] In step S103 of some embodiments, an insurance dataset is created based on the first data structure and the second data structure to build a professional database, enabling the insurance dataset to integrate various problem information, which facilitates comprehensive retrieval of problems in the future.
[0045] Step S104: Train a pre-defined natural language model using the insurance dataset to obtain an automatic question-answering language model.
[0046] In step S104 of some embodiments, a preset natural language model is trained using an insurance dataset. By repeatedly training and optimizing, an automatic question-and-answer text generation model can be obtained, thus reducing human intervention.
[0047] Specifically, in this embodiment of the application, the insurance dataset is divided into a training set, a validation set, and a test set, and then the natural language model is trained using the training set, the validation set, and the test set.
[0048] Step S105: In response to the user's search command, perform text analysis on the information of the question to be answered to obtain the analysis results.
[0049] In step S105 of some embodiments, in response to the user's retrieval instruction, text analysis is performed on the information of the question to be answered, so as to better understand the grammatical structure and semantic content of the sentence, improve the relevance of the search results, obtain the analysis results, and facilitate subsequent accurate retrieval in the insurance dataset, thereby improving the accuracy of the retrieval.
[0050] Step S106: Based on the automatic question-answering language model and the insurance dataset, the analysis results are used to retrieve answers and obtain the retrieval set.
[0051] In step S106 of some embodiments, answers are retrieved based on the automatic question-answering language model and the insurance dataset to obtain a retrieval set, thereby achieving a comprehensive retrieval of the analysis results, improving the accuracy of the retrieval, reducing the number of times users need to perform multiple searches to find all relevant information, and improving retrieval efficiency.
[0052] Step S107: When the search set is empty or the preset answer statement does not exist in the search set, the user's search instruction is converted into a manual service instruction.
[0053] In step S107 of some embodiments, when the search set is empty or there is no preset answer statement in the search set, it means that the user's input question information has not been resolved. In this case, the user's search instruction needs to be converted into a manual service instruction so that the question information can be answered manually, avoiding the situation where the answer is inconsistent with the customer's expectations, and achieving accurate answers to the question information.
[0054] It should be noted that the answer statements in this application embodiment can be set according to the user's needs, and this application embodiment does not impose any specific limitations.
[0055] It is understood that the embodiments of this application can intelligently solve the problems of difficult and cumbersome event handling in the daily customer service and operation of insurance companies. It can effectively improve the efficiency of problem handling and service quality, and also reduce the manpower input for insurance companies' customer operations, thereby achieving the goal of cost reduction and efficiency improvement.
[0056] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S206.
[0057] Step S201: Clean the financial data information.
[0058] In step S201 of some embodiments, during the knowledge representation of financial data information, the embodiments of this application will first perform data cleaning on the financial data information. Specifically, the financial data will be cleaned by handling missing values, detecting outliers, identifying duplicate data, etc. Cleaning can remove erroneous, duplicate and inconsistent data, remove irrelevant information, thereby improving the overall quality of the data, ensuring data quality, and reducing the risk of making wrong decisions based on erroneous data.
[0059] It should be noted that data cleaning operations include, but are not limited to, data consistency checks, data accuracy verification, data integrity checks, data deduplication, format conversion, etc., and the embodiments of this application do not impose specific limitations.
[0060] Step S202 involves performing synonym expansion on the cleaned financial data to obtain a financial database.
[0061] In step S202 of some embodiments, in order to ensure the comprehensiveness of keyword retrieval, keyword synonym expansion is required. That is, the financial data information after data cleaning is expanded by synonym. Specifically, the financial data information is first identified by keyword identification to determine the keywords that need to be expanded in the financial data information, such as wealth management, insurance, etc. Then, synonym expansion is performed on different keywords to obtain a financial database, which facilitates the comprehensive retrieval of keywords in the future.
[0062] It should be noted that, taking the keywords such as "insurance" and "amount" in financial data information as an example, after synonym expansion of financial data information, "insurance" can be expanded to words such as "buying insurance" and "selecting insurance," and "amount" can be expanded to words such as "how much" and "share." This application embodiment does not impose specific limitations.
[0063] It is understood that the embodiments of this application, through synonym expansion, can ensure that information related to keywords but with different wording is not missed when searching for financial data, and can improve the relevance of search results and user satisfaction, and help identify and integrate similar concepts in different languages.
[0064] Step S203: Determine common question information and answers in the financial database.
[0065] In step S203 of some embodiments, since there may be questions in the financial database that have already appeared and have standard answers, the embodiments of this application can directly determine the information and answers of common questions in the financial database, simplifying the operation process.
[0066] It is understood that the common information in the embodiments of this application may include the car insurance application process, the eligibility for pension insurance, etc., and the embodiments of this application do not impose specific limitations.
[0067] Step S204: Establish the first data structure based on the information and answers to common questions.
[0068] In step S204 of some embodiments, since the common questions information and answers are already standard questions and answers, the answers to these questions can be found through traditional question-and-answer matching. Therefore, the embodiments of this application can directly establish question-answer pairs and directly perform structured data storage. That is, the first data structure can be directly established based on the common questions information and answers, which can simplify the data flow as a whole and improve the efficiency of problem processing.
[0069] It should be noted that for simple data structures, the embodiments of this application can be directly stored in semi-structured formats such as JSON or XML.
[0070] Step S205: Perform part-of-speech analysis on the financial database to determine the related question information and the related question answers.
[0071] In step S205 of some embodiments, since there is a clear correlation between some data, this embodiment of the application will also perform part-of-speech analysis on the financial database, analyze the specific part of speech of each word in the financial database, thereby determining whether there are extended words for the current word, and finally determine the related question information and related question answers based on the part of speech of each word and the extension results, thereby ensuring the integrity and reliability of the data.
[0072] Step S206: Establish a second data structure based on the associated question information and the associated question answers.
[0073] In step S206 of some embodiments, a second data structure is established based on the associated question information and the associated question answer, making the data structure intuitive and easy to understand, which facilitates subsequent improvement of data retrieval and analysis performance.
[0074] It is understandable that, for data with clear relationships, this application embodiment may consider using a relational database for storage, and representing the relationships between data through table structures.
[0075] In some embodiments, during traditional work order processing, faced with a large influx of work orders, operations personnel often need to manually classify, identify, and process each issue one by one. This manual approach is not only time-consuming and labor-intensive but also prone to delays or errors due to human factors. Furthermore, insurance business systems may contain numerous recurring issues, such as abnormal system logins and data query errors. These recurring problems consume significant time and energy of operations personnel without being effectively resolved or prevented. The embodiments of this application, through data cleaning, synonym expansion, and part-of-speech analysis of financial data, can avoid processing delays or errors caused by human factors, thereby improving the overall quality of the data, ensuring data quality, and reducing the risk of making incorrect decisions based on erroneous data.
[0076] Please see Figure 3 , Figure 3 The flowchart of a method for text analysis of information about a question to be answered provided in an embodiment of this application includes, but is not limited to, steps S301 to S304.
[0077] Step S301: Perform word segmentation on the information of the questions to be answered to obtain multiple question texts.
[0078] In step S301 of some embodiments, during the text analysis of the question information to be answered, this embodiment first performs word segmentation on the question information. Specifically, for languages such as English that use spaces to separate words, word segmentation can be performed directly according to spaces. For languages such as Chinese that do not have obvious word separators, word segmentation can be performed according to specific rules (such as length, character type) to obtain multiple question texts, thereby enabling a better understanding of the grammatical structure and semantic content of the sentences and improving the relevance of the search results.
[0079] Step S302: Perform part-of-speech tagging on all question texts to obtain tagging information.
[0080] In step S302 of some embodiments, part-of-speech tagging is performed on all question texts to obtain the part-of-speech tag for each word in the question text, thereby obtaining tagging information, which can improve the understanding of text content and thus improve the accuracy of classification and sentiment analysis.
[0081] It should be noted that the embodiments of this application can perform word segmentation and part-of-speech tagging using Python text annotation tools, Stanford NLP text annotation tools, etc., and the embodiments of this application do not impose specific limitations.
[0082] It is understood that the annotation information in the embodiments of this application includes questions, the location of answer keywords, and answers.
[0083] Step S303: Extract keywords from the question text based on the annotation information to determine the keyword information in the question text.
[0084] In step S303 of some embodiments, after obtaining the annotation information, this application embodiment extracts keywords from the question text based on the annotation information, extracts and identifies key entities in the question, such as business type, product name, etc., and determines the keyword information in the question text so as to perform accurate retrieval in the insurance dataset in the future.
[0085] Step S304: Generate analysis results based on the question text, annotation information, and keyword information.
[0086] In step S304 of some embodiments, analysis results are generated based on the question text, annotation information, and keyword information, which facilitates accurate retrieval in the insurance dataset and improves the accuracy of the retrieval.
[0087] Please see Figure 4 In some embodiments, step S106 may include, but is not limited to, steps S401 to S405.
[0088] It should be noted that user search instructions include first user search instructions and second user search instructions.
[0089] Step S401: Extract key entities from the analysis results based on keyword information.
[0090] Step S402: In response to the first user's retrieval instruction, the key entities are retrieved using the insurance dataset.
[0091] Step S403: If the key entity does not exist in the insurance dataset, perform synonym replacement on the key entity to obtain replacement information.
[0092] Step S404: Retrieve replacement information using the insurance dataset.
[0093] Step S405: When there is no replacement information in the insurance dataset, the analysis results are retrieved using an automatic question-answering language model to obtain a retrieval set.
[0094] In steps S401 to S405 of some embodiments, during the process of retrieving answers from the analysis results based on the automatic question-answering language model and the insurance dataset, this embodiment first extracts key entities from the analysis results based on keyword information. By identifying key entities, relevant information can be located more accurately, reducing interference from irrelevant results. In response to the first user's retrieval instruction, key entities are retrieved from the insurance dataset, quickly responding to user needs and helping to allocate service resources more effectively, thus improving work efficiency. When a key entity does not exist in the insurance dataset, it means that no data consistent with the key entity has been retrieved from the insurance dataset, and the key entity needs to be replaced with a synonym. Replacement can capture documents that express the same or similar concepts using different words, thereby improving the comprehensiveness of search results. Replacement information is obtained, and then retrieved using an insurance dataset. Synonym replacement helps find documents that may not directly use the original keywords but are related in content, improving the relevance of search results. If replacement information is not found in the insurance dataset, it means that no data consistent with the replacement information was found. In this case, an automatic question-answering language model is needed to retrieve the analysis results, obtaining a search set. This enables a comprehensive retrieval of the analysis results, improves search accuracy, reduces the number of times users need to perform multiple searches to find all relevant information, and improves search efficiency.
[0095] It is understood that this application embodiment begins by responding to a first user's search instruction to perform a question search. If no corresponding content is found by searching for keywords, then keyword synonyms are used to perform the search again. For example, "the process of purchasing insurance" and "the process of applying for insurance" are the same, thereby enabling the capture of documents that use different words to express the same or similar concepts.
[0096] It is worth noting that, in response to the first user's search instruction, this application embodiment will also perform format conversion on the analysis results, converting the format of the analysis results into a query statement or format, such as JSON or XML, so as to be searched in the insurance dataset.
[0097] Please see Figure 5 , Figure 5 The flowchart of the method for retrieving analysis results using an automatic question-answering language model provided in the embodiments of this application includes, but is not limited to, steps S501 to S503.
[0098] In step S501, in response to the second user's retrieval command, the analysis results are input into the automatic question-answering language model, so that the automatic question-answering language model performs word vectorization on the labeled information, so as to convert the labeled information into vectors of a preset length and obtain word vectors.
[0099] Step S502: Perform semantic analysis on word vectors using an automatic question-answering language model and output the analysis answer.
[0100] Step S503: Organize the analyzed answers to obtain the retrieval set.
[0101] In steps S501 to S503 of some embodiments, during the retrieval of analysis results using an automatic question-answering language model, in response to a second user's retrieval command, the analysis results are input into the automatic question-answering language model. This allows the automatic question-answering language model to vectorize the labeled information into word vectors of a preset length, facilitating subsequent precise analysis of the word vectors and improving retrieval accuracy. Then, the automatic question-answering model performs semantic analysis on the word vectors to determine the semantics expressed by the word vectors and outputs the analysis answer corresponding to the question information to be answered. Finally, the analysis answer is organized, for example, by removing duplicates and sorting, to obtain a retrieval set, achieving a comprehensive retrieval of the analyzed information and enabling the retrieval of accurate question answers.
[0102] It should be noted that the automatic question-answering language model in this application embodiment includes a Word2Vec pre-trained model and a Transformer model. Specifically, this application embodiment uses the Word2Vec pre-trained model to vectorize the labeled information and uses the Transformer model to perform semantic analysis on the word vectors.
[0103] Please see Figure 6 , Figure 6 This is a flowchart of a data processing method based on a natural language model provided in another embodiment of this application. The method includes, but is not limited to, steps S601 to S604.
[0104] It should be noted that steps S601 to S604 occur after the user's search instruction is transformed into a manual service instruction.
[0105] Step S601: Collect human response information.
[0106] Step S602: Perform sensitive information detection on the manually replied information.
[0107] Step S603: When the manual reply information passes the sensitive information detection, the target reply information corresponding to the question information to be answered is generated based on the manual reply information.
[0108] Step S604: Add the information of the question to be answered and the target answer information to the insurance dataset.
[0109] In steps S601 to S604 of some embodiments, after converting the user's search instruction into a manual service instruction, this application embodiment also collects manual response information, that is, the answers given by humans to the questions to be answered. Then, sensitive information detection is performed on the manual response information to reduce the risk of personal privacy leakage. By identifying and protecting sensitive information, the overall data security is improved. When the manual response information passes the sensitive information detection, it means that there is no sensitive information in the manual response information. The target answer information corresponding to the questions to be answered can be directly generated based on the manual response information to achieve accurate answers to the questions to be answered. Finally, the questions to be answered and the target answer information are added to the insurance dataset to achieve real-time updates of the insurance dataset, which facilitates the subsequent retrieval of other questions. Real-time updates of the dataset can ensure that the information retrieved by the user is up-to-date, improving the accuracy and relevance of question matching.
[0110] It is worth noting that, in the process of detecting sensitive information in manually replied information, this application embodiment will perform sensitive word detection on the manually replied information based on preset sensitive words, where sensitive words can be personal identity information, financial data, etc.
[0111] It should be noted that when a manually replied message fails the sensitive information detection, the identified sensitive information will be marked, and then the marked sensitive information will be de-identified, for example, through masking, encryption, etc. Finally, the de-identified information will be displayed to the user, achieving secure display of information and avoiding leakage of sensitive information.
[0112] Understandably, if users' questions can be answered quickly and these answers can be retrieved by other users, it will increase user satisfaction with the service.
[0113] Please see Figure 7 , Figure 7 This is a flowchart of a data processing method based on a natural language model provided in another embodiment of this application. Figure 7 The method may include, but is not limited to, step S701.
[0114] It should be noted that step S701 occurs after the answer retrieval is performed on the analysis results based on the automatic question-answering language model and the insurance dataset.
[0115] Step S701: When the search set is not empty and a preset answer statement exists in the search set, determine the target answer statement in the search set.
[0116] In step S701 of some embodiments, after retrieving answers from the analysis results based on the automatic question-answering language model and the insurance dataset, if the retrieval set is not empty and a preset answer statement exists in the retrieval set, it means that a statement corresponding to the question information to be answered has been retrieved through the automatic question-answering language model or the insurance dataset, and the statement is a preset answer statement. Therefore, the target answer statement can be directly determined in the retrieval set to achieve accurate retrieval of the question information to be answered.
[0117] Please see Figure 8 This application also provides a data processing apparatus based on a natural language model, the apparatus comprising: The information acquisition module 801 is used to acquire financial data information and unanswered questions information, including multiple questions; The knowledge representation module 802 is used to represent financial data information to create a first data structure and a second data structure corresponding to the question information. The first data structure is the data structure corresponding to common questions, and the second data structure is the data structure corresponding to related questions. The dataset creation module 803 is used to create an insurance dataset based on a first data structure and a second data structure. Model training module 804 is used to train a preset natural language model using an insurance dataset to obtain an automatic question-answering language model; The text analysis module 805 is used to perform text analysis on the information of the question to be answered in response to the user's search command, and obtain the analysis results; The answer retrieval module 806 is used to retrieve answers based on the analysis results using an automatic question-answering language model and an insurance dataset, and obtain a retrieval set. The module 807 is used to convert user search instructions into human assistance instructions when the search set is empty or the preset answer statement does not exist in the search set.
[0118] In some embodiments, the data processing device based on a natural language model further includes an information collection module and a target determination module. The information collection module is used to collect manually provided response information; perform sensitive information detection on the manually provided response information; when the manually provided response information passes the sensitive information detection, generate target response information corresponding to the question information to be answered based on the manually provided response information; and add the question information to be answered and the target response information to the insurance dataset. The target determination module is used to determine the target answer statement in the search set when the search set is not empty and a preset answer statement exists in the search set.
[0119] The specific implementation of this data processing device based on a natural language model is basically the same as the specific implementation of the data processing method based on a natural language model described above, and will not be repeated here.
[0120] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned data processing method based on a natural language model. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0121] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the data processing method based on the natural language model of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0122] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data processing method based on a natural language model.
[0123] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0124] The data processing method, apparatus, electronic device, and storage medium based on a natural language model provided in this application first acquire financial data information and unanswered question information, perform knowledge representation on the financial data information to achieve accurate analysis of the financial data information, and create a first data structure and a second data structure corresponding to the question information in the financial data information, thereby creating data structures for common and related questions. Then, an insurance dataset is created based on the first and second data structures to construct a professional database, enabling the insurance dataset to integrate multiple question information, facilitating comprehensive retrieval of questions later. Afterwards, a preset natural language model is trained through the insurance dataset, enabling the natural language model to automatically process large amounts of text data, resulting in an automatic question-answering language model, thereby enabling... This approach reduces manual intervention by responding to user search commands and performing text analysis on the questions to be answered. This analysis identifies keywords for retrieval, yielding precise results and improving search accuracy. Furthermore, based on an automatic question-answering language model and an insurance dataset, the analysis results are used to retrieve answers, achieving a comprehensive search of the questions and resulting in a search set. This improves efficiency and service quality. If the search set is empty or lacks a pre-defined answer, indicating the user's question has not been resolved, the user's search command is converted into a request for human assistance. This ensures a human response to the question, preventing discrepancies between the answer and the customer's expectations and ensuring accurate answers. This embodiment of the application creates multiple data structures by representing financial data, then uses these structures to create an insurance dataset and train a natural language model. This enables automatic question-answering, reducing manual input and improving efficiency. Furthermore, in cases where the search set is empty or lacks a pre-defined answer, the application will initiate a human intervention to prevent discrepancies between the answer and the customer's expectations and ensure accurate answers.
[0125] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0126] It will be understood by those skilled in the art that Figure 1-9 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0129] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0130] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0131] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0132] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0133] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0134] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A data processing method based on a natural language model, characterized in that, The method includes: Acquire financial data and unanswered questions, wherein the financial data includes multiple questions. The financial data information is represented by knowledge to create a first data structure and a second data structure corresponding to the question information, wherein the first data structure is a data structure corresponding to common questions, and the second data structure is a data structure corresponding to related questions; Create an insurance dataset based on the first data structure and the second data structure; An automatic question-answering language model is obtained by training a preset natural language model using the insurance dataset. In response to a user's search command, text analysis is performed on the information of the question to be answered to obtain the analysis results; Based on the automatic question-answering language model and the insurance dataset, answers are retrieved from the analysis results to obtain a retrieval set; When the search set is empty or the preset answer statement does not exist in the search set, the user's search instruction is converted into a manual service instruction.
2. The data processing method based on a natural language model according to claim 1, characterized in that, The step of performing knowledge representation on the financial data information to create a first data structure and a second data structure corresponding to the question information includes: The financial data information is cleaned. Synonym expansion is performed on the cleaned financial data to obtain a financial database. Determine frequently asked questions and their answers from the financial database; A first data structure is established based on the information of the frequently asked questions and the answers to the frequently asked questions. Part-of-speech analysis is performed on the financial database to determine related question information and related question answers; A second data structure is established based on the associated question information and the answers to the associated questions.
3. The data processing method based on a natural language model according to claim 1, characterized in that, The text analysis of the unanswered question information, to obtain the analysis results, includes: The unanswered question information is segmented into words to obtain multiple question texts; Part-of-speech tagging was performed on all the question texts to obtain tagging information; Based on the annotation information, keywords are extracted from the question text to determine the keyword information in the question text; Analysis results are generated based on the question text, the annotation information, and the keyword information.
4. The data processing method based on a natural language model according to claim 3, characterized in that, The user search instructions include a first user search instruction and a second user search instruction; the step of retrieving answers based on the analysis results using the automatic question-answering language model and the insurance dataset to obtain a search set includes: Extract key entities from the analysis results based on the keyword information; In response to the first user's search instruction, the key entity is retrieved using the insurance dataset; If the key entity does not exist in the insurance dataset, the key entity is replaced with a synonym to obtain replacement information; The replacement information is retrieved using the insurance dataset; If the replacement information is not found in the insurance dataset, the analysis results are retrieved using the automatic question-answering language model to obtain a retrieval set.
5. The data processing method based on a natural language model according to claim 4, characterized in that, The process of retrieving the analysis results using the automatic question-answering language model to obtain a retrieval set includes: In response to a second user's retrieval command, the analysis results are input into the automatic question-answering language model, so that the automatic question-answering language model performs word vectorization on the labeled information, thereby converting the labeled information into vectors of a preset length to obtain word vectors; The automatic question-answering language model is used to perform semantic analysis on the word vectors, and the analysis answer is output. The analyzed answers are organized to obtain a retrieval set.
6. The data processing method based on a natural language model according to claim 1, characterized in that, After converting the user's search instruction into a direct call for human assistance, the method further includes: Collect human responses; Sensitive information detection is performed on the manually replied information; When the manual response information passes the sensitive information detection, a target response information corresponding to the question to be answered is generated based on the manual response information. The information on the unanswered question and the information on the target response are added to the insurance dataset.
7. The data processing method based on a natural language model according to claim 1, characterized in that, After retrieving answers from the analysis results based on the automatic question-answering language model and the insurance dataset to obtain a retrieval set, the method further includes: When the search set is not empty and a preset answer statement exists in the search set, the target answer statement is determined in the search set.
8. A data processing device based on a natural language model, characterized in that, The device includes: The information acquisition module is used to acquire financial data information and unanswered questions information, wherein the financial data information includes multiple questions. The knowledge representation module is used to represent the financial data information to create a first data structure and a second data structure corresponding to the question information, wherein the first data structure is a data structure corresponding to common questions, and the second data structure is a data structure corresponding to related questions. A dataset creation module is used to create an insurance dataset based on the first data structure and the second data structure; The model training module is used to train a preset natural language model using the insurance dataset to obtain an automatic question-answering language model. The text analysis module is used to perform text analysis on the question information to be answered in response to the user's search command, and obtain the analysis results; The answer retrieval module is used to retrieve answers based on the automatic question-answering language model and the insurance dataset from the analysis results, and obtain a retrieval set; The module for transferring to human assistance is used to convert the user's search instruction into a human assistance service instruction when the search set is empty or the search set does not contain a preset answer statement.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the data processing method based on a natural language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data processing method based on the natural language model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Business question and answer data processing method and device, computer equipment and storage medium
CN113918692A
Knowledge base expansion method and device, equipment and medium
CN115525747A