Data information processing method, device and equipment and storage medium thereof

By employing semantic understanding-based segmentation processing and vectorized retrieval technology, this technology identifies special clauses in insurance contracts, solving the problem of low processing efficiency in existing technologies and enabling rapid identification and efficient processing of special clauses in insurance contracts.

CN121786031APending Publication Date: 2026-04-03CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

When processing special clauses in insurance contracts, conventional automated processing models result in low business processing efficiency and are unable to effectively identify and process specially agreed-upon scope of liability and compensation rules.

Method used

The method employs semantic understanding-based segmentation, text information extraction, vectorization, and comparative vector retrieval to identify the text type of the target text and then process it in conjunction with the corresponding business processing strategies.

Benefits of technology

It enables rapid identification of clause types in financial business contracts, improves the efficiency of financial business processing, and allows for corresponding business processing for different clause types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786031A_ABST
    Figure CN121786031A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of text processing, and relates to a data information processing method and device, equipment and a storage medium thereof. Performing semantic understanding type segmentation processing to obtain segmented texts; determining a text type of the target text; and performing business processing on the target text in combination with the text type and the business processing strategy corresponding to the text type. When the data information processing method is applied to processing the liability term text in the financial business contract, whether the current liability term text is a term text of a special agreed type or not can be identified by combining a preset vector retrieval library, so that a corresponding business processing strategy is screened out; and performing business processing on the target text. According to the method, the clause types in the financial business contract can be quickly identified, business processing is performed according to different clause types, and the financial business processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing technology and is applied to scenarios where business execution is performed based on agreed terms and conditions. It relates to a data information processing method, apparatus, device, and storage medium. Background Technology

[0002] When an insurance contract is signed, there may be special clauses that specify the scope of liability and compensation rules to the client, in order to meet the client's personalized insurance needs.

[0003] Currently, there are relatively mature information extraction and recognition schemes for identifying and processing textual information in insurance contracts, particularly for fixed-format or template-based clauses. However, for these special clauses, which often specify the scope of liability and compensation rules, the use of conventional automated processing models for handling business liability after a liability event occurs is inefficient because the specifically agreed-upon scope of liability and compensation rules often differ from the template-based clauses. Summary of the Invention

[0004] The purpose of this application is to provide a data information processing method, apparatus, device and storage medium to solve the technical problem that the existing technology uses conventional automated processing models to process special terms and conditions, which reduces business processing efficiency.

[0005] In a first aspect, embodiments of this application provide a data information processing method, which adopts the following technical solution: A data information processing method includes the following steps: Obtain the target text; The target text is subjected to semantic understanding-based segmentation processing to obtain segmented text; Text information is extracted from all segments of text, and the extracted scenario conditions and target business information are used as intermediate information. The target business information includes target business attribute type and target business attribute value. The intermediate information is vectorized to obtain a set of vector representations to be identified; The vector representation set to be identified is input into a preset vector retrieval library for comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized; Based on the results of the vector retrieval, the text type of the target text is determined; The target text is processed in accordance with the text type and the corresponding business processing strategy.

[0006] Secondly, embodiments of this application also provide a data information processing apparatus, which adopts the technical solution described below: A data information processing device, comprising: The target text acquisition module is used to acquire target text. The text segmentation module is used to perform semantic understanding-based segmentation on the target text to obtain segmented text. The text information extraction module is used to extract text information from all segmented texts separately, and to use the extracted scenario conditions and target business information as intermediate information. The target business information includes target business attribute type and target business attribute value. The vectorization processing module is used to perform vectorization processing on the intermediate information to obtain a set of vector representations to be identified; The comparison vector retrieval module is used to input the vector representation set to be identified into a preset vector retrieval library for comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized; The text type determination module is used to determine the text type of the target text based on the comparison vector retrieval results; The business processing module is used to perform business processing on the target text by combining the text type and the business processing strategy corresponding to the text type.

[0007] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data information processing method described above.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the data information processing method described above.

[0009] Compared with the prior art, the embodiments of this application have the following main advantages: The data processing method described in this application involves: acquiring target text; performing semantic understanding-based segmentation to obtain segmented text; extracting text information, including scenario conditions and target business information; performing vectorization processing and comparative vector retrieval; determining the text type of the target text based on the comparative vector retrieval results; and performing business processing on the target text by combining the text type with the corresponding business processing strategy. Applying this data processing method to the processing of liability clauses in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition, extracting scenario conditions and target business information, and identifying whether the current liability clause text is a specially agreed-upon type by combining a pre-given vector retrieval library. This enables the selection of appropriate business processing strategies and the execution of business processing for the target text. This achieves rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, thereby improving the efficiency of financial business processing. Attached Figure Description

[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of a data information processing method according to this application; Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 202 shown; Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 203 shown; Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown; Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown; Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 206 shown; Figure 8 This is a schematic diagram of the structure of one embodiment of a data information processing device according to this application; Figure 9 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0019] It should be noted that the data information processing method provided in this application embodiment is generally executed by a server, and correspondingly, a data information processing device is generally set in the server.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0021] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a data information processing method according to this application. The data information processing method includes the following steps: Step 201: Obtain the target text.

[0022] In this embodiment, the target text includes liability clause texts in contracts under the background of financial business. These liability clause texts include specially agreed types and non-specially agreed types. Specifically, the liability clause texts of the specially agreed type stipulate the allocation of responsibilities after negotiation among the responsible parties, while the liability clause texts of the non-specially agreed type stipulate a fixed, formatted allocation of responsibilities. Generally, when signing an insurance claim contract, non-specially agreed type liability clause texts are involved, meaning that the liability clauses are fixed according to regulations, and the insurer, policyholder, or beneficiary cannot negotiate the allocation of responsibilities. Specially agreed type liability clause texts are also involved, meaning that the various parties involved in the insurance policy can independently negotiate the allocation of responsibilities and obligations. For example, the liability clause texts of the specially agreed type record the agreed content negotiated by the insurer, policyholder, or beneficiary.

[0023] By acquiring the target text, it is possible to subsequently identify the obligations and responsibilities of each subject within the target text, and then apply these obligations and responsibilities to specific business processing solutions.

[0024] Step 202: Perform semantic understanding-based segmentation on the target text to obtain segmented text.

[0025] In this embodiment, semantic understanding is used to segment the target text. For example, the Qwen2-72B semantic understanding analysis model is used to first analyze the semantic content involved in the target text. Based on the different semantic content, the target text is segmented to obtain several segmented texts. The Qwen2-72B semantic understanding analysis model is an open-source large language model that uses an 80-layer Transformer structure to optimize semantic understanding reasoning capabilities and efficiency. Specifically, assuming a target text is a special clause in an insurance claim contract, this text at least involves the insurer's claim liability and the policyholder's insurance liability. More specifically, it includes the claim amount and claim items that the insurer should bear. By performing semantic understanding-based segmentation on the target text, segmented texts are obtained, which facilitates the segmentation of the liability content involved by different responsible parties, avoids confusion, and speeds up the clarification of the allocation of responsibilities among different responsible parties.

[0026] Step 203: Extract text information from all segmented texts, and use the extracted scenario conditions and target business information as intermediate information. The target business information includes target business attribute type and target business attribute value.

[0027] In this embodiment, text information is extracted from all segmented texts. After the text information is extracted, the extracted scenario conditions and target business information are used as intermediate information. The scenario conditions include information such as the main object and application scenario involved in the business processing, while the target business information includes the target business attribute type and the target business attribute value. For example, it includes information on various claims when making a claim and the specific value of the claim amount corresponding to different claims.

[0028] By extracting text information from each segment of text, the extracted scenario conditions and target business information are used as intermediate information to clarify the business processing information and specific attribute (parameter) values ​​involved in each text segment of the target text.

[0029] Step 204: The intermediate information is vectorized to obtain the vector representation set to be identified.

[0030] In this embodiment, text data vectorization processing technology can be used to vectorize all the obtained intermediate information separately, and finally obtain the set of vector representations to be identified contained in the target text. The vectorization processing technology, for example, uses Word2Vec and GloVe models based on neural networks to vectorize the text data. Word2Vec learns the distributed representation of text data lexical units by pre-training a shallow neural network, and then uses the trained shallow neural network to identify the lexical representations involved in the intermediate information, thereby obtaining the vectorized representation result corresponding to the intermediate information.

[0031] By vectorizing the intermediate information, a set of vector representations to be identified is obtained, which realizes the transformation of text data into vector representation, making it more conducive to computer recognition and processing, and improving the computer's recognition efficiency of the intermediate information.

[0032] Step 205: Input the vector representation set to be identified into a preset vector retrieval library and perform a comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized.

[0033] In this embodiment, the preset vector retrieval library pre-organizes the vectorized processing results corresponding to different historical special texts. For example, in insurance claims, some customers and insurers have made special terms agreements in the past. The special texts of the business processing agreements agreed in the past can be set as a special text to identify whether the agreement content of subsequent customers is consistent with that of previous customers. If they are consistent, the business processing method corresponding to the historical customer agreement text can be directly adopted to process the business for the subsequent customer.

[0034] Specifically, the step of inputting the vector representation set to be identified into a preset vector retrieval library for comparison vector retrieval aims to screen whether there are highly similar comparison vectors in the preset vector retrieval library. The vector retrieval library can be the Milvus vector library, which is an open-source vector database built for the GenAI application and is mainly used for vector retrieval queries.

[0035] Step 206: Determine the text type of the target text based on the comparison vector retrieval results.

[0036] Specifically, the preset vector retrieval library is constructed based on the special terms and conditions texts agreed upon by historical customers. If a matching vector can be found by filtering the vector representation set to be identified, it indicates that the target text is also a special terms and conditions text; otherwise, it indicates that the target text is not a special terms and conditions text.

[0037] Step 207: Perform business processing on the target text in conjunction with the text type and the corresponding business processing strategy.

[0038] Specifically, for target text of non-specially agreed types, conventional business processing mode is used for business processing; while for target text of specially agreed types, unconventional business processing mode is used for business processing.

[0039] In this embodiment, the following steps are taken: First, the target text is acquired. Then, semantic understanding-based segmentation is performed to obtain segmented text. Next, text information extraction is performed, including extracted scenario conditions and target business information. This extracted information is then vectorized and compared using a vector retrieval system. Based on the vector retrieval results, the text type of the target text is determined. Finally, the target text is processed according to the text type and the corresponding business processing strategy. Applying this data processing method to the processing of liability clauses in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition. This extracts scenario conditions and target business information, and, using a pre-defined vector retrieval library, identifies whether the current liability clause is a specially agreed-upon type, thereby selecting the appropriate business processing strategy and processing the target text. This enables rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, improving the efficiency of financial business processing.

[0040] Continue to refer to Figure 3 , Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 202 shown includes: Step 301: Input the target text into the preset semantic understanding segmentation processing model; Specifically, for example, the text of liability clauses in a financial business contract is input into a preset semantic understanding segmentation processing model, wherein the semantic understanding segmentation processing model includes the Qwen2-72B semantic understanding analysis model.

[0041] Step 302: Use the semantic understanding segmentation processing model to perform semantic content recognition on the scene conditions and business attribute information covered in the target text, wherein the business attribute information includes business attribute type and business attribute value; Step 303: Based on the semantic content recognition results, classify and segment the business attribute information corresponding to different scenario conditions, so that all business attribute information corresponding to the same scenario condition is added to the same segment, and obtain the segmented text.

[0042] Specifically, after semantic understanding of the target text, the target text may involve multiple scene conditions. In this case, based on the semantic content recognition results of steps 302 and 303, the business attribute information corresponding to different scene conditions is categorized and segmented, so that all business attribute information corresponding to the same scene condition is added to the same segment, thus obtaining the segmented text. This achieves segmentation of the target text according to different scene conditions.

[0043] Continue to refer to Figure 4 , Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 203 shown includes: Step 401: Identify the One-shot task examples provided by the preset Prompt knowledge base, and identify the task extraction format and extraction rules included in the One-shot task examples; In this embodiment, the preset Prompt knowledge base, such as an enterprise-level Prompt knowledge base for financial and insurance claims business scenarios, covers various claims business knowledge, claims rules, claims strategies, calculation methods, and other knowledge content when conducting financial and insurance claims business.

[0044] Step 402: According to the task extraction format and extraction rules, extract the target business attribute type and target business attribute value corresponding to different scenario conditions in each segment of text.

[0045] Specifically, the One-shot task example can be understood as a task information extraction template. Using this One-shot task example, the target business attribute type and target business attribute value corresponding to different scenario conditions in each segment of text can be extracted into a fixed template format.

[0046] By utilizing the One-shot task examples provided by the pre-set Prompt knowledge base, scenario conditions and target business information are extracted from all segmented texts. This allows the target business attribute types and values ​​corresponding to different scenario conditions in each segmented text to be extracted into a fixed template format, thereby improving computer processing efficiency. At the same time, it also improves recognition efficiency when processing intermediate information in the subsequent process.

[0047] Continue to refer to Figure 5 , Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown includes: Step 501: Take the target business attribute type and target business attribute value corresponding to different scenario conditions in each segmented text as a piece of intermediate information to be vectorized. Step 502: Each piece of intermediate information is vectorized using a preset vectorization processing technique to obtain the vectorized representation result corresponding to each piece of intermediate information, and the vector representation set to be identified is constructed. The preset vectorization processing technique includes any one of the following: word embedding-based vectorization processing technique, bag-of-words model-based vectorization processing technique, or TF-IDF algorithm-based vectorization processing technique. The word embedding-based vectorization processing technique includes the Word2Vec and GloVe models based on neural networks.

[0048] By converting text data recognition into vector representation recognition, the computer's recognition efficiency for the intermediate information can be improved to some extent.

[0049] Continue to refer to Figure 6 , Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown includes: Step 601: Extract each element from the vector representation set to be identified; Step 602: Input each element in the vector representation set to be identified as the current search element into the preset vector retrieval library, and compare its similarity with all elements in the reference vector representation set in the preset vector retrieval library. By first segmenting the target text, extracting intermediate information from each segment, and then vectorizing it, an element is obtained from the vector representation set to be identified. Subsequently, each element from the vector representation set to be identified is sequentially input into a preset vector retrieval library as the current retrieval element, and its similarity is compared with all elements in the reference vector representation set in the preset vector retrieval library. This achieves a more refined comparison vector retrieval from the perspective of comparing the target text as a whole.

[0050] Step 603: Select the elements in the reference vector representation set that have a similarity exceeding the preset similarity threshold and are the highest values, and use them as the reference vector corresponding to the current search element. Step 604: Organize all the reference vectors matched by the vector representation set to be identified, and obtain the reference vector retrieval results.

[0051] In this embodiment, the step of organizing all the reference vectors matched in the vector representation set to be identified to obtain the reference vector retrieval result includes: if all the matched reference vectors are duplicated, then the duplicated reference vectors are deduplicated.

[0052] Continue to refer to Figure 7 , Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 206 shown includes: Step 701: Identify the reference vector retrieval result corresponding to the vector representation set to be identified; Step 702: If the retrieval result of the reference vector corresponding to the vector representation set to be identified is empty, then the target text is classified as non-contracted text for business processing. Step 703: If the reference vector retrieval result corresponding to the vector representation set to be identified is not empty, then based on all the reference vectors included in the reference vector retrieval result, determine all the special texts covered by the target text, and divide the target text into business processing special texts, wherein the business processing non-special texts correspond to the liability clause texts of the non-special agreement type, and the business processing special texts correspond to the liability clause texts of the special agreement type.

[0053] In this embodiment, the step of performing business processing on the target text by combining the text type and the corresponding business processing strategy specifically includes: if the target text is a non-specialized business processing text, then a preset conventional business processing strategy is used to process the target text; if the target text is a specialized business processing text, then all specialized texts covered by the target text are identified; and the target text is processed according to the non-conventional business processing strategies corresponding to each of the specialized texts covered by the target text. Both the conventional and non-conventional business processing strategies involve corresponding business processing steps, the processing order between all business processing steps, and the constraints and processing parameters of each business processing step.

[0054] In this embodiment, the following steps are taken: First, the target text is acquired. Then, semantic understanding-based segmentation is performed to obtain segmented text. Next, text information extraction is performed, including extracted scenario conditions and target business information. This extracted information is then vectorized and compared using a vector retrieval system. Based on the vector retrieval results, the text type of the target text is determined. Finally, the target text is processed according to the text type and the corresponding business processing strategy. Applying this data processing method to the processing of liability clauses in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition. This extracts scenario conditions and target business information, and, using a pre-defined vector retrieval library, identifies whether the current liability clause is a specially agreed-upon type, thereby selecting the appropriate business processing strategy and processing the target text. This enables rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, improving the efficiency of financial business processing.

[0055] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0056] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0057] In this embodiment, the following steps are taken: First, the target text is acquired. Then, semantic understanding-based segmentation is performed to obtain segmented text. Next, text information extraction is performed, including extracted scenario conditions and target business information. This extracted information is then vectorized and compared using a vector retrieval system. Based on the vector retrieval results, the text type of the target text is determined. Finally, the target text is processed according to the text type and the corresponding business processing strategy. Applying this data processing method to the processing of liability clauses in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition. This extracts scenario conditions and target business information, and, using a pre-defined vector retrieval library, identifies whether the current liability clause is a specially agreed-upon type, thereby selecting the appropriate business processing strategy and processing the target text. This enables rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, improving the efficiency of financial business processing.

[0058] Further reference Figure 8 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data information processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0059] like Figure 8 As shown, the data information processing device 800 described in this embodiment includes: a target text acquisition module 801, a text segmentation processing module 802, a text information extraction module 803, a vectorization processing module 804, a vector retrieval module 805, a text type determination module 806, and a business processing module 807. Wherein: The target text acquisition module 801 is used to acquire target text. The text segmentation processing module 802 is used to perform semantic understanding-based segmentation processing on the target text to obtain segmented text. The text information extraction module 803 is used to extract text information from all segmented texts separately, and to use the extracted scenario conditions and target business information as intermediate information. The target business information includes target business attribute type and target business attribute value. The vectorization processing module 804 is used to perform vectorization processing on the intermediate information to obtain a vector representation set to be identified; The comparison vector retrieval module 805 is used to input the vector representation set to be identified into a preset vector retrieval library for comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized; The text type determination module 806 is used to determine the text type of the target text based on the comparison vector retrieval results; The business processing module 807 is used to perform business processing on the target text by combining the text type and the business processing strategy corresponding to the text type.

[0060] This application obtains target text; performs semantic understanding-based segmentation to obtain segmented text; extracts text information, including scenario conditions and target business information; performs vectorization processing and performs comparative vector retrieval; determines the text type of the target text based on the comparative vector retrieval results; and performs business processing on the target text by combining the text type and the corresponding business processing strategy. Applying this data information processing method to the processing of liability clause text in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition, extracting scenario conditions and target business information, and identifying whether the current liability clause text is a specially agreed type by combining a pre-given vector retrieval library. This enables the selection of appropriate business processing strategies and the execution of business processing for the target text. This achieves rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, thereby improving the efficiency of financial business processing.

[0061] In this embodiment, the text segmentation processing module 802 includes a target text input unit, a semantic content recognition unit, and a classification and segmentation processing unit. Wherein: The target text input unit is used to input the target text into a preset semantic understanding segmentation processing model; The semantic content recognition unit is used to perform semantic content recognition on the scene conditions and business attribute information covered in the target text using the semantic understanding segmentation processing model, wherein the business attribute information includes business attribute type and business attribute value; The classification and segmentation processing unit is used to classify and segment the business attribute information corresponding to different scenario conditions according to the semantic content recognition results, so that all business attribute information corresponding to the same scenario condition is added into the same segment to obtain the segmented text.

[0062] In this embodiment, the text information extraction module 803 includes: a task sample recognition unit and an information extraction execution unit. Wherein: The task sample recognition unit is used to recognize the one-shot task samples provided by the preset Prompt knowledge base, and to identify the task extraction format and extraction rules included in the one-shot task sample. The information extraction execution unit is used to extract the target business attribute type and target business attribute value corresponding to different scenario conditions in each segment of text, according to the task extraction format and extraction rules.

[0063] In this embodiment, the vectorization processing module 804 includes an intermediate information processing unit and a vectorization processing unit. Wherein: The intermediate information processing unit is used to treat the target business attribute type and target business attribute value corresponding to different scenario conditions in each segmented text as a piece of intermediate information to be vectorized. The vectorization processing unit is used to perform vectorization processing on each piece of intermediate information using a preset vectorization processing technique, so as to obtain the vectorization representation result corresponding to each piece of intermediate information and construct the vector representation set to be identified.

[0064] In this embodiment, the comparison vector retrieval module 805 includes a vector element extraction unit, a retrieval comparison unit, a comparison vector filtering unit, and a retrieval result processing unit. Wherein: A vector element extraction unit is used to extract each element from the vector representation set to be identified. The retrieval and comparison unit is used to input each element in the vector representation set to be identified as the current retrieval element into a preset vector retrieval library, and compare the similarity with all elements in the reference vector representation set in the preset vector retrieval library respectively. The reference vector filtering unit is used to filter out elements in the reference vector representation set corresponding to the highest value when the similarity exceeds the preset similarity threshold, and use them as the reference vector corresponding to the current search element. The retrieval result processing unit is used to process all the reference vectors matched by the vector representation set to be identified, and obtain the reference vector retrieval results.

[0065] In this embodiment, the text type determination module 806 includes a vector retrieval result recognition unit, a first target text type division unit, and a second target text type division unit. Wherein: A vector retrieval result recognition unit is used to recognize the reference vector retrieval result corresponding to the vector representation set to be recognized. The first segmentation unit for target text type is used to segment the target text into business processing non-special text if the retrieval result of the reference vector corresponding to the vector representation set to be identified is null. The second segmentation unit for target text type is used to determine all special texts covered by the target text based on all the reference vectors contained in the reference vector retrieval result if the reference vector retrieval result corresponding to the vector representation set to be identified is not empty, and to segment the target text into special texts for business processing.

[0066] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0067] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0068] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0069] The computer device 9 includes a memory 9a, a processor 9b, and a network interface 9c that are interconnected via a system bus. It should be noted that... Figure 9 Only a computer device 9 with component memory 9a, processor 9b, and network interface 9c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0070] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0071] The memory 9a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 9a may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 9a may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 9. Of course, the memory 9a may include both the internal storage unit and its external storage device of the computer device 9. In this embodiment, the memory 9a is typically used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions for a data information processing method. In addition, the memory 9a can also be used to temporarily store various types of data that have been output or will be output.

[0072] In some embodiments, the processor 9b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 9b is typically used to control the overall operation of the computer device 9. In this embodiment, the processor 9b is used to execute computer-readable instructions stored in the memory 9a or to process data, for example, to execute computer-readable instructions of the data processing method described above.

[0073] The network interface 9c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.

[0074] The computer device proposed in this embodiment belongs to the field of text processing technology and is applied in scenarios where business execution processing is performed based on agreed-upon terms and conditions. This application obtains the target text; performs semantic understanding-based segmentation processing to obtain segmented text; extracts text information, including scenario conditions and target business information; performs vectorization processing and performs a vector retrieval comparison; determines the text type of the target text based on the vector retrieval results; and performs business processing on the target text by combining the text type and the corresponding business processing strategy. Applying this data information processing method to the processing of liability clause text in financial business contracts allows for the semantic understanding and segmentation of the contract text after acquisition, extraction of scenario conditions and target business information, and identification of whether the current liability clause text is a specially agreed-upon type by combining a pre-given vector retrieval library. This enables the selection of appropriate business processing strategies and the execution of business processing for the target text. This achieves rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, improving the efficiency of financial business processing.

[0075] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the data information processing method described above.

[0076] The computer-readable storage medium proposed in this embodiment belongs to the field of text processing technology and is applied to scenarios where business execution processing is performed based on agreed-upon terms and conditions. This application obtains the target text; performs semantic understanding-based segmentation processing to obtain segmented text; extracts text information, including scenario conditions and target business information; performs vectorization processing and performs a vector retrieval comparison; determines the text type of the target text based on the vector retrieval results; and performs business processing on the target text by combining the text type and the corresponding business processing strategy. Applying this data information processing method to the processing of liability clause text in financial business contracts allows for the semantic understanding and segmentation of the contract text after obtaining it, extracting scenario conditions and target business information, and identifying whether the current liability clause text is a specially agreed-upon type by combining a pre-given vector retrieval library. This enables the selection of appropriate business processing strategies and the execution of business processing for the target text. This achieves rapid identification of clause types in financial business contracts and allows for separate business processing based on different clause types, improving the efficiency of financial business processing.

[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0078] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

Claims

1. A data information processing method, characterized in that, Includes the following steps: Obtain the target text; The target text is subjected to semantic understanding-based segmentation processing to obtain segmented text; Text information is extracted from all segments of text, and the extracted scenario conditions and target business information are used as intermediate information. The target business information includes target business attribute type and target business attribute value. The intermediate information is vectorized to obtain a set of vector representations to be identified; The vector representation set to be identified is input into a preset vector retrieval library for comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized; Based on the results of the vector retrieval, the text type of the target text is determined; The target text is processed in accordance with the text type and the corresponding business processing strategy.

2. The data information processing method according to claim 1, characterized in that, The step of performing semantic understanding-based segmentation on the target text to obtain segmented text specifically includes: Input the target text into a preset semantic understanding segmentation processing model; The semantic understanding segmentation processing model is used to perform semantic content recognition on the scene conditions and business attribute information covered in the target text, wherein the business attribute information includes business attribute type and business attribute value; Based on the semantic content recognition results, the business attribute information corresponding to different scenario conditions is classified and segmented, so that all business attribute information corresponding to the same scenario condition is added to the same segment, thus obtaining the segmented text.

3. The data information processing method according to claim 1, characterized in that, The step of extracting text information from all segmented texts and using the extracted scenario conditions and target business information as intermediate information specifically includes: The system identifies the one-shot task examples provided by the preset Prompt knowledge base, and identifies the task extraction format and extraction rules included in the one-shot task examples. According to the task extraction format and extraction rules, extract the target business attribute type and target business attribute value corresponding to different scenario conditions in each segment of text.

4. The data information processing method according to claim 1, characterized in that, The step of vectorizing the intermediate information to obtain the vector representation set to be identified specifically includes: The target business attribute type and target business attribute value corresponding to different scenario conditions in each segmented text are respectively regarded as a piece of intermediate information to be vectorized; Each piece of intermediate information is vectorized using a preset vectorization processing technique to obtain the vectorized representation result corresponding to each piece of intermediate information, thereby constructing the vector representation set to be identified. The preset vectorization processing technique includes any one of the following: word embedding-based vectorization processing technique, bag-of-words model-based vectorization processing technique, or TF-IDF algorithm-based vectorization processing technique.

5. The data information processing method according to claim 1, characterized in that, The step of inputting the vector representation set to be identified into a preset vector retrieval library for comparison vector retrieval specifically includes: Extract each element from the vector representation set to be identified; Each element in the vector representation set to be identified is sequentially input into a preset vector retrieval library as the current retrieval element, and a similarity comparison is performed with all elements in the reference vector representation set in the preset vector retrieval library. Select the elements in the reference vector representation set whose similarity exceeds the preset similarity threshold and is the highest value, and use them as the reference vector corresponding to the current search element; Organize all the reference vectors matched by the vector representation set to be identified to obtain the reference vector retrieval results.

6. The data information processing method according to any one of claims 1 or 5, characterized in that, The step of determining the text type of the target text based on the comparison vector retrieval results specifically includes: Identify the reference vector retrieval results corresponding to the vector representation set to be identified; If the retrieval result of the reference vector corresponding to the vector representation set to be identified is empty, then the target text will be classified as non-contracted text for business processing. If the reference vector retrieval result corresponding to the vector representation set to be identified is not empty, then based on all the reference vectors contained in the reference vector retrieval result, all the special texts covered by the target text are determined, and the target text is classified into special texts for business processing.

7. The data information processing method according to claim 6, characterized in that, The step of performing business processing on the target text by combining the text type and the corresponding business processing strategy specifically includes: If the target text is a business processing non-special text, then a preset routine business processing strategy is adopted to process the target text. If the target text is a business processing special text, then identify all special texts covered by the target text; Based on the non-routine business processing strategies corresponding to all the special texts covered by the target text, the target text is processed. The routine business processing strategy and the non-routine business processing strategy both involve corresponding business processing steps, the processing order between all business processing steps, the constraints and processing parameters of each business processing step.

8. A data information processing device, characterized in that, include: The target text acquisition module is used to acquire target text. The text segmentation module is used to perform semantic understanding-based segmentation on the target text to obtain segmented text. The text information extraction module is used to extract text information from all segmented texts separately, and to use the extracted scenario conditions and target business information as intermediate information. The target business information includes target business attribute type and target business attribute value. The vectorization processing module is used to perform vectorization processing on the intermediate information to obtain a set of vector representations to be identified; The comparison vector retrieval module is used to input the vector representation set to be identified into a preset vector retrieval library for comparison vector retrieval, wherein the comparison vector is the vectorized processing result corresponding to different historical special texts that have been pre-organized; The text type determination module is used to determine the text type of the target text based on the comparison vector retrieval results; The business processing module is used to perform business processing on the target text by combining the text type and the business processing strategy corresponding to the text type.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data information processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data information processing method as described in any one of claims 1 to 7.