Report generation method and device based on structured information processing, computer equipment and storage medium
By using automated data collection and structured storage technologies, combined with big data models, insurance industry reports are generated, solving the problem of low efficiency in information collection and analysis, achieving efficient and accurate product development support, and enhancing the competitiveness of insurance companies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-17
AI Technical Summary
In the insurance industry, the report generation process based on structured information processing suffers from low information collection efficiency and poor accuracy. In particular, the information collection and analysis stages rely on manual operation, resulting in high time costs, information omissions, and analytical biases, making it difficult to meet the efficiency and compliance requirements of product development.
Using a structured information processing approach, regulatory policies and competitor terms are automatically collected through web crawling technology, transformed into computable text, classified and structured for storage, and combined with big data models for difference comparison and intelligent recommendation to generate intelligent recommendation reports.
Significantly reduce information gathering time, avoid missing key information, improve the accuracy and efficiency of analysis, provide multi-dimensional decision support, ensure product design compliance and market adaptability, and reduce R&D trial and error costs.
Smart Images

Figure CN121881986A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of report generation technology based on structured information processing, and more specifically to report generation methods, apparatus, computer equipment, and storage media based on structured information processing. Background Technology
[0002] In the insurance industry within the financial sector, product development capabilities are crucial for insurance companies to build core competitiveness, directly determining their market position and development potential. Report generation based on structured information processing requires comprehensive consideration of market demand, competitor dynamics, regulatory policies, and risk control, among other factors. The efficiency, accuracy, and compliance of this process are vital to the successful market launch and realization of commercial value of the product.
[0003] However, the current product development process in the insurance industry is generally dominated by manual operations, which faces many problems that urgently need to be solved in practice, specifically in the following three core aspects: First, information gathering is difficult, and efficiency and completeness are hard to guarantee. In the early stages of generating reports based on structured information processing, product managers need to extensively acquire key information such as competitor information, market trend data, and regulatory policy documents to support product positioning and design decisions. Currently, most insurance companies rely on manual information gathering, requiring staff to search for relevant materials by accessing public information channels, regulatory agency websites, and industry research report publishing platforms. However, this information is often fragmented, and much of the data exists in unstructured forms, such as PDF documents in different formats, image-based industry reports, and scattered text fragments from web pages. Manually screening, organizing, and managing this unstructured data is not only time-consuming and labor-intensive, resulting in low efficiency, but also easily leads to the omission of key information due to the limitations of manual operation, such as specific coverage clauses of a competitor's product or the latest regulatory policy details. This can negatively impact the rationality and market adaptability of subsequent product design.
[0004] Secondly, the analysis tasks are arduous, and the accuracy and intelligence levels are insufficient. In the core analysis stage of report generation based on structured information processing, R&D personnel need to conduct detailed comparisons and in-depth analyses of collected competitor terms, rate standards, and regulatory requirements to identify market gaps and product differentiation directions. Currently, R&D personnel mainly rely on basic office tools such as Excel spreadsheets and Word documents for manual comparison work, needing to check the differences in terms line by line, manually calculate rate competitiveness, and verify compliance with regulatory requirements item by item. This process is not only time-consuming and labor-intensive, significantly extending the product development cycle, but also difficult to avoid analytical biases caused by human error, such as misjudging differences in compensation ratios among different products or omitting prohibited clauses in regulatory policies. At the same time, existing analysis methods lack intelligent auxiliary insight capabilities, failing to quickly extract potential patterns and market trends from massive amounts of data, and making it difficult to provide R&D personnel with accurate and quantifiable analytical conclusions, thus restricting the scientific nature of product design decisions.
[0005] To address these issues, some insurance companies have attempted to introduce data collection technologies for information acquisition and simple reporting tools to assist in data processing. However, these technologies are limited to optimizing a single process and have failed to form an integrated intelligent solution covering the entire "information collection-analysis" process. Consequently, they cannot fundamentally solve the core pain points of low efficiency and poor accuracy in the existing report generation process based on structured information processing. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a report generation method, apparatus, computer equipment and storage medium based on structured information processing.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: Report generation methods based on structured information processing include: Obtaining structured information on R&D requirement parameters; Collect formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms, and convert them into computable text; Computable text is classified and structured for storage to obtain structured term documents; The structured storage terms and conditions documents and R&D requirements parameters are input into the big data model to generate differentiated comparison parameters for different product terms; By combining differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results, an intelligent recommendation report is generated.
[0008] The present invention also provides a report generation apparatus based on structured information processing, comprising: The acquisition unit is used to acquire structured information about R&D requirement parameters. The data collection and conversion unit is used to collect formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms and conditions, and convert them into computable text. The classification storage unit is used to classify and structure computable text to obtain a structured document. The input generation unit is used to input structured storage terms documents and R&D requirement parameters into the big data model to generate differential comparison parameters for different product terms; The combined unit is used to combine differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results to generate intelligent recommendation reports.
[0009] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0010] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0011] The advantages of this invention compared to existing technologies are as follows: Through data collection technology, key information such as regulatory policies, industry trends, and competitor clauses can be automatically captured, eliminating the need for manual searching across multiple platforms, thus shortening information collection time and significantly reducing manual labor costs. Simultaneously, format conversion technology transforms unstructured files into computable text, solving the problem of inefficiently processing unstructured data manually and preventing the omission of key information such as competitor clauses and regulatory details due to format barriers. Furthermore, data collection can focus on core information sources for insurance research and development, avoiding interference from irrelevant information, and supports dynamic adjustment of the collection scope according to research and development needs, ensuring a high degree of alignment between acquired information and research direction. Moreover, classifying and structuring computable text transforms scattered clause information into an ordered, directly accessible data format, avoiding the tedious manual process of filtering and analyzing objects sentence by sentence from massive amounts of unstructured text. At the same time, structured data facilitates big data modeling. The system directly reads and processes data, providing an efficient data input foundation for subsequent comparisons. Furthermore, the big data model can accurately pinpoint comparison dimensions based on R&D needs, automatically marking differences between product terms and generating standardized, visualized results. This not only reduces comparison time from hours to minutes but also eliminates subjective bias, ensuring the objectivity and accuracy of analysis results. This provides direct support for R&D personnel to quickly grasp the core differences of competitors. In addition, by deeply integrating term difference data, claims big data analysis tables, and market sentiment data, it can provide a basis for R&D decisions from three dimensions: risk, market, and competitor levels. This avoids the bias caused by single data points. Through comprehensive analysis of multi-dimensional data, it outputs intelligent recommendation reports containing specific optimization directions, directly guiding product term design, helping R&D personnel quickly determine the core competitiveness of the product, reducing R&D trial-and-error costs, and improving market acceptance after product launch.
[0012] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A schematic diagram illustrating an application scenario of the report generation method based on structured information processing provided in this embodiment of the invention; Figure 2 A flowchart illustrating the report generation method based on structured information processing provided in an embodiment of the present invention; Figure 3A schematic block diagram of a report generation device based on structured information processing provided in an embodiment of the present invention; Figure 4 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0017] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0018] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0019] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the report generation method based on structured information processing provided in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating a report generation method based on structured information processing provided in an embodiment of the present invention. This report generation method based on structured information processing is applied to a server, which interacts with terminals. Through a full-process technical design encompassing demand acquisition, data collection and transformation, classification and storage, difference comparison, and intelligent recommendation, it effectively reduces the workload of manual labor, improves the efficiency and accuracy of each stage of product development, and helps insurance companies enhance their product development capabilities and strengthen their core competitiveness in the market.
[0020] Figure 2 This is a flowchart illustrating the report generation method based on structured information processing provided in an embodiment of the present invention. Figure 2As shown, the method includes the following steps S110 to S150.
[0021] S110, R&D requirement parameters for obtaining structured information; Specifically, structured information refers to insurance products. The system supports R&D personnel to input requirements through multiple platforms (Web, App, and internal OA system). Input formats include "text description" (e.g., "Develop a critical illness insurance product for people aged 30-45, covering specific rare diseases"), "upload competitor documents" (e.g., directly uploading the terms and conditions of a competitor's critical illness insurance product), and "select a requirement template" (the system has built-in "critical illness insurance R&D requirement templates" and "medical insurance R&D requirement templates," containing preset fields such as coverage period, payout ratio, and target customer group; personnel only need to fill in key parameters). Furthermore, NLP technology is used to decompose the input unstructured requirement text, automatically extracting core elements and generating "requirement tags." For example, the tag "coverage scope - rare diseases" is extracted from "coverage of specific rare diseases," and the tag "target customer group - young and middle-aged people" is extracted from "people aged 30-45." Simultaneously, the extracted elements are linked to the system's knowledge base (including insurance product classification and core parameter definitions) to form a standardized "structured form for R&D requirements," avoiding subsequent R&D deviations caused by ambiguous requirement descriptions.
[0022] In other words, by structuring abstract requirements into explicit parameterized forms, the "misunderstanding of requirements" that often occurs in traditional manual communication (such as developers misinterpreting "rare disease coverage" as "some rare diseases") is avoided. This ensures that subsequent information collection and clause analysis are based on accurate requirements, reducing R&D rework costs. Furthermore, multi-platform input and templated design lower the barrier to entry for R&D personnel to submit requirements, eliminating the need to manually write complex requirement documents. Structured forms can be directly synchronized to the "intelligent search module" and "intelligent analysis module," avoiding secondary entry of requirements across modules and enabling "one-time input, full-process reuse" of R&D requirements.
[0023] In one embodiment, the R&D requirement parameters for obtaining structured information include: To obtain R&D requirement parameters, obtain structured information or specific requirements input by R&D personnel through the user interaction layer.
[0024] Specifically, the web-based input supports both structured templates and free text input. The system's web interface has a built-in "Report Generation Requirement Template Library Based on Structured Information Processing," with preset templates categorized by product type (critical illness insurance, medical insurance, life insurance, etc.). The templates include "core fields": target customer group (e.g., "young and middle-aged people aged 30-45" or "parents of newborns"), coverage (e.g., "covering 120 critical illnesses + 20 moderate illnesses" or "including inpatient medical care + outpatient surgery"), premium range (e.g., "annual premium of 500-2000 yuan"), and regulatory compliance requirements (e.g., "compliant with the new regulations for critical illness insurance in 2024"). Developers can directly fill in the parameters in the templates without manually outlining the requirement framework. Simultaneously, a "free text input box" is supported, allowing developers to describe product direction using unstructured text (e.g., "developing a small-amount medical insurance product suitable for county-level markets, focusing on cost-effectiveness and convenient claims processing"), meeting the needs of scenarios where specific parameters are not clearly defined in the initial stages of requirement development.
[0025] App Input: Adapted for mobile scenarios, supports uploading attachments in multiple formats. The app's input interface has been optimized with a "lightweight form" design, retaining only core required fields (product type, core requirement direction). Developers can quickly input preliminary requirements when conducting field research (e.g., "Research found that county-level users are concerned about 'low inpatient deductible,' which needs to be included in medical insurance R&D requirements"). It supports uploading attachments in multiple formats, such as photos of competitor promotional pages taken at the research site and user interview recordings transcribed into text documents. The content of the attachments can be automatically linked to the requirement record, providing support for subsequent requirement refinement.
[0026] Enterprise internal OA system integration: Enables linkage between requirements and internal collaborative processes, and connects with the insurance company's internal OA system (such as DingTalk and WeChat Workbench). R&D personnel can directly initiate requirement input in the "Product Development Module" of the OA system. After the input is completed, a "Requirement Work Order" is automatically generated and synchronized to the department head's approval process. Approved requirement work orders will be automatically synchronized to the report generation support system based on structured information processing, avoiding repeated data entry by R&D personnel across multiple systems and achieving integrated collaboration of "requirement submission - approval - system access".
[0027] In addition, the system automates the processing of the requirements information input by R&D personnel (whether structured templates or free text) and transforms it into standardized data that the system can recognize and reuse.
[0028] S120. Collect formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms, and convert them into computable text; Specifically, a distributed web crawling framework (Scrapy + multi-threaded scheduling) is adopted. Collection rules are configured based on tags (such as "critical illness insurance" and "regulatory policies") in the "structured form of R&D requirements." This allows for targeted crawling of official websites (regulatory policies), insurance industry association announcements (industry standards), and publicly available competitor terms and conditions (e.g., PDF files) from insurance company websites, avoiding interference from irrelevant information. Simultaneously, a scheduled task is configured to support "incremental updates" (e.g., crawling new policies from regulatory websites every day at midnight) and "full re-collection" (e.g., re-crawling all competitor terms and conditions on the 1st of each month to ensure information timeliness). "Real-time collection" can also be triggered as needed (e.g., when R&D personnel need to obtain newly released regulatory details, targeted crawling can be manually triggered). Furthermore, for editable PDF files, text content is extracted using PDF parsing tools (such as PyPDF2); for scanned documents and image-based PDFs (where text cannot be directly extracted), OCR recognition technology (such as Tesseract + an insurance industry-specific training model, optimizing the accuracy of recognizing professional terms such as "compensation ratio" and "exclusion clauses") is introduced to convert image content into editable text. Redundant spaces and garbled characters are removed from the editable text, and professional terminology is standardized (e.g., "payment ratio" and "compensation ratio" are standardized to "compensation ratio"), ultimately generating a standardized, computable text that can be directly used for subsequent analysis.
[0029] In other words, automated data collection replaces manual searching, reducing the time required to gather regulatory policies and competitor clauses from 2-3 days to 1-2 hours, significantly lowering labor costs. Collection rules ensure a high degree of alignment between information and R&D needs, avoiding the time-consuming filtering of irrelevant information and the omission of crucial information (such as missing newly released regulatory details for critical illness insurance) that often occur with manual searches. Furthermore, OCR technology and industry-specific models overcome the barriers to processing unstructured documents, solving the problem of inefficiently extracting scanned and image-based clause content using traditional manual methods. This ensures the complete and comprehensive conversion of competitor clauses and regulatory documents, providing a solid data foundation for subsequent analysis. Moreover, scheduled updates and incremental updates ensure real-time synchronization of regulatory policies and competitor information, preventing product designs from failing to meet the latest regulatory requirements due to information lag. Text preprocessing and standardized terminology eliminate analytical biases caused by different expressions of the same concept.
[0030] In one embodiment, the process of collecting publicly available regulatory websites, industry association announcements, and competitor terms in format and converting them into computable text includes: Based on the product type and market positioning of the structured information, determine the scope of publicly available regulatory websites, industry association announcements, and competitor companies that need to be collected, and form a target list for collection. Specifically, the system pre-defines associated regulatory websites, industry associations, and competitor companies based on insurance product types (critical illness insurance, medical insurance, life insurance, annuity insurance, etc.). If the development requirements explicitly state "market positioning as small-amount medical insurance in county-level markets," then in addition to the default collection scope, it supplements the "County-level Insurance Product Guidelines" announcements from local insurance industry associations (such as a provincial insurance industry association). Simultaneously, it includes small and medium-sized insurance companies focusing on county-level markets (such as a regional life insurance company) in the competitor collection list, ensuring a high degree of alignment between the collected information and the target market. Furthermore, the system automatically generates a structured collection target list, including fields such as "target type (regulatory website / industry association / competitor company), target name, specific URL, type of content to be collected (policy document / terms PDF), and update frequency (daily / weekly / monthly)." For example, "Competitor Company - A Life Insurance Company - Official Website Product Terms Page URL - Critical Illness Insurance Terms PDF - Weekly Update." If a key competitor company or a regulatory document section is missing from the list, R&D personnel can manually add the target information; if the URL of a competitor company's terms page is invalid, the system will automatically prompt and support replacement and updates to ensure the completeness and effectiveness of the target list.
[0031] Automatically access and download formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms according to the target data collection list; Specifically, an architecture of "Scrapy framework + multi-threaded scheduler + proxy IP pool" is adopted. Scrapy is responsible for parsing the target URL page structure and extracting PDF file download links. The multi-threaded scheduler supports launching 10-20 collection tasks simultaneously (adjustable according to server performance), reducing the single-target collection time from minutes to seconds. The proxy IP pool automatically switches IP addresses to avoid IP blocking issues caused by frequent access to the same website, ensuring collection continuity (especially against anti-scraping mechanisms of regulatory websites and competitor company websites). In addition, scheduled tasks are configured according to the "update frequency" of the collection target list. For example, for the "daily update" section of the regulatory website, automatic collection is set at 2:00 AM every day (avoiding peak website access times); for the "weekly update" terms and conditions pages of competitor companies, full collection is set at 3:00 AM every Monday. "Instant collection" triggering is also supported. If developers need to obtain a newly released regulatory policy PDF, they can manually select the corresponding item in the target list to trigger the crawler to immediately execute collection, meeting urgent needs. Meanwhile, the crawler automatically identifies PDF file links on the page (by parsing content containing ".pdf" in the "href" attribute of HTML tags) and stores them on the system server according to the path rule of "target type / target name / collection date / file name" (e.g., "regulatory website / 20240510 / critical illness insurance new regulations.pdf"), facilitating subsequent traceability and management. If the PDF download fails due to network interruption during the collection process, it will automatically resume downloading from the breakpoint on the next collection to avoid duplicate downloads; by verifying the file's MD5 value, if a PDF file already exists and its content has not been updated (MD5 value matches), the download will be skipped, reducing server storage usage and network bandwidth consumption.
[0032] Convert the downloaded format file into a computable text format to obtain computable text.
[0033] Specifically, the system first determines the type of the downloaded PDF file by parsing its metadata, distinguishing between "editable PDFs" (where the text layer can be directly extracted) and "non-editable PDFs" (scanned documents, image-based PDFs, without a text layer). For example, policy PDFs published on regulatory websites are mostly editable, while some competitor companies' earlier clause PDFs may be scanned documents. For editable PDFs, tools such as PyPDF2 and PDFMiner are used to directly extract the text content, preserving paragraph structure, table borders, and other formatting information (e.g., extracting the "compensation ratio table" from the clause as a text table according to its original structure), thus avoiding subsequent analysis biases caused by lost formatting. For non-editable PDFs, OCR recognition technology (such as TesseractOCR + insurance industry-specific training model) is introduced. First, each page of the PDF is converted into an image format (such as PNG), and then the text content in the image is recognized by the OCR model. For tables (such as rate tables) and formulas (such as compensation amount calculation formulas) in PDFs, "table recognition algorithms (such as TableNet)" and "formula recognition modules" are used to assist in extraction, ensuring that the row and column relationships of tables and formula symbols are accurately converted into text format (such as accurately recognizing "annual premium = sum insured × rate coefficient").
[0034] In addition, the extracted text undergoes "cleaning and optimization" to remove redundant spaces and garbled characters (such as invalid characters like "□" and " " generated during PDF conversion), and to standardize the expression of professional terminology (e.g., unifying "payment ratio" and "compensation ratio" as "compensation ratio", and unifying "deductible rate" and "deductible amount ratio" as "deductible rate") to avoid subsequent classification and comparison errors caused by inconsistent terminology, thus forming preprocessed text. Furthermore, the preprocessed text is stored in TXT or JSON format (containing metadata such as "text content, source file path, collection time, and file type"), and a mapping is established between the text and the original PDF file. When developers view the text later, they can jump to the corresponding page in the original PDF file with one click, facilitating verification of text accuracy (e.g., confirming that the "compensation ratio 80%" recognized by OCR is consistent with the original PDF).
[0035] In other words, the target data collection list generated based on "product type + market positioning" can filter out information irrelevant to R&D needs (such as when developing "county-level micro-medical insurance," there is no need to collect competitor terms for high-end medical insurance), avoiding invalid information from consuming storage resources and subsequent analysis time. In addition, by supplementing targets such as local associations and regional competitors, it is ensured that the collected information covers the specific policy requirements of the target market and the dynamics of competitors (such as premium restrictions and claims service characteristics in county-level markets), providing data support for product design to meet the needs of the target market and reducing product positioning deviations caused by incomplete information. Furthermore, addressing the pain points of traditional manual "site-by-site retrieval and manual PDF download" which is time-consuming and labor-intensive, and prone to failure due to IP blocking and network interruptions, this solution achieves "automated and efficient" collection through a distributed crawler framework. Multi-threaded scheduling and scheduled tasks shorten the original 2-3 days of manual collection work to 1-2 hours, significantly reducing manual time costs. In addition, proxy IP pools and breakpoint resume mechanisms ensure the stability of collection, avoid interruptions caused by regulatory policies and competitor terms due to IP blocking, ensure that "daily" and "weekly" updates are obtained on time, and prevent product development from failing to meet the latest regulatory requirements due to information lag.
[0036] S130. Classify and structure the computable text to obtain a structured document of terms. Specifically, a "large-scale zero-shot text classification technology" (based on a pre-trained language model such as BERT, fine-tuned for insurance scenarios) is adopted to automatically classify texts according to "terms function": competitor terms are divided into categories such as "definition terms" (e.g., "definition of critical illness"), "coverage liability terms" (e.g., "cancer compensation rules"), "exclusion terms" (e.g., "no compensation for critical illnesses caused by war"), and "claims process terms" (e.g., "reporting time limit requirements"). Regulatory documents and industry announcements are classified according to "policy type" (e.g., "product filing requirements" and "rate regulation rules"). In addition, multimodal auxiliary classification is introduced to identify the associated text content of tables (e.g., competitor rate tables), formulas (e.g., compensation amount calculation methods), and images (e.g., disease diagnosis standard diagrams) embedded in the text, and determine their classification (e.g., rate tables are associated with the "rate terms" category). At the same time, a "human feedback loop" is added, allowing developers to correct classification errors (e.g., reclassifying "coverage liability terms" that were mistakenly identified as "exclusion terms"). The correction results are used to fine-tune the classification model to improve the accuracy of subsequent classifications.
[0037] Furthermore, an insurance knowledge graph is constructed, using "nodes" to represent insurance products (e.g., "Critical Illness Insurance A from a certain company"), terms (e.g., "Critical Illness Insurance A's Exclusion Clauses"), and regulatory documents (e.g., "2024 Critical Illness Insurance Regulatory Rules"), and "edges" to represent relationships (e.g., "Critical Illness Insurance A includes Exclusion Clause X" and "Critical Illness Insurance A must comply with the 2024 Critical Illness Insurance Regulatory Rules"), thus realizing the logical connection between terms, products, and policies. Simultaneously, the categorized computable text (terms and policies) is transformed into vectors using an embedding model and stored in a vector database (e.g., Milvus); key information from the text (e.g., product name, payout ratio, regulatory document number) is extracted as "structured fields" and stored in a relational database (e.g., MySQL), forming a "vector + structured" dual storage mode. This supports "semantic retrieval" (e.g., inputting "Critical Illness Insurance Rare Disease Payout" to match relevant term vectors) and "precise field query" (e.g., querying "Critical Illness Insurance Terms with Payout Ratio ≥ 90%)", ultimately yielding a structured stored term document.
[0038] In other words, the classification and dual-storage model transforms scattered, computable text into structured assets that are "clearly categorized, well-connected, and searchable," avoiding the time-consuming search problems caused by storing clauses as "unordered folders" in traditional manual methods (e.g., opening each PDF individually to find a particular type of disclaimer clause). Developers can locate target clauses within minutes through "classification filtering + semantic retrieval," significantly improving data reuse efficiency. Furthermore, the categorized clauses are functionally broken down and can be directly used as comparison units in the "intelligent analysis module" (e.g., extracting only "coverage liability clauses" for competitor analysis), avoiding interference from irrelevant text. The knowledge graph's relationships support analysis of "whether clauses comply with regulatory requirements" (e.g., whether a compensation clause conflicts with regulatory details), and the semantic retrieval capabilities of the vector database can assist in discovering "similar clauses" (e.g., disclaimers with different wording but the same meaning in different competitors), improving the comprehensiveness of the analysis.
[0039] In one embodiment, classifying and structuring the computable text to obtain a structured term document includes: Noise removal is performed on the computable text to obtain cleaned text; Specifically, common noise types in calculable text are identified, including formatting redundancy (such as page numbers left over from PDF conversion, headers and footers like "Product Terms of an Insurance Company" or "Page X"), character interference (such as garbled characters " " or "□", extra spaces / line breaks), invalid content (such as non-core statements at the end of the terms like "Disclaimer: This document is for illustrative purposes only; the final contract shall prevail"), and duplicate text (such as "Definition of Insurance Liability" appearing repeatedly in different chapters). Furthermore, regular expressions are used to match and remove formatting redundancy (e.g., using "Page \d+" to remove page numbers), and character encoding verification (e.g., filtering non-standard characters using UTF-8 encoding) is used to remove garbled characters. Invalid content is marked based on keyword matching (e.g., "Disclaimer" or "Example") and deleted after confirmation by developers. Duplicate segments are identified using text fingerprinting algorithms (e.g., SimHash), retaining the first complete occurrence and deleting subsequent repetitions. Meanwhile, the system automatically outputs a "text comparison report before and after cleaning", marking the deleted noise content (such as "deleted page numbers: page 3, page 5; deleted garbled characters: 2 places"), which is convenient for R&D personnel to trace. For ambiguous noise (such as text that is suspected to be redundant but may contain core information, such as "this clause applies to people aged 18-60"), the system pops up a window to prompt R&D personnel to judge whether to retain it, avoiding accidental deletion of key content, and finally generating clean text without redundancy and interference.
[0040] The cleaned text is fed into a pre-trained language model to generate classification results; Specifically, a pre-trained language model fine-tuned for the insurance industry is selected. Based on BERT and GPT series pre-trained models, it is fine-tuned using insurance industry corpus (including massive clauses and regulatory documents) to optimize the semantic understanding ability of professional terms such as "insurance liability," "exclusion clauses," and "claims process." The model input is cleaned complete text, and the output is "classification label + confidence score" (e.g., "classification label: protection liability clauses; confidence score: 0.98"). In addition, a classification dimension and label system are set, and a two-level classification label system is established: the first-level labels are divided by function (e.g., "protection liability category," "exclusion category," "claims category," "definition category," "regulatory compliance category"), and the second-level labels are subdivided into specific scenarios (e.g., "protection liability category" is divided into "critical illness compensation clauses," "medical expense reimbursement clauses," and "death benefit clauses") to ensure that the classification granularity fits the needs of insurance R&D. If the cleaned text contains tables (such as "Critical Illness Payout Ratio Table") and formulas (such as "Payout Amount = Sum Insured × Payout Ratio") associated with the original PDF, a multimodal model can be invoked to analyze the table titles and formula associations to help confirm the classification labels (such as associating the "Critical Illness Payout Ratio Table" with the "Critical Illness Payout Clause" label). Developers can then correct the classification results (such as reclassifying "Inpatient Medical Reimbursement Clause" which was mistakenly classified as "Exclusion Category" as "Critical Illness Payout Clause"). The corrected data is automatically stored in the model training set, and the pre-trained model is fine-tuned periodically (such as monthly) to improve the subsequent classification accuracy, forming an adaptive closed loop of "model classification - manual correction - model optimization".
[0041] The classification results are stored in a structured format to obtain a structured document of terms.
[0042] Specifically, a three-in-one storage architecture of "relational database + knowledge graph + vector database" is constructed: Relational databases (such as MySQL): store structured field information, including "text ID, category tags (first-level + second-level), core content summary (such as "critical illness compensation clause: a one-time payment of 80% of the insured amount after diagnosis of critical illness"), source file path, and collection time", and support quick filtering by category tags, collection time and other dimensions; Knowledge Graph: Referring to the "node-edge" modeling logic, "cleaned text" is used as a node. It is associated with category tag nodes through "belongs to" edges (such as "text A - belongs to - protection liability category - critical illness compensation clauses"), regulatory document nodes through "reference" edges (such as "text A - reference - 2024 new regulations for critical illness insurance"), and competitor company nodes through "associated products" edges (such as "text A - associated product - critical illness insurance B of a certain company"), so as to realize the logical association between text and categories, regulations, and competitors. Vector databases (such as Milvus): After cleaning the text, it is converted into vectors through an embedding model (such as Sentence-BERT) and stored. It supports semantic retrieval (such as inputting "compensation ratio after diagnosis of critical illness" and matching "critical illness compensation clause" text with similar vectors).
[0043] In addition, the system automatically integrates the aforementioned stored data to generate a "structured storage terms document," with category tags as the directory (e.g., "Coverage Liability Category → Critical Illness Claim Terms"). Each directory displays the core summary, complete content, and associated knowledge graph nodes of the corresponding text (e.g., clicking "Related Product - Critical Illness Insurance B from a Certain Company" will take you to view the complete terms of that competitor's product). Researchers can quickly locate target texts through "Category Tag Filtering" (e.g., selecting "Exclusion Category - War Exclusion Terms"), "Semantic Search" (e.g., entering "Medical Insurance Deductible Rules"), and "Knowledge Graph Association Query" (e.g., querying "All Terms Citing a Certain New Regulatory Regulation"). The document supports exporting to Word / Excel format for easy subsequent analysis.
[0044] In other words, automated cleaning rules efficiently remove invalid information such as page numbers and garbled characters, preventing noise from interfering with the model's semantic understanding of core clauses. Manual correction ensures no key information is lost, providing a high-quality text foundation for subsequent model classification and clause comparison, and reducing analytical bias caused by noise. Furthermore, relational databases support rapid filtering by category tags, and vector databases support semantic retrieval, significantly improving text retrieval efficiency. Additionally, knowledge graph association enables integrated queries across "text-category-regulation-competitors," providing direct data support for subsequent clause comparisons (such as comparing deductible differences among different competitors), avoiding repeated data retrieval across systems, and promoting efficient integration of the R&D process.
[0045] S140. Input the structured storage terms document and R&D requirement parameters into the big data model to generate differentiated comparison parameters for different product terms; Specifically, target clauses (such as critical illness insurance coverage clauses from 3-5 competing products) are filtered from structured stored terms documents based on tags in the "R&D Requirements Structured Form" (e.g., "Critical Illness Insurance - Coverage") as the comparison data source. Simultaneously, "core focus dimensions" from the R&D requirements (e.g., "Rare Disease Coverage," "Payout Ratio," "Waiting Period") are input into the big data model as the comparison focus. A combined architecture of "large model + rule engine" is adopted, with a pre-trained large model (e.g., a finely tuned version of GPT-4) as the core, responsible for text semantic understanding and difference identification. Insurance industry-specific rules are embedded (e.g., "Payout ratio difference threshold ≥ 5% requires marking" and "Waiting period difference ≥ 30 days requires marking") to ensure that difference identification conforms to industry understanding.
[0046] Furthermore, the model first performs "sub-unit decomposition" on the selected clauses, further breaking down the "coverage liability clauses" into the smallest analytical units such as "rare disease coverage sub-clauses," "cancer compensation sub-clauses," and "cardiovascular and cerebrovascular disease compensation sub-clauses," ensuring precise granularity of comparison. For each sub-unit, the model performs "semantic-level difference identification." For example, comparing the "rare disease coverage sub-clauses" of competitor A and competitor B, it identifies that "A covers 20 rare diseases, while B covers 15, of which A has 5 unique diseases, namely XXX"; comparing the "compensation ratio sub-clauses," it marks that "A's cancer compensation ratio is 90%, while B's is 80%." In addition, the comparison results can automatically generate comparison tables in Excel / Word format (columns include "comparison dimension," "competitor A," "competitor B," "competitor C," and "difference summary"), with differences highlighted in red; it also generates visual charts (such as bar charts showing the differences in compensation ratios among competitors), allowing R&D personnel to export them with one click for direct use in internal reports or analysis documents.
[0047] In other words, automated segmentation and semantic comparison replace manual sentence-by-sentence verification, reducing the time required to compare clauses from three competing products (originally 1-2 hours) to 10-15 minutes. The model accurately marks differences based on industry rules, avoiding "missed differences" (e.g., overlooking coverage of a rare disease) or "misjudgments" (e.g., considering clauses with different wording but the same meaning as differences) caused by human "visual fatigue" or "professional oversight," thus improving comparison accuracy. Furthermore, the visualization results transform complex clause differences into "intuitive tables and charts," allowing developers to quickly grasp the core differences between competing products (e.g., "which competitor has the broadest rare disease coverage" or "which has the highest payout ratio") without having to read multiple clause documents in depth. This provides a direct basis for product differentiation design (e.g., "addressing the pain point of competitor B's limited rare disease coverage, this product can add coverage for 5 more rare diseases").
[0048] In one embodiment, the step of inputting structured storage terms documents and R&D requirement parameters into a big data model to generate differential comparison parameters for different product terms includes: Analyze the key parameters of the R&D requirements to generate a list of requirement keywords; Specifically, an insurance R&D demand parameter system should be established, classifying key parameter types according to four dimensions: "product attributes," "coverage requirements," "customer characteristics," and "regulatory compliance." For example: Product attributes: Product type (critical illness insurance / medical insurance / life insurance), coverage period (1 year / 20 years / lifetime), payment method (annual payment / monthly payment); Coverage requirements: reimbursement ratio (e.g., "cancer reimbursement ≥ 90%"), deductible (e.g., "medical insurance deductible ≤ 5000 yuan"), coverage scope (e.g., "covering 120 critical illnesses + 20 moderate illnesses"). Customer characteristics: target age (e.g., "30-45 years old"), occupation category (e.g., "occupations of categories 1-3"); Regulatory compliance: Applicable regulatory documents (such as "Compliant with the new regulations for critical illness insurance in 2024").
[0049] The system employs NLP semantic parsing technology to extract parameters and calls upon an insurance industry-specific NLP model to perform word segmentation and entity recognition on the R&D requirement parameters. For example, from the requirement "Develop a lifetime critical illness insurance policy for people aged 30-45, with cancer payout ≥90%, complying with the 2024 new regulations for critical illness insurance," the system automatically extracts the parameters "Product type - critical illness insurance, Coverage period - lifetime, Target age - 30-45 years old, Cancer payout ratio - ≥90%, Regulatory compliance - 2024 new regulations for critical illness insurance." Simultaneously, the system transforms the extracted key parameters into a structured keyword list of "parameter type + parameter value + priority." If the model misses parameters, developers can supplement or correct them, ultimately generating a precise list of requirement keywords as the core basis for subsequent clause matching.
[0050] Based on the list of required keywords, match the corresponding related sub-units in the structured storage terms document to form a correlation lookup table; Specifically, the structured storage terms document is split into multi-level sub-units, referencing the splitting logic of "definition—coverage liability—exclusions—claims process," and further refined to the smallest analytical unit (sub-unit). For example, the "coverage liability terms" are split into "critical illness payout sub-unit," "moderate illness payout sub-unit," and "cancer additional payout sub-unit." Each sub-unit contains "sub-unit name, core content, and parameter value" (e.g., "cancer additional payout sub-unit - core content: 50% additional payout upon diagnosis of cancer - parameter value: 50%)." In addition, a "keyword-sub-unit" index is built for the sub-units. Based on knowledge graph technology, the "sub-unit name, core content, and parameter value" of the sub-units are associated with keywords (e.g., the keyword "cancer payout ratio" is associated with "cancer additional payout sub-unit" and "critical illness payout sub-unit (including cancer payout)"). This supports quick location of sub-units based on required keywords. Simultaneously, the system matches sub-units sequentially according to the "priority" of the keyword list, prioritizing high-priority keywords (such as "product type - critical illness insurance" and "cancer payout ratio - ≥90%)" to filter out "cancer payout-related sub-units" of all critical illness insurance products from the structured terms document; then it matches medium-priority keywords (such as "target age - 30-45 years old") to further filter out critical illness insurance cancer payout sub-units targeting that age group. The correlation table is formatted with "demand keywords" as rows and "competitor terms / regulatory terms" as columns, and the information of the matched sub-units is filled in. Import the correlation comparison table, the list of demand keywords, and the sub-cell data of the structured storage terms document into the big data model to generate differentiated comparison parameters for different product terms.
[0051] Specifically, a combined model of "large model + industry rule engine" is adopted, with a fine-tuned version of GPT-4 and the BERT model as the core, responsible for semantic understanding and difference identification of sub-unit text. Insurance industry-specific rules are embedded (such as "payout ratio differences ≥10% require special marking" and "age range overlap less than 50% requires prompting") to ensure that difference identification meets industry R&D needs. The correlation comparison table (including requirement-sub-unit matching relationships), the list of requirement keywords (including parameter thresholds, such as "≥90%)", and sub-unit data (including core content and parameter values) are input into the model in a structured format (JSON), specifying the dimensions the model needs to compare (such as "parameter value differences", "text expression differences", and "compliance differences").
[0052] Furthermore, the parameter values and demand thresholds of different clause sub-units were compared. For example, "Competitor A pays out 40% for cancer and Competitor B pays out 60%, both lower than the demand threshold of 90%, with differences of 50% and 30% respectively." The semantic differences in the core content of the sub-units were identified. For example, "Competitor A states 'payment within 30 days of cancer diagnosis,' while Competitor B states 'payment after 15 days of survival following cancer diagnosis,' with the core difference being the 'survival waiting period.'" The sub-units were compared with regulatory clause sub-units. For example, "Competitor C's 'no payout for congenital diseases' exclusion sub-unit does not meet the regulatory clause requirement of 'payment is available 3 years after diagnosis of congenital diseases.'" Meanwhile, the comparison is displayed in columns of "Difference Type, Involved Clauses, Sub-unit ID, Difference Content, and Demand Matching Degree," with differences highlighted in color (e.g., parameter values below the demand threshold are marked in red, and compliance issues are marked in orange). A bar chart shows the parameter differences between different competitors and the demand threshold (e.g., differences in cancer payout ratios), and a radar chart shows the multi-dimensional demand matching degree (e.g., competitor A's matching degree in "payout ratio," "age suitability," and "compliance" is 45%, 80%, and 100%, respectively). The comparison tables and charts can be exported to Word / Excel / PDF formats with one click and used directly in R&D analysis reports, i.e., comparison parameters of different product clause differences.
[0053] In other words, the keyword list clearly defines the core R&D requirements (e.g., "cancer payout ≥ 90%)", avoiding getting bogged down in irrelevant clauses during comparison (e.g., death benefit clauses in life insurance), and reducing ineffective analysis time. The correlation comparison table directly establishes the correspondence between "requirement" and "sub-unit", eliminating the need for R&D personnel to sift through massive amounts of clauses and allowing them to quickly focus on sub-units relevant to their needs, thus improving the targeting of comparisons. Furthermore, the model not only identifies differences in parameter values (e.g., payout ratios) but also captures differences in textual expression (e.g., survival waiting period) and compliance differences (e.g., conflicts with regulatory clauses), avoiding missed discrepancies due to human error such as "visual fatigue" or "professional oversight" (e.g., overlooking the difference between "survival for 15 days" and "within 30 days"), thus improving the completeness of comparisons. In addition, the difference comparison tables and charts intuitively display "competitor differences" and "gaps with requirements" (e.g., a radar chart clearly showing competitor A's low matching degree in payout ratios), allowing R&D personnel to quickly grasp the core differences without interpreting complex text.
[0054] S150 combines differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results to generate intelligent recommendation reports.
[0055] Specifically, input "terms differentiation comparison parameters" (e.g., "competitors generally have a cancer payout ratio of 80%-90%, and cover 10-20 rare diseases"); access "claims big data analysis table" (including insurance company's historical customer claims records, extracting "payout rate", "reasons for denial", and "proportion of claims for high-incidence diseases", such as "cancer claims account for 35%, and 'not meeting the diagnostic criteria' accounts for 40% of the reasons for denial"); crawl "market sentiment data" (e.g., user reviews on social media and insurance evaluation platforms, extracting "positive demands" such as "hope to increase cancer payouts" and "negative feedback" such as "waiting period is too long") through sentiment analysis, and calculating "volume ranking". Data fusion is achieved using a combination of machine learning models and weight allocation. The XGBoost model is used for "risk prediction" of claims data (e.g., "If the cancer payout ratio is increased to 95%, the payout rate is predicted to increase by 2.3%"). The LDA topic model is used for "demand mining" of public opinion data (e.g., "The public opinion volume related to 'rare disease protection' has increased by 15% month-on-month, indicating high user attention"). Based on the importance of R&D needs (e.g., "The target customer group is young and middle-aged people, so the weight of 'cancer protection' is set to 0.4, and the weight of 'waiting period' is set to 0.2"), weights are assigned to three categories of data: "term differences, risk prediction, and public opinion needs," forming a comprehensive analytical basis.
[0056] Furthermore, the intelligent recommendation report comprises three parts: core conclusions, specific suggestions, and data support. Core conclusions include opportunities such as "insufficient coverage of rare diseases and room for improvement in cancer payout ratios" in the current critical illness insurance market. Specific suggestions include: "1. Set the cancer payout ratio to 95% (5-15 percentage points higher than the average of competing products), predicting a controllable payout rate (an increase of 2.3%); 2. Add five rare diseases with high public opinion (such as ALS and spinal muscular atrophy) to match user needs; 3. Set the waiting period to 90 days (consistent with mainstream competitors to avoid negative public opinion caused by excessively long waiting periods)." The data support section includes a "payout ratio prediction curve," a "public opinion volume chart," and a "comparison table of competitor terms," ensuring that the suggestions are quantifiable and verifiable. Simultaneously, developers can request modifications to the recommendation report (e.g., "How would the predicted payout rate change if the cancer payout ratio were set to 92%?"). The system will then call the model in real time to recalculate and update the report content, achieving a rapid closed loop of "requirement adjustment - model recalculation - report update."
[0057] In other words, multi-source data fusion avoids the one-sidedness of traditional R&D that "makes decisions solely based on competitor terms." It provides decision-making support from three dimensions: "risk (claims data), market (public opinion data), and competitors (term differences)," ensuring that recommendations "both meet market demands and control risks" (e.g., the recommendation to "increase cancer payout ratios" matches user needs in public opinion and verifies controllable payout rates through claims data). Furthermore, intelligent recommendation reports provide "directly implementable product design parameters" (e.g., specific payout ratios, rare disease types) rather than the "qualitative conclusions" of traditional analysis (e.g., "coverage should be increased"). R&D personnel can directly translate recommendations into product term design schemes, shortening the "analysis-design" cycle. Optimized report interaction supports rapid response to design adjustment needs, avoiding the time-consuming problem of "re-analysis and re-reporting" caused by parameter adjustments, thus improving R&D flexibility.
[0058] In one embodiment, the step of combining differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results to generate an intelligent recommendation report includes: The core data dimension table was selected from the comparison parameters of different product terms; Specifically, a core dimension system for report generation based on structured information processing is established, defining mandatory and optional dimensions around three core R&D requirements: "competitive differentiation, demand matching, and compliance risks." Mandatory dimensions include: Differentiation from competitors: Key clause parameters (such as "cancer payout ratio, waiting period length, and deductible amount" for critical illness insurance), coverage (such as "types of critical illnesses covered, whether it includes moderate / mild illness coverage"), and premium differences (such as "annual premium for a 30-year-old male, and payment period options"). Demand matching dimensions: the degree of matching between each competitor's product and the keywords of R&D needs (such as "the degree of matching of the demand for cancer reimbursement ratio ≥90% and the age suitability of the target customer group"), and compliance matching results (such as "whether it complies with the latest regulatory terms and conditions and whether there are any prohibited disclaimers"). Support R&D personnel to customize optional dimensions: If the R&D needs focus on "claims service", optional dimensions such as "claims application channels and claims time commitment" can be added to ensure that the selection dimensions are in line with the specific R&D scenario.
[0059] In addition, the system automatically extracts core data from differentiated comparison parameters (such as comparison tables and difference analysis charts). For example, from "difference in cancer payout ratio between competitors A and B (40% vs. 60%)", it extracts "Cancer payout ratio - competitor A: 40%, competitor B: 60%, demand threshold: ≥90%, difference value: 50% / 30%". At the same time, the extracted data is standardized in format, and the units and expressions are unified (such as unifying "waiting period of 2 months" and "waiting period of 60 days" into "waiting period: 60 days", and unifying "premium of 1500 yuan / year" and "annual premium of 1500" into "annual premium: 1500 yuan") to avoid data format confusion from affecting subsequent fusion analysis, and finally forming a "core data dimension table".
[0060] Historical data from the insurance company's internal claims database is retrieved to create a big data analysis table of claims. Specifically, the scope of data retrieval is determined based on the "product type" and "target customer group" of the R&D needs. For example, when developing "lifetime critical illness insurance for people aged 30-45", only historical data of "critical illness insurance products, claim applicant age 30-45 years old, and claim time in the past 3 years" in the internal claims database are retrieved to avoid interference from irrelevant data. R&D personnel need to pass internal identity authentication (such as OA account associated with system permissions) and can only retrieve anonymized data related to the R&D needs (such as hiding customer names and ID numbers, and only retaining analysis fields such as "age, occupation category, claim disease, and compensation amount"), which meets data security and compliance requirements. In addition, the retrieved data covers "basic claims information" (such as claims application time, case closure time, compensation amount, and whether the claim was rejected), "risk-related information" (such as the payout rate, the proportion of claims for different diseases, and the distribution of reasons for rejection, such as the proportion of "not meeting the diagnostic criteria" and "incidents occurring during the waiting period"), and "customer group characteristic-related information" (such as the average claim amount for customers aged 30-45 and the diseases with high-frequency claims). The integrated data is then converted into a structured format (such as JSON and Excel) to generate a "claims big data analysis table".
[0061] Collect public opinion information from social media, insurance evaluation platforms, and user review channels, and use text filtering technology to remove advertisements and irrelevant comments in order to obtain market public opinion analysis results; Specifically, core data collection channels should be identified, covering mainstream public opinion channels in the insurance industry, including social media platforms (such as Weibo's "Insurance Topics," Xiaohongshu's "Critical Illness Insurance Recommendations" notes, and Douyin's "Insurance Review" video comment sections), professional insurance review platforms (such as product review sections of insurance review websites and insurance industry forums), and user reviews from insurance company official channels (such as official website customer service feedback and in-app user ratings and reviews). Keyword collection rules should be configured, setting collection keywords based on R&D needs keywords (such as "critical illness insurance, cancer payout, 30-45 years old, county-level market"), and only collecting public opinion information containing target keywords, such as collecting comments / notes containing "insufficient cancer payouts in critical illness insurance" or "recommendations for buying critical illness insurance at age 30," to reduce invalid data collection. Meanwhile, a dual text filtering technology of "rule-based filtering + machine learning filtering" is adopted. Rule-based filtering removes obvious advertisements (such as content containing keywords such as "scan to get discounts" or "agent franchise") and irrelevant comments (such as content discussing "car insurance claims"). Machine learning models (such as a text classification model based on BERT) identify hidden advertisements (such as "a product with extremely high cost performance, DM to learn more") and irrelevant information to ensure that the retained information is strongly related to R&D needs. In addition, the filtered public opinion information is analyzed in multiple dimensions, including sentiment analysis (such as positive comments: "hope for higher cancer payouts in critical illness insurance", negative comments: "the waiting period for critical illness insurance is too long"), volume statistics (such as the volume of public opinion related to "cancer payout ratio": 5,000, an increase of 15% month-on-month), and core demand extraction (such as extracting potential demands such as "hope for coverage of rare diseases" and "reduction of deductible" from user comments), ultimately forming "market public opinion analysis results".
[0062] By combining core data dimension tables, claims big data analysis tables, and market sentiment analysis results, an intelligent recommendation report is generated.
[0063] Specifically, a data fusion analysis model is constructed, taking core data dimension tables (competitive product differentiation, demand matching), claims big data analysis tables (risk prediction), and market sentiment analysis results (user demand) as inputs. A weighted algorithm is used to allocate weights to each data point (e.g., "competitive product differentiation weight 30%, claims risk weight 35%, and public sentiment demand weight 35%)" to comprehensively calculate the direction for product design optimization. XGBoost or Random Forest models are called to predict the risks of different design schemes based on the claims big data analysis tables (e.g., "if the cancer payout ratio is increased to 95%, the predicted payout rate will increase by 2.3%), and to predict market acceptance based on public sentiment data (e.g., "adding coverage for 5 rare diseases predicts a 20% increase in positive public sentiment"). This ensures that recommendations are both risk-controllable and market-adaptable.
[0064] Furthermore, the intelligent recommendation report adopts a structured format of "core conclusions + specific suggestions + data support." The core conclusions clearly identify market opportunities and risks (e.g., "Currently, the cancer payout ratio in the critical illness insurance market is generally below 90%, indicating room for improvement, but the risk of payout ratio needs to be controlled"). The specific suggestions include implementable product design parameters (e.g., "Set the cancer payout ratio to 95%, the waiting period to 90 days, and add 3 rare diseases with high public opinion volume"). The data support section includes a core data dimension table, a claims risk prediction curve, and a public opinion volume chart to ensure that the suggestions are quantifiable and verifiable. If R&D personnel have doubts about a suggestion (e.g., "Concerned that a 95% payout ratio will lead to an excessively high payout ratio"), they can modify the parameters in the system (e.g., adjust it to 92%). The model will recalculate the risk and market acceptance in real time and update the report content, achieving a rapid closed loop of "parameter adjustment - model recalculation - report update."
[0065] In other words, the core data dimension table provides the basis for "matching competitors with demand," the claims big data analysis table provides the basis for "risk controllability," and the market sentiment analysis results provide the basis for "user demand and market acceptance." The combination of these three covers the entire R&D dimension of "competitors-risk-market," avoiding design biases caused by single data (such as focusing only on competitor differences while ignoring the risk of high payout ratios). The data fusion analysis model, through weighted calculations and machine learning predictions, transforms scattered data into related decision-making basis (such as "combining the public opinion of 'high demand for cancer payouts' with the claims 'controllable payout ratio,' recommending an increase in the payout ratio"), ensuring that the recommendations are scientific and reasonable.
[0066] In one embodiment, after generating the intelligent recommendation report by combining the differentiated comparison parameters of different product terms, the claims big data analysis table, and the market sentiment analysis results, the method further includes: Call the regulatory reporting template library and fill in the template with parameters based on the parameters in the intelligent recommendation report to generate a preliminary reporting document; Specifically, the system has a built-in "regulatory filing template library" covering all insurance categories, categorized by filing material type (such as clause filing templates, actuarial report templates, and product instruction manual templates). The templates strictly adhere to the latest format requirements (e.g., clause filing templates include mandatory sections such as "Basic Product Information," "Coverage Liability," "Exclusions," and "Claims Process," with section title numbering and font size conforming to regulatory standards). Based on the "Product Type" (e.g., critical illness insurance, medical insurance) in the intelligent recommendation report, the system automatically matches the corresponding template. For example, if the report points to "Lifetime Critical Illness Insurance for People Aged 30-45," it will call the "Personal Insurance Product Clause Filing Template (Critical Illness Insurance Only)," avoiding format rework due to template mismatch. Furthermore, the system extracts the core parameters to be filled from the intelligent recommendation report, covering "Basic Product Parameters" (e.g., product name, underwriting company, target customer age range), "Core Clause Parameters" (e.g., cancer payout ratio 95%, waiting period 90 days, deductible 5000 yuan), and "Actuarial Related Parameters" (e.g., predicted payout ratio 28%, basis for rate calculation). The template has preset parameter variable identifiers, and the system automatically matches and fills the extracted parameters with the variable identifiers. For parameters that need to be filled in batches, "one-click global fill" is supported to avoid repeated operations and finally generate a preliminary reporting document with complete structure and parameters (formats include Word and PDF).
[0067] The large model is invoked to polish and verify the consistency of the initial filing documents, and the compliance is detected by the rule engine to generate compliant filing documents; Specifically, a large-scale model for fine-tuning insurance regulatory texts (such as an insurance compliance model based on GPT-4) was selected to refine the texts, focusing on "standardized regulatory language" and "accuracy of expression." Rigid parameter expressions in initial filings (such as "95% compensation for cancer") were optimized into regulatory-compliant expressions (such as "If the insured is diagnosed with cancer that meets the criteria stipulated in this contract, the insurance company will compensate 95% of the basic sum insured"). The system includes provisions for cancer insurance benefits; ambiguous statements (such as "'young people' can apply") are optimized to precise statements (such as "insured persons aged 30 to 45 can apply"); in addition, consistency checks cover "parameter consistency" and "statement consistency": at the parameter level, the system checks whether the values of the same parameter are consistent throughout the text (such as whether "waiting period" is always 90 days, without contradictory statements such as "60 days" or "2 months"); at the statement level, the system checks whether the terminology for the same concept is consistent (such as whether "basic sum insured" does not use abbreviations such as "sum insured" or "insurance amount" interchangeably). If inconsistencies are found, the system automatically marks them and corrects them according to the optimal statement, and the correction record is synchronously saved to the document version log.
[0068] In addition, the system has a built-in "Insurance Regulatory Rules Knowledge Base," covering legal provisions and industry self-regulatory norms. The rules are broken down into executable detection logic (such as "Exclusion clauses must not include 'Congenital diseases are not covered under any circumstances,' and must be supplemented with the exception of 'Payment is available 3 years after diagnosis'" and "The clauses must not contain unrealistic promises such as 'Guaranteed renewal until age 100'"). The rule engine scans the preliminary filing document segment by segment using a "chapter-by-chapter detection" method: for clause chapters, it checks for prohibited exclusions; for actuarial report chapters, it checks whether the rate calculation basis complies with regulatory requirements; for product brochure chapters, it checks for misleading statements (such as "'zero risk' 'guaranteed profit'"). When compliance risks are detected, the system automatically marks the risk location (such as "Chapter 5, Article 3 contains a prohibited exclusion statement") and pushes alternative statement suggestions (such as "It is recommended to modify it to 'If the insured is diagnosed with a congenital disease 3 years after diagnosis and meets the requirements of this contract, the insurance company shall assume the responsibility of paying insurance benefits'"). After the R&D personnel confirm the correction, a compliance filing document is generated.
[0069] Based on NLP models, logical consistency and format checks are performed on compliance reporting documents to output tiered quality inspection reports.
[0070] Specifically, a "textual logical association analysis model" (built based on BERT + graph neural network) is used to break down the logical association points in the document: detecting "term and rule logic" (e.g., "If the terms stipulate 'no compensation for accidents during the waiting period,' does the subsequent 'claims process' clearly specify 'the handling method for reports during the waiting period' to avoid logical gaps") and "parameter derivation logic" (e.g., "Is the 'predicted payout rate of 28%' in the actuarial report consistent with the derivation of basic data such as 'cancer payout ratio of 95%' and 'target customer group claim rate of 35%', without any calculation contradictions"). For detected logical contradictions (e.g., "The terms stipulate '90% reimbursement ratio for inpatient medical expenses,' but the 'exclusion clause' states 'only 80% reimbursement for inpatient expenses'"), the system automatically locates the contradictory chapters, analyzes the contradiction type (e.g., "rule conflict" or "derivation gap"), and generates a logical problem explanation.
[0071] In addition, the system strictly follows the regulatory requirements for the format of the filed documents. The inspection dimensions include: chapter format (such as whether it is arranged in the mandatory order of "Chapter 1 General Provisions → Chapter 2 Insurance Liability → Chapter 3 Exclusion Clauses", and whether the chapter number is in the standard hierarchy of "I, (I), 1, (1)"), text format (such as the title font is "Bold No. 4", the body text is "SimSun No. 4", and the line spacing is "fixed 22 points"), attachment format (such as whether the actuarial report is accompanied by a "rate calculation table", and whether the table is marked with a standard title such as "Table 1-1 Rate Calculation Details"); for format deviations (such as "Chapter 3 title font is SimSun No. 4" "the table is not marked with a standard title"), the system automatically takes a screenshot to mark the deviation position and provides format correction guidance (such as "it is recommended to change the title font of Chapter 3 to bold No. 4").
[0072] Furthermore, the inspection results are categorized according to the classification criteria of "serious issues, general issues, and suggestive issues": serious issues (such as logical contradictions rendering clauses invalid or violations of mandatory regulatory provisions), general issues (such as format deviations in non-core expressions or inconsistent terminology), and suggestive issues (such as "it is recommended to supplement the 'Claim Application Materials List' to improve user clarity, which is not a mandatory regulatory requirement"). The report includes four parts: "Issue Statistics, Issue Details, Correction Suggestions, and Quality Inspection Conclusion." The issue statistics use a pie chart to show the percentage of the three types of issues; the issue details indicate the location, type, and screenshot of the issue; the correction suggestions provide directly adoptable modification solutions; and the quality inspection conclusion clarifies "whether the preliminary quality inspection has been passed" (such as "no serious issues, and after correcting the two general issues, it can be submitted for regulatory filing"). The report can be exported as a PDF file and retained as an attachment to the filing materials.
[0073] In other words, the built-in template library ensures that templates comply with the latest regulatory formats, avoiding rejection of reports due to the use of old templates (such as failure to add a "Consumer Rights Notice" section according to the new regulations); variable mapping and global filling shorten the manual filling work that originally took 1-2 days to within 30 minutes, and the parameter filling accuracy rate reaches 100%, avoiding manual omissions (such as forgetting to fill in "the basis for rate calculation in the actuarial report") and errors (such as filling in "60 days" instead of "90-day waiting period"), thus improving the efficiency and accuracy of initial document generation. Furthermore, the logical consistency check covers clause rules and parameter derivation, avoiding deep logical contradictions that are difficult for humans to detect (such as "hidden conflicts between compensation ratios and claims processes"); the format check strictly matches regulatory standards, eliminating minor deviations (such as fonts and table titles); the tiered quality inspection report accurately marks the location, type, and correction suggestions of problems, allowing R&D personnel to quickly locate and correct problems without having to check page by page (such as directly finding "format deviation in Article 3 of Chapter 5" through screenshots), while the tiered conclusions provide clear basis for the submission of reports (such as "no serious problems, can be submitted"), ensuring that the final submitted reports fully comply with regulatory requirements and improving the approval rate of reports.
[0074] The aforementioned report generation method based on structured information processing, through data collection technology, can automatically capture key information such as regulatory policies, industry trends, and competitor clauses, eliminating the need for manual access to multiple platforms for retrieval, thus shortening information collection time and significantly reducing manual labor costs. Simultaneously, by using format conversion technology to transform unstructured PDF files into computable text, it solves the problem of traditional manual inefficient processing of unstructured data, avoiding the omission of key information such as competitor special clauses and regulatory details due to format barriers. Furthermore, data collection can focus on core information sources for insurance research and development, avoiding interference from irrelevant information, and supports dynamic adjustment of the collection scope according to research and development needs, ensuring a high degree of alignment between the acquired information and the research direction. Moreover, classifying and structured storage of computable text can transform scattered clause information into an ordered, directly accessible data format, avoiding the tedious operation of manually filtering and analyzing objects sentence by sentence from massive amounts of unstructured text. At the same time, structured data is convenient for big data processing. The model directly reads and processes data, providing an efficient data input foundation for subsequent comparisons. Furthermore, the big data model can accurately pinpoint comparison dimensions based on R&D needs, automatically mark differences between product terms, and generate standardized visual results. This not only reduces term comparison time from hours to minutes but also eliminates subjective bias, ensuring the objectivity and accuracy of the analysis results. This provides direct support for R&D personnel to quickly grasp the core differences of competitors. In addition, by deeply integrating term difference data, claims big data analysis tables, and market sentiment data, it can provide a basis for R&D decisions from three dimensions: risk, market, and competitor levels. This avoids the bias caused by single data points. Through comprehensive analysis of multi-dimensional data, it outputs intelligent recommendation reports containing specific optimization directions, directly guiding product term design, helping R&D personnel quickly determine the core competitiveness of the product, reducing R&D trial-and-error costs, and improving market acceptance after product launch.
[0075] Figure 3 This is a schematic block diagram of a report generation device 300 based on structured information processing provided in an embodiment of the present invention. Figure 3 As shown, corresponding to the above-described report generation method based on structured information processing, the present invention also provides a report generation apparatus 300 based on structured information processing. This report generation apparatus 300 includes a unit for executing the above-described report generation method based on structured information processing, and the apparatus can be configured in a server. Specifically, please refer to... Figure 3 The report generation device 300 based on structured information processing includes: Acquisition unit 301 is used to acquire structured information of R&D requirement parameters; The data collection and conversion unit 302 is used to collect formatted files from publicly available regulatory websites, industry association announcements, and competitor terms, and convert them into computable text. The classification storage unit 303 is used to classify and structure the computable text to obtain a structured storage clause document; Input generation unit 304 is used to input structured storage terms documents and R&D requirement parameters into the big data model to generate different comparison parameters for different product terms; Unit 305 is used to combine different product terms and conditions with comparative parameters, claims big data analysis tables, and market sentiment analysis results to generate intelligent recommendation reports.
[0076] In one embodiment, the acquisition unit 301 is used to acquire structured information content or specific requirement content input by R&D personnel through the user interaction layer, so as to obtain R&D requirement parameters.
[0077] In one embodiment, the acquisition and conversion unit 302 includes: The module is used to determine the scope of publicly available regulatory websites, industry association announcements, and competitor companies to be collected based on the product type and market positioning of the structured information, and to generate a list of collection targets. The download module is used to automatically access and download formatted files from publicly available regulatory websites, industry association announcements, and competitor terms according to the list of data collection targets. The conversion module is used to convert downloaded format files into a computable text format to obtain computable text.
[0078] In one embodiment, the classification storage unit 303 includes: The cleaning module is used to clean up noise in computable text to obtain cleaned text. The input generation module is used to input the cleaned text into a pre-trained language model to generate classification results; The storage module is used to store the classification results in a structured manner to obtain a structured storage terms document.
[0079] In one embodiment, the input generation unit 304 includes: The parsing module is used to parse the key parameters of the R&D requirements to generate a list of requirement keywords; The matching module is used to match corresponding related sub-units in the structured stored terms document based on the list of required keywords, so as to form a correlation lookup table; The import generation module is used to import sub-cell data from the correlation comparison table, the list of demand keywords, and the structured storage terms document into the big data model to generate differentiated comparison parameters for different product terms.
[0080] In one embodiment, the bonding unit 305 includes: The filtering module is used to filter out the core data dimension table from the differential comparison parameters of different product terms; The data retrieval module is used to retrieve historical data from the insurance company's internal claims database and generate a big data analysis table for claims. The data collection and filtering module is used to collect public opinion information from social media, insurance evaluation platforms, and user review channels, and uses text filtering technology to remove advertisements and irrelevant comments in order to obtain market sentiment analysis results. The module combines core data dimension tables, claims big data analysis tables, and market sentiment analysis results to generate intelligent recommendation reports.
[0081] In one embodiment, the device further includes: The inspection unit is used to call the regulatory reporting template library and fill in the parameters of the template according to the parameters in the intelligent recommendation report to generate a preliminary reporting document; the large model is called to polish and verify the consistency of the preliminary reporting document, and the compliance is detected by the rule engine to generate a compliant reporting document; the compliant reporting document is checked for logical consistency and format based on the NLP model to output a graded quality inspection report.
[0082] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned report generation device 300 based on structured information processing and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0083] The aforementioned report generation device 300 based on structured information processing can be implemented as a computer program, which can, for example... Figure 4 It runs on the computer device shown.
[0084] Please see Figure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0085] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0086] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a report generation method based on structured information processing.
[0087] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0088] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a report generation method based on structured information processing.
[0089] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0090] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps: The process involves: acquiring structured information on R&D requirements parameters; collecting formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms, and converting them into computable text; classifying and structuring the computable text to obtain structured terms documents; inputting the structured terms documents and R&D requirements parameters into a big data model to generate comparative parameters for different product terms; and combining these comparative parameters with claims big data analysis tables and market sentiment analysis results to generate an intelligent recommendation report.
[0091] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0092] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0093] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps: The process involves: acquiring structured information on R&D requirements parameters; collecting formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms, and converting them into computable text; classifying and structuring the computable text to obtain structured terms documents; inputting the structured terms documents and R&D requirements parameters into a big data model to generate comparative parameters for different product terms; and combining these comparative parameters with claims big data analysis tables and market sentiment analysis results to generate an intelligent recommendation report.
[0094] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0096] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0097] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0099] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A report generation method based on structured information processing, characterized in that, include: Obtaining structured information on R&D requirement parameters; Collect formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms, and convert them into computable text; Computable text is classified and structured for storage to obtain structured term documents; The structured storage terms and conditions documents and R&D requirements parameters are input into the big data model to generate differentiated comparison parameters for different product terms; By combining differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results, an intelligent recommendation report is generated.
2. The report generation method based on structured information processing according to claim 1, characterized in that, The R&D requirement parameters for obtaining structured information include: To obtain R&D requirement parameters, obtain structured information or specific requirements input by R&D personnel through the user interaction layer.
3. The structured information processing based report generation method of claim 1, wherein, The process involves collecting publicly available regulatory websites, industry association announcements, and competitor terms in their formatted documents, and converting them into computable text, including: Based on the product type and market positioning of the structured information, determine the scope of publicly available regulatory websites, industry association announcements, and competitor companies that need to be collected, and form a target list for collection. Automatically access and download formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms according to the target data collection list; Convert the downloaded format file into a computable text format to obtain computable text.
4. The structured information processing based report generation method of claim 1, wherein, The process of classifying and structuring computable text to obtain a structured document includes: Noise removal is performed on the computable text to obtain cleaned text; The cleaned text is fed into a pre-trained language model to generate classification results; The classification results are stored in a structured format to obtain a structured document of terms.
5. The structured information processing based report generation method of claim 1, wherein, The process of inputting structured storage terms documents and R&D requirement parameters into a big data model to generate differentiated comparison parameters for different product terms includes: Analyze the key parameters of the R&D requirements to generate a list of requirement keywords; Based on the list of required keywords, match the corresponding related sub-units in the structured storage terms document to form a correlation lookup table; Import the correlation comparison table, the list of demand keywords, and the sub-cell data of the structured storage terms document into the big data model to generate differentiated comparison parameters for different product terms.
6. The structured information processing based report generation method of claim 1, wherein, The method combines differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results to generate an intelligent recommendation report, including: The core data dimension table was selected from the comparison parameters of different product terms; Historical data from the insurance company's internal claims database is retrieved to create a big data analysis table of claims. Collect public opinion information from social media, insurance evaluation platforms, and user review channels, and use text filtering technology to remove advertisements and irrelevant comments in order to obtain market public opinion analysis results; By combining core data dimension tables, claims big data analysis tables, and market sentiment analysis results, an intelligent recommendation report is generated.
7. The structured information processing based report generation method of claim 1, wherein, After generating the intelligent recommendation report by combining differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results, it also includes: Call the regulatory reporting template library and fill in the template with parameters based on the parameters in the intelligent recommendation report to generate a preliminary reporting document; The large model is invoked to polish and verify the consistency of the initial filing documents, and the compliance is detected by the rule engine to generate compliant filing documents; Based on NLP models, logical consistency and format checks are performed on compliance reporting documents to output tiered quality inspection reports.
8. A report generation apparatus based on structured information processing, characterized by include: The acquisition unit is used to acquire structured information about R&D requirement parameters. The data collection and conversion unit is used to collect formatted documents from publicly available regulatory websites, industry association announcements, and competitor terms and conditions, and convert them into computable text. The classification storage unit is used to classify and structure computable text to obtain a structured document. The input generation unit is used to input structured storage terms documents and R&D requirement parameters into the big data model to generate differential comparison parameters for different product terms; The combined unit is used to combine differentiated comparison parameters of different product terms, claims big data analysis tables, and market sentiment analysis results to generate intelligent recommendation reports.
9. A computer device, comprising: The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1 to 7.
10. A storage medium, characterized by The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.