Automatic text filling method and device, storage medium and computer equipment

By combining large models with preset filling rules, it automatically parses and generates content to be filled, solving the problems of low information processing efficiency and format diversity in the financial and medical fields, achieving efficient and compliant text processing, and supporting data sharing between different business systems.

CN120688465APending Publication Date: 2025-09-23PING AN INT FINANCIAL LEASING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510788074.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the fields of financial technology and healthcare, untimely information processing leads to inefficient business operations, repeated filling and parsing of information are difficult, traditional parsing technology has difficulty understanding semantic associations, format diversity leads to high error rates, and there is a lack of a unified data sharing solution.

Method used

By using a large model to identify the original text and combining it with preset filling rules, it automatically parses and generates the content to be filled. It supports parsing of mixed content in formats such as PDF and Word, realizes cross-modal information fusion, and improves text processing efficiency and compliance.

Benefits of technology

It has improved text processing efficiency and compliance in the financial and medical fields, reduced the instability and subjectivity of manual operations, and supported seamless connection and sharing of data between different business systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688465A_ABST
    Figure CN120688465A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computers, and discloses an automatic text filling method and device, a storage medium and computer equipment, which can be applied to an automatic text filling business scene of financial science and technology and medical health, and the method comprises the following steps: obtaining an original text; identifying the original text to obtain multiple pieces of text information; based on preset filling rules corresponding to fields needing to be filled, the text information is analyzed, to-be-filled field content conforming to the preset filling rules is generated, the number of the fields needing to be filled is multiple, and the fields needing to be filled correspond to the preset filling rules; the to-be-filled field content is filled to the position of a field needing to be filled in a preset text template, a filled compliant text is obtained, and the preset text template comprises the field needing to be filled. By automatically analyzing the original text, the to-be-filled content is generated according to the rule and is accurately filled, so that the text processing efficiency and compliance in the financial and medical fields can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer technology, financial technology, and medical health technology, and in particular to a method and apparatus for automatic text filling, a storage medium, and a computer device. Background Art

[0002] Currently, in the fields of financial technology and medical health technology, untimely information processing will affect business efficiency to a certain extent, specifically:

[0003] In the financial sector, businesses involve filling out a large amount of complex and repetitive information. When individuals apply for multiple businesses, they need to submit a lot of information. Because different financial institutions have independent systems, users need to fill out the forms repeatedly on different platforms, which is time-consuming and labor-intensive and prone to errors due to manual input, affecting business processing. At the same time, when financial institutions process user files, traditional parsing technologies have limitations. PDF or Word files have complex layouts and contain professional terminology. OCR (Optical Character Recognition) technology has difficulty understanding semantic associations. Rule engines rely on predefined templates and are unable to flexibly process diverse and complex files. In addition, the file formats of different institutions vary greatly, and traditional parsing tools have poor compatibility, increasing processing difficulty and error rates. In addition, data silos exist in the financial sector, and users lack a unified data sharing solution, which increases workload and risk management difficulty.

[0004] Similarly, in the medical field, the efficiency of patient information entry and transmission is low. Patients must repeatedly fill in information at different hospitals, which is prone to omissions and errors, affecting diagnostic decisions. When medical institutions conduct business and scientific research, compiling and filling in information is time-consuming, labor-intensive, and error-prone. Inconsistent data formats across institutions lead to difficulties in information transmission. Traditional parsing technology for medical documents also faces limitations similar to those in the financial sector. Summary of the Invention

[0005] In view of this, the present application provides a text automatic filling method and device, storage medium, and computer equipment, which can automatically parse the original text, generate content to be filled according to rules and fill it accurately, thereby improving the text processing efficiency and compliance in the financial and medical fields.

[0006] According to one aspect of the present application, a text automatic filling method is provided, the method comprising:

[0007] Get the original text;

[0008] Recognizing the original text to obtain multiple text messages;

[0009] Based on the preset filling rules corresponding to the required filling fields, the text information is parsed to generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple fields, each of which corresponds to a preset filling rule;

[0010] Fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field.

[0011] According to another aspect of the present application, a text automatic filling device is provided, the device comprising:

[0012] Original text acquisition module, used to obtain original text;

[0013] An original text recognition module, used to recognize the original text and obtain multiple text messages;

[0014] A filling content parsing module is used to parse the text information based on the preset filling rules corresponding to the required filling fields, and generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple fields, and each required filling field corresponds to a preset filling rule;

[0015] The compliant text filling module is used to fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field.

[0016] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the above-mentioned text automatic filling method is implemented.

[0017] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned text automatic filling method when executing the program.

[0018] Through the above technical solution, the present application provides a text automatic filling method and device, storage medium, and computer equipment that can automatically parse the original text, generate content to be filled according to rules and fill it accurately, thereby improving the text processing efficiency and compliance in the financial and medical fields.

[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 A flowchart of a text automatic filling method provided by an embodiment of the present application is shown;

[0022] Figure 2 A schematic diagram of a process for identifying content of a field to be filled provided by an embodiment of the present application is shown;

[0023] Figure 3 A schematic diagram illustrating a flow chart of another text automatic filling method provided in an embodiment of the present application is shown;

[0024] Figure 4 A flowchart of another text automatic filling method provided by an embodiment of the present application is shown;

[0025] Figure 5 A schematic structural diagram of a text automatic filling device provided in an embodiment of the present application is shown;

[0026] Figure 6 A structural diagram of another text automatic filling device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0028] In this embodiment, a text automatic filling method is provided, such as Figure 1 As shown, the method includes:

[0029] Step 101: Obtain the original text.

[0030] Step 102: Identify the original text to obtain multiple text messages.

[0031] Currently, users in the fields of finance and medical technology face similar and severe information processing challenges.

[0032] In the financial industry, whether individuals applying for financial services (such as loans, credit cards, and investment account openings) or businesses submitting applications for various financial services (such as corporate loans, project financing, and listing documents), they are required to fill out a large amount of complex and repetitive information. When applying for services at different financial institutions, individuals must repeatedly fill out basic and financial-related fields such as name, ID number, contact information, income, asset status, and occupation across different platforms. Not only is this information voluminous, but some fields (such as income composition and asset details) require detailed breakdown and precise entry, making manual entry time-consuming and error-prone. When businesses apply for financial services, they face complex fields such as basic company information (such as company name, registered address, and legal representative), financial information (such as balance sheet, income statement, and cash flow statement data), and project information (such as project name, project size, and expected project returns). These complex fields require significant time and effort to manually organize and complete, severely impacting business processing efficiency and creating a poor user experience.

[0033] Furthermore, when processing financial documents (such as financial statements, contracts, and agreements), while OCR technology can recognize text content, it cannot understand semantic relationships. For example, in financial statements, OCR technology cannot distinguish whether time expressions such as "first quarter of 2023" and "second quarter of 2023" describe the time range of financial data or other business-related time information, resulting in inaccurate information extraction. Consequently, rule engines that rely on predefined templates are unable to process financial documents. Financial documents come in a variety of formats, including complex tables, charts, and specialized terminology. Rule engines struggle to adapt to this unstructured data and are unable to accurately extract key information. Furthermore, the layout of financial documents submitted by different financial institutions varies significantly. Some use standardized tables generated by professional financial software, while others are free-form Word documents or PDF files that may also include charts, images, and other elements. This diverse format results in extremely high error rates when extracting information using traditional parsing tools, seriously impacting the efficiency and quality of financial transactions.

[0034] Therefore, when applying for services from different financial institutions, individual and corporate users must repeatedly upload various financial documents on each platform, lacking a unified data export and sharing solution. This not only increases user workload but can also lead to inconsistent data across platforms, making financial risk management more difficult.

[0035] In the healthcare industry, the challenges of duplicate information entry and parsing difficulties arise when patients seek medical care, medical institutions apply for projects, and researchers apply for research funding. Patients must repeatedly fill in fields such as personal information (such as name, age, gender, and contact information), medical history (such as previous illnesses, allergies, and family medical history), and symptom descriptions across different hospital information systems. Patients may struggle to accurately describe complex symptoms and medical histories, and manual entry is not only time-consuming but can also lead to inaccurate information, impacting doctors' diagnoses and treatment. When applying for various medical projects (such as research projects and designated medical insurance programs), medical institutions must fill in a large number of complex fields, including institutional information (such as hospital name, address, and departmental structure), personnel information (such as physician qualifications and researcher resumes), and project information (such as project title, research content, and expected outcomes). Organizing and filling in this information requires significant manpower and time, and is prone to errors. Similarly, when processing medical documents (such as medical records, examination reports, and research papers), optical character recognition (OCR) technology cannot understand semantic associations. For example, in a medical record, "2023-05-10" may indicate a consultation date, or it may indicate an examination date or a surgery date. OCR technology cannot accurately distinguish between the two, resulting in information extraction errors. Medical file formats are complex and contain a large number of professional terms, charts, images, and other elements. Rule engines rely on predefined templates and are unable to process this unstructured medical data, making it difficult to accurately extract key information, which affects the application of medical projects and the development of scientific research. In addition, the medical documents generated by different medical institutions have different layouts. Some use standardized documents generated by electronic medical record systems, while others are PDF files or images scanned from handwritten medical records. This diversity of formats makes traditional parsing tools face huge challenges when extracting information, and the error rate remains high.

[0036] As a result, patients seeking medical treatment must repeatedly upload medical records, examination reports, and other documents between hospitals, lacking a unified data sharing platform. This not only increases the burden on patients but can also lead to disjointed medical information, impacting doctors' diagnostic and treatment decisions. Furthermore, when applying for projects, medical institutions must repeatedly submit documents across different departments and platforms, lacking efficient data export and integration solutions.

[0037] In the above embodiment of the present application, by supporting mixed content parsing (text, tables, charts) in formats such as PDF and Word, and realizing cross-modal information fusion through a large model (such as the Transformer architecture), and then automatically filling in text, it is possible to improve information processing efficiency and can be applied to the fields of financial technology and medical health technology. Specifically, the original text is first obtained, and then the original text can be recognized by the large model to obtain multiple text information, preparing for subsequent automatic text filling.

[0038] Furthermore, in the financial sector, for example, in securities companies' securities account opening scenarios, securities companies can use rule engines to extract text information from account opening forms. For example, according to preset rules, text information such as investment years and risk tolerance level are extracted from the investment experience information (i.e., the original text) filled in by investors, such as "Investment years: 5 years" and "Risk tolerance: Medium". Named entity recognition technology is also used to identify entity information such as names and bank card numbers from the form. For example, "Name: Zhang San" and "Bank card number: 622848001234567890" are accurately extracted to form independent text information.

[0039] In the field of medical technology, for example, in hospital settings, this can be applied to electronic medical record systems. When a patient visits a hospital, the doctor records their symptoms, diagnosis, treatment plan, and other information in the electronic medical record system. This information is stored as electronic documents on the hospital's server, becoming raw text. From the text in the electronic medical record, hospitals can use semantic analysis technology to extract patient symptom information. For example, they can identify symptoms such as "fever," "cough," and "headache" from the doctor's diagnosis note and convert them into structured text information, such as "Symptoms: fever, cough, headache." Next, regular expressions can be used to extract the patient's vital signs, such as temperature, blood pressure, and heart rate, from the medical record. For example, they can extract text information such as "temperature: 38.5°C" and "blood pressure: 120 / 80 mmHg."

[0040] Step 103 , based on the preset filling rules corresponding to the required filling fields, the text information is parsed to generate the content of the fields to be filled that conforms to the preset filling rules, wherein the required filling fields include multiple fields, each of which corresponds to a preset filling rule.

[0041] Next, different required fields have specific semantics and formatting requirements. Preset filling rules ensure that only content that meets the field requirements is accurately extracted from text information. For example, in the financial loan business, if the required field is "loan amount," the preset rules can be set to only recognize numerical values ​​in the text, and the values ​​must be within a reasonable loan limit. This prevents other irrelevant values ​​(such as date numbers) from being mistakenly entered into the field, thereby improving information accuracy.

[0042] Therefore, manual processing of text information and filling in fields is susceptible to subjective factors, leading to oversights or errors. However, pre-set filling rules can automatically perform parsing and filling operations, filtering information based on pre-defined logic and conditions. This can avoid the instability and subjectivity of manual operations, thereby improving the accuracy of filled content. Furthermore, traditional text information processing and field filling rely on manual labor, which is time-consuming and labor-intensive. Pre-set filling rules automate the process, quickly extracting and generating content for fields to be filled from large amounts of text information. For example, in medical record processing, key information such as symptoms and diagnoses can be automatically extracted from patient records and filled into the corresponding fields, saving significant time and labor costs. They can also process multiple text messages simultaneously and generate content for multiple required fields. In financial batch loan application review scenarios, they can quickly process large volumes of application documents, improving processing speed and efficiency and adapting to large-scale business needs. In particular, the required fields and text information characteristics vary across different business scenarios, and pre-set filling rules can be flexibly customized to suit specific scenarios. In the financial sector, loan applications and insurance claims require different fields and text information to be filled, and rules can be set separately. In the medical field, the processing rules for outpatient and inpatient medical records also differ, allowing for flexible adjustment and adaptation. By parsing text information based on preset filling rules corresponding to the required filling fields, the content of the fields to be filled that meets the preset filling rules is generated. This enables seamless data connection and sharing between different business systems of financial institutions and different department systems of medical institutions, improving data utilization efficiency and value.

[0043] Alternatively, as Figure 2 As shown, the required filling field includes a keyword field, a character field, an amount field, and a time difference field. The keyword field corresponds to a keyword, and the content of the field to be filled includes the content of the keyword field, the content of the character field, the content of the amount field, and the content of the time difference field. Regarding step 103, "based on the preset filling rule corresponding to the required filling field, parsing the text information to generate the content of the field to be filled that complies with the preset filling rule", specifically includes:

[0044] Step 1031 : For the keyword field, identify the keyword field content in the text information that meets the keyword semantic features through semantic analysis.

[0045] Step 1032 : for the character field, identify the character sequence in the text information that complies with the preset character sequence format, remove the non-compliant characters in the character sequence, and obtain the character field content.

[0046] Step 1033, for the amount field, identify the amount in the text information. If the identified amount is a fixed value, the identified fixed value is determined as the amount field content. If the identified amount is a numerical range, the maximum value in the numerical range is taken as the amount field content.

[0047] Step 1034: for the time difference field, identify the time in the text information for which the time difference calculation is required, and obtain the content of the time difference field based on the difference between the current time and the identified time.

[0048] In the above embodiment of the present application, for the keyword field:

[0049] In the medical field, for example, when parsing patient complaints, we can obtain text information such as "I have been having headaches recently, especially when I am stressed at work, and I also have mild nausea and poor sleep at night."

[0050] Using semantic analysis techniques from the DIKWP semantic mathematical model, the text can be parsed to identify keywords such as "headache," "work stress," "nausea," and "poor sleep." "Headache" matches the semantic characteristics of a symptom keyword, indicating the location and type of discomfort; "work stress" is a trigger keyword, suggesting factors that may contribute to the symptoms; "nausea" and "poor sleep" are also symptom-related keywords, further describing the patient's physical reactions. By identifying these keywords, doctors can quickly understand the patient's primary symptoms and possible related factors, providing important clues for diagnosis.

[0051] For character fields:

[0052] In the healthcare field, for example, when extracting patient ID information, hospitals use optical character recognition (OCR) technology to scan paper documents and convert them into electronic format. OCR technology accurately recognizes character sequences for character fields such as patient ID numbers and medical record numbers. For example, the OCR system can recognize character sequences such as "110101199001011234" for ID numbers. The system also removes any non-compliant characters, such as stains, blurred areas, and other interfering information, to accurately capture the content of the character fields. This facilitates patient information management and querying by medical institutions, ensuring the accuracy and completeness of patient information.

[0053] In the financial sector, for example, when entering ID information for bank account opening, a user submits their ID, and bank staff scans it using OCR technology. OCR technology quickly recognizes and extracts character information from the ID, such as the ID number "110101199001011234" or the name "Zhang San." During the recognition process, verification is performed against a preset character sequence format to remove any potentially non-compliant characters, ensuring the accuracy of the entered information, improving business processing efficiency, and meeting compliance requirements in the financial industry.

[0054] For the Amount field:

[0055] In the healthcare field, for example, a medical bill might include an examination fee of 200-300 yuan. Semantic analysis identifies this as a range of amounts. Based on pre-set rules, the maximum value in the range, 300 yuan, is used as the amount field. This helps medical institutions accurately calculate expenses, facilitating patient understanding of costs, and providing accurate cost basis for medical insurance reimbursement and other processes.

[0056] For the time difference field:

[0057] In the medical field, for example, a patient's hospitalization duration can be calculated. For example, if the patient's medical record records "Admission Date: 2025-05-28," and the current time is 2025-05-30, the time difference is calculated, resulting in "2 days" as the time difference field.

[0058] Optionally, in step 1031, "identifying keyword field content in the text information that meets the keyword semantic features through semantic analysis" specifically includes:

[0059] Step 10311: extract the semantic features of the keywords corresponding to the keyword field, and perform lexical analysis on the text information to obtain words or phrases.

[0060] Step 10312, matching the word or phrase obtained by lexical analysis with the semantic features of the keyword. If the semantics of the word or phrase matches the semantic features of the keyword, the matched word or phrase is determined as the keyword field content of the keyword field.

[0061] In the above embodiment of the present application, keyword fields such as name fields or valid academic qualifications fields need to extract unique semantic features for different keyword fields:

[0062] 1. Semantic features of the name field, for example:

[0063] The Name field typically contains a person's name, consisting of a surname and a given name. In the Chinese context, a surname is usually a single character (such as Zhang, Wang, and Li), while a given name can be one or two characters (such as Wei, Fang, and Jianguo). The Name field may also contain compound surnames (such as Ouyang and Sima) and names of ethnic minorities.

[0064] 2. Semantic features of valid education fields, for example:

[0065] Valid education fields typically include information such as educational level (e.g., bachelor's, master's, doctoral), school name, and major. Educational level is typically expressed in a fixed format. School name may be the name of an institution or organization, while major name refers to knowledge in a specific field.

[0066] Next, the text is subjected to lexical analysis, for example using natural language processing (NLP) techniques to extract words or phrases. Lexical analysis is a fundamental task in natural language processing, involving segmenting text into meaningful units (such as words, phrases, or subwords) and classifying and labeling these units.

[0067] Next, the lexical analysis results are matched with the semantic features, that is, the words or phrases obtained by the lexical analysis are matched with the semantic features of the keywords in the keyword field. If the semantics of the word or phrase matches the semantic features of the keyword, it is determined to be the keyword field content.

[0068] Furthermore, in a loan application scenario in the financial sector, for example, a bank needs to verify the applicant's basic information, including name and education. For example, the applicant's information might be: "Applicant's name: Li Fang, Education: Master's degree, School of Economics and Management, Tsinghua University." The processing process might be:

[0069] 1. Extracting semantic features:

[0070] Name field: expect to extract "Li Fang" as the name.

[0071] Valid education field: It is expected to extract "Tsinghua University School of Economics and Management" and "Master" as education information.

[0072] 2. Lexical analysis:

[0073] Use NLP tools to perform lexical analysis on the text and obtain words or phrases: "applicant", "name", ":", "Li Fang", ",", "education", ":", "School of Economics and Management, Tsinghua University", "Master", ".".

[0074] 3. Semantic matching:

[0075] Name field: In the lexical analysis results, "Li Fang" matches the semantic feature (personal name) of the name field, so "Li Fang" is determined to be the content of the name field.

[0076] Valid academic qualification field: In the lexical analysis results, "Tsinghua University School of Economics and Management" matches the semantic features of the school name, and "Master" matches the semantic features of the academic qualification level. Therefore, "Tsinghua University School of Economics and Management" and "Master" are determined to be valid academic qualification field contents.

[0077] Optionally, after “performing lexical analysis on the text information to obtain words or phrases” in step 10311, the method further includes:

[0078] Step 10313: Based on the keyword corresponding to the keyword field, determine multiple keyword variants corresponding to the keyword, associate the keyword and the keyword variant to obtain the keyword group corresponding to the keyword field, and build a keyword variant library based on the keyword group of each keyword field.

[0079] Step 10314: for any keyword group in the keyword variant library, when a word or phrase obtained through lexical analysis hits any keyword or keyword variant in the keyword group, the hit word or phrase is determined as the keyword field content of the keyword field.

[0080] In the above embodiments of the present application, not only can keywords be matched semantically, but keyword variants can also be identified, which is applicable to a variety of language scenarios, especially non-standardized, colloquial language environments. At the same time, context analysis can also be combined to identify the content of the field to be filled. Specifically, in the field of medical technology, taking symptom identification in the electronic health record (EHR) system as an example, when the case text is lexically analyzed, the first step is to determine the corresponding multiple keyword variants based on the keyword field "fever". For example, expressions such as "fever", "elevated body temperature" and "high fever" are all variant forms of "fever". Subsequently, "fever" is associated with these variants to construct a "fever" keyword group. In the same way, corresponding keyword groups are also constructed for other symptom keyword fields such as "chest pain" and "headache", and finally all keyword groups are integrated to form a symptom keyword variant library.

[0081] When a new case text is subjected to lexical analysis and words or phrases are analyzed, if a keyword group in the keyword variant library is hit, such as the phrase "high fever", since it is a variant of the "fever" keyword group, "high fever" will be determined as the content of the "fever" keyword field, thereby providing key information basis for subsequent disease diagnosis, treatment plan formulation, etc., which can improve the accuracy of "keyword" recognition, be applicable to a variety of "language environments", and be more flexible.

[0082] Optionally, the character field includes a communication method character field and an email address character field. Step 1032 identifies a character sequence in the text message that conforms to a preset character sequence format, removes non-compliant characters in the character sequence, and obtains the character field content, specifically including:

[0083] Step 10321, for the communication method character field, identify the character sequence in the text information that conforms to the preset communication method format, remove the non-numeric characters in the identified character sequence, and obtain the content of the communication method character field, wherein the preset communication method format includes a numeric sequence format with an international code prefix, a numeric sequence format without an international code prefix, and a pure numeric sequence format, and the non-numeric characters include at least one of a space, a bracket, and a connector.

[0084] Step 10322: for the email character field, identify the character sequence in the text message that conforms to the preset email address format, and remove the spaces in the identified character sequence to obtain the content of the email character field.

[0085] In the above embodiment of the present application, the text message is such as: "Please contact me at +86 (123) 456-7890 or 1234567890." First, the communication method character field is identified. The preset communication method format covers a variety of situations, such as a digital sequence format with an international code prefix (for example, +861234567890), a digital sequence format without an international code prefix (for example, 1234567890), and a pure digital sequence format. Character sequences that conform to these formats are identified from the text, namely "+86 (123) 456-7890" and "1234567890". Subsequently, non-digital characters (such as spaces, brackets, connectors, etc.) in these character sequences are removed, and finally the purely digital communication method character field content is obtained: "+861234567890" and "1234567890".

[0086] Similarly, for email character fields, if a text message contains "My email address is username@example.com (mailto:name@example.com) or user_name@example.com.cn (mailto:user_name@example.com.cn), please check.", the system can identify character sequences that conform to the preset email address format, namely "username@example.com (mailto:name@example.com)" and "user_name@example.com.cn (mailto:user_name@example.com.cn)." Subsequently, spaces in these character sequences are removed to obtain the correct email character field content: "username@example.com (mailto:username@example.com)" and "user_name@example.com.cn (mailto:user_name@example.com.cn)."

[0087] Through such a processing flow, communication methods and email information can be accurately extracted and standardized, providing reliable support for subsequent communication or data management.

[0088] Step 104 , fill the content of the field to be filled into the required filling field position in the preset text template to obtain a filled compliant text, wherein the preset text template includes the required filling field.

[0089] In the above embodiment of the present application, the automatically parsed content of the to-be-filled field is automatically filled into the required filling field position, and finally the filled compliant text is obtained. By automatically parsing the original text, the content to be filled is generated according to the rules and accurately filled, thereby improving the text processing efficiency and compliance in the financial and medical fields.

[0090] Alternatively, as Figure 3 As shown, after obtaining the filled compliant text in step 104, the method further includes:

[0091] Step 105: Display the filled-in compliant text on a preset information display platform.

[0092] Step 106, when a text modification instruction is received, the filled-in compliant text is modified based on the to-be-modified requirement filling field corresponding to the text modification instruction and the to-be-modified field content corresponding to the to-be-modified requirement filling field, and the modified filled-in compliant text is redisplayed on the preset information display platform, wherein the text modification instruction is triggered by the user based on the preset information display platform, the text modification instruction corresponds to the to-be-modified requirement filling field and the to-be-modified field content corresponding to the to-be-modified requirement filling field.

[0093] In the above-described embodiment of the present application, the computer device sends the filled-in, compliant text to the front-end webpage via the back-end server, which then renders it on the display interface of a preset information display platform using, for example, HTML, CSS, and JavaScript. The front-end page design includes a "Modify Information" button for the user to trigger the text modification instruction.

[0094] When the user clicks the "Modify Information" button to trigger the text modification instruction, the text modification refers to the fields to be modified and the contents of the fields to be modified. After receiving the text modification instruction, the back-end server parses the fields to be modified and the corresponding contents of the fields to be modified. It uses string processing technology (such as regular expressions or string replacement) to find the original value corresponding to the "field to be modified" in the filled-in compliant text and replaces it with the "content of the field to be modified". Then, the back-end server sends the modified text to the front-end, which dynamically updates the page content through JavaScript without reloading the entire page, improving the user experience. To this end, it can facilitate users to check in real time whether the automatically parsed content is accurate, and provide the function of modification at any time to improve the accuracy of the final text formation.

[0095] Optionally, the original text supports multiple preset text formats, wherein the preset text formats include PDF and Word, and the original text includes a single content form and a mixed content form, the single content form includes text, tables and charts, and the mixed content form includes any combination of text, tables and charts.

[0096] In the above-mentioned embodiment of the present application, the original text display and modification functions of multiple preset text formats (such as PDF and Word) are supported, especially for the case where the original text contains a single content form (text, table, chart) and mixed content form (a combination of text, table, chart), it can comprehensively consider the processing methods of different formats, content analysis and reconstruction, and user interaction, is applicable to various original texts, is convenient for users, and can improve convenience.

[0097] In another specific embodiment, it can also be applied to the application scenario of automatic resume parsing. Currently, when users apply for positions on major recruitment websites, they need to repeatedly fill in the same information on different websites, which takes a lot of time to enter one by one, and the experience is extremely poor. Although some websites support users to upload files in PDF or Word format, the traditional parsing tools have a low success rate in identifying and matching due to the complex layout of the files uploaded by users. The problems caused by this are as follows:

[0098] 1. Manual filling is inefficient. Users must repeatedly fill out resume fields (such as name, education background, and work experience) across different recruitment platforms, which is time-consuming and prone to errors. Complex fields (such as project experience and skill descriptions) must be manually split and categorized, resulting in a poor user experience.

[0099] 2. Limitations of traditional parsing technology and insufficient OCR technology: it can only recognize text content but cannot understand semantic associations (for example, it cannot distinguish whether "2019-2021" is working time or project time).

[0100] 3. The rule engine is rigid and relies on predefined templates, which cannot handle unstructured resumes (such as free-form PDF or Word documents).

[0101] 4. Poor format compatibility: Different resume layouts (tables, columns, mixed text and graphics) lead to a high error rate in information extraction.

[0102] 5. Data silo problem: users need to repeatedly upload resume files on multiple recruitment platforms and lack a unified data export solution.

[0103] In the above embodiments of the present application, the natural language reasoning ability of a large model (such as a Transformer architecture) can be used to identify implicit relationships in resumes (such as the matching degree of work experience, skills and positions corresponding to the time range). Example: Identify the "start and end time", "school" and "major" fields in "2019.09-2023.06, XX University Computer Science and Technology Major". When using a large model, you can set the corresponding required filling fields and preset filling rules to obtain the resume format text that ultimately meets the requirements of the recruitment platform, that is, the "compliant text after filling". When using a large model for automatic text filling, you can also set some prompt words (required filling fields) to extract key fields from Chinese resumes and output them in a structured JSON format. In this process, specific rules (preset filling rules) need to be followed, for example:

[0104] 1. Data extraction requirements, including extraction of basic information and supplementary information:

[0105] Basic information:

[0106] Name (keyword field): directly extract the standard name.

[0107] Contact information (communication method character field): also known as mobile phone number, identifies phone numbers that match the "+86-" format, have no international codes, or are in simple numeric format, and removes all non-numeric characters (such as spaces, brackets, and hyphens).

[0108] Email (Email character field): Extracts the full email address, removing leading and trailing spaces or other symbols.

[0109] Expected Monthly Salary (Amount field): If a fixed value (12000) is entered, the value is directly extracted. If a range (10000-20000) is entered, the maximum value is taken and returned.

[0110] Additional information:

[0111] Age (difference field): Calculated based on the year of birth, using the current date as a benchmark for estimation (retain integers).

[0112] Highest Education Level (Keyword field): Supports the following options: "PhD," "Master," "Bachelor's," "College," "High School," or other valid educational level identifiers. If not specified, the default value is null.

[0113] In particular, when identifying "mobile phone numbers", it is necessary to identify mobile phone numbers in a specific format, that is, the preset communication method formats include three types, so that the large model can recognize three formats of telephone numbers: one is a number with a "+86-" prefix, "+86" is China's international telephone area code, and this format clearly indicates that the number belongs to China; the second is a number without international codes, that is, it may be a number representation used purely domestically; the third is a number in a simple digital format, that is, a number consisting only of numbers.

[0114] Next, after identifying a phone number that matches the above format, it will be further processed to remove all non-numeric characters contained in the number. Characters such as spaces, brackets (such as the common brackets used to separate different parts of the number, such as (123)), and connectors (such as the "-" used to separate numeric segments) will be removed, and the final result is a pure phone number consisting only of numbers.

[0115] For example, for the number "+86-13812345678", "+86-" will be removed after recognition, and "13812345678" will be obtained; for the number "(139)1234-5678", the brackets and connectors will be removed, and it will become "13912345678".

[0116] 2. Format requirements:

[0117] The output must be in JSON format without any additional content or comments.

[0118] 3. Handling of exceptions:

[0119] If a field is not provided in the resume, then "null" is returned.

[0120] If the field format is ambiguous (e.g. a phone number lacking a country code), the most common number format is prioritized.

[0121] 4. Language and cultural adaptability:

[0122] It is necessary to strictly follow the common format and expression habits of Chinese resumes (such as "expected salary", "highest education", etc.).

[0123] At the same time, it can identify variations of field names (such as "contact number", "mobile phone", "e-mail", etc.) and uniformly map them to standard field names.

[0124] 5. Performance requirements:

[0125] Ensure the parsing process is efficient and accurate, avoiding costly errors caused by inconsistent formats or missing information.

[0126] Finally, according to the form structure of the target recruitment system, the field name and hierarchy of the output JSON are adaptively adjusted (for example, mapping "work experience" to "professionalExperience" or "workHistory").

[0127] Furthermore, in the process of using large models to automatically parse resumes, you can also Figure 4 As shown, the user first uploads the resume file, which supports PDF or Word format. Then, the file format is standardized, such as PDF to text conversion and Word parsing, and then content cleaning is performed, such as denoising, segmentation, key information marking, etc. Then, the large model parsing engine is started to perform semantic analysis and field recognition, such as time range, company name and skill classification, etc., and then context relationship modeling is performed, such as associating project experience with time, position with company, and generating structured data (JSON format). Finally, field mapping adaptation is performed to match the API requirements of the target system, so that it can be automatically exported to the recruitment system through the APU or database interface.

[0128] By applying the technical solution of this embodiment, an end-to-end automated process can be achieved without human intervention, and JSON data that meets the requirements of the recruitment system API interface can be directly generated and automatically submitted through a secure protocol.

[0129] Further, as Figure 1The specific implementation of the method, the embodiment of the present application provides a text automatic filling device, such as Figure 5 As shown, the device includes:

[0130] The original text acquisition module 201 is used to acquire the original text;

[0131] The original text recognition module 202 is used to recognize the original text and obtain multiple text messages;

[0132] The filling content parsing module 203 is used to parse the text information based on the preset filling rules corresponding to the required filling fields, and generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple required filling fields, each of which corresponds to a preset filling rule;

[0133] The compliance text filling module 204 is used to fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field.

[0134] Optionally, the required filling fields include a keyword field, a character field, an amount field, and a time difference field. The keyword field corresponds to a keyword, and the contents of the fields to be filled include the contents of the keyword field, the character field, the amount field, and the time difference field. The filling content parsing module 203 is further used to:

[0135] For keyword fields, semantic analysis is performed to identify keyword field content in text information that meets the keyword semantic characteristics;

[0136] For a character field, identify a character sequence in the text information that conforms to a preset character sequence format, remove non-compliant characters in the character sequence, and obtain the character field content;

[0137] For the amount field, identify the amount in the text information. If the identified amount is a fixed value, the fixed value is determined as the content of the amount field. If the identified amount is a numerical range, the maximum value in the numerical range is taken as the content of the amount field.

[0138] For the time difference field, identify the time in the text information for which the time difference calculation is required, and obtain the content of the time difference field based on the difference between the current time and the identified time.

[0139] Optionally, the filling content parsing module 203 is further configured to:

[0140] Extracting semantic features of keywords corresponding to the keyword field, and performing lexical analysis on text information to obtain words or phrases;

[0141] The words or phrases obtained by lexical analysis are matched with the semantic features of the keywords. If the semantics of the words or phrases match the semantic features of the keywords, the matched words or phrases are determined as the keyword field content of the keyword field.

[0142] Optionally, the filling content parsing module 203 is further configured to:

[0143] Based on the keywords corresponding to the keyword fields, determining multiple keyword variants corresponding to the keywords, associating the keywords with the keyword variants to obtain keyword groups corresponding to the keyword fields, and constructing a keyword variant library based on the keyword groups of each keyword field;

[0144] For any keyword group in the keyword variant library, when a word or phrase obtained by lexical analysis hits any keyword or keyword variant in the keyword group, the hit word or phrase is determined as the keyword field content of the keyword field.

[0145] Optionally, the character field includes a communication method character field and an email address character field, and the filled content parsing module 203 is further configured to:

[0146] For the communication method character field, identifying a character sequence in the text message that conforms to a preset communication method format, removing non-numeric characters from the identified character sequence, and obtaining content of the communication method character field, wherein the preset communication method format includes a numeric sequence format with an international code prefix, a numeric sequence format without an international code prefix, and a purely numeric sequence format, and the non-numeric characters include at least one of a space, a bracket, and a connector;

[0147] For the email character field, a character sequence that conforms to a preset email address format in the text message is identified, and spaces in the identified character sequence are removed to obtain the email character field content.

[0148] Optionally, the original text supports multiple preset text formats, wherein the preset text formats include PDF and Word, and the original text includes a single content form and a mixed content form, the single content form includes text, tables and charts, and the mixed content form includes any combination of text, tables and charts.

[0149] Further, as Figure 3 The specific implementation of the method, the embodiment of the present application provides another text automatic filling device, such as Figure 6 As shown, the device includes:

[0150] The original text acquisition module 201 is used to acquire the original text;

[0151] The original text recognition module 202 is used to recognize the original text and obtain multiple text messages;

[0152] The filling content parsing module 203 is used to parse the text information based on the preset filling rules corresponding to the required filling fields, and generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple required filling fields, each of which corresponds to a preset filling rule;

[0153] The compliant text filling module 204 is used to fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field;

[0154] The compliance text display module 205 is used to display the filled compliance text on the preset information display platform; when a text modification instruction is received, the filled compliance text is modified based on the to-be-modified requirement filling field corresponding to the text modification instruction and the to-be-modified field content corresponding to the to-be-modified requirement filling field, and the modified filled compliance text is redisplayed on the preset information display platform, wherein the text modification instruction is triggered by the user based on the preset information display platform, and the text modification instruction corresponds to the to-be-modified requirement filling field and the to-be-modified field content corresponding to the to-be-modified requirement filling field.

[0155] It should be noted that for other corresponding descriptions of the functional units involved in the text automatic filling device provided in the embodiment of the present application, please refer to Figures 1 to 2 The corresponding description in the method will not be repeated here.

[0156] Based on the above Figures 1 to 4 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, the embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned operation is performed. Figures 1 to 4 The text auto-fill method shown.

[0157] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0158] Based on the above Figures 1 to 4 The method shown, and Figure 5 、 Figure 6In order to achieve the above-mentioned purpose, the embodiment of the present application further provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figures 1 to 4 The text auto-fill method shown.

[0159] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a Wi-Fi interface), etc.

[0160] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0161] The storage medium may also include an operating system and a network communication module. An operating system is a program that manages and stores the hardware and software resources of a computer device, supporting the execution of information processing programs and other software and / or programs. The network communication module facilitates communication between components within the storage medium, as well as with other hardware and software within the physical device.

[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or by hardware to obtain the original text; identify the original text to obtain multiple text messages; parse the text messages based on the preset filling rules corresponding to the required filling fields, and generate the content of the fields to be filled that conforms to the preset filling rules, wherein the required filling fields include multiple, and each required filling field corresponds to a preset filling rule; fill the content of the fields to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field. It can automatically parse the original text, generate the content to be filled according to the rules and fill it accurately, thereby improving the efficiency and compliance of text processing in the financial and medical fields.

[0163] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0164] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosures are only a few specific implementation scenarios of the present application, but the present application is not limited thereto, and any changes that can be made by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A text automatic filling method, characterized in that: The method comprises: Get the original text; Recognizing the original text to obtain multiple text messages; Based on the preset filling rules corresponding to the required filling fields, the text information is parsed to generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple fields, each of which corresponds to a preset filling rule; Fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field.

2. The method according to claim 1, characterized in that The required filling fields include a keyword field, a character field, an amount field, and a time difference field. The keyword field corresponds to a keyword, and the contents of the fields to be filled include the contents of the keyword field, the character field, the amount field, and the time difference field. Based on the preset filling rules corresponding to the required filling fields, the text information is parsed to generate the contents of the fields to be filled that comply with the preset filling rules, including: For keyword fields, semantic analysis is performed to identify keyword field content in text information that meets the keyword semantic characteristics; For a character field, identify a character sequence in the text information that conforms to a preset character sequence format, remove non-compliant characters in the character sequence, and obtain the character field content; For the amount field, identify the amount in the text information. If the identified amount is a fixed value, the fixed value is determined as the content of the amount field. If the identified amount is a numerical range, the maximum value in the numerical range is taken as the content of the amount field. For the time difference field, identify the time in the text information for which the time difference calculation is required, and obtain the content of the time difference field based on the difference between the current time and the identified time.

3. The method according to claim 2, characterized in that The semantic analysis is used to identify keyword field contents in the text information that meet the keyword semantic characteristics, including: Extracting semantic features of keywords corresponding to the keyword field, and performing lexical analysis on text information to obtain words or phrases; The words or phrases obtained by lexical analysis are matched with the semantic features of the keywords. If the semantics of the words or phrases match the semantic features of the keywords, the matched words or phrases are determined as the keyword field content of the keyword field.

4. The method according to claim 3, characterized in that After performing lexical analysis on the text information to obtain words or phrases, the method further includes: Based on the keywords corresponding to the keyword fields, determining multiple keyword variants corresponding to the keywords, associating the keywords with the keyword variants to obtain keyword groups corresponding to the keyword fields, and constructing a keyword variant library based on the keyword groups of each keyword field; For any keyword group in the keyword variant library, when a word or phrase obtained by lexical analysis hits any keyword or keyword variant in the keyword group, the hit word or phrase is determined as the keyword field content of the keyword field.

5. The method according to claim 2, characterized in that The character fields include a communication method character field and an email address character field. For the character fields, character sequences that conform to a preset character sequence format in the text message are identified, and after removing non-compliant characters from the identified character sequences, the character field content is obtained, including: For the communication method character field, identifying a character sequence in the text message that conforms to a preset communication method format, removing non-numeric characters from the identified character sequence, and obtaining content of the communication method character field, wherein the preset communication method format includes a numeric sequence format with an international code prefix, a numeric sequence format without an international code prefix, and a purely numeric sequence format, and the non-numeric characters include at least one of a space, a bracket, and a connector; For the email character field, a character sequence that conforms to a preset email address format in the text message is identified, and spaces in the identified character sequence are removed to obtain the email character field content.

6. The method according to any one of claims 1 to 5, characterized in that After obtaining the filled compliant text, the method further includes: Displaying the filled-in compliance text on a preset information display platform; When a text modification instruction is received, the filled-in compliant text is modified based on the to-be-modified requirement filling field corresponding to the text modification instruction and the to-be-modified field content corresponding to the to-be-modified requirement filling field, and the modified filled-in compliant text is redisplayed on the preset information display platform, wherein the text modification instruction is triggered by the user based on the preset information display platform, and the text modification instruction corresponds to the to-be-modified requirement filling field and the to-be-modified field content corresponding to the to-be-modified requirement filling field.

7. The method according to claim 6, characterized in that The original text supports multiple preset text formats, wherein the preset text formats include PDF and Word, and the original text includes a single content form and a mixed content form. The single content form includes text, tables and charts, and the mixed content form includes any combination of text, tables and charts.

8. A text automatic filling device, characterized in that: The device comprises: Original text acquisition module, used to obtain original text; An original text recognition module, used to recognize the original text and obtain multiple text messages; A filling content parsing module is used to parse the text information based on the preset filling rules corresponding to the required filling fields, and generate the content of the field to be filled that complies with the preset filling rules, wherein the required filling fields include multiple fields, and each required filling field corresponds to a preset filling rule; The compliant text filling module is used to fill the content of the field to be filled into the required filling field position in the preset text template to obtain the filled compliant text, wherein the preset text template includes the required filling field.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for automatic text filling in any one of claims 1 to 7 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method for automatic text filling in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Advertisement data analysis method and device, equipment and storage medium

    CN121352882A