Automatic form filling method based on word vector retrieval model
Through the automatic form filling method based on the word vector search model, the problems of limitations in the prior art of information entry and insufficient compatibility of multi-format form processing are solved, and efficient and accurate form filling is achieved.
Patent Information
- Application Number
- CN202510114241.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing automatic form filling method has limitations in information entry and lacks compatibility for multi-format form data processing, resulting in low form filling efficiency.
The automatic form filling method based on the word vector search model is adopted to analyze the key information of the historical form and the requirement form by constructing a keyword analysis model, obtain the target key-value pair, and obtain the target key-value information that needs to be filled by matching the key fields.
It improves form compatibility and information filling accuracy, realizes unified management and application of information, and significantly improves the efficiency of automated form filling and data accuracy.
Smart Images

Figure CN120197599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to an automatic form filling method based on a word vector retrieval model. Background Art
[0002] In work, there will be various application form filling tasks involving personal information, such as resume forms, employee registration forms, etc. For the parts involving personal information among different application forms, such as the same content like name, mobile phone number, etc., need to be filled in repeatedly. The filling formats of different types of application forms are not exactly the same. When encountering various application forms at work, it will lead to a large amount of repetitive input work. At the same time, the traditional manual input method has problems such as low efficiency and easy errors.
[0003] Currently, the automatic form filling tools used in the prior art usually include an initialization information entry stage and a new form filling stage. However, in the initialization information entry stage, information is manually entered into the database, and the error rate of information entry is relatively high and the user experience is poor; in the new form filling stage, first, the word document is parsed to extract the table information, and traditional regular expressions are used to extract all keywords in the table, the keywords are matched with the personal information database to obtain the corresponding personal information, and finally the information is filled in the corresponding blanks; the method of relying on regular expressions for keyword extraction has poor compatibility. When facing application forms with different formats, there will be different expressions for the same keyword fields, and special logical processing is required to avoid information filling errors, resulting in low efficiency in filling new forms.
[0004] The patent "A Form Mapping Method, Device, Computer Device and Storage Medium", publication number: CN110895533B, publication date: January 17, 2023, discloses: obtaining an original form and extracting all the original field names included in the original form; performing word segmentation processing on each original field name to obtain an original field word segmentation set and an original field word vector set; respectively matching the original field word vector set with the standard field word vector sets of at least two standard form templates, and obtaining the standard form template that meets the matching conditions as the form mapping template corresponding to the original form. However, this solution mainly processes the key data of the initial table and matches it with the standard vectors of the preset form templates, so as to store the new table data in the form templates with standard formats, in order to structure multi-source forms and improve the form query efficiency, and cannot be applied to automatically filling the corresponding keyword field values based on the key information of the new form. Summary of the Invention
[0005] The object of the present invention is to address the problem that the existing automatic form-filling method has limitations in information entry and lacks compatibility for processing multi-format form data, resulting in low form-filling efficiency. An automatic form-filling method based on a word vector retrieval model is proposed. By using the word vector retrieval model to construct a keyword parsing model containing a built-in keyword list, key information in historical form files and requirement form files is parsed respectively to obtain the first target key-value pairs and the second target key-value pairs. Taking advantage of the high semantic understanding performance of the word vector retrieval model, it can accurately identify the manifestation forms of synonymous information in different forms, taking into account the format diversity of different forms, improving the data parsing efficiency of form unit files, and ensuring the accuracy of key information at the same time. By querying and matching the keyword fields of the second target key-value pairs with those of the first target key-value pairs, the target key-value information to be filled is obtained, ensuring the accuracy of the key information in the target form and realizing reliable automation of form filling. It overcomes the limitations of conventional automatic form-filling methods due to manual limitations in initial information entry and lack of compatibility for processing multi-format form data, significantly improving the efficiency of automatic form filling while taking into account data accuracy.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is: an automatic form-filling method based on a word vector retrieval model, including the following steps: Establish a built-in keyword list and use it as the benchmark keyword of the word vector retrieval model to construct a keyword parsing model; Parse the key information in the historical form file through the keyword parsing model to obtain the first target key-value pairs; Parse the key information in the requirement form file through the keyword parsing model to obtain the second target key-value pairs; Match the keyword fields of the second target key-value pairs with those of the first target key-value pairs to obtain the target key-value information to be filled; Write the target key-value information into the target form to complete the automatic filling of the new form information.
[0007] In this solution, a keyword parsing model is constructed through a word vector retrieval model, which can more accurately identify the semantic relationships in form files, including identifying various variants of keyword fields, so as to extract all key information more efficiently, reduce misidentification, improve parsing efficiency and the compatibility of multi-format forms; by comparing the keywords in two key-value pairs, it is not necessary to screen and match all the text content of the form, reducing the amount of data processing and improving data processing efficiency. At the same time, according to the keywords matched between the first form and the second form, the key values corresponding to the keywords in the second form are obtained, and the target key-value information can be obtained, ensuring that only the correct key values are filled into the target form and guaranteeing the accuracy of the filling position; through the automatic filling of key values, the form filling process is automated, avoiding the time cost of manual operations and the error rate of manual filling, and taking into account the form filling efficiency while improving data accuracy.
[0008] Preferably, establishing the built-in keyword list and using it as the benchmark keyword of the word vector retrieval model to construct the keyword parsing model includes: Based on the business requirement information and historical form statistical information, obtain the keyword field information in the form; According to the keyword field information, establish a built-in keyword list and use it as the benchmark keyword of the word vector retrieval model; based on the benchmark keyword, construct the keyword parsing model by loading the word vector retrieval model library.
[0009] Preferably, parsing the key information of the historical form file through the keyword parsing model to obtain the first target key-value pair includes: Extract the first table information in the historical form file, and convert it into a first word vector based on the keyword parsing model; arrange and combine the first word vector and the benchmark word vector of the benchmark keyword to generate a number of first vector key pairs; calculate the vector correlation between the first word vector and the benchmark word vector in the first vector key pair; According to the vector correlation, screen the first vector key pairs to obtain the first target key-value pair.
[0010] Preferably, screening the first vector key pairs according to the vector correlation to obtain the first target key-value pair includes: Compare the vector correlation with the similarity threshold, and take the first vector key pairs with the vector correlation greater than or equal to the similarity threshold as the first screening result; Based on the first screening result, perform a second screening on the remaining first vector key pairs, where: When the vector correlations generated by combining the same first word vector with different reference word vectors all meet the primary screening, sort the vector correlations, and use the first vector key pair with the largest vector correlation as the first target key value pair; When the vector correlations of the same reference word vector corresponding to different first word vectors all meet the primary screening, perform a secondary screening based on the priority of the reference word vector, and use the first vector key pair with the highest priority of the reference word vector as the first target key value pair.
[0011] Preferably, the step of parsing the key information of the requirement form file through the keyword parsing model to obtain the second target key value pair includes: Extract the second table information in the requirement form file, and convert it into a second word vector based on the keyword parsing model; perform permutation and combination on the second word vector and the reference word vector of the reference keyword to generate a number of second vector key pairs; calculate the vector correlation between the second word vector and the reference word vector in the second vector key pairs; Screen the second vector key pairs according to the vector correlation to obtain the second target key value pair.
[0012] Preferably, the step of screening the second vector key pairs according to the vector correlation to obtain the second target key value pair includes: Compare the vector correlation with the similarity threshold, and use the second vector key pairs with the vector correlation greater than or equal to the similarity threshold as the primary screening result; Perform a secondary screening on the remaining second vector key pairs based on the primary screening result, where: When the vector correlations generated by combining the same second word vector with different reference word vectors all meet the primary screening, sort the vector correlations, and use the second vector key pair with the largest vector correlation as the second target key value pair; When the vector correlations of the same reference word vector corresponding to different second word vectors all meet the primary screening, perform a secondary screening based on the priority of the reference word vector, and use the second vector key pair with the highest priority of the reference word vector as the second target key value pair.
[0013] Preferably, both the first target key value pair and the second target key value pair at least include the cell coordinate information of the target keyword.
[0014] Preferably, the step of matching the keyword fields of the second target key value pair with the keyword fields of the first target key value pair to obtain the target key value information to be filled includes: Match the second target keyword with the first target keyword stored in the database, and perform a relationship mapping between the second target key-value pair containing the same target keyword and the first target key-value pair to obtain the target key-value information.
[0015] Preferably, the target key-value information further includes the second table information in the requirement form file that is different from the historical form file; If there is no first target keyword in the database that is the same as the second target keyword, recall the key value corresponding to the second target keyword and add it to the target key-value information for filling the new table.
[0016] Preferably, the method further includes: dynamically updating the built-in keyword list based on the vector correlation of the first vector key pair and the second vector key pair, specifically including: If the first word vector is equal to the second word vector and the vector correlations between the word vectors in the first vector key pair and the second vector key pair and the reference word vector both satisfy the low similarity threshold range, use the keyword corresponding to the first word vector or / and the second word vector as a new keyword and add it to the built-in keyword list.
[0017] Advantages of the present invention: Automatically parse the historical table and the requirement table through the keyword parsing model to extract keywords, thereby improving the table compatibility and the accuracy of information filling; By matching the target keywords, the key information of the two tables can be effectively integrated, the keywords with different expressions and their corresponding key values can be associated, realizing the unified management and application of information, and avoiding information isolation caused by differences in keyword expressions; improving the efficiency and accuracy of automatic filling of table information; at the same time, through the matching of target keywords, the information in the new table that is different from the historical table can be quickly obtained and automatically filled into the new table template as the information to be filled, significantly improving the automation of form filling and the integrity and accuracy of the final table information. Description of the Drawings
[0018] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0019] Figure 1 It is a flowchart of an automatic form filling method based on a word vector retrieval model according to an embodiment of the present application.
[0020] Figure 2Schematic diagram of an initialization information entry process according to an embodiment of the present application.
[0021] Figure 3 Schematic diagram of a keyword parsing process according to an embodiment of the present application.
[0022] Figure 4 Schematic diagram of a new form filling process according to an embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only the best embodiments of the present invention, which are only used to explain the present invention and do not limit the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Embodiment 1: As Figure 1 shown, an automatic form filling method based on a word vector retrieval model includes steps S1 - S5, where: S1. Establish a built - in keyword list and construct a keyword parsing model using it as the benchmark keyword of the word vector retrieval model; specifically including: Based on the business requirement information and historical form statistical information, obtain the keyword field information in the form; According to the keyword field information, establish a built - in keyword list and use it as the benchmark keyword of the word vector retrieval model; based on the benchmark keyword, construct the keyword parsing model by loading the word vector retrieval model library.
[0025] As an implementation manner, the built - in keyword list can be established in combination with specific business scenarios and form types. Analyze the key information required by the business through specific business scenarios to facilitate setting keyword fields; it is also possible to analyze past form data, count the keywords with higher frequencies and analyze them through expert experience, and use them as the built - in keywords to ensure the reliability of the built - in keyword list.
[0026] It is understandable that by constructing a keyword parsing model through a word vector retrieval model, various variants and synonyms of keyword fields in form files can be effectively processed, so as to associate different expressions of tabular text with the same key information, improve the integrity of form information extraction, reduce information omission or incorrect matching caused by expression differences, and improve the accuracy of key information parsing. For example, different expressions such as "date of birth", "birthday", and "year and month of birth" appear in the historical form and the new form respectively. By covering these keywords in the built-in keyword list, the accuracy of data extraction is ensured. It helps to achieve more accurate and flexible parsing of key information, provides reliable data support for automatic filling of the new form, and ensures the accuracy and integrity of the new form information.
[0027] Among them, the word vector retrieval model can be the text representation model bge-large. The core principle of the text representation model bge-large is based on the Embedding technology, which converts text into a high-dimensional vector representation. In this way, the model can capture the semantic information of the text, so as to achieve efficient text retrieval and semantic similarity calculation. S2. Parse the key information of the historical form file through the keyword parsing model to obtain the first target key-value pair; specifically including: Extract the first table information in the historical form file, and convert it into a first word vector based on the keyword parsing model; combine and arrange the first word vector and the benchmark word vector of the benchmark keyword to generate a number of first vector key pairs; calculate the vector correlation between the first word vector and the benchmark word vector in the first vector key pair; Filter the first vector key pairs according to the vector correlation to obtain the first target key-value pair.
[0028] Specifically, the filtering of the first vector key pairs according to the vector correlation to obtain the first target key-value pair includes: Compare the vector correlation with the similarity threshold, and take the first vector key pairs with the vector correlation greater than or equal to the similarity threshold as the first screening result; Perform a secondary screening on the remaining first vector key pairs based on the first screening result, where: When the vector correlations generated by combining the same first word vector with different benchmark word vectors all meet the first screening, sort the vector correlations, and take the first vector key pair with the largest vector correlation as the first target key-value pair; When the vector correlations between the same reference word vector and different first word vectors all meet the first - stage screening, a second - stage screening is performed based on the priority of the reference word vector, and the first - vector key pair with the highest priority of the reference word vector is used as the first target key - value pair.
[0029] As an implementation, as Figure 2 shown and Figure 3 shown, by reading the historical table provided by the user, supporting multiple - format files, such as excel or word format; after importing the historical table, the table data therein is parsed through a keyword parsing model, and the steps are as follows: First, extract the text information in the table cell, generate its word vector as the first word vector, and generate the word vector of the built - in keyword list as the reference word vector, to obtain a word - vector model representing the text in the table cell and the built - in keyword. Second, associate the first word vector and the reference word vector one by one to generate a number of key pairs; then calculate the vector correlations of the number of key pairs in sequence, compare the correlation with the similarity threshold, and screen out the key pairs that meet the threshold conditions. If there are multiple pairs of data among these key pairs, the similarities of these key pairs can be sorted, and the key pair with the greatest similarity is selected as the target key - value pair; or according to the keyword - field priority of these key pairs, the key pair with the highest priority is selected and used as the target key - value pair, and it is stored in the database for subsequent new - table filling. For example, in an inventory - management table, although "inventory quantity" and "stock quantity" are expressed differently, their word - vector similarities are relatively high. By calculating the similarity, they can be associated with the built - in keyword "inventory quantity" to ensure that the information is correctly extracted. The key information appearing in the table, including the cell - position coordinate information, can be extracted according to the target key - value pair.
[0030] Furthermore, the calculation of vector correlation can use cosine similarity to measure the similarity of the directions of two vectors, and obtain the similarity of the two vectors. The calculation formula is: cos(θ)=(A·B) / (||A||||B||); where: A·B represents the dot product (inner product) of the first word vector and the reference word vector, ||A|| and ||B|| respectively represent the norms (length or Euclidean norm) of the first word vector and the reference word vector, and the result range of cosine similarity is between [-1, 1]; If the directions of two vectors are exactly the same, then cos(θ)=1, that is, the maximum similarity; if the directions of two vectors are exactly opposite, then cos(θ)= - 1; if two vectors are orthogonal (i.e., independent of each other), then cos(θ)=0, meaning no similarity; calculate the vector correlations between the built - in keyword fields and the text in the table cell in sequence. If the cosine similarity is greater than the threshold (e.g., 0.7), then the two vectors are considered similar.
[0031] Further, the similarity threshold can be statistically analyzed according to the distribution of the similarity of key pairs. For example, during the testing process of the model, after calculating the similarity of word vector key pairs, a histogram or other statistical chart can be used to view the distribution range of similarity values, analyze the concentration area and dispersion degree of similarity, so as to preliminarily judge a more appropriate threshold range; then, by testing multiple thresholds, calculate the accuracy rate and recall rate under different similarity thresholds. At the same time, combined with the professional knowledge of the business field, determine the acceptable matching error range according to experience, and further obtain a more accurate similarity threshold.
[0032] Further, the setting of the keyword field priority can be set according to business knowledge and experience. For example, for the application form type related to individuals, the priority of keyword fields such as identity number, phone number, name, gender, etc. can be set in sequence for personal basic information.
[0033] In this embodiment, the word vector retrieval model can calculate the correlation of vectors, and build a keyword parsing model to accurately identify and associate information for different formats of tables and different expressions of keyword fields, ensuring the accuracy and reliability of the new form information filling; by calculating the correlation of vectors, the keywords that are semantically closest can be found, effectively processing synonyms and near-synonyms that appear in the table text, realizing precise matching of the table content, avoiding information that may be missed by only based on exact string matching, and helping to more accurately match the information in the source table to the correct fields in the target table when automatically filling in the form. By comprehensively using similarity calculation, similarity threshold, and keyword field priority, the process of information extraction can be optimized, and the efficiency and accuracy of extracting key information from the table can be improved.
[0034] S3. Parse the key information of the requirement form file through the keyword parsing model to obtain the second target key-value pair; specifically including: Extract the second table information in the requirement form file, and convert it into a second word vector based on the keyword parsing model; permute and combine the second word vector and the reference word vector of the reference keyword to generate a number of second vector key pairs; calculate the vector correlation between the second word vector and the reference word vector in the second vector key pairs; Filter the second vector key pairs according to the vector correlation to obtain the second target key-value pair.
[0035] Specifically, the filtering of the second vector key pairs according to the vector correlation to obtain the second target key-value pair includes: Compare the vector correlation with the similarity threshold, and use the second vector key pairs whose vector correlation is greater than or equal to the similarity threshold as the first screening result; Perform a secondary screening on the remaining second vector key pairs based on the first screening result, where: When the vector correlations generated by combining the same second word vector with different reference word vectors all meet the first screening criteria, sort the vector correlations, and use the second vector key pair with the largest vector correlation as the second target key-value pair; When the vector correlations of the same reference word vector corresponding to different second word vectors all meet the first screening criteria, perform a secondary screening based on the priority of the reference word vector, and use the second vector key pair with the highest priority of the reference word vector as the second target key-value pair.
[0036] Specifically, both the first target key-value pair and the second target key-value pair at least include the cell coordinate information of the target keyword.
[0037] S4. Match the keyword fields of the second target key-value pair with the keyword fields of the first target key-value pair to obtain the target key-value information to be filled; specifically including: Match the second target keyword with the first target keyword stored in the database, and perform a relationship mapping between the second target key-value pair and the first target key-value pair containing the same target keyword to obtain the target key-value information.
[0038] Specifically, the target key-value information further includes the second table information in the requirement form file that is different from the historical form file; If there is no first target keyword in the database that is the same as the second target keyword, recall the key value corresponding to the second target keyword and add it to the target key-value information for filling the new table.
[0039] As an implementation, such as Figure 4As shown, after parsing the new table according to the keyword parsing model, the effective keyword information in the current table is obtained, including the cell coordinates of keyword fields, etc., which is the second target key-value pair in this application and can be simplified to database keyword - keyword in the table - cell coordinates; read the key-value pairs of the historical table in the database, which is the first target key-value pair here and can be simplified to database keyword - database key value; perform keyword field query matching based on the two target key-value pairs. For example, the second target key-value pair is expressed as Name - First Name - A1, Mobile Phone - Phone Number - A2; the first target key-value pair is expressed as Name - Zhang San, Mobile Phone - 123456789; use the database keyword for matching, and the matching result can be simplified to database keyword - keyword in the table - cell coordinates - database key value. For example, Name - First Name - A1 - Zhang San, Mobile Phone - Phone Number - A2 - 123456789, so as to obtain the target key value to be filled as First Name - Zhang San.
[0040] In this embodiment, by matching the target keywords, the key information of the two tables can be effectively integrated, the keywords with different expressions and their corresponding key values are associated, and the unified management and application of information are realized, avoiding information isolation caused by differences in keyword expressions; through the matching logic, the same information in the new table and the historical table is accurately extracted, data reuse is realized, and data omission or incorrect filling caused by inconsistent keywords is effectively reduced, thereby improving the efficiency and accuracy of automatic filling of table information. At the same time, through the matching of target keywords, the information in the new table that is different from the historical table can be quickly obtained and automatically filled into the new table template as the information to be filled, significantly improving the automation of form filling and the integrity and accuracy of the final table information.
[0041] S5. Write the target key value information into the target form to complete the automatic filling of the new table information.
[0042] In this embodiment, through the automatic filling of the target key value information, manual intervention is reduced, the speed and efficiency of data processing are improved, the reuse efficiency of data between different format forms is improved, errors caused by manual searching and copying and pasting are avoided, the time cost of form filling is saved, and strong support is provided for improving the efficiency and accuracy of business processes.
[0043] Specifically, the method further includes: dynamically updating the built-in keyword list based on the vector correlation of the first vector key pair and the second vector key pair, specifically including: If the first word vector is equal to the second word vector and the vector correlations between the word vectors in the first vector key pair and the second vector key pair and the reference word vector both satisfy the low similarity threshold range, then the keyword corresponding to the first word vector or / and the second word vector is used as a new keyword and added to the built-in keyword list.
[0044] In some alternative embodiments, words with vector correlations lower than a set low similarity threshold can be used as new keywords. The low similarity threshold can be set through methods such as expert experience summary after the analysis of historical table data by the parsing model. For example, a word vector with a similarity of 0 indicates that it does not exist in the built-in list but appears in the requirement table or historical table, and it can be directly added as a new keyword to ensure the data integrity of table parsing. When adding new keywords, keyword retrieval can be performed to ensure that the newly added keywords do not exist in the built-in keyword list to avoid duplicate addition. For similar keywords, they can be merged into a compromise keyword segment through semantic analysis to avoid multiple redundant keywords in the built-in keyword list.
[0045] In this embodiment, new word vectors are obtained through real-time vector correlation analysis and added to the built-in keyword list to update the built-in keyword list, ensuring that the built-in keyword list is more complete, improving adaptability and reliability, and flexibly coping with the processing requirements of multi-format and multi-type form data.
[0046] The above specific embodiments are the preferred embodiments of the present invention, and do not limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to this specific embodiment. All equivalent changes made according to the shape, structure, and method of the present invention are within the protection scope of the present invention.
Claims
1. The automatic form filling method based on the word vector retrieval model is characterized by: The steps include: Create a built-in keyword list and use it as the benchmark keyword of the word vector retrieval model to build a keyword parsing model; Parsing key information of the historical form file by using the keyword parsing model to obtain a first target key-value pair; Parsing key information of the demand form file by using the keyword parsing model to obtain a second target key-value pair; Match the key field of the second target key-value pair with the key field of the first target key-value pair to obtain the target key-value information to be filled; The target key value information is written into the target form to complete the automatic filling of the new table information.
2. The automatic form filling method based on the word vector retrieval model according to claim 1 is characterized in that: The method of establishing a built-in keyword list and using it as a benchmark keyword of a word vector retrieval model to construct a keyword parsing model includes: Based on business demand information and historical form statistics, obtain key field information in the form; Establish a built-in keyword list based on the key field information and use it as the benchmark keyword of the word vector retrieval model; Based on the benchmark keyword, the keyword parsing model is constructed by loading a word vector retrieval model library.
3. The automatic form filling method based on the word vector retrieval model according to claim 1 is characterized in that: The step of performing key information parsing by using the keyword parsing model in response to the historical form file imported by the user to obtain a first target key-value pair includes: Extracting first table information from the historical form file, and converting it into a first word vector based on the keyword parsing model; arranging and combining the first word vector with a reference word vector of the reference keyword to generate a plurality of first vector key pairs; calculating the vector correlation between the first word vector and the reference word vector in the first vector key pair; The first vector key-value pairs are screened according to the vector relevance to obtain the first target key-value pairs.
4. The automatic form filling method based on the word vector retrieval model according to claim 3 is characterized in that: The step of screening the first vector key pair according to the vector correlation to obtain the first target key-value pair includes: Compare the vector correlation with a similarity threshold, and take the first vector key pair whose vector correlation is greater than or equal to the similarity threshold as a screening result; Based on the first screening result, the remaining first vector key pairs are screened a second time, wherein: When the correlations of vectors generated by combining the same first word vector with different reference word vectors all satisfy a screening, sorting the correlations of the vectors, and taking the first vector key pair with the largest vector correlation as the first target key-value pair; When the vector correlations of the same benchmark word vector corresponding to different first word vectors all satisfy the first screening, a second screening is performed based on the priority of the benchmark word vector, and the first vector key pair with the highest priority of the benchmark word vector is used as the first target key-value pair.
5. The automatic form filling method based on the word vector retrieval model according to claim 3 is characterized in that: The step of parsing key information according to the demand form file imported by the user through the keyword parsing model to obtain a second target key-value pair includes: Extracting the second table information from the demand form file, and converting it into a second word vector based on the keyword parsing model; arranging and combining the second word vector with the benchmark word vector of the benchmark keyword to generate a plurality of second vector key pairs; calculating the vector correlation between the second word vector and the benchmark word vector in the second vector key pair; The second vector key-value pairs are screened according to the vector correlation to obtain the second target key-value pairs.
6. The automatic form filling method based on the word vector retrieval model according to claim 5 is characterized in that: The step of screening the second vector key pair according to the vector correlation to obtain the second target key-value pair includes: Comparing the vector correlation with a similarity threshold, and taking a second vector key pair whose vector correlation is greater than or equal to the similarity threshold as a first screening result; Based on the first screening result, the remaining second vector key pairs are screened a second time, wherein: When the vector correlations generated by combining the same second word vector with different reference word vectors all satisfy a screening, sorting the vector correlations, and taking the second vector key pair with the largest vector correlation as the second target key-value pair; When the vector correlations of the same benchmark word vector corresponding to different second word vectors all satisfy the first screening, a second screening is performed based on the priority of the benchmark word vector, and the second vector key pair with the highest priority of the benchmark word vector is used as the second target key-value pair.
7. The automatic form filling method based on the word vector retrieval model according to claim 5 is characterized in that: The first target key-value pair and the second target key-value pair both include at least cell coordinate information of the target keyword.
8. The automatic form filling method based on the word vector retrieval model according to claim 5 is characterized in that: The step of matching the key field of the second target key-value pair with the key field of the first target key-value pair to obtain target key-value information to be filled includes: The target key value information is obtained by matching the second target keyword with the first target keyword stored in the database and performing relationship mapping between the second target key-value pair containing the same target keyword and the first target key-value pair.
9. The automatic form filling method based on the word vector retrieval model according to claim 1, characterized in that: The target key value information also includes the second table information in the demand form file which is different from the historical form file; If the first target keyword identical to the second target keyword does not exist in the database, the key value corresponding to the second target keyword is recalled and added to the target key value information to fill in a new table.
10. The automatic form filling method based on the word vector retrieval model according to claim 5, characterized in that: The method further includes: dynamically updating the built-in keyword list based on the vector correlation between the first vector keyword pair and the second vector keyword pair, specifically including: If the first word vector is equal to the second word vector and the vector correlations between the word vectors in the first vector key pair and the second vector key pair and the benchmark word vector both meet the low similarity threshold range, the keywords corresponding to the first word vector and / or the second word vector are taken as new keywords and added to the built-in keyword list.
Citation Information
Patent Citations
A form mapping method, apparatus, computer device, and storage medium
CN110895533B