Method and apparatus for adaptive matching of engineering file data based on textual similarity
By automatically matching engineering document data using text similarity algorithms, the problem of inability of process procedures and signature processes to adapt to migration and changes has been solved. This has enabled adaptive matching of construction project data, reduced manual processing, and improved the system's adaptability and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the procedures and signature processes for construction project documents cannot adapt to migration and changes, resulting in errors and omissions in procedures and the inability to generate signatures properly. This is especially true when there are cross-regional or standard upgrades, which require a lot of manual intervention.
By employing a text similarity-based method, the similarity of strings to be compared is matched using a preset algorithm. The system automatically identifies and generates a result table with the highest similarity, enabling adaptive matching of processes and signature procedures, and reducing manual processing.
It improves the adaptability of processes and signing procedures, reduces manual intervention, enhances the system's flexibility and accuracy, and adapts to changes in different regions and standards.
Smart Images

Figure CN116340589B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital information technology, and in particular to an adaptive matching method and system for construction engineering data procedures and signature processes based on text similarity. Background Technology
[0002] During the compilation of construction project data, in order to facilitate the compilation by professionals, the system will provide some commonly used procedures based on different engineering specialties. Different procedures need to be associated with relevant templates and tables of different specifications in order to realize the creation and compilation of tables according to procedures.
[0003] In addition, when submitting electronic signatures for construction project documents, it is also necessary to generate the signature process according to the content set in the form. Although most signature positions and related processes can be accurately generated by manual means and various rules, there will still be omissions. For example, there may be errors in the process when generating cross-regional processes, and signatures may not be generated normally due to process mismatch.
[0004] Therefore, the existing technology has at least the following technical drawbacks: the construction process and signing procedures cannot adapt to the needs of migration and change. Summary of the Invention
[0005] This application provides an adaptive matching method and system for construction engineering data procedures and signing processes based on text similarity, enabling construction procedures and signing processes to adapt to migration and change requirements.
[0006] An adaptive matching method for engineering file data based on text similarity includes:
[0007] The step is to obtain the name of the target specification table in the original process.
[0008] The extraction step involves extracting the string to be compared from the name of the target specification table.
[0009] The matching step involves using a preset algorithm to match the result table with the highest similarity to the string to be compared.
[0010] Preferably, the method further includes:
[0011] Obtain the process tree diagram of the original process;
[0012] The preset algorithm is specifically configured as follows: collect the string dataset corresponding to the original process position in the process tree diagram, and perform matching in the string dataset.
[0013] Preferably, the method further includes:
[0014] When submitting electronic signatures for project documents, check for missing signatures;
[0015] Obtain the process option string for the missing signature;
[0016] The preset algorithm is used to match the result table with the highest similarity to the process option string, and the result table includes chapter descriptions.
[0017] Preferably, the preset algorithm is a similarity algorithm based on edit distance and word vector similarity.
[0018] Preferably, the similarity algorithm is specifically implemented as follows:
[0019] Get the string to be compared;
[0020] Preprocessing of the string by replacing or transforming word vectors;
[0021] Calculate the similarity and take the maximum value;
[0022] Calculate whether the similarity meets the standard, and process and output the matching results for the target table that meets the standard.
[0023] Preferably, the method includes: assigning different weights to the corresponding word vectors of the two strings to be compared.
[0024] Preferably, the post-matching processing includes: table association and / or chapter position settings.
[0025] An apparatus for adaptive matching of engineering file data based on text similarity includes: the aforementioned method for adaptive matching of engineering file data based on text similarity.
[0026] A system for adaptive matching of engineering file data based on text similarity includes: the aforementioned apparatus for adaptive matching of engineering file data based on text similarity.
[0027] A device for adaptive matching of engineering file data based on text similarity, comprising:
[0028] At least one processing device; and
[0029] A memory communicatively connected to the at least one processing device; wherein,
[0030] The memory stores instructions that can be executed by the at least one processing device to enable the at least one processing device to perform the above-described method.
[0031] The present invention provides a method for adaptive matching of engineering document data based on text similarity, comprising: an acquisition step, acquiring the name of a target specification table in the original process; an extraction step, extracting a string to be compared from the target specification table name; and a matching step, using a preset algorithm to match a result table with the highest similarity to the string to be compared. When submitting an electronic signature for an engineering document, missing signature items are retrieved; the process option string of the missing signature item is acquired; and the preset algorithm is used to match a result table with the highest similarity to the process option string, wherein the result table includes a signature description. In the preparation of construction engineering documents, by utilizing the analysis of data in a specific field, a similarity matching method is used to generate process data and adapt missing signatures and processes. Attached Figure Description
[0032] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0033] Figure 1 This is a flowchart illustrating the adaptive matching method for engineering file data based on text similarity in an embodiment of this application.
[0034] Figure 2 This is a schematic diagram of the method for adaptive matching of engineering file data based on text similarity in an embodiment of this application.
[0035] Figure 3 This is a schematic diagram of the method for adaptive matching of engineering file data based on text similarity in an embodiment of this application.
[0036] Figure 4 This is a schematic diagram of the method for adaptive matching of engineering file data based on text similarity in an embodiment of this application.
[0037] Figure 5 This is a schematic diagram of a system for adaptive matching of engineering file data based on text similarity in an embodiment of this application;
[0038] Figure 6 This is a schematic diagram of the structure of the processing device based on the embodiments of this application. Detailed Implementation
[0039] This application provides an adaptive matching method and system for construction engineering data procedures and signing processes based on text similarity, enabling construction procedures and signing processes to adapt to migration and change requirements.
[0040] Currently, the procedures of common systems rely on manual settings, either provided by default in the product or set by the user. When there are updates to the standards or when they are applied to other similar provinces, similar professions, or similar fields, a lot of manual intervention is still required, resulting in a large amount of time investment and an inability to meet migration and change requirements in a timely manner.
[0041] This invention aims to match or generate new, usable data based on existing data for similar application scenarios. Typical scenarios are as follows:
[0042] Process generation: If process data already exists for a national standard, but no process is available for a specific provincial standard, or if provincial standard data already exists for one province, but another province has similar usage habits, then generating new data based on existing data to meet the new requirements offers significant economic advantages and ease of use.
[0043] Missing signature generation: When the signature set in the file does not match the process set by the user, the missing signature can be supplemented based on the similarity between the process and the information in the table, which can make up for the insufficient coverage of the signature position identified by manual and automatic identification.
[0044] The present invention discloses an adaptive matching method and system for construction engineering data procedures and signature processes based on text similarity, which enables construction procedures and signature processes to adapt to migration and change requirements.
[0045] Terminology definition:
[0046] Process adaptation, also known as process generation adaptation, refers to the process of compiling engineering data that is related to the construction process of the construction engineering profession. For example, the process can be divided into the project preparation stage, the construction implementation stage, and the completion and acceptance stage. Furthermore, the construction implementation stage includes different sub-sections, which in turn include sub-sub-sections, which in turn include sub-items, etc.
[0047] Each level or process requires the filling in of relevant engineering data, thus the process and the relevant forms in the specifications can be related.
[0048] During the initial process organization, the relationships between processes and relevant tables in a specific standard are established manually (or with semi-automatic assistance) according to certain rules. However, when the application scope shifts from national standards to local standards, from one province to another, or when standards are upgraded, continuing to organize the process manually would require a huge amount of work and would not allow for timely updates.
[0049] Processes are relatively stable data. Based on the correlation between past or similar processes and existing standard tables, they can be applied to new or similar standards in a certain way. String similarity matching is used to achieve matching degree, accuracy and convenience.
[0050] like Figure 1 As shown, an adaptive matching method for engineering file data based on text similarity includes:
[0051] S11: Obtain the step, obtain the target specification table name in the original process;
[0052] S12: Extraction step, extract the string to be compared from the name of the target specification table;
[0053] For example, the original process may have a connection to a certain standard form, such as "010906_Fireproof Coating Engineering Inspection Batch Quality Acceptance Record Form", but the new standard does not have a completely identical form name that corresponds to it. By matching strings similarity, the most similar name "010906 Steel Structure Fireproof Coating Engineering Inspection Batch Quality Acceptance Record" is found, and the correct connection result is obtained.
[0054] Furthermore, based on the similarity after matching, the matching data with lower similarity is further reviewed. Using this method, the workload of manual processing is greatly reduced.
[0055] S13: Matching step, using a preset algorithm, matching the result table with the highest similarity to the string to be compared.
[0056] Based on the names of the process and the associated templates, the maximum similarity matching is performed using text similarity in similar specifications (if the similarity exceeds a certain threshold).
[0057] Obtain the process tree diagram of the original process;
[0058] The process is represented in data structure as a tree (tree storage structures vary; only a tree representation is given here for ease of understanding), the typical form is shown in the reference. Figure 2 :
[0059] The preset algorithm is specifically configured as follows: collect the string dataset corresponding to the original process position in the process tree diagram, and perform matching in the string dataset.
[0060] In the actual adaptation process, in addition to considering template similarity, we also need to consider the similarity of their positions in the tree, which is a comprehensive similarity judgment.
[0061] The above functions can be packaged into a tool for users to use directly. If a user has a certain process division that they are used to, an operation entry can be provided so that the user can import their personal process data into the system, specify the specifications of the associated application, and use the method provided by this invention to match the data that is closest to the user's personal process to meet more flexible data application needs.
[0062] refer to Figure 3 The document illustrates the signature process adaptation process, including:
[0063] S31: When submitting electronic signatures for project documents, retrieve any missing signatures;
[0064] When submitting project documents with electronic signatures, a check for omissions is performed based on the default signature configuration in the form and the signature process set by the user. Omissions are usually signature positions that were missed during manual binding or cannot be accurately identified according to existing rules.
[0065] S32: Obtain the process option string for the missing signature item;
[0066] S33: Using the preset algorithm, a result table is matched with the string that has the highest similarity to the process option string. The result table contains chapter descriptions.
[0067] Similarly, based on the options descriptions in the user flow, the corresponding descriptions in the table can be found through similarity matching, and the position where the signature should be placed can be determined according to certain rules to set the signature position, so as to ensure that the signature position matches the flow when signing.
[0068] For example, as shown in the table below,
[0069] Construction unit project stamp Chapter Position Construction Project Manager: Project stamp of supervision unit Chapter Position Chief Supervising Engineer: Chapter Position
[0070] As shown above, three corresponding stamp positions have been set in the table. When setting up the signing process, in addition to the three signing processes generated by the system by default, the user adds a "Construction Technical Manager" position. When submitting, the system finds that the "Construction Technical Manager" stamp position is missing by comparison. At this time, the similarity matching rules of this invention are applied, and the system adds the stamp position at "Construction Professional Manager" and submits, thus achieving a correct match.
[0071] In this case, even if the similarity is low, the assigned chapter position can usually compensate for the impact of missing chapters.
[0072] In summary, the above defines two typical scenarios for using similarity matching to solve data adaptation problems. In these scenarios, matching with the highest similarity usually produces better adaptation results. By applying this invention, manual input is reduced and the system's adaptability is improved.
[0073] Preferably, the preset algorithm is a similarity algorithm based on edit distance and word vector similarity.
[0074] For example, in the front end of a web application, using edit distance requires fewer dependencies and can achieve a better balance between performance and resources, while in the back end or when processing large amounts of data, using word vectors can achieve better results and batch processing capabilities.
[0075] The similarity range obtained in this invention is 0 to 1. A suitable threshold, such as 0.7, can be selected according to the data and test results, and is not limited thereto.
[0076] refer to Figure 4 The similarity algorithm is specifically implemented as follows:
[0077] S41: Obtain the string to be compared;
[0078] S42: Perform preprocessing on the string, such as substitution or transformation of word vectors;
[0079] S43: Calculate the similarity and take the maximum value;
[0080] S44: Calculate whether the similarity meets the standard, and process and output the matching results for the target table that meets the standard.
[0081] Preferably, the method includes: assigning different weights to the corresponding word vectors of the two strings to be compared.
[0082] When there are many identical words, such as "AAA inspection batch acceptance record" and "BBB inspection batch acceptance record", their similarity is very high if no processing is done, but they may actually be two completely different concepts. In such descriptions, the weight of the words "AAA" and "BBB" is much higher than the rest. In this case, when doing edit distance, it may be necessary to remove the parts that cannot express effective meaning. In word vector methods, the keyword weight can be increased through bag-of-words processing or other methods.
[0083] Preferably, the post-matching processing includes: table association and / or chapter position settings.
[0084] Once the similarity reaches the required level, further processing is performed, such as table association and chapter positioning.
[0085] refer to Figure 5 A system 5 for adaptive matching of engineering document data based on text similarity, comprising: a device 51 for adaptive matching of engineering document data based on text similarity. The operating principle of the adaptive matching device for engineering document data based on text similarity is described below. Figure 1-4 The methods and solutions illustrated and explained will not be repeated here.
[0086] Figure 6 The computing device 60 shown is a method for adaptive matching of engineering document data based on text similarity, which can be configured as an apparatus for adaptive matching of engineering document data based on text similarity.
[0087] like Figure 6As shown, the processing device 60 is presented in the form of a general-purpose processing device. The components of the processing device 60 may include, but are not limited to: at least one processing device 61, at least one memory 62, and a bus 63 connecting different system components (including memory 62 and processing device 61).
[0088] Bus 63 represents one or more of several bus structures, including a memory bus or memory controller, peripheral bus, processing device, or a local bus using any of the various bus structures.
[0089] The memory 62 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 621 and / or cache memory 622, and may further include read-only memory (ROM) 623.
[0090] The memory 62 may also include a program / utility 625 having a set (at least one) of program modules 624, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0091] Processing device 60 can also communicate with one or more external devices 64 (e.g., keyboard, pointing device, etc.), and with one or more devices that enable a user to interact with processing device 60, and / or with any device that enables processing device 60 to communicate with one or more other processing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 65. Furthermore, processing device 60 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 66. As shown, network adapter 66 communicates with other modules used in processing device 60 via bus 63. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with processing device 60, including but not limited to: microcode, device drivers, redundant processing devices, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0092] In summary:
[0093] The present invention provides a method for adaptive matching of engineering document data based on text similarity, comprising: an acquisition step, acquiring the name of a target specification table in the original process; an extraction step, extracting a string to be compared from the target specification table name; and a matching step, using a preset algorithm to match a result table with the highest similarity to the string to be compared. When submitting an electronic signature for an engineering document, missing signature items are retrieved; the process option string of the missing signature item is acquired; and the preset algorithm is used to match a result table with the highest similarity to the process option string, wherein the result table includes a signature description. In the preparation of construction engineering documents, by utilizing the analysis of data in a specific field, a similarity matching method is used to generate process data and adapt missing signatures and processes.
[0094] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0095] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for adaptive matching of engineering file data based on literal similarity, characterized in that, The method comprises the following steps: an acquisition step of acquiring a target specification table name in a previous process; an extraction step of extracting a to-be-compared string in the target specification table name; a matching step of matching, by a preset algorithm, a result table with the highest similarity to the to-be-compared string; when an engineering file is submitted for electronic signature, retrieving a signature omission according to a signature set by default in a table and a signature process set by a user; acquiring a process option string of the signature omission; matching, by the preset algorithm, a result table with the highest similarity to the process option string, the result table being provided with a chapter position description.
2. The method for text similarity based engineering file data adaptive matching according to claim 1, characterized in that, The method further comprises the following steps: acquiring a process tree diagram in which the previous process is located; the preset algorithm is specifically configured to collect a string data set corresponding to a position of the previous process in the process tree diagram, and matching is performed in the string data set.
3. The method for text similarity based engineering file data adaptive matching as claimed in claim 1 wherein, The preset algorithm is a similarity algorithm of an edit distance and a word vector similarity.
4. The method for text similarity based engineering file data adaptive matching according to claim 3, characterized in that, The similarity algorithm is specifically implemented as follows: acquiring a to-be-compared string; performing preprocessing on the string by giving a replacement or converting a word vector; calculating a similarity and taking a maximum item; calculating whether the similarity meets a standard, and performing processing and output after matching for a target table that meets the standard.
5. The method for text similarity based engineering file data adaptive matching as claimed in claim 1 wherein, The method further comprises the following steps: assigning different weights to word vectors corresponding to two to-be-compared strings.
6. The method for text similarity based engineering file data adaptive matching as claimed in claim 4 wherein, The processing after matching includes table association and / or chapter position setting.
7. An apparatus for adaptive matching of engineering file data based on literal similarity, the apparatus comprising: Embedding the method for engineering file data adaptive matching based on character similarity according to any one of claims 1-6.
8. A system for adaptive matching of engineering file data based on literal similarity, the system comprising: The device comprises: at least one processing device; 9. An apparatus for adaptive matching of engineering file data based on literal similarity, the apparatus comprising: and a memory in communication connection with the at least one processing device; wherein the memory stores instructions executable by the at least one processing device, and the instructions are executed by the at least one processing device to enable the at least one processing device to perform the method according to any one of claims 1-6. The device comprises: at least one processing device; and a memory in communication connection with the at least one processing device; wherein the memory stores instructions executable by the at least one processing device, and the instructions are executed by the at least one processing device to enable the at least one processing device to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Construction service table management method, system, storage medium and electronic terminal
CN110472103A
Disease code matching method and device of non-standard disease name and computer equipment
CN114386397A
Electronic signature positioning method and device
CN114708186A