Engineering file data adaptive matching method and system
By standardizing the engineering file data, schema matching degree calculation, keyword filtering and similarity coefficient analysis, feature file data mining and semantic correlation calculation, the problem of low accuracy of engineering file data matching in the existing technology is solved, and a more efficient data matching effect is achieved.
Patent Information
- Application Number
- CN202410841851.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-06-27
AI Technical Summary
The prior art is difficult to achieve accurate data matching in complex engineering file data and diversity engineering projects, resulting in reduced matching accuracy.
By obtaining the engineering file data to be matched, standardized processing and data architecture recognition, calculate the architecture matching degree; extracting file keywords, performing filtering and similarity coefficient calculations; extracting feature file data, mining environmental information and data logic; calculating semantic correlation degree, determining the matching file data and performing sequence adjustments.
Improve the accuracy of adaptive matching of engineering file data and can more effectively process complex and diverse engineering file data.
Smart Images

Figure CN118410011B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data matching, and in particular to a method and system for adaptively matching engineering file data. Background Art
[0002] Engineering file data refers to various file types used in the engineering field to describe, record and store engineering-related information and data. Depending on different engineering fields and project requirements, engineering file data may cover different information, such as structural design, mechanical devices, electrical and electronic, civil engineering, etc. These files usually contain various details, parameters, constraints, calculation results and other important information of the engineering project. They are the basis and basis for the realization of engineering projects. In order to improve the integrity of the data, the data needs to be matched.
[0003] The existing file data matching adopts the rule matching method, which uses pre-defined rules and patterns to match engineering file data. These rules are formulated based on specific metadata, attributes or formats and can be implemented through regular expressions. However, this method requires manual definition and maintenance of matching rules, which may be very difficult for complex engineering file data and diverse engineering projects, thereby reducing the matching accuracy of engineering file data. Summary of the invention
[0004] The present invention provides a method and system for adaptively matching engineering file data, the main purpose of which is to improve the accuracy of adaptively matching engineering file data.
[0005] To achieve the above object, the present invention provides a method for adaptively matching engineering file data, comprising:
[0006] Acquire engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify the data architecture corresponding to the target file data, and calculate the architecture matching degree between the data architectures;
[0007] Extracting file keywords from the target file data, screening out the file keywords to obtain preferred keywords, and calculating similarity coefficients between the preferred keywords;
[0008] Extracting characteristic file data from the target file data, mining environmental information corresponding to the characteristic file data, locating data positions corresponding to the characteristic file data, and analyzing data logic between the characteristic file data in combination with the data positions and the environmental information;
[0009] The semantic association between the target file data is calculated, and the corresponding matching file data is determined from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and the matching file data is sequenced according to the data logic to obtain the target matching file data.
[0010] Optionally, the step of performing standardization processing on the engineering file data to obtain target file data includes:
[0011] Performing data cleaning on the engineering file data to obtain cleaned file data;
[0012] Identifying the document image data in the cleaned document data, and analyzing the image connotation corresponding to the document image data;
[0013] According to the image connotation, the document image data is converted into image text data;
[0014] According to the image text data, the cleaning file data is updated to obtain target file data.
[0015] Optionally, analyzing the image connotation corresponding to the file image data includes:
[0016] Performing noise reduction processing on the image in the document image data to obtain a noise-reduced document image;
[0017] Performing edge detection on the denoised document image to obtain the document image edge;
[0018] According to the edge of the document image, segmenting the denoised document image to obtain a document image area;
[0019] Identify the region body in the document image region, and analyze the image connotation corresponding to the document image data based on the region body.
[0020] Optionally, calculating the architecture matching degree between the data architectures includes:
[0021] Analyze the architecture attributes corresponding to the data architecture, and calculate the attribute weights corresponding to the architecture attributes;
[0022] Classifying the data architecture according to the architecture attributes to obtain a classified data architecture;
[0023] Calculating the architecture overlap between the classified data architectures;
[0024] The architecture matching degree between the data architectures is calculated by combining the architecture overlap degree and the attribute weight.
[0025] Optionally, the calculating the similarity coefficient between the preferred keywords includes:
[0026] Performing vectorization processing on the preferred keyword to obtain a keyword vector;
[0027] The vector similarity between the keyword vectors is calculated by the following formula:
[0028] ;
[0029] Among them, Q represents the vector similarity between keyword vectors, and They represent the ath vector and a+1th vector in the keyword vector respectively. a and a+1 are the sequence numbers of the keyword vectors respectively. and They represent the norms of the ath vector and the a+1th vector respectively,
[0030] The vector similarities are normalized to obtain similarity coefficients between the preferred keywords.
[0031] Optionally, extracting the characteristic file data from the target file data includes:
[0032] Querying the quality evaluation index corresponding to the target file data;
[0033] Extracting the index data corresponding to the target file data according to the quality evaluation index;
[0034] Quantitatively process the indicator data to obtain a quantitative value of the indicator;
[0035] Calculate the quality score value corresponding to the quality evaluation indicator according to the indicator quantification value;
[0036] According to the quality score value, feature file data in the target file data is extracted.
[0037] Optionally, combining the data location and the environment information to analyze the data logic between the feature file data includes:
[0038] Determining the data order corresponding to the feature file data according to the data position;
[0039] Calculating an information entropy value corresponding to the environmental information, and screening out key environmental information from the environmental information according to the information entropy value;
[0040] Constructing a visualization chart corresponding to the key environmental information;
[0041] Analyzing information dependencies between the key environmental information according to the visualization chart and the data sequence;
[0042] According to the information dependency, the data logic between the feature file data is determined.
[0043] Optionally, the calculating the semantic association between the target file data includes:
[0044] Identify data content corresponding to the target file data;
[0045] Extracting features from the data content to obtain content features;
[0046] The Gini coefficient corresponding to the content feature is calculated by the following formula:
[0047] ;
[0048] Among them, D represents the Gini coefficient corresponding to the content feature, represents the proportion of the i-th feature in the content features, i represents the sequence number corresponding to the content feature, and t represents the number of content features;
[0049] According to the Gini coefficient, the data content is filtered to obtain target data content;
[0050] Performing semantic analysis on the target data content to obtain content semantics;
[0051] Calculating the relevance between the content semantics to obtain content relevance;
[0052] Combined with the content relevance, the semantic relevance between the target file data is obtained.
[0053] Optionally, calculating the relevance between the content semantics to obtain the content relevance includes:
[0054] Performing vector processing on the content semantics to obtain a semantic vector;
[0055] Calculating the semantic covariance between the content semantics and the semantic standard deviation of the content semantics according to the semantic vector;
[0056] Combining the semantic covariance and the semantic standard deviation, the correlation between the content semantics is calculated using the following formula:
[0057] ;
[0058] Among them, F represents the correlation between content semantics, represents the semantic covariance between the bth and b+1th semantics in the content semantics, and They respectively represent the semantic standard deviations corresponding to the b-th and b+1-th semantics in the content semantics, and b and b+1 respectively represent the serial numbers corresponding to the content semantics.
[0059] An engineering file data adaptive matching system, characterized in that the system comprises:
[0060] A matching degree calculation module is used to obtain engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify the data architecture corresponding to the target file data, and calculate the architecture matching degree between the data architectures;
[0061] A similarity coefficient calculation module is used to extract file keywords from the target file data, screen out the file keywords, obtain preferred keywords, and calculate similarity coefficients between the preferred keywords;
[0062] A data logic analysis module, used to extract characteristic file data from the target file data, mine the environment information corresponding to the characteristic file data, locate the data position corresponding to the characteristic file data, and analyze the data logic between the characteristic file data in combination with the data position and the environment information;
[0063] A data matching module is used to calculate the semantic association between the target file data, determine the corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and adjust the sequence of the matching file data according to the data logic to obtain the target matching file data.
[0064] The present invention can remove invalid data in the engineering file data by standardizing the engineering file data, and convert image data into text data, so as to facilitate the subsequent recognition processing of the subsequent data architecture. The present invention can remove unimportant keywords in the file keywords by screening the file keywords, reduce the redundant information in the file keywords, calculate the similarity coefficient between the preferred keywords, analyze the similarity between the preferred keywords, and provide a basis for the subsequent determination of matching file data. The present invention extracts the feature file data in the target file data, and can obtain representative data in the feature file data, which is convenient for the subsequent mining of environmental information and provides a basis for the analysis of data logic. The present invention can understand the semantic association between the target file data by calculating the semantic association between the target file data, and provides a basis for the subsequent improvement of the matching accuracy of the target file data. Therefore, the method and system for adaptive matching of engineering file data provided by the embodiment of the present invention can improve the accuracy of adaptive matching of engineering file data. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 A schematic diagram of a flow chart of a method for adaptively matching engineering file data provided by an embodiment of the present invention;
[0066] Figure 2 A functional module diagram of an engineering file data adaptive matching system provided by one embodiment of the present invention.
[0067] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0068] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0069] The embodiment of the present application provides a method for adaptively matching engineering file data. In the embodiment of the present application, the execution subject of the method for adaptively matching engineering file data includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for adaptively matching engineering file data can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0070] Reference Figure 1 FIG. 1 is a flow chart of a method for adaptively matching engineering file data provided by an embodiment of the present invention. In this embodiment, the method for adaptively matching engineering file data includes steps S1 to S4.
[0071] S1. Acquire engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify a data architecture corresponding to the target file data, and calculate an architecture matching degree between the data architectures.
[0072] The present invention can remove invalid data in the engineering file data and convert image data into text data by standardizing the engineering file data, so as to facilitate the subsequent improvement of recognition processing of subsequent data architecture, wherein the engineering file data is relevant file data used in the engineering field.
[0073] As an embodiment of the present invention, the standardized processing of the engineering file data to obtain target file data includes: data cleaning of the engineering file data to obtain cleaned file data, identifying file image data in the cleaned file data, analyzing the image connotation corresponding to the file image data, performing data conversion on the file image data according to the image connotation to obtain image text data, and updating the cleaned file data according to the image text data to obtain target file data.
[0074] Among them, the cleaned file data is the data obtained after removing duplicate data and invalid data in the engineering file data, the file image data is the image type data in the cleaned file data, the image connotation is the image content corresponding to the file image data, and the image text data is the data expressed by the file image data through the image content, which converts the image into text-type data.
[0075] Furthermore, data cleaning of the engineering file data can be achieved by a cleaning tool, which is compiled by a scripting language, such as JS language; data conversion of the file image data can be achieved by replacing it according to the image content.
[0076] Further, as an optional embodiment of the present invention, the analysis of the image content corresponding to the file image data includes: performing denoising processing on the image in the file image data to obtain a denoised file image, performing edge detection on the denoised file image to obtain a file image edge, performing segmentation processing on the denoised file image based on the file image edge to obtain a file image region, identifying a region body in the file image region, and analyzing the image content corresponding to the file image data based on the region body.
[0077] The document image edge is an image contour in the denoised document image, the document image region is a region obtained by segmenting an image region formed by the document image edge, and the region body is an object in the document image region.
[0078] Furthermore, the edge detection of the denoised document image can be achieved through the Canny edge detection algorithm; the segmentation processing of the denoised document image can be achieved through the threshold segmentation method; the regional subject in the document image area can be realized through the target detection algorithm, such as the Fast R-CNN algorithm; the image connotation corresponding to the document image data is obtained by analyzing the subject semantics corresponding to the regional subject.
[0079] By calculating the architecture matching degree between the data architectures, the present invention can understand the matching degree between the data architectures, so as to facilitate the subsequent matching processing of the target file data.
[0080] As an embodiment of the present invention, the calculation of the architectural matching degree between the data architectures includes: analyzing the architectural attributes corresponding to the data architectures, calculating the attribute weights corresponding to the architectural attributes, classifying the data architectures according to the architectural attributes to obtain classified data architectures, calculating the architectural overlap between the classified data architectures, and calculating the architectural matching degree between the data architectures by combining the architectural overlap and the attribute weights.
[0081] Among them, the architectural attribute is the type corresponding to the data architecture, such as a file name or a data type, and the architectural overlap indicates the degree of overlap between each architecture in the classification architecture. Furthermore, the analysis of the architectural attributes corresponding to the data architecture can be implemented through an object-oriented model, and the structure and relationship of the data are analyzed through the model to reveal the attributes of the data architecture; the attribute weights corresponding to the architectural attributes can be implemented through a weight calculator; the classification processing of the data architecture can be implemented through a K-means clustering algorithm, and the calculation of the architectural overlap between the classified data architectures can be implemented through a clustering evaluation index, such as a silhouette coefficient index; the architectural matching degree between the data architectures can be obtained by summing the architectural overlap corresponding to each architecture, and multiplying the sum by the corresponding attribute weight to obtain the architectural matching degree.
[0082] S2. extracting file keywords from the target file data, screening out the file keywords to obtain preferred keywords, and calculating similarity coefficients between the preferred keywords.
[0083] The present invention can remove unimportant keywords from the file keywords and reduce redundant information in the file keywords by screening the file keywords, calculate the similarity coefficients between the preferred keywords, analyze the similarity between the preferred keywords, and provide a basis for determining subsequent matching file data. Optionally, the screening of the file keywords can be achieved through a TF-IDF algorithm.
[0084] As an embodiment of the present invention, the calculating the similarity coefficient between the preferred keywords includes: performing vectorization processing on the preferred keywords to obtain keyword vectors, and calculating the vector similarity between the keyword vectors by the following formula:
[0085] ;
[0086] Among them, Q represents the vector similarity between keyword vectors, and They represent the ath vector and a+1th vector in the keyword vector respectively. a and a+1 are the sequence numbers of the keyword vectors respectively. and They represent the norms of the ath vector and the a+1th vector respectively,
[0087] The vector similarities are normalized to obtain similarity coefficients between the preferred keywords.
[0088] Optionally, the normalization process of the vector similarity is to eliminate the influence of differences between vectors, so as to facilitate the subsequent comparison of similarity coefficients. The normalization process of the vector similarity can be achieved by a Z-Score normalization method.
[0089] S3, extracting characteristic file data from the target file data, mining environmental information corresponding to the characteristic file data, locating data positions corresponding to the characteristic file data, and analyzing data logic between the characteristic file data in combination with the data positions and the environmental information.
[0090] The present invention extracts characteristic file data from the target file data, and can obtain representative data from the characteristic file data, thereby facilitating subsequent mining of environmental information and providing a basis for the analysis of data logic.
[0091] As an embodiment of the present invention, the extracting of characteristic file data from the target file data includes: querying a quality evaluation indicator corresponding to the target file data, extracting the indicator data corresponding to the target file data based on the quality evaluation indicator, quantizing the indicator data to obtain an indicator quantization value, calculating a quality score value corresponding to the quality evaluation indicator based on the indicator quantization value, and extracting the characteristic file data from the target file data based on the quality score value.
[0092] Among them, the quality evaluation index is the quality evaluation standard of the target file data, the indicator data is the data related to the quality evaluation index, such as the value of the data or the data time series, the indicator quantization value is the expression value corresponding to the indicator data, and the quality score value represents the score value corresponding to the quality evaluation index.
[0093] Furthermore, the query of the quality evaluation index corresponding to the target file data can be obtained from the Internet through human-computer interaction; the extraction of the index data corresponding to the target file data can be achieved through an extraction function, and the extraction function is compiled by a programming language; the quantization processing of the index data can be achieved through a one-hot encoding method; the quality score value corresponding to the quality evaluation index can be obtained by calculating the sum of the quantization values of the index.
[0094] The present invention analyzes the data logic between the feature file data by combining the data position and the environmental information, and can understand the logical relationship between the feature file data to facilitate the subsequent sequence adjustment processing of the matching file data, wherein the data logic is a description of the logical relationship between the feature file data, and further, the mining of the environmental information corresponding to the feature file data can be achieved through a decision tree algorithm.
[0095] As an embodiment of the present invention, the combination of the data position and the environmental information to analyze the data logic between the feature file data includes: determining the data order corresponding to the feature file data according to the data position, calculating the information entropy value corresponding to the environmental information, screening out key environmental information from the environmental information according to the information entropy value, constructing a visualization chart corresponding to the key environmental information, analyzing the information dependency between the key environmental information according to the visualization chart and the data order, and determining the data logic between the feature file data according to the information dependency.
[0096] The data order indicates the order of the characteristic file data, the information entropy value indicates the amount of information contained in the environmental information, and the information dependency indicates the degree of dependency between the key environmental information.
[0097] Furthermore, the information entropy value corresponding to the environmental information can be calculated by the Shannon entropy formula; the information entropy value can be compared with a preset threshold, and when the information entropy value is greater than the preset threshold, the key environmental information is screened out from the environmental information, and the preset threshold can be 0.8, or can be set according to the actual application scenario; the visualization chart corresponding to the key environmental information can be realized by the Visio drawing tool; the causal relationship between the key environmental information can be determined based on the visualization chart, and the information dependency between the key environmental information can be analyzed by combining the data order and the causal relationship.
[0098] S4. Calculate the semantic association between the target file data, determine corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and adjust the sequence of the matching file data according to the data logic to obtain target matching file data.
[0099] By calculating the semantic association between the target file data, the present invention can understand the semantic association relationship between the target file data, and provide a basis for subsequently improving the matching accuracy of the target file data.
[0100] As an embodiment of the present invention, the calculation of the semantic association between the target file data includes: identifying the data content corresponding to the target file data, performing feature extraction on the data content to obtain content features, calculating the Gini coefficient corresponding to the content features, filtering the data content according to the Gini coefficient to obtain target data content, performing semantic analysis on the target data content to obtain content semantics, calculating the association between the content semantics to obtain content association, and combining the content association to obtain the semantic association between the target file data.
[0101] Among them, the content feature is the representation corresponding to the data content, the Gini coefficient represents the importance corresponding to the content feature, the content semantics is the meaning expressed by the target data content, and the content relevance represents the association relationship between the content semantics.
[0102] Optionally, the recognition of the data content corresponding to the target file data can be achieved through OCR recognition technology; the feature extraction of the data content can be achieved through a bag-of-words model; the data content can be filtered according to the numerical value of the Gini coefficient; the semantic analysis of the target data content can be achieved through a semantic analysis method; the content associations corresponding to the target file data are summed to obtain the semantic associations between the target file data.
[0103] Optionally, as an optional embodiment of the present invention, the calculating the Gini coefficient corresponding to the content feature includes:
[0104] The Gini coefficient corresponding to the content feature is calculated by the following formula:
[0105] ;
[0106] Among them, D represents the Gini coefficient corresponding to the content feature, It represents the proportion of the i-th feature in the content features, i represents the serial number corresponding to the content feature, and t represents the number of content features.
[0107] Optionally, as an optional embodiment of the present invention, the calculating the correlation between the content semantics to obtain the content correlation includes: performing vector processing on the content semantics to obtain a semantic vector, calculating the semantic covariance between the content semantics according to the semantic vector, and calculating the semantic standard deviation of the content semantics, combining the semantic covariance and the semantic standard deviation, and calculating the correlation between the content semantics by the following formula:
[0108] ;
[0109] Among them, F represents the correlation between content semantics, represents the semantic covariance between the bth and b+1th semantics in the content semantics, and They respectively represent the semantic standard deviations corresponding to the b-th and b+1-th semantics in the content semantics, and b and b+1 respectively represent the serial numbers corresponding to the content semantics.
[0110] The present invention determines corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association degree, so as to improve the accuracy of data matching. Optionally, the architecture matching degree, the similarity coefficient and the semantic association degree are normalized, the normalized results are summed, and the corresponding matching file data is determined from the target file data according to the summed value; the sequence adjustment of the matching file data can be implemented according to the data logic to obtain the target matching file data.
[0111] The present invention can remove invalid data in the engineering file data by standardizing the engineering file data, and convert image data into text data, so as to facilitate the subsequent recognition processing of the subsequent data architecture. The present invention can remove unimportant keywords in the file keywords by screening the file keywords, reduce the redundant information in the file keywords, calculate the similarity coefficient between the preferred keywords, analyze the similarity between the preferred keywords, and provide a basis for the subsequent determination of matching file data. The present invention extracts the feature file data in the target file data, and can obtain representative data in the feature file data, which is convenient for the subsequent mining of environmental information and provides a basis for the analysis of data logic. The present invention can understand the semantic association between the target file data by calculating the semantic association between the target file data, and provides a basis for the subsequent improvement of the matching accuracy of the target file data. Therefore, the adaptive matching method for engineering file data provided by the embodiment of the present invention can improve the accuracy of adaptive matching of engineering file data.
[0112] like Figure 2 , which is a functional module diagram of an engineering file data adaptive matching system provided by an embodiment of the present invention.
[0113] The engineering file data adaptive matching system 100 of the present invention can be installed in an electronic device. According to the functions to be implemented, the engineering file data adaptive matching system 100 can include a matching degree calculation module 101, a similarity coefficient calculation module 102, a data logic analysis module 103 and a data matching module 104. The module of the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0114] In this embodiment, the functions of each module / unit are as follows:
[0115] The matching degree calculation module 101 is used to obtain engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify the data architecture corresponding to the target file data, and calculate the architecture matching degree between the data architectures;
[0116] The similarity coefficient calculation module 102 is used to extract the file keywords in the target file data, screen out the file keywords to obtain preferred keywords, and calculate the similarity coefficients between the preferred keywords;
[0117] The data logic analysis module 103 is used to extract the characteristic file data in the target file data, mine the environment information corresponding to the characteristic file data, locate the data position corresponding to the characteristic file data, and analyze the data logic between the characteristic file data in combination with the data position and the environment information;
[0118] The data matching module 104 is used to calculate the semantic association between the target file data, determine the corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and adjust the sequence of the matching file data according to the data logic to obtain the target matching file data.
[0119] In detail, each module described in the engineering file data adaptive matching system 100 described in the embodiment of the present application adopts the same Figure 1 The same technical means as the engineering file data adaptive matching method described in the text and can produce the same technical effects are not described here.
[0120] In the several embodiments provided by the present invention, it should be understood that the provided methods and systems can be implemented in other ways. For example, the method embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for adaptively matching engineering file data, characterized in that: The method comprises: Acquire engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify the data architecture corresponding to the target file data, and calculate the architecture matching degree between the data architectures; Extracting file keywords from the target file data, screening out the file keywords to obtain preferred keywords, and calculating similarity coefficients between the preferred keywords; Extracting characteristic file data from the target file data, mining environmental information corresponding to the characteristic file data, locating data positions corresponding to the characteristic file data, and analyzing data logic between the characteristic file data in combination with the data positions and the environmental information, wherein the combining the data positions and the environmental information to analyze the data logic between the characteristic file data includes: Determining the data order corresponding to the feature file data according to the data position; Calculating an information entropy value corresponding to the environmental information, and screening out key environmental information from the environmental information according to the information entropy value; Constructing a visualization chart corresponding to the key environmental information; Analyzing information dependencies between the key environmental information according to the visualization chart and the data sequence; Determining the data logic between the feature file data according to the information dependency; Calculate the semantic association between the target file data, determine corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and adjust the matching file data in sequence according to the data logic to obtain target matching file data.
2. The method for adaptively matching engineering file data according to claim 1, characterized in that: The step of performing standardization processing on the engineering file data to obtain target file data includes: Performing data cleaning on the engineering file data to obtain cleaned file data; Identifying the document image data in the cleaned document data, and analyzing the image connotation corresponding to the document image data; According to the image connotation, the document image data is converted into image text data; According to the image text data, the cleaning file data is updated to obtain target file data.
3. The method for adaptively matching engineering file data according to claim 2, characterized in that: The analyzing the image connotation corresponding to the file image data includes: Performing noise reduction processing on the image in the document image data to obtain a noise-reduced document image; Performing edge detection on the denoised document image to obtain the document image edge; According to the edge of the document image, segmenting the denoised document image to obtain a document image area; Identify the region body in the document image region, and analyze the image connotation corresponding to the document image data based on the region body.
4. The method for adaptively matching engineering file data according to claim 1, characterized in that: The calculating the architecture matching degree between the data architectures includes: Analyze the architecture attributes corresponding to the data architecture, and calculate the attribute weights corresponding to the architecture attributes; Classifying the data architecture according to the architecture attributes to obtain a classified data architecture; Calculating the architecture overlap between the classified data architectures; The architecture matching degree between the data architectures is calculated by combining the architecture overlap degree and the attribute weight.
5. The method for adaptively matching engineering file data according to claim 1, characterized in that: The calculating the similarity coefficient between the preferred keywords includes: Performing vectorization processing on the preferred keyword to obtain a keyword vector; The vector similarity between the keyword vectors is calculated by the following formula: ; Among them, Q represents the vector similarity between keyword vectors, and They represent the ath vector and a+1th vector in the keyword vector respectively. a and a+1 are the sequence numbers of the keyword vectors respectively. and They represent the norms of the ath vector and the a+1th vector respectively, The vector similarities are normalized to obtain similarity coefficients between the preferred keywords.
6. The method for adaptively matching engineering file data according to claim 1, characterized in that: The extracting the characteristic file data from the target file data comprises: Querying the quality evaluation index corresponding to the target file data; Extracting the index data corresponding to the target file data according to the quality evaluation index; Quantitatively process the indicator data to obtain a quantitative value of the indicator; Calculate the quality score value corresponding to the quality evaluation indicator according to the indicator quantification value; According to the quality score value, feature file data in the target file data is extracted.
7. The method for adaptively matching engineering file data according to claim 1, characterized in that: The calculating the semantic association between the target file data includes: Identify data content corresponding to the target file data; Extracting features from the data content to obtain content features; The Gini coefficient corresponding to the content feature is calculated by the following formula: ; Among them, D represents the Gini coefficient corresponding to the content feature, represents the proportion of the i-th feature in the content features, i represents the sequence number corresponding to the content feature, and t represents the number of content features; According to the Gini coefficient, the data content is filtered to obtain target data content; Performing semantic analysis on the target data content to obtain content semantics; Calculating the relevance between the content semantics to obtain content relevance; Combined with the content relevance, the semantic relevance between the target file data is obtained.
8. The method for adaptively matching engineering file data according to claim 7, characterized in that: The calculating the relevance between the content semantics to obtain the content relevance includes: Performing vector processing on the content semantics to obtain a semantic vector; Calculating the semantic covariance between the content semantics and the semantic standard deviation of the content semantics according to the semantic vector; Combining the semantic covariance and the semantic standard deviation, the correlation between the content semantics is calculated using the following formula: ; Among them, F represents the correlation between content semantics, represents the semantic covariance between the bth and b+1th semantics in the content semantics, and They respectively represent the semantic standard deviations corresponding to the b-th and b+1-th semantics in the content semantics, and b and b+1 respectively represent the serial numbers corresponding to the content semantics.
9. An engineering file data adaptive matching system, characterized in that: The system comprises: A matching degree calculation module is used to obtain engineering file data to be matched, perform standardization processing on the engineering file data to obtain target file data, identify the data architecture corresponding to the target file data, and calculate the architecture matching degree between the data architectures; A similarity coefficient calculation module is used to extract file keywords from the target file data, screen out the file keywords, obtain preferred keywords, and calculate similarity coefficients between the preferred keywords; A data logic analysis module is used to extract the characteristic file data in the target file data, mine the environment information corresponding to the characteristic file data, locate the data position corresponding to the characteristic file data, and analyze the data logic between the characteristic file data in combination with the data position and the environment information, wherein the combining the data position and the environment information to analyze the data logic between the characteristic file data includes: Determining the data order corresponding to the feature file data according to the data position; Calculating an information entropy value corresponding to the environmental information, and screening out key environmental information from the environmental information according to the information entropy value; Constructing a visualization chart corresponding to the key environmental information; Analyzing information dependencies between the key environmental information according to the visualization chart and the data sequence; Determining the data logic between the feature file data according to the information dependency; A data matching module is used to calculate the semantic association between the target file data, determine the corresponding matching file data from the target file data according to the architecture matching degree, the similarity coefficient and the semantic association, and adjust the sequence of the matching file data according to the data logic to obtain the target matching file data.
Citation Information
Patent Citations
Computer file similarity identification system and method based on image analysis
CN111666928A
Similar text matching method and device, electronic equipment and computer storage medium
CN112541338A