Project matching method and system based on automatic data identification
Through automatic identification of bill data and multi-dimensional data analysis, the problem of insufficient accuracy of bill identification and matching in the existing system is solved, efficient project matching and recommendation is achieved, and the accuracy and efficiency of the system are improved.
Patent Information
- Application Number
- CN202410988597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-07-23
AI Technical Summary
When facing the diverse bill processing and project matching systems, the identification technology is insufficient, and it is difficult to conduct in-depth correlation analysis, resulting in low correlation and accuracy of matching results.
The bill data is preprocessed, type recognition, layout analysis, character recognition and information extraction through automatic recognition methods. Combined with multi-dimensional data analysis, multiple candidate project information are obtained and feature extraction and vectorization are performed. Collaborative filtering, association rule mining and deep learning are used for intelligent matching to generate a recommended project list.
It improves the accuracy of ticket type recognition and region division, enhances the accuracy of character recognition and the depth of information extraction, improves the accuracy and efficiency of project matching, provides rich data support, and enhances the interpretability and credibility of recommended results.
Smart Images

Figure CN118761591B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data identification and matching, and in particular to a project matching method and system based on automatic data identification. Background Art
[0002] With the rapid development of information technology, businesses and organizations are facing the challenge of processing and managing massive amounts of bill data. Traditional manual processing methods are not only time-consuming and labor-intensive, but also prone to errors and omissions, making them unable to meet the efficiency and accuracy requirements of the modern business environment. Furthermore, with a vast amount of project information within companies, how to quickly identify the most suitable project from among numerous candidate projects becomes a key issue in improving resource utilization efficiency and decision-making quality.
[0003] However, current bill processing and project matching systems often suffer from several issues. Existing recognition technologies lack accuracy for diverse bill types and formats, making them incapable of handling complex bill layouts and character recognition tasks. Furthermore, insufficient analysis of the correlation between bill data and project information leads to low relevance and accuracy in matching results. Traditional matching algorithms lack the ability to comprehensively analyze multi-dimensional data and are unable to fully utilize the rich information contained within bills. Summary of the Invention
[0004] The present application provides a project matching method and system based on automatic data identification, which is used to automatically identify bill data, thereby improving the accuracy of project matching.
[0005] In a first aspect, the present application provides a project matching method based on automatic data identification, the project matching method based on automatic data identification comprising:
[0006] Acquiring original bill data and preprocessing the original bill data to obtain target bill data;
[0007] Performing type recognition and layout analysis on the target bill data to obtain structured area data;
[0008] Performing character recognition and information extraction on the structured area data to obtain bill structured data;
[0009] Summarizing and analyzing the bill structured data to obtain multi-dimensional data analysis results;
[0010] Acquire information of multiple candidate projects, perform feature extraction and vectorization on the information of multiple candidate projects, and obtain a project feature library;
[0011] Intelligently match the multi-dimensional data analysis results with the project feature library to obtain a recommended project list.
[0012] In a second aspect, the present application provides a project matching system based on automatic data identification, the project matching system based on automatic data identification comprising:
[0013] An acquisition module, configured to acquire original bill data and pre-process the original bill data to obtain target bill data;
[0014] An analysis module, configured to perform type recognition and layout analysis on the target bill data to obtain structured area data;
[0015] A recognition module, configured to perform character recognition and information extraction on the structured area data to obtain bill structured data;
[0016] A summary module, used to summarize and analyze the bill structured data to obtain multi-dimensional data analysis results;
[0017] The extraction module is used to obtain information of multiple candidate projects, perform feature extraction and vectorization on the information of multiple candidate projects, and obtain a project feature library;
[0018] The matching module is used to intelligently match the multi-dimensional data analysis results with the project feature library to obtain a list of recommended projects.
[0019] The third aspect of the present application provides a project matching device based on automatic data recognition, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the project matching device based on automatic data recognition to execute the above-mentioned project matching method based on automatic data recognition.
[0020] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned project matching method based on automatic data identification.
[0021] The technical solution provided in this application improves the quality of original bill data through image processing. Multi-scale feature extraction and cluster analysis, combined with bill template matching, achieve accurate bill type identification and regional segmentation, effectively improving the extraction accuracy of structured data. Combining feature extraction, classification recognition, and contextual analysis significantly improves character recognition accuracy, while semantic analysis and entity recognition enhance the depth of information extraction. By comprehensively exploring the potential value of bill data, rich data support is provided for project matching. Comprehensive feature extraction of candidate project information, including word frequency statistics, word vector training, and topic modeling, forms a multi-faceted project feature representation, improving matching accuracy. A combination of collaborative filtering, association rule mining, and deep learning matching methods effectively improves the coverage and accuracy of project recommendations. Through frequent item set generation, multiple filtering, and rule optimization, high-quality association rules are extracted, enhancing the interpretability and credibility of recommendation results. A high-performance deep learning recommendation model is constructed, significantly improving matching accuracy and efficiency. Through weighted fusion and multi-objective optimization, the advantages of multiple recommendation methods are effectively integrated, balancing multiple recommendation objectives and improving the accuracy of project matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a schematic diagram of an embodiment of a project matching method based on automatic data identification in an embodiment of the present application;
[0024] Figure 2 This is a schematic diagram of an embodiment of a project matching system based on automatic data identification in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The embodiments of the present application provide a method and system for project matching based on automatic data identification. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.
[0026] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiment of the present application, an embodiment of the project matching method based on automatic data identification includes:
[0027] Step S101: obtaining original bill data and preprocessing the original bill data to obtain target bill data;
[0028] It is understandable that the execution subject of this application can be a project matching system based on automatic data recognition, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0029] Specifically, raw bill data is acquired through scanning or other image capture devices, and image conversion and image enhancement are performed on the raw bill data to obtain an enhanced bill image. The enhanced bill image is then subjected to denoising to remove any noise and impurities, resulting in a de-noised bill image. The de-noised bill image is then binarized to convert it to black and white for ease of subsequent image processing and analysis. The binarization process results in a binarized bill image. The binarized bill image is then subjected to tilt correction to correct any tilt that may have occurred during scanning or capturing, resulting in a corrected bill image. Edge detection is then performed on the corrected bill image to extract edge contour data, resulting in bill edge contour data. The bill edge contour data is then cropped to remove any excess background from the image, resulting in a cropped bill image. The cropped bill image is then subjected to resolution standardization to ensure uniform resolution for all bill images, resulting in a standardized bill image. The standardized bill image is then subjected to color space conversion, converting the image from one color space to another, resulting in a converted bill image. Color space conversion can improve the visual quality of an image and enhance the visibility of key information within it. Lighting is then balanced on the converted bill image to create a uniform lighting distribution. This balanced bill image is then converted to its data format to obtain the final target bill data.
[0030] Step S102: performing type recognition and layout analysis on the target bill data to obtain structured region data;
[0031] Specifically, multi-scale feature extraction is performed on the target bill data to obtain multi-scale feature maps. These feature maps capture key information about bill images at different scales. Cluster analysis is then performed on the multi-scale feature maps. Using a clustering algorithm, similar features are classified and aggregated to produce bill type feature clustering results. Based on the clustering results, a bill type recognition model is constructed, which can learn and recognize the characteristic patterns of different bill types. The bill type recognition model is then used to predict the type of the target bill data, obtaining a bill type recognition result and enabling preliminary classification of the bill type. Bill template matching is then performed on the bill type recognition result. By comparing the recognition result with a pre-stored bill template, matching bill template data is obtained. Based on the matching bill template data, the target bill data is segmented into regions, obtaining preliminary segmented bill region data. Edge detection and contour analysis are performed on the preliminary segmented bill region data to identify the edges and contours of the bill region and obtain bill region contour data. Based on the contour data, region optimization is performed on the preliminary segmented bill region data to adjust and refine the bill region, obtaining optimized bill region data. The optimized bill region data is analyzed for layout structure, analyzing the layout structure of these regions within the bill layout to obtain the bill layout structure data. Semantic annotation is then performed on the optimized bill region data based on the bill layout structure data. This semantic annotation identifies the information in different regions as specific content, such as amount, date, invoice number, etc., to obtain structured region data.
[0032] Step S103: performing character recognition and information extraction on the structured area data to obtain bill structured data;
[0033] Specifically, character segmentation is performed on the structured region data. Using character segmentation techniques, the text region in the bill is divided into a set of individual character images. Feature extraction is performed on each character image to generate a set of character feature vectors. These feature vectors represent the morphological and structural information of the characters. Classification and recognition are then performed on the set of character feature vectors. A pre-trained character classification model is used to perform preliminary recognition of each character, resulting in a preliminary recognition result. Because the character recognition process may contain errors, contextual analysis is performed on the preliminary recognition results. By considering the semantic and positional relationships of the characters within the context, the preliminary recognition results are optimized to obtain an optimized character recognition result. Text semantic analysis is then performed on the optimized character recognition results. Natural language processing techniques are used to perform semantic analysis on the recognized text, resulting in semantically annotated data. Based on the semantically annotated data, key entities are identified to obtain entity information within the bill, including important information such as the amount, date, and invoice number. Relationship extraction is performed on the bill entity information to identify associations and relationships between entities and generate an entity-relationship diagram. The entity-relationship diagram illustrates the connections and interactions between entities within the bill. Based on the entity-relationship diagram, a knowledge graph is constructed. By organizing entities and relationships into a knowledge graph, a bill knowledge graph is generated. The knowledge graph can systematically represent the knowledge structure and information relationships in bills, improving the interpretability and application value of the information. A structured transformation is performed on the bill knowledge graph, converting the information in the knowledge graph into a standardized structured data format to obtain bill structured data.
[0034] Step S104: Summarize and analyze the bill structured data to obtain multi-dimensional data analysis results;
[0035] Specifically, the structured invoice data is cleansed to remove redundant, erroneous, and incomplete data, resulting in cleaned invoice data. The cleaned invoice data is then standardized to unify the measurement units and formats, resulting in standardized invoice data. Standardization eliminates data bias and enables data from different sources and formats to be processed and compared within the same analytical framework. Time series analysis is performed on the standardized invoice data to analyze the temporal trends and patterns of data change, resulting in temporal analysis results. Time series analysis reveals the temporal dynamics of data and helps predict future trends and changes. Spatial cluster analysis is also performed on the standardized invoice data to identify the distribution patterns and clustering characteristics of the data in geographic space, resulting in spatial analysis results. Spatial cluster analysis reveals geographic spatial relationships in the data and helps identify regional characteristics and patterns. Multivariate correlation analysis is performed on the standardized invoice data to examine the interrelationships and dependencies between different variables, resulting in variable association analysis results. Multivariate correlation analysis reveals the mutual influence and inherent connections between variables. Based on the results of the variable association analysis, principal component analysis is performed on the standardized invoice data to extract key features using dimensionality reduction techniques, resulting in reduced feature data. Principal component analysis can simplify data structure and reduce data dimensionality while retaining key information and features. Anomaly detection is performed on the reduced feature data, detecting outliers and patterns in the data to generate anomaly data analysis results. Anomaly detection can identify anomalies and potential problems in the data, helping to promptly detect and address anomaly data. Cluster analysis can also be performed on the reduced feature data, grouping data with similar characteristics to generate data clustering results. Cluster analysis can reveal the inherent structure and distribution characteristics of the data, helping to identify distinct groups and patterns within the data. Decision tree analysis is performed on the data clustering results, building a decision tree model to derive rules and decision paths within the data, generating rule derivation results. Decision tree analysis provides clear decision rules and logical relationships, helping to understand the decision-making process and logic within the data. Multidimensional fusion of the temporal, spatial, variable association, anomaly, clustering, and rule derivation results integrates the analysis results from different dimensions to generate multidimensional data analysis results, comprehensively considering the various characteristics and relationships of the data.
[0036] Step S105: Acquire multiple candidate project information, perform feature extraction and vectorization on the multiple candidate project information, and obtain a project feature library;
[0037] Specifically, multiple candidate project information is obtained and processed using data cleaning techniques to remove redundant, erroneous, and incomplete data. This cleansed project data is then processed to ensure data quality and consistency. Text segmentation is performed on the cleaned project data, breaking the text content into individual words or phrases to obtain a set of project keywords. A word frequency count is performed on the project keyword set to calculate the frequency of each keyword in the text and obtain a word frequency feature vector. The word frequency feature vector reflects the importance and frequency of each keyword. Simultaneously, word vector training is performed on the project keyword set. By training a word vector model, each keyword is converted into a fixed-dimensional vector representation to obtain a word vector model. A topic model is then trained on the cleaned project data to mine potential topics and topic distribution within the text, obtaining a project topic distribution. The topic model reveals the topic structure and relationships between topics within the project text, facilitating understanding and analysis of the project's content and characteristics. Named entity recognition is also performed on the cleaned project data to identify named entities within the text, such as names of people, places, and organizations, and obtain entity feature vectors. The word frequency feature vector, word vector model, project topic distribution, and entity feature vector are fused to produce a fused feature vector. Feature fusion comprehensively considers multiple feature information of the text, improving the comprehensiveness and accuracy of feature representation. The fused feature vector is then subjected to dimensionality reduction. Using dimensionality reduction techniques such as principal component analysis, high-dimensional feature vectors are reduced to lower dimensions to produce the reduced-dimensional project feature vector. Dimensionality reduction can reduce feature redundancy and noise, improving feature computational efficiency and model generalization. Vectors are clustered and indexed based on the reduced-dimensional project feature vectors. Similar feature vectors are grouped and organized using a clustering algorithm to produce a project feature library. The project feature library efficiently stores and retrieves project feature information.
[0038] Step S106: Intelligently match the multi-dimensional data analysis results with the project feature library to obtain a recommended project list.
[0039] Specifically, feature extraction is performed on the results of multi-dimensional data analysis to obtain a bill feature vector. This feature vector represents the key information and features of the bill data. Similarity calculation is performed between the bill feature vector and the project feature library. This calculation yields a preliminary matching result. This similarity calculation effectively measures the similarity between bill and project features. Collaborative filtering is then applied to the preliminary matching results. Using collaborative filtering technology, a recommendation list is generated based on the preferences of similar users or projects, resulting in a collaborative filtering recommendation list. Collaborative filtering recommendation methods can identify potential matching projects based on historical data and user behavior, improving the accuracy and personalization of recommendations. Cluster analysis is also performed on the bill feature vectors. Clustering algorithms are used to group similar bill features, resulting in bill clustering results. Association rule mining is then performed on the bill clustering results and the project feature library. By mining the correlations between bill and project features, a recommendation list is generated, resulting in an association rule recommendation list. Furthermore, a deep learning model is trained using the bill feature vectors and the project feature library. Deep learning technology is used to construct a recommendation model, resulting in a deep learning recommendation model. Deep learning models can learn complex feature relationships and patterns through training on large amounts of data, improving the accuracy and generalization of recommendations. A deep learning recommendation model performs predictions and matches the bill feature vectors with the item feature library to generate a recommendation list, resulting in a deep learning recommendation list. A weighted fusion of the collaborative filtering recommendation list, the association rule recommendation list, and the deep learning recommendation list combines the results of different recommendation methods to create a fused recommendation list. Weighted fusion comprehensively considers the strengths and weaknesses of different methods, improving the diversity and accuracy of recommendation results. Multi-objective optimization is performed on the fused recommendation list. An optimization algorithm balances and optimizes different objectives to produce an optimized recommendation list. Multi-objective optimization considers multiple evaluation metrics to generate more relevant recommendation results. Based on the optimized recommendation list, recommended items are sorted and filtered, ranking items according to priority and importance to select the best recommended items and create a recommended item list.
[0040] In the embodiments of this application, the quality of original bill data is improved through image processing. Multi-scale feature extraction and cluster analysis, combined with bill template matching, achieve accurate bill type identification and regional division, effectively improving the extraction accuracy of structured data. The combination of feature extraction, classification recognition, and context analysis significantly improves the accuracy of character recognition. At the same time, semantic analysis and entity recognition enhance the depth of information extraction. By comprehensively exploring the potential value of bill data, rich data support is provided for project matching. Comprehensive feature extraction of candidate project information, including word frequency statistics, word vector training, and topic models, forms a multi-angle project feature representation, improving matching accuracy. The integrated application of multiple matching methods, such as collaborative filtering, association rule mining, and deep learning, effectively improves the coverage and accuracy of project recommendations. Through frequent item set generation, multiple filtering, and rule optimization, high-quality association rules are extracted, enhancing the interpretability and credibility of recommendation results. A high-performance deep learning recommendation model is constructed, significantly improving matching accuracy and efficiency. Through weighted fusion and multi-objective optimization, the advantages of multiple recommendation methods are effectively integrated, balancing multiple recommendation objectives and improving the accuracy of project matching.
[0041] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0042] (1) Obtaining original bill data, and performing image conversion and image enhancement on the original bill data to obtain an enhanced bill image;
[0043] (2) Denoising the enhanced bill image to obtain a de-noised bill image, and binarizing the de-noised bill image to obtain a binarized bill image;
[0044] (3) Perform tilt correction on the binary bill image to obtain a corrected bill image, and perform edge detection on the corrected bill image to obtain bill edge contour data;
[0045] (4) Cropping the edge contour data of the bill to obtain a cropped bill image, and normalizing the resolution of the cropped bill image to obtain a standardized bill image;
[0046] (5) Performing color space conversion on the standardized bill image to obtain the converted bill image;
[0047] (6) Performing illumination equalization on the converted bill image to obtain an illumination-balanced bill image, and performing data format conversion on the illumination-balanced bill image to obtain target bill information data.
[0048] Specifically, the original bill data is obtained through a scanning device or camera. The original bill data is converted and enhanced to improve the clarity and contrast of the image, making the content on the bill easier to identify. Image conversion includes operations such as adjusting the brightness, contrast and sharpening of the image, while image enhancement uses algorithms to improve image quality. The enhanced bill image is denoised. Denoising can be performed using a variety of methods, such as Gaussian filtering and median filtering. Gaussian filtering is a common denoising method that smoothes the image using a Gaussian function. The formula for Gaussian filtering is:
[0049] ;
[0050] in, represents the pixel value after Gaussian filtering, represents the standard deviation of the Gaussian function, and Represents the coordinates of the pixel. By adjusting The value of controls the degree of denoising. After denoising, the denoised bill image is obtained. The denoised bill image is binarized to convert the image into a black and white format. The purpose of binarization is to highlight the text and key areas on the bill to facilitate subsequent processing and recognition. The commonly used binarization method is the Otsu method, which determines the optimal threshold by maximizing the inter-class variance. The formula of the Otsu method is:
[0051] ;
[0052] in, represents the between-class variance, and denote the probabilities of foreground and background pixels respectively, and denote the average values of foreground and background pixels, respectively. Represents the overall average. Perform skew correction on the binary bill image. This is done by detecting text lines in the image and rotating the image to align it with the horizontal line. A commonly used skew correction method is the Hough transform, which calculates the skew angle by detecting straight lines in the image. Perform edge detection on the corrected bill image to extract the bill's edge contour data. Edge detection can use the Canny edge detection algorithm, which detects edges by calculating the gradient of the image. The formula for Canny edge detection is:
[0053] ;
[0054] Among them, Gradient Represents the gradient strength of the pixel point, and Representing the gradients along the x-axis and y-axis, respectively. After obtaining edge contour data, the image is cropped to remove excess background, resulting in a cropped bill image. The cropped bill image then undergoes resolution normalization, adjusting the image to the specified resolution through interpolation to ensure uniform resolution across all bill images. The normalized bill image undergoes color space conversion, converting the image from one color space (such as RGB) to another (such as grayscale or HSV) for further processing and analysis. The converted bill image undergoes illumination equalization to achieve a uniform illumination distribution. This can be achieved using adaptive histogram equalization, which adjusts the image's histogram to achieve a more uniform brightness distribution. The illuminated bill image undergoes data format conversion to the format required by the target bill data. The purpose of data format conversion is to ensure compatibility and processability of the image data for subsequent processing and analysis.
[0055] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0056] (1) Perform multi-scale feature extraction on the target bill data to obtain a multi-scale feature map of the bill, and perform cluster analysis on the multi-scale feature map of the bill to obtain the bill type feature clustering result;
[0057] (2) Construct a bill type recognition model based on the bill type feature clustering results, and use the bill type recognition model to predict the type of target bill data to obtain the bill type recognition results;
[0058] (3) Perform bill template matching on the bill type recognition result to obtain matching bill template data, and divide the target bill data into regions based on the matching bill template data to obtain preliminary divided bill region data;
[0059] (4) Performing edge detection and contour analysis on the initially divided bill area data to obtain bill area contour data, and performing regional optimization on the initially divided bill area data based on the bill area contour data to obtain optimized bill area data;
[0060] (5) Performing layout analysis on the optimized bill area data to obtain bill layout structure data, and performing semantic annotation on the optimized bill area data based on the bill layout structure data to obtain structured area data.
[0061] Specifically, multi-scale feature extraction is performed on the target bill data to capture the characteristic information of the bill image at different scales. Multi-scale feature extraction can be achieved through convolutional neural networks. Convolutional neural networks can extract local features of images at different levels and obtain feature maps of different scales through multi-layer convolution operations. The formula is as follows:
[0062] ;
[0063] in, Indicates the Tier feature maps, Indicates the Tier convolution kernels, * represents the convolution operation, Indicates the Tier The bias of the feature map, Denotes the activation function. Through multi-layer convolution operations, a multi-scale feature map of the bill is obtained. Cluster analysis is performed on the multi-scale feature map of the bill. Similar feature maps are grouped using a clustering algorithm to obtain the bill type feature clustering results. Common clustering algorithms include K-means and DBSCAN. The goal of the K-means algorithm is to minimize the sum of squared distances of samples within a cluster. Its objective function is:
[0064] ;
[0065] in, represents the objective function value, represents the number of clusters, Indicates the clusters, Indicates the samples, Indicates the The center of each cluster. Construct a bill type recognition model based on the clustering results of bill type features. Use support vector machines or deep neural networks to build recognition models. Learn the relationship between bill type features and categories by training the model. The trained model can perform type prediction on the target bill data to obtain bill type recognition results. Perform bill template matching on the bill type recognition results, and obtain matching bill template data by comparing with the pre-defined bill template. Based on the matching bill template data, perform regional division on the target bill data to obtain preliminary divided bill area data. Region division can use template matching algorithms, such as Hough transform or template matching method, to identify key areas in the bill and divide it into different functional areas. Perform edge detection and contour analysis on the preliminary divided bill area data to extract the edge and contour information of the bill area to obtain bill area contour data. Edge detection can use the Canny edge detection algorithm, and its formula is:
[0066] ;
[0067] Among them, Gradient Represents the gradient strength of the pixel point, and Represent the gradients in the x-axis and y-axis directions respectively. The gradient intensity of each pixel is calculated by the above formula, and the final edge image is obtained by non-maximum suppression and double threshold detection. Contour analysis is performed based on the edge image to optimize the preliminarily divided bill area data to obtain the optimized bill area data. Layout analysis is performed on the optimized bill area data, and the bill layout structure data is obtained by analyzing the positional relationship of each area in the bill. Layout analysis can use connected domain analysis and morphological processing methods to identify and mark each area in the bill. Semantic annotation is performed on the optimized bill area data based on the bill layout structure data, and the content in different areas is annotated to obtain structured area data. For example, the invoice number, date, amount and other information in the bill are annotated and extracted.
[0068] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0069] (1) Perform character segmentation on the structured region data to obtain a set of single character images, and perform feature extraction on the set of single character images to obtain a set of character feature vectors;
[0070] (2) Classify and recognize the character feature vector set to obtain preliminary recognition results, and perform context analysis on the preliminary recognition results to obtain optimized character recognition results;
[0071] (3) Perform text semantic analysis on the optimized character recognition results to obtain semantic annotation data, and perform entity recognition on the semantic annotation data to obtain bill entity information;
[0072] (4) Extract the relationship between the bill entity information to obtain the entity relationship graph, and construct the knowledge graph of the entity relationship graph to obtain the bill knowledge graph;
[0073] (5) Perform structured transformation on the bill knowledge graph to obtain bill structured data.
[0074] Specifically, character segmentation is performed on structured area data, and the text area in the bill is divided into a set of individual character images through character segmentation technology. Common character segmentation methods include projection method and connected component analysis. The projection method determines the boundaries of the characters by calculating the vertical projection of the image, and connected component analysis segments the characters by marking the connected areas in the image. Feature extraction is performed on a single character image set. Each character image is converted into a feature vector for subsequent classification and recognition. Common feature extraction methods include pixel value method, shape feature method and gradient feature method. The pixel value method directly uses the pixel value of the character image as a feature, the shape feature method extracts the geometric shape features of the character (such as edges, contours, etc.) as features, and the gradient feature method calculates the gradient direction and amplitude of the character image as features. For example, the directional gradient histogram of the character image is calculated as a feature vector. The calculation formula of the HOG feature is:
[0075] ;
[0076] Among them, HOG Indicates the location HOG feature values, and Respectively indicated at position horizontal and vertical gradients, Represents the gradient direction. The HOG features of each character image are calculated using the above formula to obtain a set of character feature vectors. Classification and recognition are then performed on the set of character feature vectors. Each character is classified using a pretrained character classification model to obtain a preliminary recognition result. A convolutional neural network is used for character classification and recognition. Multiple layers of convolution and pooling operations are used to extract high-level features from the character image, and classification is performed using a fully connected layer. Contextual analysis is performed on the preliminary recognition results. Character recognition results are optimized by incorporating contextual information about the characters in the text. For example, a recurrent neural network is used for contextual analysis to improve character recognition accuracy by capturing contextual dependencies within character sequences. The optimized character recognition results are then subjected to text semantic analysis. This extracts meaningful information from the text and annotates it. For example, named entity recognition methods are used for text semantic analysis to identify named entities (such as names of people, places, and organizations) in the text to obtain semantically annotated data. Entity recognition is then performed on the semantically annotated data. Specific entity information, such as amount, date, and invoice number, is extracted from the text. For example, a deep learning model is used for entity recognition. By training the model to learn the characteristics and patterns of entities in the text, entity information from invoices can be extracted. Relationship extraction is performed on bill entity information to identify relationships between entities and construct an entity-relationship graph. For example, a supervised learning-based approach can be used to train a model to learn the relationship patterns between entities and obtain an entity-relationship graph. A knowledge graph is constructed based on the entity-relationship graph. By organizing entities and relationships into a graph structure, complex associations between entities can be displayed. The knowledge graph can systematically represent the knowledge structure in bills, improving the interpretability and application value of the information. The knowledge graph is structured and converted. The information in the knowledge graph is converted into a standardized structured data format to facilitate subsequent data processing and analysis. For example, a template-based approach can be used to convert the information in the knowledge graph into structured data using a predefined conversion template to obtain structured bill data.
[0077] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0078] (1) Cleaning the structured bill data to obtain cleaned bill data, and standardizing the cleaned bill data to obtain standardized bill data;
[0079] (2) Perform time series analysis on the standardized bill data to obtain the results of time dimension analysis; perform spatial clustering on the standardized bill data to obtain the results of spatial dimension analysis; perform multivariate correlation analysis on the standardized bill data to obtain the results of variable association analysis;
[0080] (3) Based on the results of variable association analysis, principal component analysis is performed on the standardized bill data to obtain the feature data after dimensionality reduction;
[0081] (4) Perform anomaly detection on the feature data after dimensionality reduction to obtain the abnormal data analysis results, and perform cluster analysis on the feature data after dimensionality reduction to obtain the data clustering results; perform decision tree analysis on the data clustering results to obtain the rule deduction results;
[0082] (5) The time dimension analysis results, space dimension analysis results, variable association analysis results, abnormal data analysis results, data clustering results and rule derivation results are multi-dimensionally integrated to obtain multi-dimensional data analysis results.
[0083] Specifically, the structured data of bills is cleaned to remove redundancy, errors and missing values in the data, and to improve data quality and consistency. Data cleaning includes operations such as processing missing values, removing duplicate records and correcting outliers. For example, if a field has missing values, it can be filled by interpolation or using the mean or median of the field. The cleaned bill data is standardized to unify the measurement unit and range of the data to obtain standardized bill data, eliminate scale differences in the data, and enable different variables to be compared and processed under the same analysis framework. Time series analysis is performed on the standardized bill data to obtain the results of the time dimension analysis. By analyzing the changing trends and patterns of data in the time dimension, the time dynamic characteristics of the data are revealed. Common time series analysis methods include moving average, exponential smoothing and autoregressive models. For example, the autoregressive model (AR model) can be expressed as:
[0084] ;
[0085] in, Indicates time The value of represents a constant term, represents the autoregressive coefficient, Represents the error term. Spatial clustering is performed on the standardized bill data to obtain spatial dimension analysis results. The purpose of spatial clustering is to reveal the geospatial relationships of data by identifying the distribution patterns and clustering characteristics of data in geographic space. For example, using the K-means clustering algorithm to minimize the sum of squared distances of samples within a cluster, its objective function is:
[0086] ;
[0087] in, represents the objective function value, represents the number of clusters, Indicates the clusters, Indicates the samples, Indicates the The center of each cluster. Perform multivariate correlation analysis on the standardized bill data to obtain the variable association analysis results. Multivariate correlation analysis studies the relationship and dependence between different variables. Commonly used methods include Pearson correlation coefficient and Spearman correlation coefficient. The formula for Pearson correlation coefficient is:
[0088] ;
[0089] in, represents the correlation coefficient, and Represents the first Sample values, and Representing the mean of the two variables, respectively. Based on the results of the variable association analysis, principal component analysis is performed on the standardized bill data to obtain the feature data after dimensionality reduction. The original data is converted to a new feature space through linear transformation, retaining the main feature information while reducing the data dimension. Principal component analysis finds the eigenvectors and eigenvalues of the data covariance matrix and determines the principal components based on the size of the eigenvalues. The formula for principal component analysis is:
[0090] ;
[0091] in, Represents the feature data after dimensionality reduction, represents the original data matrix, Represents the eigenvector matrix. Anomaly detection is performed on the reduced feature data to identify outliers and unusual patterns in the data, generating anomaly data analysis results. Anomaly detection can utilize statistical methods and machine learning algorithms, such as distance-based methods and the isolation forest algorithm. The isolation forest algorithm constructs multiple random trees to detect anomalies, based on the principle that outliers have short paths within the tree. Cluster analysis is performed on the reduced feature data to generate data clustering results. By grouping similar data, the inherent structure and distribution characteristics of the data are revealed. Decision tree analysis is performed on the data clustering results. By constructing a decision tree model, rules and decision paths within the data are derived, resulting in rule derivation results. Decision trees recursively partition the data into subsets by selecting optimal split points until a stopping condition is met. The results of temporal analysis, spatial analysis, variable association analysis, anomaly analysis, data clustering, and rule derivation are integrated to generate multidimensional data analysis results that comprehensively consider the various characteristics and relationships of the data.
[0092] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0093] (1) Obtaining information on multiple candidate projects and performing data cleaning on the information to obtain cleaned project data;
[0094] (2) Perform text segmentation on the cleaned project data to obtain a set of project keywords, and perform word frequency statistics on the project keyword set to obtain a word frequency feature vector;
[0095] (3) Perform word vector training on the project keyword set to obtain a word vector model, and perform topic model training on the cleaned project data to obtain the project topic distribution;
[0096] (4) Perform named entity recognition on the cleaned project data to obtain entity feature vectors;
[0097] (5) Fusing the word frequency feature vector, word vector model, project topic distribution, and entity feature vector to obtain a fused feature vector;
[0098] (6) Reduce the dimension of the fused feature vector to obtain the reduced-dimensional project feature vector, and perform vector clustering and indexing based on the reduced-dimensional project feature vector to obtain the project feature library.
[0099] Specifically, data is obtained from the text information of multiple candidate projects, including different aspects of project information, such as project name, description, goals, etc. Data cleaning is performed to remove redundant, erroneous and incomplete data. The cleaning process may include processing missing values, removing duplicate records and correcting erroneous values. Text segmentation is performed on the cleaned project data to obtain a set of project keywords. The text in the project description is divided into individual words or phrases. Word segmentation can be performed using word segmentation tools or libraries such as NLTK, SpaCy, etc. After word segmentation, word frequency statistics are performed on the project keyword set, and the frequency of each keyword in the text is calculated to obtain a word frequency feature vector. The word frequency feature vector can reflect the importance and frequency of occurrence of the keyword. Word vector training is performed on the project keyword set to obtain a word vector model. The purpose of the word vector model is to convert each keyword into a vector representation of a fixed dimension, thereby capturing the semantic relationship and contextual information between words. Commonly used word vector training methods include Word2Vec, GloVe, etc. The basic formula of the word vector model is:
[0100] cosine similarity ;
[0101] in, and Represents two word vectors respectively, represents the vector dot product, Represents the modulus of the vector. The word vector model is trained by maximizing the cosine similarity of similar words. At the same time, the topic model is trained on the cleaned project data to obtain the project topic distribution. By analyzing the co-occurrence relationship of words in the text, the potential topics in the text are identified. Common topic model methods include latent Dirichlet distribution. The latent Dirichlet distribution model performs topic analysis of the text by assuming that each document is a mixture of several topics, and each topic is a mixture of several words. The basic formula of the latent Dirichlet distribution model is:
[0102] ;
[0103] in, Represents a document The word set in Represents a document The theme collection in represents the topic distribution of words, Represents a document The number of words in the text. Named entity recognition is performed on the cleaned project data to obtain entity feature vectors. Named entities in the text, such as names of people, places, and organizations, are identified and converted into feature vectors. Named entity recognition can be performed using a pretrained model or by training a custom model. Feature fusion is performed on the word frequency feature vector, word vector model, project topic distribution, and entity feature vector to obtain a fused feature vector. This comprehensively considers multiple feature information from the text to improve the comprehensiveness and accuracy of feature representation. Dimensionality reduction is performed on the fused feature vector to obtain a reduced project feature vector, reducing feature redundancy and noise, improving feature computational efficiency and model generalization. Common dimensionality reduction methods include principal component analysis (PCA) and t-SNE. PCA finds the eigenvectors and eigenvalues of the data covariance matrix and determines the principal components based on the magnitude of the eigenvalues. Vectors are clustered and indexed based on the reduced project feature vectors to obtain a project feature library. The goal of vector clustering is to reveal the inherent structure and distribution characteristics of projects by grouping similar project feature vectors. Common clustering algorithms include K-means and DBSCAN. The purpose of indexing is to establish an efficient data structure to facilitate subsequent rapid retrieval and matching.
[0104] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0105] (1) Extract features from the multi-dimensional data analysis results to obtain the bill feature vector, and calculate the similarity between the bill feature vector and the project feature library to obtain preliminary matching results;
[0106] (2) Perform collaborative filtering on the preliminary matching results to obtain a collaborative filtering recommendation list, and perform cluster analysis on the bill feature vectors to obtain the bill clustering results;
[0107] (3) Perform association rule mining on the bill clustering results and the project feature library to obtain a list of recommended association rules;
[0108] (4) Deep learning model training is performed through the bill feature vector and project feature library to obtain a deep learning recommendation model;
[0109] (5) Predictions are made through the deep learning recommendation model to obtain a deep learning recommendation list;
[0110] (6) Perform weighted fusion on the collaborative filtering recommendation list, association rule recommendation list, and deep learning recommendation list to obtain a fused recommendation list;
[0111] (7) Perform multi-objective optimization on the fused recommendation list to obtain an optimized recommendation list, and sort and filter the items in the optimized recommendation list to obtain a recommended item list.
[0112] Specifically, feature information related to bills is extracted from the multi-dimensional data analysis results and converted into bill feature vectors. Common methods for converting complex multi-dimensional data into vector representations that can be used for similarity calculations include principal component analysis, linear discriminant analysis, and feature selection. Similarity is calculated between the bill feature vector and the project feature library to obtain a preliminary matching result. Common similarity calculation methods include cosine similarity, Euclidean distance, and Manhattan distance. Collaborative filtering is performed on the preliminary matching results to obtain a collaborative filtering recommendation list. Items of potential interest are recommended by analyzing the user's historical behavior and preferences. Collaborative filtering is divided into user-based collaborative filtering and item-based collaborative filtering. For example, item-based collaborative filtering recommends items similar to the preliminary matching results by calculating the similarity between items. Cluster analysis is performed on the bill feature vectors to obtain bill clustering results. By grouping similar bills, the inherent structure and distribution characteristics of the bills are revealed. Common clustering algorithms include K-means and DBSCAN. The goal of the K-means clustering algorithm is to minimize the sum of squared distances of samples within a cluster. Its objective function is:
[0113] ;
[0114] in, represents the objective function value, represents the number of clusters, Indicates the clusters, Indicates the samples, Indicates the The center of each cluster. Perform association rule mining on the bill clustering results and the project feature library to obtain a list of recommended association rules. The purpose of association rule mining is to discover statistically significant associations in the data. Commonly used methods include the Apriori algorithm and the FP-Growth algorithm. The basic formula for association rule mining is:
[0115] confidence ;
[0116] Among them, confidence Representation Rules The confidence level, Representing item sets support Representing item sets The support level of a recommendation is calculated using a deep learning model trained on bill feature vectors and a project feature library to generate a deep learning recommendation model. Through training on a large amount of data, the deep learning recommendation model learns complex feature relationships and patterns, improving recommendation accuracy. The trained model can be fed with bill feature vectors to predict project recommendations, generating a deep learning recommendation list. A weighted fusion of the collaborative filtering recommendation list, the association rule recommendation list, and the deep learning recommendation list is performed to generate a fused recommendation list. By comprehensively considering the advantages and disadvantages of different recommendation methods, different weights are assigned to the multiple recommendation results and a weighted average is calculated. Multi-objective optimization is performed on the fused recommendation list to generate an optimized recommendation list. Items in this optimized recommendation list are then sorted and filtered to obtain a list of recommended items. The goal of multi-objective optimization is to generate recommendation results that meet various requirements based on multiple evaluation metrics. For example, using the particle swarm optimization algorithm, the optimized recommendation list is sorted and filtered to ultimately obtain the optimal recommended items.
[0117] In a specific embodiment, the execution step performs association rule mining on the bill clustering results and the project feature library to obtain the association rule recommendation list may specifically include the following steps:
[0118] (1) Generate frequent item sets for the bill clustering results to obtain bill frequent item sets, and generate frequent item sets for the project features based on the project feature library to obtain project frequent item sets;
[0119] (2) Generate candidate association rules for bill frequent item sets and project frequent item sets to obtain candidate association rule sets;
[0120] (3) Calculate the rule support based on the candidate association rule set to obtain a rule set after support filtering, and calculate the confidence of the rule set after support filtering to obtain a rule set after confidence filtering;
[0121] (4) Calculate the rule lift based on the confidence filtered rule set to obtain the lift filtered rule set, and optimize the rule set to obtain the optimized association rule set;
[0122] (5) Perform rule interpretability analysis based on the optimized association rule set to obtain an association rule set with explanations, and sort the rules of the association rule set with explanations to obtain a sorted association rule set;
[0123] (6) Perform item recommendation matching based on the sorted association rule set to obtain an association rule recommendation list.
[0124] Specifically, the Apriori or FP-Growth algorithm is applied to the bill clustering results and project features respectively to generate frequent item sets. First, a minimum support threshold is set, and candidate item sets of length k are iteratively generated. Then, the database is scanned to calculate the support, and the item sets that meet the minimum support are retained. This process will obtain bill frequent item sets and project frequent item sets respectively. Based on the generated frequent item sets, candidate association rules of the form X->Y are constructed. For each frequent item set, all its non-empty true subsets are generated as the antecedent X, and the remaining part is used as the consequent Y. In this way, a candidate set containing all possible association rules can be obtained. The support of each candidate rule is calculated: support(X->Y)=count(X∪Y) / N, where N is the total number of transactions. Rules with support greater than the preset threshold are retained. Then the confidence is calculated: confidence(X->Y)=support(X∪Y) / support(X), and rules with confidence greater than the preset threshold are also retained. Calculate the lift of the rule: lift(X->Y) = confidence(X->Y) / support(Y). A lift greater than 1 indicates a positive correlation, and these rules are retained. The rules can then be optimized through methods such as pruning and merging to remove redundant or low-quality rules. Add an explainable description to each rule. A comprehensive scoring function can be constructed based on indicators such as support, confidence, and lift to sort the rules. Business indicators (such as profit margin) can also be introduced for weighted sorting. For a given bill or project, find all applicable association rules. Generate a recommendation list based on the consequence of the rule, and the recommended items can be sorted based on the rule score.
[0125] In a specific embodiment, the execution step performs deep learning model training using the bill feature vector and the project feature library to obtain a deep learning recommendation model may specifically include the following steps:
[0126] (1) Preprocess the bill feature vector and project feature library to obtain a standardized training data set;
[0127] (2) Perform data negative sampling based on the standardized training data set to obtain a balanced training sample set, and perform feature crossover on the balanced training sample set to obtain a cross feature set;
[0128] (3) Perform feature autoencoder training on the cross feature set to obtain a feature compression model, and perform attention mechanism on the compressed features output by the feature compression model to obtain attention weighted features;
[0129] (4) Perform multi-layer neural network training based on the attention-weighted features to obtain deep feature representation, and perform residual network processing on the deep feature representation to obtain optimized feature representation;
[0130] (5) Perform multi-task learning based on the optimized feature representation to obtain a multi-objective prediction model, and then perform integrated learning on the multi-objective prediction model to obtain an integrated recommendation model;
[0131] (6) Perform knowledge distillation on the integrated recommendation model to obtain a deep learning recommendation model.
[0132] Specifically, the bill feature vectors and project feature database are first standardized. This may include filling missing values, handling outliers, normalizing or standardizing numerical features, and one-hot encoding categorical features. The goal is to convert all features to the same scale to facilitate model learning. To address the imbalance between positive and negative samples, oversampling is performed on minority class samples or undersampling is performed on majority class samples. Feature crossover is then performed, for example, combining different features pairwise to generate new features, enhancing the model's nonlinear representation capabilities. An autoencoder is used to reduce the dimensionality of high-dimensional features and extract the most important feature information. An attention mechanism is then applied to allow the model to automatically learn the importance weights of different features, highlighting the role of key features. A multi-layer neural network is constructed to learn the complex nonlinear relationships between features. Residual connections are introduced to alleviate the vanishing gradient problem in deep networks, facilitating the training of deeper network structures. Multiple related tasks (such as click-through rate and conversion rate prediction) are simultaneously optimized, sharing the underlying feature representations. Ensemble learning methods (such as bagging and boosting) are then used to combine multiple base models to improve overall prediction performance. The knowledge of a large and complex ensemble model (teacher model) is transferred to a smaller single model (student model). The student model not only learns the true labels, but also learns the soft labels of the teacher model to obtain richer knowledge.
[0133] The above describes the project matching method based on automatic data recognition in the embodiment of the present application. The following describes the project matching system based on automatic data recognition in the embodiment of the present application. Figure 2 In one embodiment of the present application, a project matching system based on automatic data identification includes:
[0134] The acquisition module 201 is used to acquire original bill data and pre-process the original bill data to obtain target bill data;
[0135] Analysis module 202, for performing type recognition and layout analysis on target bill data to obtain structured area data;
[0136] Recognition module 203, used to perform character recognition and information extraction on the structured area data to obtain bill structured data;
[0137] The aggregation module 204 is used to aggregate and analyze the structured data of the bill to obtain multi-dimensional data analysis results;
[0138] Extraction module 205, used to obtain multiple candidate project information, and perform feature extraction and vectorization on the multiple candidate project information to obtain a project feature library;
[0139] The matching module 206 is used to intelligently match the multi-dimensional data analysis results with the project feature library to obtain a list of recommended projects.
[0140] Through the collaborative efforts of these components, image processing improves the quality of raw bill data. Multi-scale feature extraction and cluster analysis, combined with bill template matching, achieve accurate bill type identification and region segmentation, effectively improving the accuracy of structured data extraction. Combining feature extraction, classification recognition, and contextual analysis significantly enhances character recognition accuracy, while semantic analysis and entity recognition enhance the depth of information extraction. By comprehensively exploring the potential value of bill data, this provides rich data support for project matching. Comprehensive feature extraction of candidate project information, including word frequency statistics, word vector training, and topic modeling, forms a multi-faceted project feature representation, enhancing matching accuracy. A comprehensive application of multiple matching methods, including collaborative filtering, association rule mining, and deep learning, effectively improves the coverage and accuracy of project recommendations. Through frequent item set generation, multiple filtering, and rule optimization, high-quality association rules are extracted, enhancing the interpretability and credibility of recommendation results. A high-performance deep learning recommendation model is constructed, significantly improving matching accuracy and efficiency. Through weighted fusion and multi-objective optimization, the advantages of various recommendation methods are effectively integrated, balancing multiple recommendation objectives and improving project matching accuracy.
[0141] The present application also provides a project matching device based on automatic data identification, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the project matching method based on automatic data identification in the above-mentioned embodiments.
[0142] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the project matching method based on automatic data identification.
[0143] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0145] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A project matching method based on automatic data identification, characterized in that: The method comprises: Acquiring original bill data and preprocessing the original bill data to obtain target bill data; Performing type identification and layout analysis on the target bill data to obtain structured region data; specifically comprising: performing multi-scale feature extraction on the target bill data to obtain a multi-scale feature map of the bill, and performing cluster analysis on the multi-scale feature map of the bill to obtain a bill type feature clustering result; constructing a bill type identification model based on the bill type feature clustering result, and performing type prediction on the target bill data using the bill type identification model to obtain a bill type identification result; performing bill template matching on the bill type identification result to obtain matched bill template data, and performing region division on the target bill data based on the matched bill template data to obtain preliminarily divided bill region data; performing edge detection and contour analysis on the preliminarily divided bill region data to obtain bill region contour data, and performing region optimization on the preliminarily divided bill region data based on the bill region contour data to obtain optimized bill region data; performing layout analysis on the optimized bill region data to obtain bill layout structure data, and performing semantic annotation on the optimized bill region data based on the bill layout structure data to obtain structured region data; Performing character recognition and information extraction on the structured area data to obtain bill structured data; specifically comprising: performing character segmentation on the structured area data to obtain a set of single character images, and performing feature extraction on the set of single character images to obtain a set of character feature vectors; performing classification recognition on the set of character feature vectors to obtain a preliminary recognition result, and performing context analysis on the preliminary recognition result to obtain an optimized character recognition result; performing text semantic analysis on the optimized character recognition result to obtain semantic annotation data, and performing entity recognition on the semantic annotation data to obtain bill entity information; performing relationship extraction on the bill entity information to obtain an entity relationship graph, and performing knowledge graph construction on the entity relationship graph to obtain a bill knowledge graph; performing structured conversion on the bill knowledge graph to obtain bill structured data; The bill structured data is aggregated and analyzed to obtain multi-dimensional data analysis results; specifically including: performing data cleaning on the bill structured data to obtain cleaned bill data, and standardizing the cleaned bill data to obtain standardized bill data; performing time series analysis on the standardized bill data to obtain time dimension analysis results; performing spatial clustering on the standardized bill data to obtain spatial dimension analysis results; performing multivariate correlation analysis on the standardized bill data to obtain variable association analysis results; performing principal component analysis on the standardized bill data based on the variable association analysis results to obtain feature data after dimensionality reduction; performing anomaly detection on the feature data after dimensionality reduction to obtain anomaly data analysis results, and performing cluster analysis on the feature data after dimensionality reduction to obtain data clustering results; performing decision tree analysis on the data clustering results to obtain rule deduction results; performing multi-dimensional fusion on the time dimension analysis results, the spatial dimension analysis results, the variable association analysis results, the anomaly data analysis results, the data clustering results and the rule deduction results to obtain multi-dimensional data analysis results; Acquire information of multiple candidate projects, perform feature extraction and vectorization on the information of multiple candidate projects, and obtain a project feature library; Intelligently match the multi-dimensional data analysis results with the project feature library to obtain a recommended project list.
2. The project matching method based on automatic data identification according to claim 1 is characterized in that: The step of obtaining original bill data and preprocessing the original bill data to obtain target bill data includes: Obtaining original bill data, and performing image conversion and image enhancement on the original bill data to obtain an enhanced bill image; Denoising the enhanced bill image to obtain a denoised bill image, and binarizing the denoised bill image to obtain a binarized bill image; Performing tilt correction on the binarized bill image to obtain a corrected bill image, and performing edge detection on the corrected bill image to obtain bill edge contour data; Cropping the bill edge contour data to obtain a cropped bill image, and standardizing the resolution of the cropped bill image to obtain a standardized bill image; Performing color space conversion on the standardized bill image to obtain a converted bill image; Performing illumination balancing on the converted bill image to obtain an illumination-balanced bill image, and performing data format conversion on the illumination-balanced bill image to obtain target bill information data.
3. The project matching method based on automatic data identification according to claim 1 is characterized in that: The method of obtaining information of multiple candidate projects and performing feature extraction and vectorization on the information of multiple candidate projects to obtain a project feature library includes: Acquire multiple candidate project information, and perform data cleaning on the multiple candidate project information to obtain cleansed project data; Performing text segmentation on the cleaned project data to obtain a project keyword set, and performing word frequency statistics on the project keyword set to obtain a word frequency feature vector; Performing word vector training on the project keyword set to obtain a word vector model, and performing topic model training on the cleaned project data to obtain project topic distribution; Performing named entity recognition on the cleaned project data to obtain entity feature vectors; Performing feature fusion on the word frequency feature vector, the word vector model, the project topic distribution, and the entity feature vector to obtain a fused feature vector; The fusion feature vector is subjected to dimensionality reduction to obtain a project feature vector after dimensionality reduction, and vector clustering and indexing are performed based on the project feature vector after dimensionality reduction to obtain a project feature library.
4. The project matching method based on automatic data identification according to claim 1 is characterized in that: The intelligent matching of the multi-dimensional data analysis results and the project feature library to obtain a recommended project list includes: Performing feature extraction on the multi-dimensional data analysis results to obtain a bill feature vector, and performing similarity calculation on the bill feature vector and the project feature library to obtain a preliminary matching result; Performing collaborative filtering on the preliminary matching results to obtain a collaborative filtering recommendation list, and performing cluster analysis on the bill feature vectors to obtain bill clustering results; Performing association rule mining on the bill clustering results and the project feature library to obtain an association rule recommendation list; Performing deep learning model training using the bill feature vector and the project feature library to obtain a deep learning recommendation model; Perform predictions using the deep learning recommendation model to obtain a deep learning recommendation list; Performing weighted fusion on the collaborative filtering recommendation list, the association rule recommendation list, and the deep learning recommendation list to obtain a fused recommendation list; Multi-objective optimization is performed on the fused recommendation list to obtain an optimized recommendation list, and the items in the optimized recommendation list are sorted and screened to obtain a recommended item list.
5. The project matching method based on automatic data identification according to claim 4 is characterized in that: The association rule mining is performed on the bill clustering result and the project feature library to obtain an association rule recommendation list, including: Performing frequent item set generation on the bill clustering result to obtain bill frequent item set, and performing frequent item set generation on the project feature according to the project feature library to obtain project frequent item set; Generating candidate association rules for the bill frequent item set and the project frequent item set to obtain a candidate association rule set; Performing rule support calculation on the candidate association rule set to obtain a support-filtered rule set, and performing confidence calculation on the support-filtered rule set to obtain a confidence-filtered rule set; Calculating rule lifts based on the confidence-filtered rule set to obtain a lift-filtered rule set, and optimizing the lift-filtered rule set to obtain an optimized association rule set; Performing rule interpretability analysis on the optimized association rule set to obtain an association rule set with explanations, and sorting the rules of the association rule set with explanations to obtain a sorted association rule set; Item recommendation matching is performed according to the sorted association rule set to obtain an association rule recommendation list.
6. The project matching method based on automatic data identification according to claim 4 is characterized in that: The deep learning model training is performed using the bill feature vector and the project feature library to obtain a deep learning recommendation model, including: Performing data preprocessing on the bill feature vector and the project feature library to obtain a standardized training data set; Performing data negative sampling according to the standardized training data set to obtain a balanced training sample set, and performing feature crossover on the balanced training sample set to obtain a crossover feature set; Performing feature autoencoder training on the cross feature set to obtain a feature compression model, and performing an attention mechanism on the compressed features output by the feature compression model to obtain attention-weighted features; Performing multi-layer neural network training based on the attention-weighted features to obtain a deep feature representation, and performing residual network processing on the deep feature representation to obtain an optimized feature representation; Performing multi-task learning based on the optimized feature representation to obtain a multi-objective prediction model, and performing ensemble learning on the multi-objective prediction model to obtain an ensemble recommendation model; Perform knowledge distillation on the integrated recommendation model to obtain a deep learning recommendation model.
7. A project matching system based on automatic data identification, characterized in that: A system for performing an item matching method based on automatic data identification according to any one of claims 1 to 6, comprising: An acquisition module, configured to acquire original bill data and pre-process the original bill data to obtain target bill data; An analysis module, configured to perform type recognition and layout analysis on the target bill data to obtain structured area data; A recognition module, configured to perform character recognition and information extraction on the structured area data to obtain bill structured data; A summary module, used to summarize and analyze the bill structured data to obtain multi-dimensional data analysis results; The extraction module is used to obtain information of multiple candidate projects, perform feature extraction and vectorization on the information of multiple candidate projects, and obtain a project feature library; The matching module is used to intelligently match the multi-dimensional data analysis results with the project feature library to obtain a list of recommended projects.
Citation Information
Patent Citations
Public opinion information recommendation method based on knowledge graph
CN118193850A
Support ticket summarizer, similarity classifier, and resolution forecaster
US20210328888A1