BERT-based method and system for accurately docking data requirements

By using a pre-trained BERT model to accurately match government data with demand, the problem of high error rate in matching government data supply and demand has been solved, and efficient and accurate matching of data services has been achieved.

CN120994878APending Publication Date: 2025-11-21SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511055277.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies suffer from high error rates in matching government data supply and demand, making it difficult to achieve accurate matching and resulting in insufficient accuracy of data services.

Method used

A pre-trained BERT model is used for data preprocessing and semantic similarity calculation. Through the semantic similarity calculation module and the demand matching module, the data catalog is accurately matched with user needs. By leveraging the powerful semantic understanding capability of the BERT model, user needs can be responded to quickly.

Benefits of technology

It has enabled precise matching of government data, reduced the error rate of supply and demand matching, and improved the accuracy and efficiency of data services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994878A_ABST
    Figure CN120994878A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a BERT-based accurate data demand docking method and system, and the method comprises the following steps: data preprocessing, model input, similarity calculation, matching display, demand submission, demand check and demand summarization. The method has the beneficial effects that the semantic information required by the user can be accurately captured by utilizing the strong semantic understanding capability of the BERT model, and the accurate matching of the requirements is realized; the pre-trained BERT model is adopted, complex model training is not needed, and user requirements can be quickly responded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, specifically to a method and system for precise matching of data requirements based on BERT. Background Technology

[0002] With the development of artificial intelligence technology, semantic analysis and comparison technology has evolved from early rule-based methods to today's deep learning and large-scale pre-trained models.

[0003] With the construction of Digital China and Digital Government, the government data catalog is gradually being improved, and data services based on the government data catalog are becoming more and more common. In order to solve the problem of data supply, demand and quality, artificial intelligence technology should be incorporated, and the BERT model with more accurate semantic analysis should be used to reduce the error rate of supply and demand matching.

[0004] Therefore, a method and system based on BERT is needed to accurately match data requirements, accurately determine the similarity between information items, and recommend more accurate government data catalogs or data requirements with higher similarity to users. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for precise data demand matching based on BERT, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for precise data demand matching based on BERT, comprising the following steps:

[0007] Data preprocessing steps: Organize the existing data catalog information items and requirement information items in the platform, perform preprocessing operations on the data, and convert the text into an input format acceptable to the BERT model;

[0008] Model input steps: Input the processed data into the pre-trained BERT model;

[0009] Similarity calculation steps: The user inputs data requirement information items based on the requirement submission module, calls the semantic similarity calculation module, and outputs the similarity between the requirement information items and the information items in each data directory;

[0010] Matching and display steps: Call the demand matching module, sort according to similarity, output a directory of data with similarity in descending order and within a certain threshold, and display it to the user;

[0011] Request submission steps: Users select a data directory with high similarity, fill in other basic information about their requests, and then submit their requests. The result saving module is then called to save the matching results.

[0012] Preferably, the specific system operations in the similarity calculation step and the matching display step are as follows: The user enters the original requirement sorting menu, performs requirement sorting operation, and enters the requirement name, requirement content, and requirement information item content; the user clicks to select a directory, and according to the minimum matching degree set by the user, the system displays data directory information with a matching degree higher than the minimum matching degree to the user, and sorts them in reverse order by matching degree, and the user selects the corresponding high matching degree data directory.

[0013] Preferably, the specific system operation in the requirement submission step is as follows: After the user completes the selection of the highly matching directory, selects the sharing method, update frequency, fills in the input and output parameters of the requirement interface, and clicks the "Submit" button to complete the requirement sorting and submission.

[0014] Preferably, the requirement verification step is also included: the verifier enters the original requirement verification menu, selects the requirement to be verified, and performs the verification operation; on the verification page, the verifier can view the basic information of the requirement and preview the matching degree of the resource provision directory selected when submitting the requirement; by clicking "View Matching", the verifier can view the directory matching results saved when submitting the requirement and highlight the selected directory in red.

[0015] Preferably, the system also includes a requirement aggregation and distribution step: After the aggregator selects a requirement, it clicks "Next". The system displays data requirements that match the selected requirement within a certain threshold and lists them in descending order of matching degree. The aggregator selects the requirement with the highest matching degree and clicks "Next". In the final step, the aggregator sets the basic information of the aggregated requirement, selects the data providing directory and the data providing department, and then completes the requirement aggregation and distribution. When selecting the data providing directory, matching degree-related processing and display are performed.

[0016] A system for a method of accurately matching data requirements based on BERT includes:

[0017] Data preprocessing module: Used to organize the existing data catalog information items and requirement information items in the platform, and perform data preprocessing operations to convert the text into an input format acceptable to the BERT model;

[0018] BERT model input module: Connected to the data preprocessing module, it is used to input the processed data into the pre-trained BERT model;

[0019] Semantic similarity calculation module: This module is called when the user inputs data requirement information items based on the requirement submission module. It is used to output the similarity between the requirement information items and the information items in each data directory.

[0020] Demand matching module: connected to the semantic similarity calculation module, used to sort data according to similarity and output a directory of data with similarity in descending order and within a certain threshold.

[0021] Request submission module: Allows users to select a data directory with high similarity, fill in other basic information about their requests, and complete the request submission. When submitting the request, the result saving module is called to save the matching results.

[0022] Result saving module: Connected to the requirement submission module, it is used to save the matching results.

[0023] Preferably, the requirement submission module includes:

[0024] Requirements Gathering Operation Unit: After entering the original requirements gathering menu, users can perform requirements gathering operations through this unit, and enter the requirement name, requirement content, and requirement information items.

[0025] Directory selection display unit: When the user clicks to select a directory, the system displays data directory information with a matching degree higher than the minimum matching degree set by the user, and sorts them in reverse order by matching degree, allowing the user to select the corresponding high matching degree data directory;

[0026] Information Submission Section: After selecting a highly compatible directory, users can use this section to choose the sharing method, update frequency, and fill in the input and output parameters of the required interface. Then, they can click the "Submit" button to complete the requirement review and submission.

[0027] Preferably, it also includes a requirement verification module, which includes:

[0028] Verification Operation Selection Unit: After entering the original requirement verification menu, the verifier can select the requirement to be verified through this unit to perform the verification operation;

[0029] Basic Information Viewing Unit: On the verification page, verifiers can view the basic information required.

[0030] Matching degree preview unit: On the verification page, it is possible to preview the matching degree of the resource provision directory selected when submitting the request;

[0031] Matching Result Viewing Unit: When the checker clicks "Matching View," they can view the directory matching results saved when the requirement was submitted through this unit, and the selected directory will be highlighted in red.

[0032] Preferably, it also includes a demand aggregation and distribution module, the demand aggregation and distribution module comprising:

[0033] Requirement Selection Unit: The aggregator selects the requirements to be aggregated in this unit and clicks "Next".

[0034] Similar Requirement Display Unit: The system displays data requirements that match the selected requirement within a certain threshold, and lists them in descending order of matching degree, allowing the aggregator to select the requirement with the highest matching degree and click "Next";

[0035] Summary and Distribution Settings Unit: In the final step, the aggregator sets the basic information of the summary requirements through this unit, selects the data provision directory and the data provision department, and then completes the summary and distribution of the requirements. When selecting the data provision directory, matching degree related processing and display are performed.

[0036] Preferably, when the demand aggregation and distribution module calls the semantic similarity calculation module and the demand matching module, the semantic similarity calculation module is used to output the similarity between demand information items and other demand information items, and the demand matching module is used to obtain data demands with reverse similarity and similarity within a certain threshold, so as to assist the aggregator in completing the operations of adding new data demands with high similarity, filling in the basic information of the aggregated demand, and selecting the data providing department.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] The present invention proposes a method and system for precise data demand matching based on BERT. By leveraging the powerful semantic understanding capabilities of the BERT model, it can accurately capture the semantic information of user needs and achieve precise demand matching. By using a pre-trained BERT model, there is no need for complex model training, and it can quickly respond to user needs. Attached Figure Description

[0039] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of the present invention clear and complete, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only some, not all, embodiments of the present invention, and are merely illustrative of the embodiments of the present invention. They are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1, please refer to Figure 1 This invention provides a technical solution: a method for precise data demand matching based on BERT, comprising the following steps:

[0042] Step 1: Organize the existing data catalog items and requirement information items in the platform, perform data preprocessing operations, and convert the text into an input format acceptable to the BERT model.

[0043] Step 2: Input the processed data into the pre-trained BERT model.

[0044] Step 3: The user inputs data requirement information items through the requirement submission module, which then calls the semantic similarity calculation module to output the similarity between the requirement information items and the information items in each data directory. The system operation is as follows:

[0045] Users can access the original requirements sorting menu to sort out requirements, and enter the requirement name, requirement content, requirement information items, etc.

[0046] Step 4: Invoke the demand matching module, sort the data according to similarity, and output a directory of data arranged in descending order of similarity with similarity within a certain threshold, which is then displayed to the user. The system operation is as follows:

[0047] When a user clicks to select a directory, the system displays data directories with matching scores higher than the minimum set by the user, and sorts them in descending order of matching score. The user can then select the corresponding high-matching data directories.

[0048] Step 5: The user selects a data directory with high similarity, fills in other basic information about their requirements, and submits the request. Upon submission, the results saving module is invoked to save the matching results. The system operation is as follows:

[0049] After selecting a highly compatible directory, users choose the sharing method, update frequency, and fill in the required interface input and output parameters, then click the "Submit" button to complete the requirement review and submission.

[0050] Step 6: The reviewer examines the requirements submitted by the requester and the matching results, and performs a review of the requirements. The system operation is as follows:

[0051] The verifier enters the original requirement verification menu and selects the requirement to be verified to perform the verification operation.

[0052] On the verification page, you can view the basic information of the requirements and preview the matching degree of the resource provision directory selected when submitting the requirements.

[0053] Clicking "Match View" allows you to view the directory matching results saved when submitting your request, and highlights the selected directories in red.

[0054] Step 7: After verification, the aggregator, based on the demand aggregation module, aggregates demands from multiple departments. It selects the demands to be aggregated, calls the semantic similarity calculation module to output the similarity between demand information items and other demand information items, and then calls the demand matching module to obtain data demands with similarity sorted in descending order and within a certain threshold. The aggregator adds new data demands with high similarity, fills in the basic information of the aggregated demand, selects the data providing department, and completes the aggregation and distribution. The system operation is as follows:

[0055] After selecting a requirement, the aggregator clicks "Next".

[0056] The system displays data requirements that match the selected needs within a certain threshold, and shows them in a list in descending order of matching degree. Users select the requirements with the highest matching degree and click "Next".

[0057] In the final step, the basic information of the aggregated requirements is set, and the data provider directory and data provider department are selected to complete the aggregated and distributed requirements. The data provider directory selection process in this step is similar to the requirement submission process, with matching-related processing and display to facilitate user selection of the target directory and improve the efficiency of requirement coordination.

[0058] Example 2, based on Example 1, proposes a system for precise data demand matching based on BERT, including:

[0059] Data preprocessing module: Used to organize the existing data catalog information items and requirement information items in the platform, and perform data preprocessing operations to convert the text into an input format acceptable to the BERT model;

[0060] BERT model input module: Connected to the data preprocessing module, it is used to input the processed data into the pre-trained BERT model;

[0061] Semantic similarity calculation module: This module is called when the user inputs data requirement information items based on the requirement submission module. It is used to output the similarity between the requirement information items and the information items in each data directory.

[0062] Demand matching module: connected to the semantic similarity calculation module, used to sort data according to similarity and output a directory of data with similarity in descending order and within a certain threshold.

[0063] Request Submission Module: This module allows users to select a data directory with high similarity, fill in other basic request information, and submit their request. Upon submission, it calls the result saving module to save the matching results. The request submission module includes:

[0064] Requirements Gathering Operation Unit: After entering the original requirements gathering menu, users can perform requirements gathering operations through this unit, and enter the requirement name, requirement content, and requirement information items.

[0065] Directory selection display unit: When the user clicks to select a directory, the system displays data directory information with a matching degree higher than the minimum matching degree set by the user, and sorts them in reverse order by matching degree, allowing the user to select the corresponding high matching degree data directory;

[0066] Information Submission Section: After selecting a highly matched directory, users can use this section to choose the sharing method, update frequency, and fill in the input and output parameters of the required interface. Then, they can click the "Submit" button to complete the requirement review and submission.

[0067] The requirements verification module includes:

[0068] Verification Operation Selection Unit: After entering the original requirement verification menu, the verifier can select the requirement to be verified through this unit to perform the verification operation;

[0069] Basic Information Viewing Unit: On the verification page, verifiers can view the basic information required.

[0070] Matching degree preview unit: On the verification page, it is possible to preview the matching degree of the resource provision directory selected when submitting the request;

[0071] Matching Result Viewing Unit: When the checker clicks "Matching View," they can view the directory matching results saved when the requirement was submitted through this unit, and the selected directory will be highlighted in red.

[0072] Result saving module: Connected to the requirement submission module, it is used to save the matching results.

[0073] The demand aggregation and distribution module includes:

[0074] Requirement Selection Unit: The aggregator selects the requirements to be aggregated in this unit and clicks "Next".

[0075] Similar Requirement Display Unit: The system displays data requirements that match the selected requirement within a certain threshold, and lists them in descending order of matching degree, allowing the aggregator to select the requirement with the highest matching degree and click "Next";

[0076] The aggregation and distribution settings unit: In the final step, the aggregator uses this unit to set the basic information of the aggregation requirements, select the data provision directory and data provision department, and then complete the aggregation and distribution of the requirements. During the selection of the data provision directory, matching-related processing and display are performed. When the requirement aggregation and distribution module calls the semantic similarity calculation module and the requirement matching module, the semantic similarity calculation module outputs the similarity between requirement information items and other requirement information items, and the requirement matching module obtains data requirements in descending order of similarity that are within a certain similarity threshold. This assists the aggregator in adding new data requirements with high similarity, filling in the basic information of the aggregation requirements, and selecting the data provision department.

[0077] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for precise data demand matching based on BERT, characterized in that: Includes the following steps: Data preprocessing steps: Organize the existing data catalog information items and requirement information items in the platform, perform preprocessing operations on the data, and convert the text into an input format acceptable to the BERT model; Model input steps: Input the processed data into the pre-trained BERT model; Similarity calculation steps: The user inputs data requirement information items based on the requirement submission module, calls the semantic similarity calculation module, and outputs the similarity between the requirement information items and the information items in each data directory; Matching and display steps: Call the demand matching module, sort according to similarity, output a directory of data with similarity in descending order and within a certain threshold, and display it to the user; Request submission steps: Users select a data directory with high similarity, fill in other basic information about their requests, and then submit their requests. The result saving module is then called to save the matching results.

2. The method for precise data demand matching based on BERT according to claim 1, characterized in that: The specific system operations for the similarity calculation and matching display steps are as follows: The user enters the original requirement sorting menu, performs requirement sorting operations, and enters the requirement name, requirement content, and requirement information items; the user clicks to select a directory, and based on the minimum matching degree set by the user, the system displays data directory information with a matching degree higher than the minimum matching degree, and sorts them in reverse order by matching degree, allowing the user to select the corresponding high matching degree data directory.

3. The method for precise data demand matching based on BERT according to claim 2, characterized in that: The specific system operations for submitting requirements are as follows: After selecting a highly compatible directory, the user chooses the sharing method, update frequency, fills in the input and output parameters of the requirement interface, and clicks the "Submit" button to complete the requirement review and submission.

4. The method for precise data demand matching based on BERT according to claim 3, characterized in that: It also includes a requirement verification step: the verifier enters the original requirement verification menu, selects the requirement to be verified, and performs the verification operation; on the verification page, the verifier can view the basic information of the requirement and preview the matching degree of the resource provision directory selected when submitting the requirement; by clicking "View Matching", the verifier can view the directory matching results saved when submitting the requirement and highlight the selected directory in red.

5. The method for precise data demand matching based on BERT according to claim 4, characterized in that: It also includes a requirement aggregation and distribution step: After the aggregator selects a requirement, it clicks "Next". The system displays data requirements that match the selected requirement within a certain threshold and lists them in descending order of matching degree. The aggregator selects the requirement with the highest matching degree and clicks "Next". In the final step, the aggregator sets the basic information of the aggregated requirement, selects the data providing directory and the data providing department, and then completes the requirement aggregation and distribution. When selecting the data providing directory, matching degree-related processing and display are performed.

6. A system for the method of precise data demand matching based on BERT according to claim 5, characterized in that: include: Data preprocessing module: Used to organize the existing data catalog information items and requirement information items in the platform, and perform data preprocessing operations to convert the text into an input format acceptable to the BERT model; BERT model input module: Connected to the data preprocessing module, it is used to input the processed data into the pre-trained BERT model; Semantic similarity calculation module: This module is called when the user inputs data requirement information items based on the requirement submission module. It is used to output the similarity between the requirement information items and the information items in each data directory. Demand matching module: connected to the semantic similarity calculation module, used to sort data according to similarity and output a directory of data with similarity in descending order and within a certain threshold. Request submission module: Allows users to select a data directory with high similarity, fill in other basic information about their requests, and complete the request submission. When submitting the request, the result saving module is called to save the matching results. Result saving module: Connected to the requirement submission module, it is used to save the matching results.

7. The system according to claim 6, characterized in that: The requirement submission module includes: Requirements Gathering Operation Unit: After entering the original requirements gathering menu, users can perform requirements gathering operations through this unit, and enter the requirements name, requirements content, and requirements information items; Directory selection display unit: When the user clicks to select a directory, the system displays data directory information with a matching degree higher than the minimum matching degree set by the user, and sorts them in reverse order by matching degree, allowing the user to select the corresponding high matching degree data directory; Information Submission Section: After selecting a highly compatible directory, users can use this section to choose the sharing method, update frequency, and fill in the input and output parameters of the required interface. Then, they can click the "Submit" button to complete the requirement review and submission.

8. The system according to claim 7, characterized in that: It also includes a requirements verification module, which includes: Verification Operation Selection Unit: After entering the original requirement verification menu, the verifier can select the requirement to be verified through this unit to perform the verification operation; Basic Information Viewing Unit: On the verification page, verifiers can view the basic information required. Matching degree preview unit: On the verification page, it is possible to preview the matching degree of the resource provision directory selected when submitting the request; Matching Result Viewing Unit: When the checker clicks "Matching Viewing", they can view the directory matching results saved when the requirement was submitted through this unit, and the selected directory will be highlighted in red.

9. A system according to claim 8, characterized in that: It also includes a demand aggregation and distribution module, which includes: Requirement Selection Unit: The aggregator selects the requirements to be aggregated in this unit and clicks "Next". Similar Requirement Display Unit: The system displays data requirements that match the selected requirement within a certain threshold, and lists them in descending order of matching degree, allowing the aggregator to select the requirement with the highest matching degree and click "Next"; Summary and Distribution Settings Unit: In the final step, the aggregator sets the basic information of the summary requirements through this unit, selects the data provision directory and the data provision department, and then completes the summary and distribution of the requirements. When selecting the data provision directory, matching degree related processing and display are performed.

10. A system according to claim 9, characterized in that: When the demand aggregation and distribution module calls the semantic similarity calculation module and the demand matching module, the semantic similarity calculation module is used to output the similarity between demand information items and other demand information items, and the demand matching module is used to obtain data demands with similarity sorted in reverse order and within a certain threshold, so as to assist the aggregation party in completing the operations of adding new data demands with high similarity, filling in the basic information of the aggregation demand, and selecting the data providing department.

Citation Information

Patent Citations

  • Demand matching method and device, storage medium and terminal

    CN108595506A

  • Text semantic retrieval method in military scene

    CN116150335A

  • Government affair intelligent customer service implementation method based on artificial intelligence large model

    CN118093844A

  • Information matching method and system based on large language model

    CN118484510A

  • Intelligent drug recommendation method and system based on voice recognition

    CN118897891A