A method and system for data traceability
The method improves data sourcing accuracy by employing semantic understanding and query tracing to establish precise associations between functional points and data tables, automating the process and enhancing data management efficiency.
Patent Information
- Application Number
- CN202011328158.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-11-24
AI Technical Summary
The existing data traceability methods are insufficient in the big data platform, and it is difficult to accurately match the source and processing of data, resulting in decision-making errors and losses.
A variety of matching methods are adopted, including semantic understanding matching, data analysis matching and query and tracking matching, combining functional point lists, display text and data table call records, and determining the association relationship between data table and functional point through semantic understanding and vector similarity calculation.
It improves the accuracy of data traceability, realizes the automatic correlation between the system's front-end business functions and the back-end database tables, reduces manual workload, and improves data asset management efficiency and software development capabilities.
Smart Images

Figure CN114547231B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of databases, and particularly relates to a method and a system for data traceability. Background Art
[0002] With the rapid development of computer mobile Internet and information storage capabilities, various information has shown an explosive exponential growth. Along with the advent of the cloud computing and big data era, people have gradually realized the importance of data. However, due to the huge complexity of data, various data-related problems will inevitably arise, such as data loss, data inconsistency, and data reliability. When people obtain some data, they often consider whether these data are true and reliable, otherwise it may lead to wrong decisions for us. These data information can usually be divided into two categories, one is the most original input data, and the other is the data derived from these data. However, what is usually exposed to users more is the derived data, that is, the data that has been processed in various ways. These data are often stored through various complex conversions or editing methods. Since we do not know the conversion process, people often have a skeptical attitude towards this result data. In fact, sometimes the result data has no relation to the original data, so we must care about the generation process and their sources of these result data.
[0003] Traceability technology has been widely applied in many fields, such as archaeology, physics, astronomy, archives, etc. In recent years, data traceability has also developed in the computer field, mainly existing in research directions such as databases, scientific experiments, and workflows. However, the research in the big data field is relatively less. With the increasingly wide application of big data platforms in enterprises, various data will be processed through a series of big data models and obtain results. Decision-makers use this result data for analysis and make decisions. If the result data is inaccurate or the source is unreliable, it will lead to wrong decisions and even cause immeasurable losses to the enterprise. Therefore, data traceability under the big data platform becomes more and more important. Users often need to combine historical information such as the source and processing process of data to determine whether the data is reliable, and data traceability can describe the origin and processing process of data, and can provide users with audit mechanisms, locate errors, and debug the processing process, etc.
[0004] The existing data traceability methods include matching and analyzing the displayed content with all the content in the database based on string matching or data analysis matching. Since the data in a data table may be duplicated, the data in different tables may be similar, and each piece of data may contain multiple parts, there will be some probability models to analyze the matching probability of two pieces of data after content matching analysis, and then obtain the source information of the data. However, due to the possibility of data duplication, the results of many data sources may be similar, and at the same time, there will be some additional processing before many data are displayed, resulting in a large gap between the data in the results and the original data, causing errors and omissions in the data matching method, and the data traceability results are not very accurate. Therefore, how to improve the accuracy of data traceability is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a data traceability method, including:
[0006] Obtain the function point list corresponding to the function point that generates the data, the display text after function point operation, the data table call record involved in the function point operation, and all data table lists and data tables;
[0007] Based on the function point list, the display text after function point operation, the data table call record involved in the function point operation, and all data table lists and data tables, use multiple matching methods to obtain the correlation coefficients between each data table and the function point under each matching method;
[0008] Based on the correlation coefficients between each data table and the function point under each matching method, determine the data tables associated with the function point.
[0009] Preferably, the matching methods include: semantic understanding matching, data analysis matching, and query tracking matching methods.
[0010] Preferably, using the semantic understanding matching method to obtain the correlation coefficients between each data table and the function point under the semantic understanding matching method includes:
[0011] Based on the function point, determine the function point list corresponding to the function point;
[0012] Based on each data table, determine the data table list corresponding to each data table;
[0013] Based on the natural language processing method, perform semantic understanding annotation on the text in the function point list and each data table list respectively;
[0014] Based on the function point list and each data table list after understanding annotation, construct the vector corresponding to the function point list and the vector corresponding to each data table list in turn;
[0015] Calculate the cosine similarity between the vector corresponding to each data table list and the vector corresponding to the function point list in sequence;
[0016] Take the cosine similarity as the correlation coefficient between each data table and the function point under the semantic understanding matching method.
[0017] Preferably, obtain the correlation coefficients between each data table and the function point under the data analysis matching method by using the data analysis matching method, including:
[0018] Determine the display text after the function point operation based on the function point;
[0019] Determine the content in each data table based on each data table respectively;
[0020] Based on the display text after the function point operation and the content in each data table, construct the vector corresponding to the display text after the function point operation and the vector corresponding to the content in each data table in sequence;
[0021] Calculate the cosine similarity between the vector corresponding to the content in each data table and the vector corresponding to the display text after the function point operation in sequence;
[0022] Take the cosine similarity as the correlation coefficient between each data table and the function point under the data analysis matching method.
[0023] Preferably, obtain the correlation coefficients between each data table and the function point under the query tracking matching method by using the query tracking matching method, including:
[0024] Obtain the data table call record involved in the operation of the function point;
[0025] Based on the data table call record involved in the operation of the function point, determine the data tables associated with the function point and the data tables not associated with the function point;
[0026] Assign preset values to the data tables associated with the function point and the data tables not associated with the function point respectively, as the correlation coefficients between each data table and the function point under the query tracking matching method.
[0027] Preferably, based on the correlation coefficients between each data table and the function point under each method, determine the data tables associated with the function point, including:
[0028] Based on the correlation coefficients between each data table and the function point obtained under the semantic understanding matching, data analysis matching and query tracking matching methods and the pre-set correlation coefficient weights corresponding to each method, calculate the comprehensive association value between each data table and the function point;
[0029] Arrange the comprehensive correlation values of the respective data tables with the function points in descending order, and set data tables with relatively high comprehensive correlation values with the function points as the data tables associated with the function points.
[0030] Preferably, after determining the data tables associated with the function points, it further includes reviewing the data tables associated with the function points to obtain the final data tables associated with the function points.
[0031] Based on the same concept, the present invention also provides a data traceability system, including:
[0032] A data acquisition module for obtaining a function point list corresponding to a function point that generates data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables;
[0033] A correlation coefficient calculation module for obtaining the correlation coefficients between each data table and the function point under each matching method by using multiple matching methods based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables;
[0034] A result output module for determining the data tables associated with the function point based on the correlation coefficients between each data table and the function point under each matching method.
[0035] Preferably, the matching methods include: semantic understanding matching, data analysis matching, and query tracking matching methods.
[0036] Preferably, the result output module includes:
[0037] A comprehensive correlation value calculation unit for calculating the comprehensive correlation value between each data table and the function point based on the correlation coefficients between each data table and the function point obtained under the semantic understanding matching, data analysis matching, and query tracking matching methods and the pre-set correlation coefficient weights corresponding to each method;
[0038] A screening unit for arranging the comprehensive correlation values between each data table and the function point in descending order, and setting a data table with a relatively high comprehensive correlation value with the function point as the data table associated with the function point.
[0039] Compared with the closest prior art, the beneficial effects of the present invention are as follows:
[0040] The present invention provides a method and system for data traceability, including: obtaining a function point list corresponding to a function point that generates data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables; based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables, using multiple matching methods to obtain the correlation coefficients between each data table and the function point under each matching method; based on the correlation coefficients between each data table and the function point under each matching method, determining the data tables associated with the function point. The present invention takes into account the data matching relationships between function points and each data table in multiple dimensions, can accurately find the association relationship between function points and corresponding data tables, realize the automatic association between the front-end business functions of the system and the back-end database tables, improve the efficiency of data asset management, reduce manual workload, assist software application development capabilities, and simplify the cost of later maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. is a schematic diagram of a method for data traceability provided by the present invention;
[0042] Figure 2 FIG. is a schematic diagram of a system for data traceability provided by the present invention;
[0043] Figure 3 FIG. is a data traceability flow chart provided in an embodiment of the present invention;
[0044] Figure 4 FIG. is a program framework diagram of data traceability provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0046] Embodiment 1:
[0047] A method for data traceability provided in an embodiment of the present invention is as Figure 1 shown and includes:
[0048] S1 Obtaining a function point list corresponding to a function point that generates data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables;
[0049] S2 Based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables, using multiple matching methods to obtain the correlation coefficients between each data table and the function point under each matching method;
[0050] S3 determines the data tables associated with the function point based on the correlation coefficients between each data table and the function point under each matching method.
[0051] Specifically, the data traceability flow chart is as Figure 3 shown, including three parts: input, operation processing, and output;
[0052] Input part: The user can import data. The function point list, data table list, etc. can be imported in the form of manual input, Excel sheet import, etc. After import, it will be stored in the server database for unified management. When the data in the database is utilized after storage, it can be judged and searched through the underlying call path, or judged and searched through the similarity of the data itself. In the case of other information (database table information provided by the administrator, description content of each function, etc.), this information can also be used for auxiliary judgment.
[0053] Operation processing part: Using the obtained function point list, the display text after function point operation, the data table call records involved in function point operation, and all data table lists and data tables, multiple matching methods are used to obtain the correlation coefficients between each data table and the function point under each matching method. Specifically, it includes three steps that run simultaneously:
[0054] S2-1 Semantic understanding matching: The content of the function point list and the data table list is semantically understood and annotated using natural language processing tools, and the text similarity measurement is performed on the function point list and the data table list after understanding and annotation to obtain the correlation coefficients between each data table and the function point;
[0055] Text similarity measurement refers to regarding the text as a set of words, analyzing the number of times each word appears in the text and the number of times it appears in the entire text set, and then using this word frequency information to model the text as a vector, and calculating the similarity between texts using methods such as the cosine distance between vectors and neural network methods for sentence embedding. Text similarity measurement is widely used in many fields, such as: information retrieval field, text classification, automatic generation of text summaries, and duplicate detection of texts. The existing TF-IDF method (Term Frequency-Inverse Document Frequency) mainly models the text as a word frequency vector and then uses cosine similarity to calculate the similarity between two texts. The specific steps in this embodiment are as follows:
[0056] S2-1-1 Determine the function point list corresponding to the function point based on the function point;
[0057] S2-1-2 Based on each data table, determine the data table list corresponding to each data table;
[0058] S2-1-3 Perform semantic understanding and annotation on the text in the function point list and each data table list respectively based on natural language processing methods;
[0059] S2-1-4 Based on the function point list and each data table list after understanding and annotation, construct the vector corresponding to the function point list and the vector corresponding to each data table list in sequence;
[0060] S2-1-5 Calculate the cosine similarity between the vector corresponding to each data table list and the vector corresponding to the function point list in sequence;
[0061] S2-1-6 Use the cosine similarity as the correlation coefficient between each data table and the function point under the semantic understanding matching method.
[0062] S2-2 Data analysis matching. Compare the display text after function point operations with the specific content of each data table, and mine the degree of data correlation based on the vector space matching algorithm to obtain the correlation coefficient between each data table and the function point. The specific steps are as follows:
[0063] S2-2-1 Determine the display text after function point operations based on the function point;
[0064] S2-2-2 Determine the content in each data table based on each data table respectively;
[0065] S2-2-3 Based on the display text after function point operations and the content in each data table, construct the vector corresponding to the display text after function point operations and the vector corresponding to the content in each data table in sequence;
[0066] S2-2-4 Calculate the cosine similarity between the vector corresponding to the content in each data table and the vector corresponding to the display text after function point operations in sequence;
[0067] S2-2-5 Use the cosine similarity as the correlation coefficient between each data table and the function point under the data analysis matching method.
[0068] S2-3 Query tracking matching. Match the operation behavior of each function point of the user with the call record of the specific data table, find the data table used in the call process of each function point, and then obtain the correlation coefficient between each data table and the function point. The specific steps are as follows:
[0069] S2-3-1 Obtain the data table call record involved in the operation of the function point;
[0070] S2-3-2 Based on the data table call record involved in the function point operation, determine the data table associated with the function point and the unassociated data table;
[0071] S2-3-3 sets the correlation coefficient between the data table associated with the function point and the function point to 1, and sets the correlation coefficient between the data table not associated with the function point and the function point to 0, which serves as the correlation coefficient between each data table and the function point.
[0072] In the output part, the correlation coefficients between each data table and the function point calculated under the three matching analysis methods are subjected to corresponding weighted average calculations to obtain the comprehensive correlation coefficients between each data table and the function point, and they are sorted in descending order. The N data tables with the highest comprehensive correlation values with the function point are used as the data tables associated with the function point, and the final matching results are obtained through manual review and saved in the system for convenient export.
[0073] The program architecture for implementing the above data traceability is as Figure 4 shown, including: the front-end I / O layer, the operation processing layer, and the data management layer;
[0074] The front-end I / O layer is the user interaction interface, including: the interface UI components, the message processing module, the exception handling module, and the front-back end interface;
[0075] Various UI components are used to obtain information through various input methods such as accepting user clicks and keyboard inputs;
[0076] The message processing module and the exception handling module are used to process normal information and exception information for the information obtained by various UI components;
[0077] The front-back end interface is used to interact the processed information with the operation processing layer and the data management layer.
[0078] The operation processing layer includes a semantic understanding module, a data mining module, a query tracking module, a cloud service module, a session control module, and a data interaction interface;
[0079] The named entity recognition part in the semantic understanding module can be used to analyze the meaning of each word, and the word embedding part converts each word into a corresponding vector respectively, and then calculates the similarity in the semantic measurement component;
[0080] The data mining module uses the target information collection component to extract the results obtained after each function point user operation, and uses the resource information collection component to collect the specific content of each data table, and then uses the algorithm based on vector space matching in the data mining component to mine the data correlation degree, so as to obtain the matching degree;
[0081] The query tracking module uses the action tracking module to capture the user's input and click operations, calls the tracking module to obtain the call records of specific data tables, and performs matching in the tracking information mining component to obtain the corresponding relationship of the data tables;
[0082] The cloud service module, session service module, and data interaction interface are designed to enable the operation processing layer to store information, interact with the outside world, and obtain corresponding data content.
[0083] The data management layer includes: a data storage module, a data update module, a backup and guarantee module, and a data interaction interface;
[0084] The data management layer is used to store data. The data storage module is used to save data, the data update module is used to maintain and update data, the backup and guarantee module is used to import, export, and periodically back up data, and the data interaction interface is used to interact with the operation processing layer.
[0085] Based on the three - matching method for data traceability, taking into account the similarity between the data itself and the database data, database call records, human knowledge, etc., data matching analysis is carried out, which can accurately find the association relationship between the function points and the corresponding data tables. Even if the content of the data tables is not directly displayed in the user window, the corresponding association can be found, realizing the automated generation of the association relationship recommendation between the front - end business functions of the system and the back - end database tables, improving the automation and intelligence of the manual inventory work, and being able to effectively improve the inventory efficiency, optimize the inventory effect, and reduce the manual workload.
[0086] Embodiment 2:
[0087] An embodiment of the present invention discloses a data traceability system, as Figure 2 shown, including:
[0088] A data acquisition module, which is used to obtain the function point list corresponding to the function points generating data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables;
[0089] A correlation coefficient calculation module, which is used to obtain the correlation coefficients between each data table and the function points under each matching method by using a variety of matching methods based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables;
[0090] A result output module, which is used to determine the data tables associated with the function points based on the correlation coefficients between each data table and the function points under each matching method.
[0091] Among them, the matching methods include: semantic understanding matching, data analysis matching, and query tracking matching methods.
[0092] Furthermore, the correlation coefficient calculation module includes: a semantic understanding matching calculation unit, a data analysis matching calculation unit, and a query tracking matching calculation unit;
[0093] A semantic understanding matching calculation unit, which is used to obtain the correlation coefficients between each data table and the function point under the semantic understanding matching method by using the semantic understanding matching method;
[0094] A data analysis matching calculation unit, which is used to obtain the correlation coefficients between each data table and the function point under the data analysis matching method by using the data analysis matching method;
[0095] A query tracking matching calculation unit, which is used to obtain the correlation coefficients between each data table and the function point under the query tracking matching method by using the query tracking matching method.
[0096] Further, the semantic understanding matching calculation unit includes:
[0097] A function point list determination subunit, which is used to determine the function point list corresponding to the function point based on the function point;
[0098] A data table list determination subunit, which is used to determine the data table list corresponding to each data table based on each data table respectively;
[0099] A natural language processing subunit, which is used to perform semantic understanding annotation on the texts in the function point list and each data table list respectively based on the natural language processing method;
[0100] A vector construction subunit 1, which is used to construct the vector corresponding to the function point list and the vectors corresponding to each data table list in sequence based on the function point list and each data table list after understanding annotation;
[0101] A cosine similarity calculation subunit 1, which is used to calculate the cosine similarity between the vector corresponding to each data table list and the vector corresponding to the function point list in sequence;
[0102] A correlation coefficient calculation subunit 1, which is used to use the cosine similarity as the correlation coefficient between each data table and the function point under the semantic understanding matching method.
[0103] Further, the data analysis matching calculation unit includes:
[0104] A function point display text determination subunit, which is used to determine the display text after the operation of the function point based on the function point;
[0105] A data table content determination subunit, which is used to determine the content in each data table based on each data table respectively;
[0106] A vector construction subunit 2, which is used to construct the vector corresponding to the display text after the operation of the function point and the vectors corresponding to the content in each data table in sequence based on the display text after the operation of the function point and the content in each data table;
[0107] The cosine similarity calculation subunit 2 is used to calculate the cosine similarity between the vectors corresponding to the content in each data table and the vector corresponding to the displayed text after the function point operation in sequence;
[0108] The correlation coefficient calculation subunit 2 is used to use the cosine similarity as the correlation coefficient between each data table and the function point under the data analysis matching method.
[0109] Furthermore, the query tracking matching calculation unit includes:
[0110] The call record determination subunit is used to obtain the data table call record involved in the operation of the function point;
[0111] The call result determination subunit is used to determine the data tables associated with and not associated with the function point based on the data table call record involved in the function point operation;
[0112] The correlation coefficient calculation subunit 3 is used to assign preset values to the data tables associated with and not associated with the function point respectively as the correlation coefficient between each data table and the function point under the query tracking matching method.
[0113] Furthermore, the result output module includes:
[0114] The comprehensive association value calculation unit is used to calculate the comprehensive association value between each data table and the function point based on the correlation coefficients between each data table and the function point obtained under the semantic understanding matching, data analysis matching, and query tracking matching methods and the preset correlation coefficient weights corresponding to each method;
[0115] The screening unit is used to sort the comprehensive association values between each data table and the function point in descending order, and set data tables with higher sorted comprehensive association values with the function point as the data tables associated with the function point.
[0116] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0117] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or one or more blocks Figure 1 of the functions specified in the block.
[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or one or more blocks Figure 1 of the functions specified in the block.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks in the flow. Figure 1 one or more flows and / or one or more blocks Figure 1 of the functions specified in the block.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than to limit the scope of its protection. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present application, various changes, modifications, or equivalent substitutions can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent substitutions are all within the scope of the protection of the pending claims of the application.
Claims
1. A method for data traceability, characterized in that, Including: Obtaining a function point list corresponding to the function points that generate data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables; Based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables, using multiple matching methods to obtain the correlation coefficients between each data table and the function points under each matching method; Based on the correlation coefficients between each data table and the function points under each matching method, determining the data tables associated with the function points; The matching methods include: semantic understanding matching, data analysis matching, and query tracking matching methods; Obtaining the correlation coefficients between each data table and the function points under the semantic understanding matching method by using the semantic understanding matching method, including: Determining the function point list corresponding to the function points based on the function points; Respectively determining the data table list corresponding to each data table based on each data table; Respectively performing semantic understanding annotation on the texts in the function point list and each data table list based on natural language processing methods; Based on the function point list and each data table list after understanding annotation, sequentially constructing the vector corresponding to the function point list and the vector corresponding to each data table list; Sequentially calculating the cosine similarity between the vector corresponding to each data table list and the vector corresponding to the function point list; Taking the cosine similarity as the correlation coefficient between each data table and the function points under the semantic understanding matching method; Obtaining the correlation coefficients between each data table and the function points under the data analysis matching method by using the data analysis matching method, including: Determining the display text after function point operations based on the function points; Respectively determining the content in each data table based on each data table; Based on the display text after function point operations and the content in each data table, sequentially constructing the vector corresponding to the display text after function point operations of the function points and the vector corresponding to the content in each data table; Sequentially calculating the cosine similarity between the vector corresponding to the content in each data table and the vector corresponding to the display text after function point operations of the function points; Taking the cosine similarity as the correlation coefficient between each data table and the function points under the data analysis matching method; Obtaining the correlation coefficients between each data table and the function points under the query tracking matching method by using the query tracking matching method, including: Obtaining the data table call records involved in the operations of the function points; Based on the data table call records involved in the function point operations, determining the data tables associated with the function points and the data tables not associated with the function points; Respectively assigning preset values to the data tables associated with the function points and the data tables not associated with the function points as the correlation coefficients between each data table and the function points under the query tracking matching method.
2. The method according to claim 1, wherein The determining the data tables associated with the function points based on the correlation coefficients between each data table and the function points under each matching method includes: Calculating the comprehensive association value between each data table and the function points based on the correlation coefficients between each data table and the function points obtained under the semantic understanding matching, data analysis matching, and query tracking matching methods, and the preset correlation coefficient weights corresponding to each method; Arrange the comprehensive correlation values of the respective data tables with the function points in descending order, and set a data table with a relatively high comprehensive correlation value with the function point as the data table associated with the function point.
3. The method according to claim 1, characterized in that, After determining the data table associated with the function point, it further includes reviewing the data table associated with the function point to obtain the final data table associated with the function point.
4. A system for data traceability, characterized in that, It includes: A data acquisition module for obtaining a function point list corresponding to a function point that generates data, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables; A correlation coefficient calculation module for obtaining the correlation coefficients between each data table and the function point under each matching method by using multiple matching methods based on the function point list, the display text after function point operations, the data table call records involved in function point operations, and all data table lists and data tables; A result output module for determining the data table associated with the function point based on the correlation coefficients between each data table and the function point under each matching method; The matching methods include: semantic understanding matching, data analysis matching, and query tracking matching methods; Obtaining the correlation coefficients between each data table and the function point under the semantic understanding matching method by using the semantic understanding matching method includes: Determining the function point list corresponding to the function point based on the function point; Determining the data table list corresponding to each data table based on each data table respectively; Semantically understanding and annotating the text in the function point list and each data table list respectively based on natural language processing methods; Successively constructing the vector corresponding to the function point list and the vector corresponding to each data table list based on the function point list and each data table list after understanding and annotation; Successively calculating the cosine similarity between the vector corresponding to each data table list and the vector corresponding to the function point list; Taking the cosine similarity as the correlation coefficient between each data table and the function point under the semantic understanding matching method; Obtaining the correlation coefficients between each data table and the function point under the data analysis matching method by using the data analysis matching method includes: Determining the display text after function point operations based on the function point; Determining the content in each data table based on each data table respectively; Successively constructing the vector corresponding to the display text after function point operations of the function point and the vector corresponding to the content in each data table based on the display text after function point operations and the content in each data table; Successively calculating the cosine similarity between the vector corresponding to the content in each data table and the vector corresponding to the display text after function point operations of the function point; Taking the cosine similarity as the correlation coefficient between each data table and the function point under the data analysis matching method; Obtaining the correlation coefficients between each data table and the function point under the query tracking matching method by using the query tracking matching method includes: Obtaining the data table call records involved in the operations of the function point; Determining the data tables associated with the function point and the data tables not associated with the function point based on the data table call records involved in the operations of the function point; Preset values are respectively assigned to the data tables associated with the function points and the data tables not associated with the function points, serving as the correlation coefficients between each data table and the function points under the query tracking matching method.
5. The system according to claim 4, wherein The result output module includes: A comprehensive correlation value calculation unit, configured to calculate the comprehensive correlation value between each data table and the function point based on the correlation coefficients between each data table and the function point obtained under the semantic understanding matching, data analysis matching, and query tracking matching methods, as well as the pre-set correlation coefficient weights corresponding to each method; A screening unit, configured to sort the comprehensive correlation values between each data table and the function point in descending order, and set a data table with a relatively high ranking of the comprehensive correlation value with the function point as the data table associated with the function point.
Citation Information
Patent Citations
Data table blood relationship analysis method and system
CN110990429A
Data blood relationship detection method and system
CN111563103A