A data identification method based on blood relationship and data identification
By constructing a data lineage graph and relationship identifiers, the problems of high difficulty and resource waste in custom development of Apache Atlas were solved, enabling cross-platform data identification and transformation and improving data identification efficiency.
Patent Information
- Application Number
- CN202311600015.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-11-28
AI Technical Summary
In existing technologies, Apache Atlas faces challenges in data lineage analysis, including difficulties in custom development, slow iteration to implement core functionalities, and the inability to recognize key information in specific formats after data conversion during cross-platform transfers, resulting in resource waste.
By acquiring data flow information to construct a data lineage graph, determining relationship identifiers, and generating a key information conversion mechanism, data identification is performed using lineage association and data identification methods, avoiding the duplication of content recognition networks on each platform.
It enables rapid identification of key information in data of different formats during data flow, reducing resource waste and improving data identification efficiency and flexibility.
Smart Images

Figure CN117688191B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data identification, and in particular relates to a data identification method based on blood relation and data identification. BACKGROUND
[0002] Currently, the technology used for data blood analysis is mainly Apache Atlas, but data blood tracing is only one of the functions of Atlas. Atlas also has a large number of functions such as metadata management, data governance, data asset directory classification and management, which is easy to overlap with the functions of the existing data asset management platform. Moreover, the rich functions of Atlas mean that it is difficult to customize and develop, and it is difficult to quickly iterate the core function of data blood tracing.
[0003] In the patent file 202110187130.0 data governance method and system based on data blood analysis, it is proposed that in the data governance process, data tracing, data quality verification and help for governance personnel to trace, verify, standardize and improve data governance capability can be realized.
[0004] However, through blood analysis, it lacks data identification function and can only build data relationship network and data tracing. In the field of data identification, it is difficult to identify data only according to blood relationship. Only through data identification, data type and data key items can be determined, and a new content identification network needs to be set up for data identification. However, in the prior art, data may be converted in the process of flowing through different platforms. After data conversion, the internal identification network can only identify the key information of data in a specific format, and cannot be used for all content platforms. Each content platform needs to set up a content identification network, and the content identification network of each platform needs to be trained, so a large amount of resources may be wasted. SUMMARY
[0005] The present application provides a data identification method based on blood relation and data identification, which is used to solve the problem that Apache Atlas technology is difficult to customize and develop, and it is difficult to quickly iterate the core function of data blood tracing. In the prior art, data may be converted in the process of flowing through different platforms. After data conversion, the internal identification network can only identify the key information of data in a specific format, and cannot be used for all content platforms. Each content platform needs to set up a content identification network, and the content identification network of each platform needs to be trained, so a large amount of resources may be wasted.
[0006] The present application provides a data identification method based on blood relation and data identification, which is used to solve the problem that Apache Atlas technology is difficult to customize and develop, and it is difficult to quickly iterate the core function of data blood tracing. In the prior art, data may be converted in the process of flowing through different platforms. After data conversion, the internal identification network can only identify the key information of data in a specific format, and cannot be used for all content platforms. Each content platform needs to set up a content identification network, and the content identification network of each platform needs to be trained, so a large amount of resources may be wasted.
[0007] A data flow direction of to-be-identified data is acquired, and a data bloodline graph is constructed;
[0008] Based on the data bloodline graph and content recognition, a relationship identifier of the to-be-identified data is determined; wherein,
[0009] The relationship identifier is a relationship identifier of key information in the to-be-identified data
[0010] According to the relationship identifier, a key information conversion mechanism is generated; wherein,
[0011] The key information conversion mechanism is used to determine the key information of the to-be-identified data in different formats.
[0012] Preferably, the acquisition of the data flow direction of the to-be-identified data comprises:
[0013] The to-be-identified data is acquired, and a data address and a data path are determined;
[0014] Based on the data address and the data path, the to-be-identified data is discretely processed, and a target value interval is determined;
[0015] Based on the target value interval, interval frame height information corresponding to each index value interval of each target index is determined;
[0016] Based on the target index arrangement order, the interval frame height information and the target index value interval, a target flow direction graph of the to-be-identified data is constructed;
[0017] According to the target flow direction graph, data flow direction statistics are performed.
[0018] Preferably, the construction of the data bloodline graph comprises:
[0019] The to-be-identified data is classified and parsed, and table fields of each to-be-identified data are determined, and table field dependency information is determined;
[0020] The table field dependency information is imported into a preset target data table;
[0021] A preset graph database is called;
[0022] The table field dependency information in the target data table is imported into the graph database;
[0023] Different types of data are processed by the graph database to generate corresponding data bloodline graphs.
[0024] Preferably, the content recognition comprises:
[0025] Image recognition, audio recognition and text recognition.
[0026] Preferably, the image recognition comprises:
[0027] Setting an automatic image recognition mode and a manual recognition mode;
[0028] According to the automatic recognition mode, the static element selection and the dynamic element selection in the image are performed, and the background color is cut off;
[0029] According to the element contour of the selected static element and dynamic element, a calibration frame is set;
[0030] According to the size of the element contour, a maximum area threshold and a minimum area threshold are determined;
[0031] According to the maximum area threshold and the minimum area threshold, the static element and the dynamic element are identified by the calibration frame.
[0032] Preferably, the audio recognition includes:
[0033] Obtaining an audio segment to be recognized;
[0034] Collecting keywords in the audio segment to determine the occurrence scene corresponding to the audio segment;
[0035] According to the occurrence scene, the corresponding sound event identification is performed through the corresponding sound recognition model;
[0036] According to the sound event identification, the text recognition of the audio segment is performed.
[0037] Preferably, the text recognition includes:
[0038] Through automatic feature extraction and character coding of the text image by moving the cursor, the text image feature and character position coding feature are determined;
[0039] According to the character position coding feature, the text image feature is sampled to obtain the sampling feature in the text image;
[0040] The sampling feature is input into a large language model to output corresponding text information.
[0041] Preferably, the determination of the relationship identification of the data to be recognized includes:
[0042] According to the data bloodline map, the relationship vector between different data in the data to be recognized is determined;
[0043] According to the relationship vector, the correlation depth between different data to be recognized is calculated;
[0044] According to the correlation depth, the relationship identification between different data to be recognized is calculated.
[0045] Preferably, the determination of the key information in the data to be recognized according to the relationship identification includes:
[0046] Obtain relationship identifiers and perform serialization and sorting on different data to be identified;
[0047] Based on the serialization sorting, key data are extracted from each piece of data to be identified in sequence, and data nodes are set;
[0048] Calculate the data weight of each piece of data to be identified using data nodes;
[0049] Based on data weights, key data are serialized and sorted to generate multiple datasets;
[0050] The dataset is used as input to the information recognition model to obtain a set of key information about the data to be recognized.
[0051] The key information set is converted into text data to generate the key information data text.
[0052] Preferably, the method further includes:
[0053] Identify key information and generate information features based on that key information;
[0054] Based on the information characteristics, a logical feature representation diagram is constructed; where...
[0055] The logical feature representation diagram contains logical vectors between information features;
[0056] Based on the logical feature representation diagram, key information is sorted to determine the content sequence of key information.
[0057] The beneficial effects of this invention are as follows:
[0058] This application uses a bloodline map to collect bloodline relationships. In this process, the bloodline relationships enable logical and sequential data sorting, and key information can be identified through the sorting method.
[0059] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0060] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0061] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0062] Figure 1This is a flowchart illustrating a data identification method based on blood relation and data identification in an embodiment of the present invention.
[0063] Figure 2 This is a data flow flowchart of a data identification method based on blood relation and data identification in an embodiment of the present invention;
[0064] Figure 3 This is a flowchart illustrating the generation of key information content sequences in a data identification method based on blood ties and data identification, as described in an embodiment of the present invention. Detailed Implementation
[0065] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0066] This invention proposes a data identification method based on kinship and data identification, comprising:
[0067] Obtain the data flow direction of the data to be identified and construct a data lineage map;
[0068] Based on data lineage mapping and content recognition, the relationship identifiers of the data to be identified are determined; among them,
[0069] The relationship identifier is the relationship identifier of key information in the data to be identified;
[0070] Based on the relationship identifier, a key information conversion mechanism is generated; among which...
[0071] The key information conversion mechanism is used to determine the key information of the data to be identified in different formats.
[0072] The principle behind the above technical solution is as follows:
[0073] This application first determines the data flow direction of the data to be identified. Data flow direction refers to the direction and path of data transmission within the system and network. It primarily involves the data generated by the data source and the process of changes in the relationships and connections between data items during transmission to the business end. This is used to determine the changes, direction, relationships, and hierarchy of the data. Data flow direction includes circular and sequential flows. This process also helps determine the data type, data direction, and data attributes.
[0074] This application's data lineage mapping identifies data transmission paths, transformation processes, and processing procedures, and can be applied to various data domains. For example, it can help financial institutions such as banks and securities companies better understand customer needs and risk preferences, thereby developing more precise marketing strategies and risk control measures. It can also help doctors better understand patient conditions, diagnostic results, and treatment methods, improving diagnostic accuracy and treatment effectiveness.
[0075] Content recognition identifies keywords or key information in the data to be recognized, while lineage can only determine the data relationship, data type, and data flow.
[0076] However, existing content recognition technologies, if they only identify key information such as keywords, can only determine the core content at the data source generation stage or after processing. That is, they can determine the core content of the data at any step in the initial stage or at the final processing result. At different processing stages, a new content recognition network is needed to perform content recognition, resulting in a large waste of resources. The method of this application can achieve core content recognition of data at the initial stage. In the subsequent transmission process, only the data conversion through relation identifiers is needed to directly determine the key information, instead of setting up a content recognition mechanism again to identify the corresponding key information.
[0077] After determining the lineage of the data and the specific key content, the relationship between the key information can be determined through relationship identifiers. Then, when key content needs to be identified during the continued data transmission, the data is converted through the encoding of the relationship identifiers, rather than through data identification.
[0078] As attached Figure 1 As shown, in the process of data identification, this application first determines the flow relationship between data based on the data flow direction of the data to be identified, and determines the data lineage map. Through the data lineage map, under the condition that the relationship between different data is determined, the content of the data is identified, the relationship identifier is determined, the data is sorted into serialized sequences according to the data relationship identifier, and then the key information is directly converted out through the relationship identifier, so as to realize the extraction of key information from the serialized data.
[0079] The beneficial effects of the above technical solution are as follows:
[0080] This application uses a bloodline map to collect bloodline relationships. In this process, the bloodline relationships enable logical and sequential data sorting, and key information can be identified through the sorting method.
[0081] To avoid the need for separate key information identification technology when processing critical information on different data processing platforms, a significant amount of network resources can be saved.
[0082] Specifically, the data flow for acquiring the data to be identified includes:
[0083] Acquire the data to be identified, and determine the data address and data path;
[0084] Based on the data address and data path, the data to be identified is discretized to determine the target value range;
[0085] Based on the target value range, determine the height information of the bounding box corresponding to each value range of each target indicator;
[0086] Based on the order of target indicators, the height of the bounding box, and the range of target indicator values, a target flow diagram of the data to be identified is constructed.
[0087] Based on the target flow diagram, perform data flow statistics.
[0088] The principle behind the above technical solution is as follows:
[0089] In determining the data flow direction, this application can construct a relationship graph in the data transmission process by using data addresses and data paths, thereby determining the data transformation relationships and data change information at different nodes in the data transmission process.
[0090] Discretization can reduce data complexity and improve data processing efficiency. The target value range determined in this application is the range of values of the target variable between the independent variable and the target variable of the data to be identified during data processing. The target variable is the range of values of the change in the data during transmission and conversion, including but not limited to format conversion, data filtering, data splitting, etc.
[0091] The target indicator is the range of indicator intervals for different data nodes (e.g., after format conversion or different data platforms) during the data lineage processing. The range of intervals can determine the magnitude of data changes, thereby determining the size of the lineage at the current node. Then, by using the target indicator arrangement order, interval box height information, and target indicator value range, a target flow diagram of the data to be identified is generated. The flow target is not unique, realizing global statistics on the data flow of the data to be identified.
[0092] The beneficial effects of the above technical solution are as follows:
[0093] As attached Figure 2 As shown, this application can perform data discretization calculation through data address and data transmission path. During the discretization calculation process, the target flow graph of the data is constructed through the index range and height information of each data, thereby realizing data flow statistics.
[0094] Specifically, the data lineage maps to be constructed include:
[0095] The data to be identified is classified and parsed, and the table fields for each piece of data to be identified are determined, as well as the dependency information of the table fields.
[0096] Import table field dependency information into the preset target data table;
[0097] Call the preset graphics database;
[0098] Import the field dependency information of the target data table into the graph database;
[0099] By performing graph calculations on different types of data using a graph database, corresponding data lineage maps are generated.
[0100] The principle behind the above technical solution is as follows:
[0101] By classifying and analyzing the data to be identified, we can determine the different types of data. Based on the data type, we can determine the dependency information of different table fields in the data to be identified. The table field dependency information is the logical relationship between different fields after the data to be identified is tabulated. For example, in an order table, when different product types are selected, the values of the order amount or discount fields will change accordingly. This can help reduce user input errors and improve data quality.
[0102] The target data table is a classification chart of the data to be identified, used to determine the relationships between different data within the data to be identified;
[0103] The pre-defined graph database is a type of graph data, which is used to determine the relationships between data in a graphical form through nodes (data entities) and edges (data associations, such as adjacency matrices or adjacency tables).
[0104] Graph computation uses graph algorithms (such as graph centering algorithms, graph community discovery algorithms, graph path discovery algorithms, graph similarity algorithms, etc.) to determine the relationships between different data to be identified, thereby identifying important information and features, and generating corresponding kinship graphs based on the important information and features.
[0105] The beneficial effects of the above technical solution are as follows:
[0106] The kinship graph of this application is based on the classification and parsing of data, which realizes the processing of information from the table fields of the data, thereby constructing a corresponding graphical database. Through the processing of the graphical database, the kinship relationship between different data can be calculated.
[0107] Specifically, the content recognition includes:
[0108] Image recognition, audio recognition, and text recognition.
[0109] The principle behind the above technical solution is as follows:
[0110] In terms of identification methods, this application mainly includes image recognition, audio recognition, and text recognition of data; the fields of image recognition include, but are not limited to, remote sensing images, real-time communication video television and teleconference, as well as cells, X-ray images, etc. in clinical diagnosis in biological hospitals;
[0111] Audio recognition includes, but is not limited to, voice control in smart home technology, voice interaction in automotive, voice control in industrial control and social fields, music recognition, language translation, and so on.
[0112] Text recognition includes, but is not limited to, text recognition on images, text retrieval, and public opinion monitoring and caption recognition in the social media field.
[0113] Specifically, the image recognition;
[0114] Set automatic image recognition mode and manual recognition mode;
[0115] Based on the automatic recognition mode, static elements and dynamic elements are selected in the image, and the background color and foreground color are separated.
[0116] Based on the element outlines of the selected static and dynamic elements, set the bounding box; where,
[0117] The bounding box is used not only for element labeling but also for phylogenetic encoding of kinship maps;
[0118] The maximum and minimum area thresholds are determined based on the dimensions of the element outlines.
[0119] Static and dynamic elements are identified and calibrated using calibration boxes based on the maximum and minimum area thresholds.
[0120] The principle behind the above technical solution is as follows:
[0121] In this application, image recognition involves setting up both an image recognition model and a manual recognition model. The automatic recognition model filters static elements (elements that remain stationary in the background image) and dynamic elements (non-background elements and dynamically moving elements) to remove background color and identify the corresponding dynamic elements. Corresponding calibration boxes are then set for calibration, using the maximum and minimum areas of the calibration boxes to identify static and dynamic elements. In this application, due to the existence of a kinship graph, in addition to content recognition calibration, kinship graph calibration is also required. Calibration boxes can be used for both content recognition and kinship graph calibration, with each box assigned a code for rapid kinship relationship localization. In manual recognition mode, corresponding elements are identified and extracted according to user instructions.
[0122] The beneficial effects of the above technical solution are as follows:
[0123] This application can recognize the content of the data to be recognized through image recognition, and can also encode the vascular relationship map through the bounding box. Through the map encoding, it can also perform the fixed-point recognition of the blood relationship map.
[0124] Specifically, the audio recognition includes:
[0125] Obtain the audio segment to be identified;
[0126] Collect keywords from audio clips to determine the context in which the audio clips occur;
[0127] Based on the context in which the event occurs, the corresponding sound event is identified using the appropriate sound recognition model.
[0128] Based on the sound event identifier, perform text-based recognition of audio segments.
[0129] The principle behind the above technical solution is as follows:
[0130] In terms of audio data recognition, this application first collects keywords from the audio data, then determines the occurrence scenario of the audio segment, and identifies the audio data based on the occurrence scenario to generate corresponding sound recognition tags. Sound event tags include, but are not limited to, occurrence scenario tags, tone tags, and key audio tags. Thus, when generating a kinship map, the audio data can use these tags to determine the corresponding event information in the kinship map, such as audio from smart homes, security monitoring, and intelligent transportation. Furthermore, through the processing of event information, i.e., audio characteristics and other information, the location where the data was generated is determined. During transmission, the transmission method is determined, and text-based recognition is used to make audio recognition more accurate.
[0131] Specifically, the text recognition includes:
[0132] Automatic feature extraction and character encoding of text images by moving the cursor determine text image features and character position encoding features;
[0133] The text image features are sampled based on the character position encoding features to obtain the sampled features in the text image;
[0134] The sampled features are input into the text large language model, and the corresponding text information is output.
[0135] The principle behind the above technical solution is as follows:
[0136] This application employs a moving cursor technique for automatic sampling in text content recognition. This moving cursor is an automated tool used for automatic feature extraction and character encoding of text information. Its function is to guide text characters within the text image. Based on this guidance, features of the text region and content are identified, automatically extracting and determining the corresponding text image features and character position codes. These features are then sampled and input into a Large Language Model (LLM) to output the text. A Large Language Model (LLM) is a deep learning model trained on a large amount of text data that can generate natural language text or understand the meaning of spoken text. In this application, not only can the meaning of the text be determined, but text translation and marking of corresponding key text information are also achieved.
[0137] Specifically, determining the relationship identifier of the data to be identified includes:
[0138] Based on the data lineage map, determine the relationship vector between different data in the data to be identified;
[0139] Calculate the association depth between different data to be identified based on the relationship vector;
[0140] Based on the association depth, calculate the relationship identifier between different data to be identified.
[0141] The principle behind the above technical solution is as follows:
[0142] In calculating relationship identifiers, kinship maps can determine relationship vectors within the data to be identified. These relationship vectors are used to determine the similarity and association depth between different data within the data to be identified, and are primarily based on the calculation of feature vectors.
[0143] Association depth is an indicator that measures the degree of correlation between different data in the data to be identified. It can be used as a guiding indicator for key information. When identifying key information at different nodes, key information can be transformed by using relation identifiers and corresponding codes. As long as they have the same relation characteristics, key information can be identified through these relation characteristics instead of re-identifying the data.
[0144] The beneficial effects of the above technical solution are as follows:
[0145] The relationship identifier in this application can be encoded to identify key information related to data that have the same relationship characteristics, thereby enabling the identification of key information related to the relationship.
[0146] Specifically, the key information in the data to be identified based on the relationship identifier includes:
[0147] Obtain relationship identifiers and perform serialization and sorting on different data to be identified;
[0148] Based on the serialization sorting, key data are extracted from each piece of data to be identified in sequence, and data nodes are set;
[0149] Calculate the data weight of each piece of data to be identified using data nodes;
[0150] Based on data weights, key data are serialized and sorted to generate multiple datasets;
[0151] The dataset is used as input to the information recognition model to obtain a set of key information about the data to be recognized.
[0152] The key information set is converted into text data to generate the key information data text.
[0153] The principle behind the above technical solution is as follows:
[0154] When performing serialization and sorting through relationship identifiers, this application sorts the serialization based on key data in the data to be identified. This allows for the determination of different data nodes in the vascular relationship. Furthermore, by using these data nodes, the data weights are determined during the processing of the data to be identified, enabling the identification, processing, and text conversion of key information. This generates corresponding data text, facilitating the analysis of blood relations and the textual display of key content.
[0155] The beneficial effects of the above technical solution are as follows:
[0156] This application uses relation identifiers to construct key data and data nodes, and then uses data nodes and data weights to serialize and sort the key data of the data to be identified, thereby realizing the transformation and integration of key data and generating a data set.
[0157] Specifically, the method further includes:
[0158] Identify key information and generate information features based on that key information;
[0159] Based on the information characteristics, a logical feature representation diagram is constructed; where...
[0160] The logical feature representation diagram contains logical vectors between information features;
[0161] Based on the logical feature representation diagram, key information is sorted to determine the content sequence of key information.
[0162] The principle behind the above technical solution is as follows:
[0163] As attached Figure 3As shown, this application will determine the logical features between different key information, sort the key information through logical features, and thus determine whether there is a problem of insufficient correlation in the logical relationship of the data to be identified.
[0164] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A data identification method based on kinship and data identification, characterized in that, include: Obtain the data flow direction of the data to be identified and construct a data lineage map; Based on data lineage mapping and content recognition, the relationship identifiers of the data to be identified are determined; among which, content recognition includes: image recognition, audio recognition and text recognition, and the relationship identifiers are the relationship identifiers of key information in the data to be identified; The image recognition includes: setting a calibration box, which is used for element calibration and determining the map code of the kinship map; and identifying and calibrating static and dynamic elements through the calibration box. The audio recognition includes: determining a sound event identifier based on the occurrence scenario of the audio segment to be recognized, and performing text-based recognition of the audio segment using the sound event identifier; The text recognition includes: automatically extracting features and encoding characters from a text image by moving the cursor, determining the sampling features of the text image, inputting the sampling features into a large text language model, and outputting the corresponding text information; The relationship identifier includes: calculating the association depth between different data to be identified based on the relationship vector; Based on the association depth, calculate the relationship identifiers between different data to be identified; Calculate the data weight of each piece of data to be identified using data nodes; Based on data weights, key data are serialized and sorted to generate multiple datasets; The dataset is used as input to the information recognition model to obtain a set of key information about the data to be recognized. Based on the relationship identifier, a key information conversion mechanism is generated; among which... The key information conversion mechanism is used to determine the key information of the data to be identified in different formats; The step of determining the relationship identifier of the data to be identified based on the data lineage map and content recognition includes: determining the flow relationship between data based on the data flow direction of the data to be identified, and determining the data lineage map; determining the relationship between different data through the data lineage map, and performing content recognition of the data to be identified to determine the relationship identifier; The relationship vector represents the relationship between different data points of the data to be identified, as determined by the data lineage map. The method of determining the key information in the data to be identified based on the relationship identifier also includes: Obtain relationship identifiers and perform serialization and sorting on different data to be identified; Based on the serialization sorting, key data are extracted from each piece of data to be identified in sequence, and data nodes are set; The key information set is converted into textual data to generate data text of the key information; The method further includes: Identify key information and generate information features based on that key information; Based on the information characteristics, a logical feature representation diagram is constructed; where... The logical feature representation diagram contains logical vectors between information features; Based on the logical feature representation diagram, key information is sorted to determine the content sequence of key information.
2. The data identification method based on kinship and data identification as described in claim 1, characterized in that, The data flow for acquiring the data to be identified includes: Acquire the data to be identified, and determine the data address and data path; Based on the data address and data path, the data to be identified is discretized to determine the target value range; Based on the target value range, determine the height information of the bounding box corresponding to each value range of each target indicator; Based on the order of target indicators, the height of the bounding box, and the range of target indicator values, a target flow diagram of the data to be identified is constructed. Based on the target flow diagram, perform data flow statistics.
3. The data identification method based on kinship and data identification as described in claim 1, characterized in that, The construction of the data lineage map includes: The data to be identified is classified and parsed, and the table fields for each piece of data to be identified are determined, as well as the dependency information of the table fields. Import table field dependency information into the preset target data table; Call the preset graphics database; Import the field dependency information of the target data table into the graph database; By performing graph calculations on different types of data using a graph database, corresponding data lineage maps are generated.
4. The data identification method based on kinship and data identification as described in claim 1, characterized in that, The image recognition also includes: Set automatic image recognition mode and manual recognition mode; Based on the automatic recognition mode, static and dynamic elements are selected in the image, and the background and foreground colors are separated; among them, The calibration box is set by the element outlines of the selected static and dynamic elements; The dimensions of the element outline are determined by the maximum and minimum area thresholds of the static elements and the dynamic elements.
5. The data identification method based on kinship and data identification as described in claim 1, characterized in that, The determination of the sound event identifier includes: Collect keywords from audio clips to determine the context in which the audio clips occur; Based on the context of the event, the corresponding sound event identifier is determined using the appropriate sound recognition model.
6. The data identification method based on kinship and data identification as described in claim 1, characterized in that, The determination of sampling features in the text image includes: Determine text image features and character position encoding features; The text image features are sampled based on the character position encoding features to obtain the sampled features in the text image.
Citation Information
Patent Citations
Data management method and system based on data consanguinity analysis
CN112800149A
Method and system for processing transaction information
CN103679435A
Data flow direction display method and device and electronic equipment
CN115862802A
Method, device and equipment for identifying abnormal privacy attribute information
CN116756762A
Blood relationship analysis method and device, computer equipment and storage medium
CN116842011A