Big data management traceability method and system based on ElasticSearch
By automatically judging data types and optimizing data push process on the ElasticSearch platform, the problem of accurate restoration of a single data history in the prior art is solved, and data access efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202510145110.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Existing data traceability technologies are difficult to meet the precise restoration needs of a single data history in dynamic data access and traceability scenarios, especially in scenarios such as multi-source data integration and large-scale data flow.
The big data governance traceability method based on ElasticSearch is adopted. By connecting the original data to ElasticSearch, and selecting full access, incremental access or timestamp-trigger access according to the data type, combining algorithms such as the ratio of difference and similarity and field overlap index, the data type is automatically judged and the data push process is optimized.
It realizes accurate restoration of a single data history in complex data flow scenarios, improves the efficiency of data access, reduces system complexity, and ensures the flexibility and reliability of data updates and access.
Smart Images

Figure CN120067087A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data management, and in particular, to a big data governance traceability method and system based on ElasticSearch. Background Art
[0002] With the advent of the data-driven era, data has become the core resource for enterprise decision-making, operation, and innovation. With the continuous development of big data technology, the volume of data generated, storage requirements, and processing complexity have increased rapidly. At the same time, the complexity of data management and governance has gradually increased. Especially in terms of data security, compliance requirements, and data flow transparency, more and more enterprises have begun to attach importance to data traceability. To meet the industry's needs for data governance, data traceability technology has emerged. It ensures the transparency of the source and destination of data by tracking the flow path and change history of data, providing important support for enterprises in data management and other aspects.
[0003] The existing technology mainly adopts the method of tracking data lineage. By recording the flow trajectories between data types, it provides a visual display of the source and destination of data. This technology has certain application value in scenarios such as data governance and data exchange and is widely used in traditional data management systems. By constructing a lineage graph, the existing technology can help enterprises understand the data flow path and support basic data governance and decision-making analysis requirements.
[0004] However, although traditional data traceability methods can track data lineage at the type level, they rely on the architecture design of relational databases or data warehouses, usually focusing on static analysis at the macro level and lacking the ability to dynamically and accurately monitor the changes of individual data records. This is because the lineage graph relied on by traditional technologies only records the node states of data at different stages and fails to record and track the real-time changes of data between nodes in detail. For example, in scenarios of multi-source data integration or large-scale data flow, the dynamic changes of data may be frequent and complex, and the existing technology cannot capture every subtle change process, making it difficult to meet the requirement of accurately restoring the history of individual data. Summary of the Invention
[0005] This application provides a big data governance traceability method and system based on ElasticSearch, which can meet the requirement of accurately restoring the history of individual data in complex data flow scenarios. This application provides the following technical solutions:
[0006] In a first aspect, this application provides a big data governance traceability method based on ElasticSearch, and the method includes:
[0007] Connect the original data to ElasticSearch, and prepare to access the latest pushed data in response to the subsequent update requirements of the original data;
[0008] Select an access method based on the type of the latest pushed data, and the access methods include full-volume access method, incremental access method, and timestamp-trigger access method;
[0009] Clean the latest pushed data based on the access method of the latest pushed data;
[0010] Connect the latest pushed data after data cleaning to ElasticSearch according to the selected access method;
[0011] In response to the request for data traceability, perform data traceability operations based on ElasticSearch.
[0012] In a specific implementable solution, the selection of the access method based on the type of the latest pushed data includes:
[0013] Automatically determine whether the type of the latest pushed data is full-volume data or incremental data, and select a suitable access method according to the different types of the latest pushed data. The access methods include the following three:
[0014] Full-volume access method: When the latest pushed data is full-volume data, including complete records of business data, use the full-volume access method;
[0015] Incremental access method: When the latest pushed data is incremental data, including only newly added or updated records, use the incremental access method;
[0016] Timestamp-trigger access method: When the latest pushed data is incremental data updated based on the timestamp or trigger mechanism.
[0017] In a specific implementable solution, the automatic determination of whether the type of the latest pushed data is full-volume data or incremental data includes:
[0018] Introduce the ratio of difference degree and similarity degree, and consider the change of data volume for judgment. Use the following formula to calculate the change index V of the latest pushed data and the original data change :
[0019]
[0020] where N origin is the number of fields of the original data, N push is the number of fields of the latest pushed data, and S sim is the cosine similarity value between the latest pushed data and the original data;
[0021] Compare the calculated change index V change with a preset threshold. If the change index is greater than or equal to the preset threshold, it indicates that the difference between the original data and the latest pushed data is greater, and the latest pushed data is considered incremental data; if the change index is less than the preset threshold, it indicates that the difference between the original data and the latest pushed data is smaller, and the latest pushed data is considered full - volume data.
[0022] In a specific feasible implementation, the automatic determination of whether the latest pushed data is full - volume data or incremental data further includes:
[0023] Use the following formula to calculate the field overlap index V of the latest pushed data and the original data type :
[0024]
[0025] where N total is the total number of fields in the latest pushed data, N same is the number of identical fields in the latest pushed data and the original data, and the calculation formula of N same is as follows:
[0026] N same = |{F push} ∩ {F origin}|
[0027] where F push is the field set of the latest pushed data, F origin is the field set of the original data, and |·| represents the number of elements in the set;
[0028] The larger the calculated field overlap index V type , the higher the field overlap degree. When the field overlap index is greater than or equal to the preset threshold, it is determined that the latest pushed data is full - volume data; the smaller the calculated field overlap index V type , the lower the field overlap degree. When the field overlap index is less than the preset threshold, it is determined that the latest pushed data is incremental data.
[0029] In a specific feasible implementation, calculating the change index V of the latest pushed data and the original data change is suitable for scenarios where both the data content and the field structure have complex changes, and is applicable to environments that require a comprehensive consideration of data change situations; calculating the field overlap index V of the latest pushed data and the original data type is applicable to scenarios that focus on changes in the field set structure rather than content changes; select an appropriate judgment method according to different scenarios.
[0030] In a specific implementable embodiment, the data cleaning of the latest push data based on the access method of the latest push data includes:
[0031] For the latest push data of the full-volume access method, the latest push data is full-volume data, and the parts of the data that are added, deleted, or modified in the latest push data compared with the original data are identified as the cleaned latest push data;
[0032] For the incremental access method or the timestamp-trigger access method, the latest push data is incremental data, and there is no need to perform data cleaning on the latest push data.
[0033] In a specific implementable embodiment, after the original data is accessed to ElasticSearch, it further includes:
[0034] Create a temporary backup table, store the current original data in the temporary backup table, and keep the temporary backup table for comparison with the subsequent latest push data.
[0035] In a second aspect, the present application provides a big data governance traceability system based on ElasticSearch, adopting the following technical solutions:
[0036] A big data governance traceability system based on ElasticSearch includes:
[0037] An original data access module, configured to access the original data to ElasticSearch and prepare to perform the access operation of the latest push data in response to the subsequent update requirement of the original data;
[0038] An access method selection module, configured to select an access method based on the type of the latest push data, and the access methods include a full-volume access method, an incremental access method, and a timestamp-trigger access method;
[0039] A data cleaning module, configured to perform data cleaning on the latest push data based on the access method of the latest push data;
[0040] A push data access module, configured to access the latest push data after data cleaning to ElasticSearch according to the selected access method;
[0041] A data traceability module, configured to perform data traceability operations based on ElasticSearch in response to a data traceability request.
[0042] In a third aspect, the present application provides an electronic device, which includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a big data governance traceability method based on ElasticSearch as described in the first aspect.
[0043] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and when the program is executed by a processor, it is used to implement a big data governance traceability method based on ElasticSearch as described in the first aspect.
[0044] In summary, the beneficial effects of the present application at least include:
[0045] 1) By selecting different access methods (full-volume access, incremental access, timestamp trigger access), combining algorithms such as similarity analysis and field overlap index, automatically judge the data type and optimize the data push process, thereby improving the data access efficiency, reducing system complexity, and ensuring more flexible and reliable data update and access.
[0046] 2) Based on the powerful query and retrieval capabilities of ElasticSearch, when the system receives a data traceability request, it can quickly locate the specific content of data changes through an accurate query mechanism and generate a clear traceability report, ensuring the transparency and integrity of the data flow path and change history, and supporting the efficient execution of enterprises in data management and compliance requirements.
[0047] Through the optimization of the data access judgment method, the flexible design of the cleaning mechanism, and the accurate implementation of traceability queries, a traceability solution that can efficiently handle dynamic and complex data environments is gradually derived. Facing full-volume and incremental data, as well as data updated by timestamp or trigger mechanisms, existing solutions are difficult to efficiently and accurately judge the data type and select the appropriate access method, which easily leads to incorrect access or redundant data processing. The present application combines the change index and the field overlap index to achieve accurate judgment of the data type. Adjust the cleaning strategy according to the data type to avoid unnecessary calculations. Introduce the collaborative operation of the temporary backup table and ElasticSearch to ensure the accuracy of the traceability data. Gradually overcome the deficiencies of the prior art in dynamic data access and traceability scenarios, and provide a comprehensive, flexible and efficient solution.
[0048] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly and implement it in accordance with the content of the specification, the following takes the preferred embodiments of the present application and combines the accompanying drawings to describe in detail as follows. Description of the Drawings
[0049] Figure 1It is a schematic flowchart of the big data governance traceability method based on ElasticSearch in the embodiments of the present application.
[0050] Figure 2 It is the overall flowchart of the big data governance traceability method based on ElasticSearch in the embodiments of the present application.
[0051] Figure 3 It is a structural block diagram of the big data governance traceability system based on ElasticSearch in the embodiments of the present application.
[0052] Figure 4 It is a block diagram of the electronic device for big data governance traceability based on ElasticSearch in the embodiments of the present application. Detailed implementation manners
[0053] The following will further describe in detail the specific implementation manners of the present application with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0054] Optionally, the present application takes the big data governance traceability method based on ElasticSearch provided in each embodiment as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.
[0055] Refer to Figure 1 , which is a schematic flowchart of the big data governance traceability method based on ElasticSearch provided in an embodiment of the present application. The method at least includes the following steps:
[0056] Step S101: Connect the original data to ElasticSearch and prepare to perform the access operation of the latest push data in response to the update requirement of the subsequent original data.
[0057] In step S101, it is first necessary to comprehensively identify the original data, which includes the source, format, structure, and specific information contained in the original data. Subsequently, before the original data is officially connected to ElasticSearch, data preprocessing operations are performed, including data cleaning, format conversion, and data mapping. Among them, data cleaning aims to remove redundant, incorrect, or inconsistent data to ensure the accuracy and integrity of the data. Format conversion is to convert the original data into a format that ElasticSearch can efficiently store and query. Data mapping defines the correspondence between the original data fields and the index fields in ElasticSearch. The preprocessed original data will be directly connected to ElasticSearch.
[0058] In implementation, after the access of the original data is completed, the dynamic update requirements of the original data are further supported. As the business data continuously changes, such as adding, modifying, or deleting records, the access operation of the latest push data is prepared, and the changed content is synchronized to ElasticSearch.
[0059] Step S102: Select an access method based on the type of the latest push data. The access methods include the full-volume access method, the incremental access method, and the timestamp-trigger access method.
[0060] In step S102, during the data access process, select an appropriate access method according to the different types of the latest push data. The access methods include the following three:
[0061] Full-volume access method: When the latest push data is full-volume data, including the complete records of the business data, the full-volume access method is used. This method is applicable to the situation where the data volume is small or the overall update is required.
[0062] Incremental access method: When the latest push data is incremental data, including only the newly added or updated records, the incremental access method is used. This method is applicable to scenarios where the data volume is large and the change frequency is low, such as log data, real-time sensor data, etc.
[0063] Timestamp-trigger access method: When the latest push data is incremental data updated based on the timestamp or trigger mechanism. This method is applicable to scenarios where data updates are triggered according to time or events.
[0064] In implementation, use the following method to automatically determine whether the type of the latest push data is full-volume data or incremental data:
[0065] Introduce the ratio of the difference degree and the similarity degree, and at the same time consider the change of the data volume to make a final judgment. Use the following formula to calculate the change index V of the latest push data and the original data change :
[0066]
[0067] where N origin is the number of fields of the original data, N push is the number of fields of the latest push data, and S sim is the cosine similarity value between the latest push data and the original data. If all the data pushed is full-volume data, it will overwrite the original data. Therefore, although the number of fields may change, the similarity is high and the difference degree is small. If it is incremental data, only the changed data is pushed, and the number of fields may be quite different, so the difference degree will be large. The calculated change index V changeCompare with the preset threshold. If the change index is greater than or equal to the preset threshold, it indicates that the difference between the original data and the latest pushed data is greater, and the latest pushed data is considered incremental data. If the change index is less than the preset threshold, it indicates that the difference between the original data and the latest pushed data is smaller, and the latest pushed data is considered full-volume data.
[0068] In the above process, by combining the similarity, the difference in the number of fields, and the ratio of the number of fields into a single change index, the judgment process of data access is simplified. This method avoids complex multi-dimensional calculations and instead considers multiple factors together, making the data judgment more efficient. The formula not only depends on the similarity (measuring the change in data content), but also combines the difference in the number of fields (measuring the structural change) and the ratio of the number of fields (measuring the scale difference of the data). This comprehensive consideration makes the change index more representative and can comprehensively reflect the change of the data, thus making the judgment more accurate. Relying solely on similarity or the difference in the number of fields may lead to misjudgment. For example, full-volume data may have a large difference in the number of fields, but due to the complete replacement of the original data by the data content, the similarity is high and the difference is small; incremental data may have a relatively unchanged number of fields, but a large difference in data content. Therefore, by combining the ratio of similarity and the difference in the number of fields, the judgment error caused by relying on only one of them can be avoided.
[0069] In another feasible embodiment, the following method can also be used to automatically determine whether the type of the latest pushed data is full-volume data or incremental data:
[0070] Specifically, use the following formula to calculate the field overlap index V of the latest pushed data and the original data type :
[0071]
[0072] where N total is the total number of fields of the latest pushed data, N same is the number of the same fields in the latest pushed data and the original data, and the calculation formula of N same is as follows:
[0073] N same =|{F push}∩{F origin}|
[0074] where F push is the field set of the latest pushed data, F origin is the field set of the original data, and |·| represents the number of elements in the set. The calculated field overlap index V typeThe larger it is, the higher the field overlap degree. When the field overlap index is greater than or equal to the preset threshold, it is determined that the latest pushed data is full - volume data. The calculated field overlap index V type The smaller it is, the lower the field overlap degree. When the field overlap index is less than the preset threshold, it is determined that the latest pushed data is incremental data.
[0075] In the above process, the formula directly starts from the intersection of the field sets and determines the data type through the overlap degree of the fields, avoiding complex multi - dimensional parameter combinations and making the calculation logic simple and easy to understand. The field overlap index V type can clearly reflect the essential differences in the data push types (full - volume or incremental). The formula takes into account the proportional relationship of the number of fields, and can maintain the stability and universality of the field overlap index even when the data scales are different.
[0076] In summary, the first embodiment considers similarity (measuring data content changes), field quantity differences (measuring structural changes), and field ratios (measuring data scale differences). For full - volume data with large field quantity differences but highly similar data content (such as full - volume push during field expansion), it can be correctly identified as full - volume data. For incremental data with small field quantity changes but large content differences (such as partial update push), it can be effectively distinguished. The applicable scenarios are those where field quantity, data content, and structural changes need to be considered, such as cross - platform data synchronization or data version control, or scenarios where high accuracy is required for judgment when the data content changes significantly. The second embodiment only relies on the overlap relationship of the field sets, avoiding content - level calculations, with simple logic and easy to implement. Even when the data scales are different, the formula maintains stability through the field ratio relationship and does not depend on field values or their specific content. The applicable scenarios are those such as scenarios where interface fields are frequently updated or data push only involves structural changes, and in systems with limited computing resources where quick data type judgment is required. That is, the first embodiment is more suitable for scenarios with complex changes in both data content and field structure, and is applicable to high - requirement environments that need to comprehensively consider data change situations. The second embodiment is more applicable to scenarios that focus on structural changes in the field set rather than content changes, and is suitable for simple and efficient judgment requirements. This application can select appropriate judgment methods according to different scenarios.
[0077] In implementation, if it is determined that the latest pushed data is incremental data, it is judged whether there is incremental data updated based on the timestamp or trigger mechanism in the incremental data. Specifically, if the data table contains a timestamp field, the value of the timestamp can be compared to determine whether the data is updated through the timestamp mechanism. Usually, the timestamp reflects the time when the data was last modified. When there is a significant difference between the timestamp of the pushed data and the timestamp of the original data, it indicates that the data is updated through the timestamp mechanism. Or, it is identified by checking the operation type and source of the updated data. For example, if the update operation of a certain piece of data is triggered by a specific trigger (such as a "data change trigger"), it can be determined as data updated by the trigger mechanism.
[0078] Step S103: Perform data cleaning on the latest pushed data based on the access method of the latest pushed data.
[0079] In step S103, for the access method selected in step S102, the full-volume access method, the incremental access method, or the timestamp-trigger access method, different cleaning operations are performed on the latest pushed data.
[0080] In implementation, for the latest pushed data with the full-volume access method, the latest pushed data is full-volume data. At this time, the parts of the data that have increased, been deleted, or modified compared with the original data are identified as the latest pushed data after cleaning. Specifically, in order to identify the added, deleted, and modified parts of the data between the latest pushed data and the original data under the full-volume access method, the fields that do not exist in the original data in the latest pushed data are checked as newly added data, the fields that do not exist in the latest pushed data in the original data are checked as deleted data, and the fields that exist in both data sets are checked, and whether their values have changed is compared. A change is a modified data.
[0081] In implementation, for the incremental access method or the timestamp-trigger access method, the latest pushed data is all incremental data. At this time, there is no need to perform cleaning processing on the latest pushed data.
[0082] Step S104: Connect the latest pushed data after data cleaning to ElasticSearch according to the selected access method.
[0083] In step S104, after the latest pushed data has undergone data cleaning, next, according to the previously selected access method (full-volume access, incremental access, or timestamp-trigger access), the cleaned data is pushed to ElasticSearch.
[0084] Step S105: In response to a data traceability request, perform a data traceability operation based on ElasticSearch.
[0085] In step S105, when a data traceability request is received, a traceability operation will be performed based on ElasticSearch. This operation will utilize the data cleaned and accessed in steps S103 and S104 to ensure the accuracy of the data and the integrity of traceability.
[0086] Specifically, a user or system will initiate a data traceability request. This request may require tracing the historical changes of a certain data item, the source of the modification, or the change history of the data. This request will be closely related to the data stored in ElasticSearch in step S104 to ensure that the cleaned data can be traced. Based on the powerful query engine of ElasticSearch, the system will perform precise retrieval in ElasticSearch according to the conditions in the traceability request (such as data identifier, timestamp, version number, etc.). In step S103, the differences (additions, deletions, modifications) between the original data and the latest pushed data have been identified through the cleaning process. The system will perform efficient queries based on this cleaned data to ensure the accuracy and integrity of the query results. After ElasticSearch completes the data query, the system will generate and return detailed traceability results, including: comparing the original data and the latest pushed data, listing all modified, deleted, and added fields. Providing a timeline of data changes, showing the timestamp and updated content of each data update.
[0087] It should be noted that since the update of the original data exists dynamically, that is, there are multiple updates, the description in this application only takes the first update as an example. After that, the latest pushed data is compared with the data of the most recent update to determine the data type, access method, or data changes.
[0088] In another feasible embodiment, after the original data is accessed to ElasticSearch, a backup of the original data is made. A temporary backup table is created, and the current original data is stored in the temporary backup table and the table is retained for comparison with the latest pushed data later. Thus, when the data is accessed, it is possible to effectively identify the added, deleted, and modified parts as much as possible, prevent data loss or errors, and at the same time provide effective historical data support for data traceability.
[0089] To sum up, combined with Figure 2, this application solves the problems of insufficient ability to track dynamic data updates and difficulty in accurate traceability in the prior art by designing an efficient data access, processing, and traceability process. The method includes multiple key steps: from raw data access to data cleaning, and then to data access and traceability operations, forming a complete closed loop. In the data access stage, this application ensures the quality of data access to ElasticSearch by comprehensively identifying the source, structure, and content of raw data and performing preprocessing (such as cleaning, format conversion, and field mapping). This design lays the foundation for subsequent data processing and querying, and helps reduce the impact of redundant data on storage performance. In terms of the access of the latest pushed data, this application proposes three methods: full-volume access, incremental access, and timestamp-trigger access, combined with two judgment methods, the data change index (V_change) and the field overlap index (V_type), to effectively distinguish data types. By comprehensively considering similarity, the difference in the number of fields and their ratio, or directly analyzing the overlap degree of the field set, the solution can accurately judge the type of pushed data, adapt to complex and changing business scenarios, and improve the efficiency and accuracy of access method selection. For data cleaning operations, this application designs different processing methods for full-volume data and incremental data. The cleaning of full-volume data ensures data integrity by comparing newly added, deleted, and modified fields; while incremental data simplifies the cleaning process and improves processing efficiency. This targeted design not only optimizes the data access process but also provides a clear record of data changes for subsequent traceability operations. Finally, in the data traceability stage, based on the powerful retrieval ability of ElasticSearch and combined with the data difference records after cleaning, it is possible to achieve accurate querying and tracing of the historical changes of data. By providing detailed traceability results (such as field change records, change timelines, etc.), the solution solves the deficiencies of the prior art in the dynamic monitoring and historical restoration of single data items.
[0090] In summary, through the optimization of the data access judgment method, the flexible design of the cleaning mechanism, and the accurate implementation of the traceability query, this application gradually derives a traceability solution that can efficiently handle dynamic and complex data environments. Facing full-volume and incremental data, as well as data updated by the timestamp or trigger mechanism, existing solutions are difficult to efficiently and accurately judge the data type and select the appropriate access method, which easily leads to incorrect access or redundant data processing. This application combines the change index and the field overlap index to achieve accurate judgment of the data type. Adjust the cleaning strategy according to the data type to avoid unnecessary calculations. Introduce the collaborative operation of the temporary backup table and ElasticSearch to ensure the accuracy of the traceability data. Gradually overcome the deficiencies of the prior art in the scenarios of dynamic data access and traceability, and provide a comprehensive, flexible, and efficient solution. By introducing a temporary backup table and a difference analysis mechanism, this application backs up the original data and generates a structured comparison result during each data push, thereby recording and tracking the detailed history of data changes. The design of the temporary backup table ensures that the add, delete, and modify operations between the original data and the latest pushed data are accurately identified. At the same time, combined with the efficient query ability of ElasticSearch, the system can dynamically record and retrieve the complete change process of a single piece of data. Through these improvements, this technology effectively overcomes the limitation of the traditional method that can only perform macroscopic and static analysis, and can meet the accurate restoration requirement of the history of a single piece of data in complex data flow scenarios.
[0091] Figure 3 FIG. 4 is a structural block diagram of a big data governance traceability system based on ElasticSearch provided by an embodiment of this application. The system at least includes the following modules:
[0092] The original data access module is used to access the original data into ElasticSearch and prepare to access the latest pushed data in response to the update requirement of the subsequent original data.
[0093] The access method selection module is used to select an access method based on the type of the latest pushed data. The access methods include full-volume access method, incremental access method, and timestamp-trigger access method.
[0094] The data cleaning module is used to clean the latest pushed data based on the access method of the latest pushed data.
[0095] The pushed data access module is used to access the latest pushed data after data cleaning into ElasticSearch according to the selected access method.
[0096] The data traceability module is used to perform data traceability operations based on ElasticSearch in response to a data traceability request.
[0097] For relevant details, refer to the method embodiments above.
[0098] Figure 4 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0099] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0100] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the big data governance traceability method based on ElasticSearch provided by the method embodiments in the present application.
[0101] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include, but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0102] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.
[0103] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the big data governance traceability method based on ElasticSearch in the above method embodiment.
[0104] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the big data governance traceability method based on ElasticSearch in the above method embodiment.
[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0106] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A big data governance and tracing method based on ElasticSearch, characterized in that: The method comprises: The original data is connected to ElasticSearch, and the latest push data is ready to be connected in response to the subsequent update requirements of the original data; Selecting an access mode based on the type of the latest pushed data, the access mode including a full access mode, an incremental access mode, and a timestamp-trigger access mode; Performing data cleaning on the latest pushed data based on the access method of the latest pushed data; The latest pushed data after data cleaning is accessed to ElasticSearch according to the selected access method; In response to the data tracing request, data tracing operations are performed based on ElasticSearch.
2. The big data governance tracing method based on ElasticSearch according to claim 1 is characterized in that: The selecting access mode based on the type of the latest pushed data includes: Automatically determine whether the latest pushed data is full data or incremental data, and select the appropriate access method based on the type of the latest pushed data. The access methods include the following three: Full access mode: When the latest pushed data is full data, including complete records of business data, use the full access mode; Incremental access mode: When the latest pushed data is incremental data, which only contains newly added or updated records, the incremental access mode is used; Timestamp-trigger access mode: when the latest pushed data is incremental data updated based on the timestamp or trigger mechanism.
3. The big data governance tracing method based on ElasticSearch according to claim 2 is characterized in that: The automatic determination of whether the type of the latest pushed data is full data or incremental data includes: Introduce the ratio of difference to similarity, and consider the change in data volume to make judgments. Use the following formula to calculate the change index V between the latest pushed data and the original data change : Among them, N origin is the number of fields in the original data, N push is the number of fields in the latest pushed data, S sim It is the cosine similarity value between the latest pushed data and the original data; The calculated change index V change Compared with the preset threshold, if the change index is greater than or equal to the preset threshold, it means that the difference between the original data and the latest pushed data is greater, and the latest pushed data is considered to be incremental data; if the change index is less than the preset threshold, it means that the difference between the original data and the latest pushed data is smaller, and the latest pushed data is considered to be full data.
4. The big data governance and tracing method based on ElasticSearch according to claim 3 is characterized in that: The automatic determination of whether the type of the latest pushed data is full data or incremental data also includes: Use the following formula to calculate the field overlap index V between the latest pushed data and the original data type : Among them, N total is the total number of fields in the latest pushed data, N same is the number of fields that are the same between the latest pushed data and the original data, N same The calculation formula is as follows: N same =|{F push }∩{F origin }| Among them, F push is the field set of the latest pushed data, F origin is the field set of the original data, and |·| represents the number of elements in the set; The calculated field overlap index V type The larger the value is, the higher the field overlap is. When the field overlap index is greater than or equal to the preset threshold, the latest pushed data is considered to be full data. The calculated field overlap index V type The smaller it is, the lower the field overlap. When the field overlap index is less than the preset threshold, the latest pushed data is judged to be incremental data.
5. The big data governance and tracing method based on ElasticSearch according to claim 4 is characterized in that: Calculate the change index V between the latest pushed data and the original data change Suitable for scenarios where both data content and field structure undergo complex changes, and for environments where data changes need to be fully considered; Calculate the field overlap index V of the latest pushed data and the original data type Applicable to scenarios that focus on changes in field set structure rather than content; choose the appropriate judgment method according to different scenarios.
6. The big data governance and tracing method based on ElasticSearch according to claim 1 is characterized in that: The performing data cleaning on the latest pushed data based on the access mode of the latest pushed data comprises: For the latest push data in full access mode, the latest push data is the full data, and the part of the data added, deleted, or modified compared with the original data is identified as the latest push data after cleaning; For the incremental access method or the timestamp-trigger access method, the latest pushed data are all incremental data, and there is no need to clean the latest pushed data.
7. The big data governance and tracing method based on ElasticSearch according to claim 1 is characterized in that: After the raw data is connected to ElasticSearch, the following steps are also included: Create a temporary backup table, store the current original data in the temporary backup table, and keep the temporary backup table for comparison with the latest pushed data.
8. A big data governance and traceability system based on ElasticSearch, characterized in that: include: The original data access module is used to access the original data to ElasticSearch and prepare to access the latest pushed data in response to the subsequent update requirements of the original data; An access mode selection module, configured to select an access mode based on the type of the latest pushed data, wherein the access mode includes a full access mode, an incremental access mode, and a timestamp-trigger access mode; A data cleaning module, used for cleaning the latest pushed data based on the access mode of the latest pushed data; A push data access module is used to access the latest push data after data cleaning to ElasticSearch according to the selected access method; The data tracing module is used to respond to data tracing requests and perform data tracing operations based on ElasticSearch.
9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a big data governance and tracing method based on ElasticSearch as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, which, when executed by the processor, is used to implement a big data governance and tracing method based on ElasticSearch as described in any one of claims 1 to 7.