A data docking method and system for a big data ecosystem joint debugging system
By abstracting the data into nodes and judging dependencies, the problems of single function of the data docking system and different data structures in the existing technology are solved, and flexible data synchronization and dependency solutions are realized.
Patent Information
- Application Number
- CN202211083203.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-09-06
AI Technical Summary
The existing data docking system has a single function and the data structures of each system vary greatly, resulting in customized development of data interactions between different systems, the synchronization time frequency is inflexible, and the data dependence problem has not been reasonably solved.
By converting the behavior of synchronized data into tasks, abstracting the data into nodes, dividing it into source nodes and target nodes, determining whether the data depends on data within the predefined range according to the map structure, obtaining the id of the dependent data through predefined dependency expressions, supplementing it into the synchronized data, and performing data checksum enumeration mapping, and finally updating the data in the intermediate table or message queue.
The synchronization of different data structures is achieved by only visual configuration and a small amount of customized development code, which improves synchronization flexibility and efficiency, and reasonably solves the dependence problem between data and data.
Smart Images

Figure CN115422276B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of industrial big data, and in particular relates to a data docking method and system for a big data ecological joint debugging system. Background Art
[0002] Data connection of various ecological joint debugging systems of industrial big data includes: obtaining basic data from ERP, OA and other systems at a certain time frequency, synchronizing to the target system after data processing, and transmitting back to the source system after data processing in the target system, opening up the data flow between different systems and solving the coupling problem of dependent data. At the same time, the real-time synchronization status can be monitored through the monitoring platform, the synchronization frequency can be controlled, and abnormal alarms can be issued.
[0003] The existing data docking system has relatively simple functions, and the data structures of various systems are different, which requires customized development when doing data interaction between different systems. The synchronization time frequency is also not flexible, and the dependency problem between data has not been reasonably solved. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes a technical solution of a data docking method of a big data ecological joint debugging system to solve the above technical problems.
[0005] The first aspect of the present invention discloses a data docking method for a big data ecological joint debugging system, the method comprising:
[0006] Step S1: convert the behavior of synchronizing data into tasks, abstract the data into nodes, and the nodes are divided into source nodes and target nodes; the source node represents the source data of the synchronizing data, and the target node represents the final destination of the synchronizing data;
[0007] Step S2: After the task is called, the source node is obtained, the task synchronization method of the target node is configured, the source data is obtained through the source node and the synchronization method, and the source data is converted into data of a map structure;
[0008] Step S3, after obtaining the data of the map structure of the source data, determine whether the data depends on the data in the predefined range according to the map structure; if there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, add the id of the dependent data to the data synchronized this time, and obtain the source data after the id is added; if the dependent data does not exist, update the data synchronization state to pause;
[0009] Step S4. Determine whether the source data after supplementing the ID exists in the database of the docking service by using the primary key. If it does not exist, add the source data after supplementing the ID to the database of the docking service; if it exists, perform a hash comparison between the source data after supplementing the ID and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed; if the hash comparison result is inconsistent, an update is performed.
[0010] Step S5. Determine the attribute of the target node. If it is in the form of an intermediate table, write the source data after supplementing the ID into the target table and update the data sending status; if it is in the form of a message queue, send the source data after supplementing the ID to the message queue and update the data sending status.
[0011] Step S6. Wait for the asynchronous feedback of the failure of writing to the table or the failure of the feedback of sending the message queue.
[0012] According to the method of the first aspect of the present invention, in the step S3, the method further includes:
[0013] Perform data verification on the source data after supplementing the ID to check whether the length, type, and value of the data conform to the configured regular expression.
[0014] According to the method of the first aspect of the present invention, in the step S3, the method further includes:
[0015] Perform enumeration mapping on the common fields in the source data after supplementing the ID.
[0016] According to the method of the first aspect of the present invention, in the step S2, the scheduling of the task is managed by the open-source framework xxl_job.
[0017] According to the method of the first aspect of the present invention, in the step S2, the method of obtaining the source data through the source node and the synchronization method and converting the source data into data in the map structure includes:
[0018] If the synchronization method is synchronization in the form of an intermediate table, first obtain the database driver of the source node, and then obtain the source data by querying the table and convert the source data into data in the map structure.
[0019] According to the method of the first aspect of the present invention, in the step S2, the method of obtaining the source data through the source node and the synchronization method and converting the source data into data in the map structure further includes:
[0020] If the synchronization method is the API interface method, first obtain the request path of the source node, call the API interface to obtain the returned message; disassemble the message through the parsing engine to obtain the source data, and convert the source data into data in a map structure; due to the differences in the message structure, the format rules of the message need to be obtained through visual configuration.
[0021] According to the method of the first aspect of the present invention, in the step S6, the asynchronous feedback method includes:
[0022] Step S61, obtain the feedback status of the sent data by listening to the message queue or the target table;
[0023] Step S62, if the feedback status of the sent data is failure, update the synchronization status of the data to failure; if the feedback status is success, update the synchronization status of the data to success;
[0024] Step S63, asynchronously determine whether the source data is dependent on data within a predefined range. If there is no dependency relationship, end the process; if it is dependent on data within the predefined range, obtain all the dependent data, complete all the dependent data with the source data, wait for the next synchronization, and end the process.
[0025] The second aspect of the present invention discloses a data docking system for a big data ecological joint debugging system, and the system includes:
[0026] The first processing module is configured to convert the behavior of synchronizing data into a task, abstract the data into nodes, and the nodes are divided into source nodes and target nodes; the source node represents the source data of the synchronized data, and the target node represents the final destination of the synchronized data;
[0027] The second processing module is configured to, after the task is called, obtain the source node, configure the task synchronization method of the target node, obtain the source data through the source node and the synchronization method, and convert the source data into data in a map structure;
[0028] The third processing module is configured to, after obtaining the data in the map structure of the source data, determine whether the data depends on data within a predefined range according to the map structure. If there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, and supplement the id of the dependent data into the data of this synchronization to obtain the source data after supplementing the id; if the dependent data does not exist, update the data synchronization status to paused;
[0029] The fourth processing module is configured to determine whether the source data after supplementing the ID exists in the database of the docking service through the primary key. If it does not exist, the source data after supplementing the ID is newly added to the database of the docking service; if it exists, a hash comparison is made between the source data after supplementing the ID and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed; if the hash comparison result is inconsistent, an update is performed.
[0030] The fifth processing module is configured to determine the attribute of the target node. If it is in the form of an intermediate table, the source data after supplementing the ID is written into the target table, and the data sending status is updated; if it is in the form of a message queue, the source data after supplementing the ID is sent to the message queue, and the data sending status is updated.
[0031] The sixth processing module is configured to wait for an asynchronous feedback of a write table failure or a failure in sending a message queue feedback.
[0032] According to the system of the second aspect of the present invention, the third processing module is configured to perform data verification on the source data after supplementing the ID to check whether the length, type, and value of the data conform to the configured regular expression.
[0033] According to the system of the second aspect of the present invention, the third processing module is configured to perform enumeration mapping on the common fields in the source data after supplementing the ID.
[0034] According to the system of the second aspect of the present invention, the second processing module is configured to manage the scheduling of tasks by the open-source framework xxl_job.
[0035] According to the system of the second aspect of the present invention, the second processing module is configured to obtain the source data through the source node and the synchronization method, and the conversion of the source data into data in the map structure includes:
[0036] If the synchronization method is synchronization in the form of an intermediate table, first obtain the database driver of the source node, and then obtain the source data by querying the table, and convert the source data into data in the map structure.
[0037] According to the system of the second aspect of the present invention, the second processing module is configured to obtain the source data through the source node and the synchronization method, and the conversion of the source data into data in the map structure further includes:
[0038] If the synchronization method is the API interface method, first obtain the request path of the source node, call the API interface to obtain the returned message; disassemble the message through the parsing engine to obtain the source data, and convert the source data into data in the map structure; due to the difference in the structure of the message, the format rule of the message needs to be obtained through visual configuration.
[0039] For the system according to the second aspect of the present invention, the sixth processing module is configured such that the asynchronous feedback includes:
[0040] Obtaining the feedback status of the sent data by listening to the message queue or the target table;
[0041] If the feedback status of the sent data is failure, updating the synchronization status of the data to failure; if the feedback status is success, updating the synchronization status of the data to success;
[0042] Asynchronously determining whether the source data is data - dependent within a predefined range. If there is no dependency relationship, ending the process; if it is dependent on the data within the predefined range, obtaining all the dependent data, completing all the dependent data with the source data, waiting for the next synchronization, and ending the process.
[0043] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, it implements the steps in any one of the data docking methods of a big data ecosystem joint debugging system in the first aspect of the present disclosure.
[0044] The fourth aspect of the present invention discloses a computer - readable storage medium. A computer program is stored on the computer - readable storage medium. When the computer program is executed by the processor, it implements the steps in any one of the data docking methods of a big data ecosystem joint debugging system in the first aspect of the present disclosure.
[0045] For the solution proposed by the present invention, for the synchronization of different data structures, only visual configuration and a small amount of customized development code are required. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a flowchart of a data docking method for a big data ecosystem joint debugging system according to an embodiment of the present invention;
[0048] Figures 2a - 2c It is a logical flowchart of the data docking method for a big data ecosystem joint debugging system according to an embodiment of the present invention;
[0049] Figure 3 It is a logical flowchart of the method for asynchronous feedback according to an embodiment of the present invention;
[0050] Figures 4a - 4d Data docking logic flowchart of the big data ecological joint debugging system of Embodiment 1 according to an embodiment of the present invention;
[0051] Figure 5 Logic flowchart of the method of asynchronous feedback of Embodiment 1 according to an embodiment of the present invention;
[0052] Figures 6a - 6b Data docking logic flowchart of the big data ecological joint debugging system of Embodiment 2 according to an embodiment of the present invention;
[0053] Figure 7 Structure diagram of a data docking system of a big data ecological joint debugging system according to an embodiment of the present invention;
[0054] Figure 8 Structure diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] The first aspect of the present invention discloses a data docking method for a big data ecological joint debugging system. Figure 1 Flowchart of a data docking method for a big data ecological joint debugging system according to an embodiment of the present invention, as Figure 1 and Figures 2a - 2c shown, the method includes:
[0057] Step S1, convert the behavior of synchronizing data into a task, abstract the data into nodes, and the nodes are divided into source nodes and target nodes; the source node represents the source data of the synchronized data, and the target node represents the final destination of the synchronized data;
[0058] Step S2, after the task is called, obtain the source node, configure the task synchronization method of the target node, obtain the source data through the source node and the synchronization method, and convert the source data into data in a map structure;
[0059] Step S3: After obtaining the data in the map structure of the source data, determine whether the data depends on the data within the predefined range according to the map structure. If there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, and supplement the id of the dependent data into the data synchronized this time to obtain the source data after supplementing the id. If the dependent data does not exist, update the data synchronization status to paused;
[0060] Step S4: Determine whether the source data after supplementing the id exists in the database of the docking service by the primary key. If it does not exist, add the source data after supplementing the id to the database of the docking service. If it exists, perform a hash comparison between the source data after supplementing the id and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed. If the hash comparison result is inconsistent, an update is performed;
[0061] Step S5: Determine the attributes of the target node. If it is in the form of an intermediate table, write the source data after supplementing the id into the target table and update the data sending status. If it is in the form of a message queue, send the source data after supplementing the id to the message queue and update the data sending status;
[0062] Step S6: Wait for the asynchronous feedback of the failure of writing to the table or the failure of the feedback of sending the message queue.
[0063] In step S2, after the task is called, the source node is obtained, the task synchronization method of the target node is configured, the source data is obtained through the source node and the synchronization method, and the source data is converted into data in the map structure.
[0064] In some embodiments, in step S2, the scheduling of the task is managed by the open-source framework xxl_job.
[0065] The method of obtaining the source data through the source node and the synchronization method and converting the source data into data in the map structure includes:
[0066] If the synchronization method is synchronization in the form of an intermediate table, first obtain the database driver of the source node, and then obtain the source data by querying the table and convert the source data into data in the map structure.
[0067] The method of obtaining the source data through the source node and the synchronization method and converting the source data into data in the map structure further includes:
[0068] If the synchronization method is the API interface method, first obtain the request path of the source node, call the API interface to obtain the return message; disassemble the message through the parsing engine to obtain the source data and convert the source data into data in the map structure; due to the difference in the structure of the message, the format rule of the message needs to be obtained through visual configuration.
[0069] In step S3, after obtaining the data in the map structure of the source data, it is determined according to the map structure whether the data depends on the data within a predefined range. For example, in the industry, product master data depends on product series master data. If there is a dependency relationship, the id of the dependent data is obtained through a predefined dependency expression. The id can uniquely identify a piece of data. The id of the dependent data is supplemented to the data synchronized this time to obtain the source data after supplementing the id. If the dependent data does not exist, the data synchronization status is updated to suspended, and the entire process ends waiting for the next synchronization.
[0070] In some embodiments, in the step S3, the method further includes:
[0071] Perform data verification on the source data after supplementing the id to check whether the length, type, and value of the data conform to the configured regular expression.
[0072] The method further includes:
[0073] Perform enumeration mapping on the common fields in the source data after supplementing the id. For example, gender will be converted to 1 - male, 2 - female.
[0074] In step S6, wait for the asynchronous feedback of the failure of writing to the table or the failure of sending the message queue feedback.
[0075] In some embodiments, as Figure 3 shown, in the step S6, the method of the asynchronous feedback includes:
[0076] Step S61: Obtain the feedback status of the sent data by listening to the message queue or the target table;
[0077] Step S62: If the feedback status of the sent data is failed, update the data synchronization status to failed; if the feedback status is successful, update the data synchronization status to successful;
[0078] Step S63: Asynchronously determine whether the source data is dependent on the data within a predefined range. If there is no dependency relationship, end the process; if it is dependent on the data within a predefined range, obtain all the dependent data, and complete all the dependent data with the source data, that is, complete the Mes system id field of the source data in all the dependent data, wait for the next synchronization, and end the process.
[0079] Specifically, if data A is the source data of the feedback and data B depends on data A, after data A obtains the successful feedback status, complete the Mes system id field of A in data B. In this way, it complements the main process of data synchronization.
[0080] Embodiment 1:
[0081] A data docking method for a big data ecosystem joint debugging system according to the first aspect of the present invention is applied in Embodiment 1. Specifically, Figures 4a - 4d The figure shows the process of synchronizing product series and product data from an ERP system in a certain factory in Wenzhou, which adopts the method of intermediate tables. These two tasks are dependent on each other, and the product data depends on the product series data. In the synchronization processes of these two tasks, the data verification and enumeration mapping processes do not perform any operations because no configuration has been done.
[0082] First, synchronize the data of the product series. The synchronization process of the product series conforms to the operating principle of the middleware and will not be specially described. The final result is that the data of the product series in the ERP is synchronized into the Mes system, and a copy of the source data is retained in the middleware service database, and a Mes system id is generated. The Mes system id has a one-to-one relationship with the ERP system id.
[0083] Figure 5 Describe the product series feedback process. After the product series data is sent to the mes system through the message queue, the middleware service will listen to the feedback message queue to obtain the feedback status of the product series data synchronization. If the feedback status is failed, update the synchronization failure status of this product series data according to the product series mes system id. If the feedback is successful, update the synchronization success status of this product series data according to the product series mes system id. Subsequently, asynchronously determine whether the product series data is dependent on other data. The result is that the product series data is dependent on the product data. Find the data in the product information data whose synchronization status is suspended and the ERP product series id is equal to the ERP system id of the product series data in this feedback, and set its product series mes system id to the product series Mes system id in this feedback. Finally, the product data with the product series mes system id completed will be pushed to the mes system in the next synchronization cycle.
[0084] After synchronizing the product series, obtain the product data from the ERP system. The product data depends on the product series data. There is an ERP product series id in the product data in the ERP source data. Query the Mes system id from the product series table in the middleware service database according to the ERP product series id. If the corresponding product series exists, write the product series Mes system id into the product series Mes system id field of the product information. If it does not exist, update the synchronization status of this product information data to suspended, stop the synchronization process, and wait to search again in the next synchronization cycle. Thus, the decoupling of product and product series data is solved.
[0085] The subsequent processes and principles are consistent with the content disclosed in the first aspect, saving the product data into the middleware service database, sending it to the message queue for other services to consume, and updating the sending status.
[0086] The above process describes the method of intermediate table synchronization. The entire process does not require any code development and only requires visual configuration. This middleware can be adapted to any currently mainstream middleware for data storage.
[0087] Embodiment 2:
[0088] A data docking method for a big data ecological joint debugging system according to the first aspect of the present invention is applied in this Embodiment 2. Specifically, Figures 6a - 6b The process of synchronizing product series in a certain factory in Beilun is described. The product series does not depend on other data, and the API synchronization method is mainly illustrated.
[0089] Due to the differences in interface formats, when using the API synchronization method, it is necessary to disassemble and assemble packets according to the format rules of the message to obtain the source data of the product series.
[0090] For example, the overall structure of the message received now is json, and the format is as follows:
[0091]
[0092]
[0093] It is necessary to obtain the product series code (cInvCCode) and the product series name (cInvCName). The configured mapping rule is that the cInvCCode field is taken from / data / created / cInvCCode, and the cInvCName field is taken from / data / created / cInvCName. Such a configuration can obtain the product series data of the above two items, generate the Mes system id, save it into the middleware database, and the other processes are the same. Finally, it is synchronized into the Mes system.
[0094] The above process describes the API synchronization method. The entire process does not require any code development and only requires visual configuration to parse any mainstream format of the message.
[0095] In summary, the solution proposed by the present invention can only perform visual configuration and a small amount of customized development code for the synchronization of different data structures.
[0096] The second aspect of the present invention discloses a data docking system for a big data ecological joint debugging system. Figure 7 It is a structural diagram of a data docking system for a big data ecological joint debugging system according to an embodiment of the present invention; as Figure 7 shown, the system 100 includes:
[0097] The first processing module 101 is configured to convert the behavior of synchronizing data into tasks and abstract the data into nodes, where the nodes are divided into source nodes and target nodes; the source nodes represent the source data of the synchronized data, and the target nodes represent the final destination of the synchronized data.
[0098] The second processing module 102 is configured to, after the task is called, obtain the source node, configure the task synchronization method of the target node, obtain the source data through the source node and the synchronization method, and convert the source data into data in map structure.
[0099] The third processing module 103 is configured to, after obtaining the data in map structure of the source data, determine whether the data depends on the data within the predefined range according to the map structure. If there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, and supplement the id of the dependent data into the data of this synchronization to obtain the source data after supplementing the id. If the dependent data does not exist, update the data synchronization status to paused.
[0100] The fourth processing module 104 is configured to determine whether the source data after supplementing the id exists in the database of the docking service through the primary key. If it does not exist, add the source data after supplementing the id to the database of the docking service. If it exists, perform a hash comparison between the source data after supplementing the id and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed. If the hash comparison result is inconsistent, an update is performed.
[0101] The fifth processing module 105 is configured to determine the attribute of the target node. If it is in the form of an intermediate table, write the source data after supplementing the id into the target table and update the data sending status. If it is in the form of a message queue, send the source data after supplementing the id to the message queue and update the data sending status.
[0102] The sixth processing module 106 is configured to wait for the asynchronous feedback of the failure of writing the table or the failure of the feedback of sending the message queue.
[0103] For the system according to the second aspect of the present invention, the third processing module 103 is configured to perform data verification on the source data after supplementing the id to check whether the length, type, and value of the data conform to the configured regular expression.
[0104] For the system according to the second aspect of the present invention, the third processing module 103 is configured to perform enumeration mapping on the common fields in the source data after supplementing the id.
[0105] For the system according to the second aspect of the present invention, the second processing module 102 is configured to manage the scheduling of tasks by the open-source framework xxl_job.
[0106] For the system according to the second aspect of the present invention, the second processing module 102 is configured such that the acquisition of source data through the source node and synchronization and the conversion of the source data into data in map structure include:
[0107] If the synchronization method is synchronization via an intermediate table, first obtain the database driver of the source node, and then obtain the source data by querying the table and convert the source data into data in map structure.
[0108] For the system according to the second aspect of the present invention, the second processing module 102 is configured such that the acquisition of source data through the source node and synchronization and the conversion of the source data into data in map structure further include:
[0109] If the synchronization method is via an API interface, first obtain the request path of the source node, call the API interface to obtain the returned message; disassemble the message through a parsing engine to obtain the source data and convert the source data into data in map structure; due to the difference in the structure of the message, the format rules of the message need to be obtained through visual configuration.
[0110] For the system according to the second aspect of the present invention, the sixth processing module 106 is configured such that the asynchronous feedback includes:
[0111] Obtain the feedback status of the sent data by listening to the message queue or the target table;
[0112] If the feedback status of the sent data is failure, update the synchronization status of the data to failure; if the feedback status is success, update the synchronization status of the data to success;
[0113] Asynchronously determine whether the source data is dependent on data within a predefined range. If there is no dependency relationship, end the process; if it is dependent on data within the predefined range, obtain all the dependent data, complete all the dependent data with the source data, wait for the next synchronization, and end the process.
[0114] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, the steps in any one of the data docking methods of the big data ecological joint debugging system in the first aspect disclosed by the present invention are implemented.
[0115] Figure 8 For the structural diagram of an electronic device according to an embodiment of the present invention, as Figure 8As shown, the electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, near field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0116] Those skilled in the art can understand that Figure 8 the structure shown in is only the structural diagram of the part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0117] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in a data docking method of a big data ecological joint debugging system in any one of the first aspects disclosed by the present invention are implemented.
[0118] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered to be within the scope described in this specification. The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A data docking method for a big data ecosystem joint debugging system, characterized in that, The method includes: Step S1: Convert the behavior of synchronizing data into tasks, and abstract the data into nodes. The nodes are divided into source nodes and target nodes. The source node represents the source data of the synchronized data, and the target node represents the final destination of the synchronized data. Step S2: After the task is called, obtain the source node, configure the task synchronization method of the target node, obtain the source data through the source node and the synchronization method, and convert the source data into data in map structure. Step S3: After obtaining the data in map structure of the source data, judge whether the data depends on the data within the predefined range according to the map structure. If there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, and supplement the id of the dependent data into the data of this synchronization to obtain the source data with the supplemented id. If the dependent data does not exist, update the data synchronization status to paused. Step S4: Judge whether the source data with the supplemented id exists in the database of the docking service through the primary key. If it does not exist, add the source data with the supplemented id to the database of the docking service. If it exists, perform a hash comparison between the source data with the supplemented id and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed. If the hash comparison result is inconsistent, an update is performed. Step S5: Judge the attribute of the target node. If it is in the form of an intermediate table, write the source data with the supplemented id into the target table and update the data sending status. If it is in the form of a message queue, send the source data with the supplemented id to the message queue and update the data sending status. Step S6: Wait for the asynchronous feedback of the write table failure or the failure feedback of sending the message queue.
2. The data docking method for a big data ecosystem joint debugging system according to claim 1, characterized in that, In the step S3, the method further includes: Perform data verification on the source data with the supplemented id to check whether the length, type, and value of the data conform to the configured regular expression.
3. The data docking method for a big data ecosystem joint debugging system according to claim 1, characterized in that, In the step S3, the method further includes: Perform enumeration mapping on the common fields in the source data with the supplemented id.
4. The data docking method for a big data ecosystem joint debugging system according to claim 1, characterized in that, In the step S2, the scheduling of the task is managed by the open-source framework xxl_job.
5. The data docking method for a big data ecosystem joint debugging system according to claim 1, characterized in that, In the step S2, the method of obtaining the source data through the source node and the synchronization method and converting the source data into data in map structure includes: If the synchronization method is synchronization in the form of an intermediate table, first obtain the database driver of the source node, and then obtain the source data by querying the table and convert the source data into data in map structure.
6. The data docking method for a big data ecosystem joint debugging system according to claim 2, characterized in that, In the step S2, the method of obtaining the source data through the source node and the synchronization method and converting the source data into data in map structure further includes: If the synchronization method is the API interface method, first obtain the request path of the source node, call the API interface to obtain the returned message; disassemble the message through the parsing engine to obtain the source data, and convert the source data into data in map structure; due to the difference in the structure of the message, the format rules of the message need to be obtained through visual configuration.
7. The data docking method for a big data ecosystem joint debugging system according to claim 1, characterized in that, In the step S6, the method of the asynchronous feedback includes: Step S61: Obtain the feedback status of the sent data by listening to the message queue or the target table; Step S62: If the feedback status of the sent data is failure, update the synchronization status of the data to failure; if the feedback status is success, update the synchronization status of the data to success; Step S63: Asynchronously determine whether the source data is dependent on the data within the predefined range. If there is no dependency relationship, end the process; if it is dependent on the data within the predefined range, obtain all the dependent data, complete all the dependent data with the source data, wait for the next synchronization, and end the process.
8. A data docking system for a big data ecosystem joint debugging system, characterized in that, The system includes: A first processing module, configured to convert the behavior of synchronizing data into a task, abstract the data into nodes, and the nodes are divided into source nodes and target nodes; the source nodes represent the source data of the synchronized data, and the target nodes represent the final destination of the synchronized data; A second processing module, configured to, after the task is called, obtain the source nodes, configure the task synchronization method of the target nodes, obtain the source data through the source nodes and the synchronization method, and convert the source data into data in a map structure; A third processing module, configured to, after obtaining the data in the map structure of the source data, determine whether the data depends on the data within the predefined range according to the map structure. If there is a dependency relationship, obtain the id of the dependent data through the predefined dependency expression, and supplement the id of the dependent data into the data of this synchronization to obtain the source data after supplementing the id; if the dependent data does not exist, update the data synchronization status to suspended; A fourth processing module, configured to determine whether the source data after supplementing the id exists in the database of the docking service through the primary key. If it does not exist, add the source data after supplementing the id to the database of the docking service; if it exists, perform a hash comparison between the source data after supplementing the id and the data in the database of the docking service. If the hash comparison result is consistent, no update is performed; if the hash comparison result is inconsistent, an update is performed; A fifth processing module, configured to judge the attribute of the target node. If it is in the form of an intermediate table, write the source data after supplementing the id into the target table and update the data sending status; if it is in the form of a message queue, send the source data after supplementing the id to the message queue and update the data sending status; A sixth processing module, configured to wait for the asynchronous feedback of the failure of writing to the table or the failure of the feedback of sending the message queue.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, it implements the steps in any one of claims 1 to 7 of a data docking method for a big data ecological joint debugging system.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps in any one of claims 1 to 7 of a data docking method for a big data ecological joint debugging system.
Citation Information
Patent Citations
Multi-application system data synchronization method and device
CN111680106A
Heterogeneous data source quasi-real-time data synchronization method and system based on CDC
CN114254040A