A data stream association method, apparatus, device, and storage medium thereof
By analyzing and mapping the multi-source data streams, the expected data content is directly obtained, which solves the delay and error problems in real-time association of multi-source data streams, and improves the speed and accuracy of data processing of financial services.
Patent Information
- Application Number
- CN202311348425.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-10-18
AI Technical Summary
The prior art has problems of real-time data delay arrival and associated data errors in real-time association of multi-source data streams, resulting in delayed completion or failure of financial business tasks.
By obtaining the multi-source data stream pushed by the target push component in real time, analyzing the data attribute fields, identifying the expected data attribute fields, using the target association result table for data mapping and splicing, building a virtual form, directly obtaining the expected data content and mapping it into the association result table, realizing the rapid association of multi-source data streams.
The correlation speed of multi-source data flow is improved, the real-time and accuracy of financial business data processing is ensured, and business task delays are avoided due to correlation delays.
Smart Images

Figure CN117216114B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of fintech, and is applied to the scenario of obtaining multi-source data in financial business data, and particularly relates to a data stream association method, device, equipment and storage medium thereof. Background Art
[0002] With the rapid development of the Internet, all industries are relying on the Internet to seek breakthrough points in the industry. In recent years, the financial industry has also been expanding its online business around the Internet. Since the financial industry involves a large amount of business volume and data volume, with the continuous improvement of users' product requirements, the time limit requirements for data processing have also increased accordingly. Currently, the data in many business scenarios requires real-time processing and real-time pushing of processing results to users; the real-time processing technology in the field of big data has also been developing rapidly in recent years, but the real-time stream technology has great differences in data processing details compared with traditional offline technologies, and there are many difficulties in migrating traditional offline processing logics to the real-time processing scenario of data streams.
[0003] A major problem in migrating traditional offline processing logics to the real-time processing scenario of data streams is the real-time association of multi-source data streams. Currently, there are still problems in the real-time association of multi-source data streams, such as real-time data arriving late, the final associated data being inconsistent with the actual business, resulting in the delay of financial business tasks or the failure of business processing due to incorrect associated data. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to propose a data stream association method, device, equipment and storage medium thereof, so as to solve the problem in the prior art that in the real-time association of multi-source data streams, it may lead to the delay of financial business tasks or the failure of business processing due to incorrect associated data.
[0005] In order to solve the above technical problems, the embodiments of the present application provide a data stream association method, which adopts the following technical solutions:
[0006] A data stream association method includes the following steps:
[0007] Obtain multi-source data streams pushed in real time by a target push component, where the target push component has set source discrimination information for the multi-source data streams in advance according to the data stream sources;
[0008] Based on a preset parsing component, parse the multi-source data streams to obtain the data attribute fields respectively included in the multi-source data streams as actual data attribute fields;
[0009] Obtain a target association result table, where the target association result table is designed in advance according to the data attribute fields finally required by the target association service;
[0010] Perform form parsing on the target association result table, and identify the data attribute fields included in the target association result table according to the form parsing result as the expected data attribute fields;
[0011] Compare the expected data attribute fields with the actual data attribute fields, and determine the target data streams respectively corresponding to all the expected data attribute fields according to the comparison result;
[0012] Obtain the business data required for the final target association business from the target data resource library according to the target data streams respectively corresponding to all the expected data attribute fields and the source difference information of the multi-source data streams;
[0013] Map the business data to the association result table according to the expected data attribute fields to complete the association of the multi-source data streams.
[0014] Further, the target push component includes a distributed data push component. Before performing the step of obtaining the multi-source data streams pushed in real time by the target push component, the method further includes:
[0015] Based on a preset push record log, record the push time when the data push component performs real-time push on each data stream in the multi-source data streams;
[0016] After performing the step of obtaining the multi-source data streams pushed in real time by the target push component, the method further includes:
[0017] Obtain the source difference information respectively corresponding to each data stream in the multi-source data streams according to the source difference information preset by the target push component for the multi-source data streams;
[0018] Parse the push record log through a preset log parsing component to identify the push time respectively corresponding to each data stream in the multi-source data streams.
[0019] Further, the preset parsing component is a data stream parsing component. The step of parsing the multi-source data streams based on the preset parsing component to obtain the data attribute fields respectively included in the multi-source data streams specifically includes:
[0020] Parse each data stream in the multi-source data streams respectively by the data stream parsing component to obtain the data content transmitted in each data stream;
[0021] Based on the data content transmitted in each data stream, identify the data attribute fields respectively included in each data stream;
[0022] After performing the step of parsing the multi-source data stream based on a preset parsing component to obtain the data attribute fields respectively included in the multi-source data stream, the method further includes:
[0023] Perform a marking process on the data attribute fields respectively included in each data stream according to the source difference information respectively corresponding to each data stream in the multi-source data stream to obtain a marking result, where the marking process specifically is to assign the source difference information respectively corresponding to each data stream as a marking field to the data attribute fields respectively included in each data stream;
[0024] Identify the source difference information respectively corresponding to all actual data attribute fields according to the marking result.
[0025] Further, the step of comparing the expected data attribute fields and the actual data attribute fields and determining the target data stream respectively corresponding to the expected data attribute fields according to the comparison result specifically includes:
[0026] Compare the expected data attribute fields and the actual data attribute fields to determine the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields;
[0027] According to the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields, and the source difference information respectively corresponding to all actual data attribute fields, determine the source difference information respectively corresponding to all expected data attribute fields;
[0028] According to the source difference information respectively corresponding to each data stream in the multi-source data stream, and the source difference information respectively corresponding to all expected data attribute fields, determine the target data stream respectively corresponding to all expected data attribute fields.
[0029] Further, the step of obtaining the service data finally required for the target associated service from the target data resource library according to the target data stream respectively corresponding to all expected data attribute fields and the source difference information of the multi-source data stream specifically includes:
[0030] Send a form acquisition request to the target data resource library according to the target data stream respectively corresponding to all expected data attribute fields and the source difference information of the multi-source data stream;
[0031] Receive the request response result returned by the target data resource library based on the form acquisition request;
[0032] By parsing the request response result, identify the target data form involved in the service data, the expected data attribute fields included in each target data form, and all data attribute fields included in each target data form.
[0033] Further, the step of mapping the service data into the associated result table according to the expected data attribute fields to complete the association of the multi-source data stream specifically includes:
[0034] Perform NULL value processing on the data content corresponding to the non-expected data attribute fields in each target data form according to the expected data attribute fields included in each target data form and all the data attribute fields included in each target data form, and obtain each processed target data form;
[0035] According to the target data forms involved in the service data and the expected data attribute fields included in each target data form, splice all the processed target data forms in the way of UNION form splicing to obtain a spliced form;
[0036] Obtain the data content corresponding to the expected data attribute fields according to the expected data attribute fields included in each target data form;
[0037] According to the push time corresponding to each data stream in the multi-source data stream, the target data stream corresponding to each of all the expected data attribute fields, and the data content corresponding to the expected data attribute fields, add the data content corresponding to all the expected data attribute fields to the spliced form to obtain a spliced form filled with data content;
[0038] Map the data content in the spliced form into the associated result table according to the data attribute fields to complete the association of the multi-source data stream.
[0039] Further, the step of adding the data content corresponding to all the expected data attribute fields to the spliced form according to the push time corresponding to each data stream in the multi-source data stream, the target data stream corresponding to each of all the expected data attribute fields, and the data content corresponding to the expected data attribute fields to obtain a spliced form filled with data content specifically includes:
[0040] According to the target data stream corresponding to each of all the expected data attribute fields and the data content corresponding to the expected data attribute fields, identify whether the data content corresponding to the same expected data attribute field is pushed successively by two or more data streams;
[0041] If the data content corresponding to the expected data attribute field is pushed only by one data stream, identify the data form corresponding to the expected data attribute field according to the source difference information of the data stream, obtain the data content corresponding to the expected data attribute field from the data form, and add the data content to the splicing form in a normal insertion manner. Specifically, the normal insertion manner is to directly add the data content corresponding to the expected data attribute field to the splicing form.
[0042] If there are two or more data streams that push the data content corresponding to the same expected data attribute field successively, filter out the data stream that is pushed last according to the push time corresponding to each data stream in the multi-source data stream, identify the data form corresponding to the expected data attribute field according to the source difference information of the data stream that is pushed last, obtain the data content corresponding to the expected data attribute field from the data form as the content to be inserted, and add the content to be inserted to the splicing form in an update insertion manner. Specifically, the update insertion manner is as follows: if the data content corresponding to the expected data attribute field pushed by the prior data stream has been added to the splicing form, first delete the data content corresponding to the expected data attribute field in the splicing form, and then add the content to be inserted to the splicing form. And
[0043] if the data content corresponding to the expected data attribute field pushed by the prior data stream has not been added to the splicing form, directly add the content to be inserted to the splicing form.
[0044] Until the data content corresponding to all expected data attribute fields is added to the splicing form, a splicing form filled with data content is obtained.
[0045] To solve the above technical problems, the embodiment of the present application further provides a data stream association device, which adopts the following technical solutions:
[0046] A data stream association device includes:
[0047] A multi-source data stream acquisition module, configured to acquire a multi-source data stream pushed by a target push component in real time, where the target push component has set source difference information for the multi-source data stream in advance according to the data stream source;
[0048] An actual data attribute field acquisition module, configured to parse the multi-source data stream based on a preset parsing component, and acquire the data attribute fields respectively included in the multi-source data stream as actual data attribute fields;
[0049] A target association result table acquisition module, configured to acquire a target association result table, where the target association result table is designed in advance according to the data attribute fields finally required by the target association service;
[0050] An expected data attribute field acquisition module, configured to perform form parsing on the target association result table, and identify the data attribute fields included in the target association result table according to the form parsing result as expected data attribute fields;
[0051] A comparison and determination module, configured to compare the expected data attribute fields with the actual data attribute fields, and determine the target data streams respectively corresponding to all the expected data attribute fields according to the comparison result;
[0052] A service data acquisition module, configured to acquire the service data finally required by the target association service from the target data resource library according to the target data streams respectively corresponding to all the expected data attribute fields and the source difference information of the multi-source data streams;
[0053] A multi-source data stream association module, configured to map the service data into the association result table according to the expected data attribute fields, and complete the association of the multi-source data streams.
[0054] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0055] A computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the data stream association method described above are implemented.
[0056] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0057] A computer-readable storage medium has computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the steps of the data stream association method described above are implemented.
[0058] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0059] The data stream association method described in the embodiments of the present application completes the association of the multi-source data stream by pushing multi-source data streams, determining expected data attribute fields, identifying actual data attribute fields, obtaining from a target data resource library in a request manner, constructing a virtual form, and mapping the data content in the virtual form into an actual association result table. Directly according to the target data stream corresponding to each expected data attribute field and the source difference information of the multi-source data stream, the entire form where the expected data attribute field is located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute field is obtained. When obtaining, there is no need to query according to the data attribute field first, only the entire form needs to be obtained, which is faster, and then the form can be processed later. To a certain extent, the association speed of multi-source data streams in financial services is improved, and the real-time nature of financial service data processing is ensured. Description of the Drawings
[0060] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following-described drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;
[0062] Figure 2 is a flowchart of an embodiment of the data stream association method according to the present application;
[0063] Figure 3 is Figure 2 a flowchart of a specific embodiment of step 205 shown;
[0064] Figure 4 is Figure 2 a flowchart of a specific embodiment of step 206 shown;
[0065] Figure 5 is Figure 2 a flowchart of a specific embodiment of step 207 shown;
[0066] Figure 6 is Figure 5 a flowchart of a specific embodiment of step 504 shown;
[0067] Figure 7 is a schematic structural diagram of an embodiment of the data stream association device according to the present application;
[0068] Figure 8It is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed implementation manners
[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0070] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0071] In order to enable those in the technical field to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings.
[0072] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0073] Users may use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0074] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.
[0075] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.
[0076] It should be noted that the data stream association method provided in the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the data stream association device is generally set in the server / terminal device.
[0077] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0078] Continuing to refer to Figure 2 , a flowchart of an embodiment of the data stream association method according to the present application is shown. The data stream association method includes the following steps:
[0079] Step 201, obtain multi-source data streams pushed in real time by a target push component, where the target push component has set source discrimination information for the multi-source data streams in advance according to the data stream sources.
[0080] In this embodiment, the multi-source data streams are data streams corresponding to financial service data from multiple data sources, that is, financial service data pushed from multiple financial service databases in the form of data streams.
[0081] In this embodiment, the target push component includes a distributed data push component. Before performing the step of obtaining the multi-source data streams pushed in real time by the target push component, the method further includes: based on a preset push record log, record the push time when the data push component performs real-time push on each data stream in the multi-source data streams.
[0082] Specifically, the distributed data push component includes a distributed data push component composed of a Flink cluster real-time task and Kafka message push as a framework. Apache Flink is a distributed processing engine framework used for data stream processing. Flink can run in all common cluster environments, and can perform calculations at memory speed and any scale, and can perform real-time stream data processing. Apache Kafka generally refers to a messaging system that tends to process real-time data. This application combines Flink cluster real-time tasks and Kafka message push, which can not only process data streams in real time, but also process data pushed in data streams in real time, thereby improving the processing speed of financial business data.
[0083] In this embodiment, after executing the step of obtaining the multi-source data stream pushed in real time by the target push component, the method also includes: obtaining the source distinction information corresponding to each data stream in the multi-source data stream according to the source distinction information pre-set by the target push component for the multi-source data stream; parsing the push record log through a preset log parsing component to identify the push time corresponding to each data stream in the multi-source data stream.
[0084] By identifying the push time corresponding to each data stream in the multi-source data stream, it is convenient to update the associated data content according to the push time during subsequent association.
[0085] Step 202: parsing the multi-source data stream based on a preset parsing component, and obtaining data attribute fields respectively contained in the multi-source data stream as actual data attribute fields.
[0086] In this embodiment, the preset parsing component is a data stream parsing component. Since Apache Flink can be directly used to perform stateful computing on the data stream, Apache Flink can be directly used as the data stream parsing component.
[0087] In this embodiment, the step of parsing the multi-source data stream based on a preset parsing component to obtain the data attribute fields respectively contained in the multi-source data stream specifically includes: parsing each data stream in the multi-source data stream according to the data stream parsing component to obtain the data content transmitted in each data stream; based on the data content transmitted in each data stream, identifying the data attribute fields respectively contained in each data stream.
[0088] Through the parsing operation, the data attribute fields included in each data stream are identified. Among them, the data attribute fields include the primary key ID of each piece of financial business data in the data stream and the data attribute fields in each piece of financial business data involved in the target business task. When pushing, according to the financial business requirements, it may not be necessary to push a complete piece of financial business data, or it may only push the data content corresponding to multiple data attribute fields in a piece of financial business data. Identifying the data attribute fields included in each data stream aims to provide a support basis for subsequent data stream association operations.
[0089] In this embodiment, after performing the step of parsing the multi-source data stream based on a preset parsing component to obtain the data attribute fields included in each data stream of the multi-source data stream, the method further includes: performing a marking process on the data attribute fields included in each data stream according to the source difference information corresponding to each data stream in the multi-source data stream to obtain a marking result, where the marking process is specifically to assign the source difference information corresponding to each data stream as a marking field to the data attribute fields included in each data stream; according to the marking result, identifying the source difference information corresponding to all actual data attribute fields.
[0090] By assigning the source difference information corresponding to each data stream as a marking field to the data attribute fields included in each data stream, the purpose is to identify the source difference information corresponding to all actual data attribute fields, which is convenient for subsequent programs to perform data stream association operations.
[0091] Step 203, obtain a target association result table, where the target association result table is designed in advance according to the data attribute fields finally required by the target association business.
[0092] By designing the target association result table in advance according to the data attribute fields finally required by the target association business, it is convenient to indicate the association direction for the association of the multi-source data stream.
[0093] Step 204, perform form parsing on the target association result table, and identify the data attribute fields included in the target association result table according to the form parsing result as the expected data attribute fields.
[0094] By performing form parsing on the target association result table and identifying the data attribute fields included in the target association result table according to the form parsing result, it is convenient to clarify the association direction and the data attribute fields required during the association when subsequently associating the multi-source data stream.
[0095] Step 205: Compare the expected data attribute fields with the actual data attribute fields, and based on the comparison result, determine the target data streams corresponding to all the expected data attribute fields respectively.
[0096] Continue to refer to Figure 3 , Figure 3 Yes Figure 2 Figure 205 shows a flowchart of a specific embodiment, including:
[0097] Step 301: Compare the expected data attribute fields with the actual data attribute fields, and determine the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields;
[0098] Step 302: Based on the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields, and the source difference information corresponding to all the actual data attribute fields respectively, determine the source difference information corresponding to all the expected data attribute fields respectively;
[0099] Step 303: Based on the source difference information corresponding to each data stream in the multi-source data streams respectively, and the source difference information corresponding to all the expected data attribute fields respectively, determine the target data streams corresponding to all the expected data attribute fields respectively.
[0100] By comparing the expected data attribute fields with the actual data attribute fields, since the actual data attribute fields respectively correspond to corresponding data streams and the source difference information of the corresponding data streams, the source difference information and the target data streams corresponding to all the expected data attribute fields are finally determined. This facilitates subsequent association operations on the multi-source data streams.
[0101] Step 206: Based on the target data streams corresponding to all the expected data attribute fields respectively and the source difference information of the multi-source data streams, obtain the service data required for the target associated service from the target data resource library.
[0102] Continue to refer to Figure 4 , Figure 4 Yes Figure 2 Figure 206 shows a flowchart of a specific embodiment, including:
[0103] Step 401: Based on the target data streams corresponding to all the expected data attribute fields respectively and the source difference information of the multi-source data streams, send a form acquisition request to the target data resource library;
[0104] Step 402: Receive the request response result returned by the target data resource library based on the form acquisition request;
[0105] Step 403: By parsing the request response result, identify the target data forms involved in the service data, the expected data attribute fields included in each target data form, and all the data attribute fields included in each target data form.
[0106] In this embodiment, actually, in addition to the data stream push and transmission processing line, a request response processing line is added. Through the request response method, the service data ultimately required for the target associated service is directly obtained from the target data resource library. In this way, not only can the obtained service data be compared with the result data associated with the subsequent data stream to identify whether the association is successful, but also when the association is not successful, the service data directly obtained from the target data resource library through the request response method can be used as the association result, which has a reference value for the multi-source data stream association operation result.
[0107] Step 207: Map the service data into the association result table according to the expected data attribute fields, and complete the association of the multi-source data stream.
[0108] Continue to refer to Figure 5 , Figure 5 Yes Figure 2 is
[0109] Step 501: According to the expected data attribute fields included in each target data form and all the data attribute fields included in each target data form, perform NULL value processing on the data content corresponding to the non-expected data attribute fields in each target data form to obtain each processed target data form.
[0110] Substantially, in this embodiment, directly based on the target data streams corresponding to all the expected data attribute fields and the source difference information of the multi-source data stream, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first, just obtain the entire form, which is faster, and then process the form later.
[0111] Step 502: According to the target data forms involved in the service data and the expected data attribute fields included in each target data form, splice all the processed target data forms using the UNION form splicing method to obtain a spliced form.
[0112] The UNION form splicing method is a usage in SQL statements. Using UNION as the splicing character, multiple related forms are spliced to generate a business comprehensive form, that is, the spliced form. The spliced form can be a virtual form cached in memory.
[0113] Step 503: Obtain the data content corresponding to the expected data attribute fields according to the expected data attribute fields included in each target data form.
[0114] Step 504: According to the push time corresponding to each data stream in the multi-source data stream, the target data stream corresponding to all expected data attribute fields, and the data content corresponding to the expected data attribute fields, add the data content corresponding to all expected data attribute fields to the spliced form to obtain a spliced form filled with data content.
[0115] Continue to refer to Figure 6 , Figure 6 Yes Figure 5 The flowchart of a specific embodiment of step 504 shown includes:
[0116] Step 601: Identify whether the data content corresponding to the same expected data attribute field is pushed successively by two or more data streams according to the target data stream corresponding to all expected data attribute fields and the data content corresponding to the expected data attribute fields.
[0117] Step 602: If the data content corresponding to the expected data attribute field is pushed only by one data stream, identify the data form corresponding to the expected data attribute field according to the source difference information of the data stream, obtain the data content corresponding to the expected data attribute field from the data form, and add the data content to the spliced form in a normal insertion manner. The normal insertion manner is specifically to directly add the data content corresponding to the expected data attribute field to the spliced form.
[0118] Step 603, if there are data contents corresponding to the same expected data attribute field pushed by two or more data streams successively, then according to the push time corresponding to each data stream in the multi-source data streams, filter out the data stream that is pushed last. According to the source difference information of the data stream that is pushed last, identify the data form corresponding to the expected data attribute field, and obtain the data content corresponding to the expected data attribute field from the data form as the content to be inserted. Then add the content to be inserted to the splicing form in an update insertion manner. The update insertion manner is specifically as follows: if the data content corresponding to the expected data attribute field pushed by the prior data stream has been added to the splicing form, first delete the data content corresponding to the expected data attribute field in the splicing form, and then add the content to be inserted to the splicing form. And
[0119] if the data content corresponding to the expected data attribute field pushed by the prior data stream has not been added to the splicing form, directly add the content to be inserted to the splicing form;
[0120] Step 604, until all the data contents corresponding to the expected data attribute fields are added to the splicing form, obtain the splicing form filled with data contents.
[0121] Step 505, map the data contents in the splicing form to the associated result table according to the data attribute fields, and complete the association of the multi-source data streams.
[0122] In this embodiment, mapping the data contents in the splicing form to the associated result table according to the data attribute fields. Specifically, the location where the associated result table is located can be in Apache Hudi. Substantially, it is to map the splicing form, that is, the data contents in the virtual form cached in the memory, to the associated result table in Apache Hudi.
[0123] In this embodiment, the association of the multi-source data stream is completed by using Apache Flink and Apache Kafka to push the multi-source data stream, determine the expected data attribute fields, identify the actual data attribute fields, obtain them from the target data repository in a request-based manner, construct a virtual form, and map the data content in the virtual form to the actual associated result table. Based directly on the target data streams corresponding to all the expected data attribute fields and the source difference information of the multi-source data stream, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first. Just obtain the entire form, which is faster, and then process the form later. To a certain extent, the association speed of the multi-source data stream in financial services is improved, and the real-time nature of financial service data processing is ensured. The virtual form ensures the accuracy of the multi-source data stream association result, and also ensures that when the association is delayed, the virtual form can be directly used to support financial service data, avoiding delays in financial service tasks.
[0124] This application completes the association of the multi-source data stream by pushing the multi-source data stream, determining the expected data attribute fields, identifying the actual data attribute fields, obtaining them from the target data repository in a request-based manner, constructing a virtual form, and mapping the data content in the virtual form to the actual associated result table. Based directly on the target data streams corresponding to all the expected data attribute fields and the source difference information of the multi-source data stream, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first. Just obtain the entire form, which is faster, and then process the form later. To a certain extent, the association speed of the multi-source data stream in financial services is improved, and the real-time nature of financial service data processing is ensured. The virtual form ensures the accuracy of the multi-source data stream association result, and also ensures that when the association is delayed, the virtual form can be directly used to support financial service data, avoiding delays in financial service tasks.
[0125] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0126] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data stream correlation technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0127] In the embodiments of the present application, by pushing multi-source data streams, determining expected data attribute fields, identifying actual data attribute fields, obtaining from a target data resource library in a request manner, constructing a virtual form, and mapping the data content in the virtual form into an actual associated result table, the association of the multi-source data streams is completed. Directly according to the target data streams corresponding to all the expected data attribute fields respectively and the source difference information of the multi-source data streams, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first, just obtain the entire form, which is faster, and then process the form later. To a certain extent, the association speed of multi-source data streams in financial operations is improved, and the real-time nature of financial business data processing is ensured.
[0128] Further referring to Figure 7 , as an implementation of the method shown above Figure 2 , the present application provides an embodiment of a data stream association device. This device embodiment corresponds to the method embodiment shown in Figure 2 , and this device can be specifically applied to various electronic devices.
[0129] As Figure 7 shown, the data stream association device 700 described in this embodiment includes: a multi-source data stream acquisition module 701, an actual data attribute field acquisition module 702, a target association result table acquisition module 703, an expected data attribute field acquisition module 704, a comparison determination module 705, a service data acquisition module 706, and a multi-source data stream association module 707. Among them:
[0130] The multi-source data stream acquisition module 701 is used to acquire the multi-source data stream pushed in real time by the target push component, where the target push component has set source difference information for the multi-source data stream in advance according to the data stream source;
[0131] The actual data attribute field acquisition module 702 is used to parse the multi-source data stream based on a preset parsing component, and acquire the data attribute fields respectively included in the multi-source data stream as actual data attribute fields;
[0132] A target association result table acquisition module 703 is configured to acquire a target association result table, where the target association result table is designed in advance according to data attribute fields finally required by a target association service;
[0133] An expected data attribute field acquisition module 704 is configured to perform form parsing on the target association result table, and identify data attribute fields included in the target association result table according to a form parsing result as expected data attribute fields;
[0134] A comparison and determination module 705 is configured to compare the expected data attribute fields with the actual data attribute fields, and determine target data streams respectively corresponding to all the expected data attribute fields according to a comparison result;
[0135] A service data acquisition module 706 is configured to acquire service data finally required by the target association service from a target data resource library according to the target data streams respectively corresponding to all the expected data attribute fields and source difference information of the multi-source data streams;
[0136] A multi-source data stream association module 707 is configured to map the service data into the association result table according to the expected data attribute fields, and complete the association of the multi-source data streams.
[0137] This application completes the association of the multi-source data streams by pushing the multi-source data streams, determining the expected data attribute fields, identifying the actual data attribute fields, acquiring from a target data resource library in a request manner, constructing a virtual form, and mapping the data content in the virtual form into an actual association result table. Directly according to the target data streams respectively corresponding to all the expected data attribute fields and the source difference information of the multi-source data streams, the entire form where the expected data attribute fields are located is directly acquired. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When acquiring, there is no need to query according to the data attribute fields first, and only the entire form needs to be acquired, which is faster, and then the form can be processed later. To a certain extent, the association speed of the multi-source data streams in financial services is improved, and the real-time nature of financial service data processing is ensured. The accuracy of the multi-source data stream association result is ensured through the virtual form, and it is also ensured that when the association is delayed, the virtual form can be directly used to support financial service data, avoiding delays in financial service tasks.
[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When this program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), etc., or a random access memory (RAM), etc.
[0139] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0140] To solve the above technical problems, the embodiments of the present application also provide a computer device. For details, please refer to Figure 8 , Figure 8 which is the basic structural block diagram of the computer device in this embodiment.
[0141] The computer device 8 includes a memory 8a, a processor 8b, and a network interface 8c that are communicatively connected to each other through a system bus. It should be noted that only the computer device 8 with components 8a - 8c is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of this technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0142] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0143] The memory 8a includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disc, etc. In some embodiments, the memory 8a may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 8a may also be an external storage device of the computer device 8, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 8. Of course, the memory 8a may also include both the internal storage unit and the external storage device of the computer device 8. In this embodiment, the memory 8a is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions of a data flow association method. In addition, the memory 8a can also be used to temporarily store various data that have been output or will be output.
[0144] In some embodiments, the processor 8b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data flow association chips. The processor 8b is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 8b is used to run the computer-readable instructions stored in the memory 8a or process data, such as running the computer-readable instructions of the data flow association method.
[0145] The network interface 8c may include a wireless network interface or a wired network interface, and the network interface 8c is generally used to establish a communication connection between the computer device 8 and other electronic devices.
[0146] The computer device proposed in this embodiment belongs to the field of fintech and is applied to the scenario of obtaining multi-source data for financial business data. This application completes the association of the multi-source data stream by pushing the multi-source data stream, determining the expected data attribute fields, identifying the actual data attribute fields, obtaining from the target data resource library in a request manner, constructing a virtual form, and mapping the data content in the virtual form into the actual associated result table. Directly based on the target data stream corresponding to each of all the expected data attribute fields and the source difference information of the multi-source data stream, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first. Just obtain the entire form, which is faster, and then process the form later. To a certain extent, the association speed of the multi-source data stream in financial business is improved, and the real-time nature of financial business data processing is ensured.
[0147] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to execute the steps of the data stream association method as described above.
[0148] The computer-readable storage medium proposed in this embodiment belongs to the field of fintech and is applied to the scenario of obtaining multi-source data for financial business data. This application completes the association of the multi-source data stream by pushing the multi-source data stream, determining the expected data attribute fields, identifying the actual data attribute fields, obtaining from the target data resource library in a request manner, constructing a virtual form, and mapping the data content in the virtual form into the actual associated result table. Directly based on the target data stream corresponding to each of all the expected data attribute fields and the source difference information of the multi-source data stream, the entire form where the expected data attribute fields are located is directly obtained. By processing the data content in the form, a form containing only the data content corresponding to the expected data attribute fields is obtained. When obtaining, there is no need to query according to the data attribute fields first. Just obtain the entire form, which is faster, and then process the form later. To a certain extent, the association speed of the multi-source data stream in financial business is improved, and the real-time nature of financial business data processing is ensured.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0150] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.
Claims
1. A data stream association method, characterized in that, Including the following steps: Obtain the multi-source data stream pushed in real time by the target push component, where the target push component has set source discrimination information for the multi-source data stream in advance according to the data stream source; Based on a preset parsing component, parse the multi-source data stream to obtain the data attribute fields respectively included in the multi-source data stream as actual data attribute fields; Obtain the target association result table, where the target association result table is designed in advance according to the data attribute fields finally required by the target association service; Perform form parsing on the target association result table, and identify the data attribute fields included in the target association result table according to the form parsing result as expected data attribute fields; Compare the expected data attribute fields with the actual data attribute fields, and determine the target data stream corresponding to each expected data attribute field according to the comparison result; According to the target data stream corresponding to each expected data attribute field and the source discrimination information of the multi-source data stream, obtain the service data finally required by the target association service from the target data resource library; Map the service data to the association result table according to the expected data attribute fields to complete the association of the multi-source data stream.
2. The data stream association method according to claim 1, wherein The target push component includes a data push component based on a distributed system. Before performing the step of obtaining the multi-source data stream pushed in real time by the target push component, the method further includes: Based on a preset push record log, record the push time when the data push component performs real-time push on each data stream in the multi-source data stream; After performing the step of obtaining the multi-source data stream pushed in real time by the target push component, the method further includes: According to the source discrimination information preset by the target push component for the multi-source data stream, obtain the source discrimination information corresponding to each data stream in the multi-source data stream; Parse the push record log through a preset log parsing component to identify the push time corresponding to each data stream in the multi-source data stream.
3. The data flow association method according to claim 1, wherein The preset parsing component is a data stream parsing component. The step of parsing the multi-source data stream based on the preset parsing component to obtain the data attribute fields respectively included in the multi-source data stream specifically includes: Parse each data stream in the multi-source data stream according to the data stream parsing component to obtain the data content transmitted in each data stream; Based on the data content transmitted in each data stream, identify the data attribute fields respectively included in each data stream; After performing the step of parsing the multi-source data stream based on the preset parsing component to obtain the data attribute fields respectively included in the multi-source data stream, the method further includes: Perform marking processing on the data attribute fields respectively included in each data stream according to the source discrimination information corresponding to each data stream in the multi-source data stream to obtain a marking result, where the marking processing method is specifically to use the source discrimination information corresponding to each data stream as a marking field and assign it to the data attribute fields respectively included in each data stream; Based on the marked results, identify the source difference information corresponding to each actual data attribute field respectively.
4. The data stream association method according to claim 3, wherein The step of comparing the expected data attribute fields with the actual data attribute fields and determining the target data streams corresponding to the expected data attribute fields respectively according to the comparison results specifically includes: Compare the expected data attribute fields with the actual data attribute fields to determine the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields; According to the one-to-one correspondence between the expected data attribute fields and the actual data attribute fields, and the source difference information corresponding to each actual data attribute field respectively, determine the source difference information corresponding to each expected data attribute field respectively; According to the source difference information corresponding to each data stream in the multi-source data streams respectively, and the source difference information corresponding to each expected data attribute field respectively, determine the target data streams corresponding to each expected data attribute field respectively.
5. The data stream association method according to claim 2, wherein The step of obtaining the business data required for the target associated service finally from the target data repository according to the target data streams corresponding to each expected data attribute field respectively and the source difference information of the multi-source data streams specifically includes: According to the target data streams corresponding to each expected data attribute field respectively and the source difference information of the multi-source data streams, send a form acquisition request to the target data repository; Receive the request response result returned by the target data repository based on the form acquisition request; By parsing the request response result, identify the target data forms involved in the business data, the expected data attribute fields included in each target data form, and all the data attribute fields included in each target data form.
6. The data flow association method according to claim 5, wherein The step of mapping the business data to the association result table according to the expected data attribute fields to complete the association of the multi-source data streams specifically includes: According to the expected data attribute fields included in each target data form and all the data attribute fields included in each target data form, perform NULL value processing on the data content corresponding to the non-expected data attribute fields in each target data form to obtain each processed target data form; According to the target data forms involved in the business data and the expected data attribute fields included in each target data form, splice all the processed target data forms by using the UNION form splicing method to obtain a spliced form; According to the expected data attribute fields included in each target data form, obtain the data content corresponding to the expected data attribute fields; According to the push time corresponding to each data stream in the multi-source data streams respectively, the target data streams corresponding to each expected data attribute field respectively, and the data content corresponding to the expected data attribute fields, add the data content corresponding to each expected data attribute field to the spliced form to obtain a spliced form filled with data content; Map the data content in the spliced form to the association result table according to the data attribute fields to complete the association of the multi-source data streams.
7. The data flow association method according to claim 6, wherein The step of adding the data content corresponding to all expected data attribute fields to the splicing form according to the push time respectively corresponding to each data stream in the multi-source data stream, the target data stream respectively corresponding to all expected data attribute fields, and the data content corresponding to the expected data attribute fields, to obtain a splicing form filled with data content, specifically includes: According to the target data stream respectively corresponding to all expected data attribute fields and the data content corresponding to the expected data attribute fields, identify whether the data content corresponding to the same expected data attribute field is pushed successively by two or more data streams; If the data content corresponding to the expected data attribute field is pushed only by one data stream, then according to the source difference information of the data stream, identify the data form corresponding to the expected data attribute field, and obtain the data content corresponding to the expected data attribute field from the data form, and add the data content to the splicing form in a normal insertion manner, where the normal insertion manner is specifically to directly add the data content corresponding to the expected data attribute field to the splicing form; If there is data content corresponding to the same expected data attribute field that is pushed successively by two or more data streams, then according to the push time respectively corresponding to each data stream in the multi-source data stream, filter out the data stream that is pushed last, and according to the source difference information of the data stream that is pushed last, identify the data form corresponding to the expected data attribute field, and obtain the data content corresponding to the expected data attribute field from the data form as the content to be inserted, and add the content to be inserted to the splicing form in an update insertion manner, where the update insertion manner is specifically that if the data content corresponding to the expected data attribute field pushed by the prior data stream has been added to the splicing form, then first delete the data content corresponding to the expected data attribute field in the splicing form, and then add the content to be inserted to the splicing form, and if the data content corresponding to the expected data attribute field pushed by the prior data stream has not been added to the splicing form, directly add the content to be inserted to the splicing form; Until the data content corresponding to all expected data attribute fields is added to the splicing form, a splicing form filled with data content is obtained.
8. A data flow association device, characterized in that, Including: A multi-source data stream acquisition module, configured to acquire a multi-source data stream pushed in real time by a target push component, where the target push component has set source difference information for the multi-source data stream in advance according to the data stream source; An actual data attribute field acquisition module, configured to parse the multi-source data stream based on a preset parsing component, and acquire the data attribute fields respectively included in the multi-source data stream as actual data attribute fields; A target associated result table acquisition module, configured to acquire a target associated result table, where the target associated result table is designed in advance according to the data attribute fields finally required by the target associated service; An expected data attribute field acquisition module, configured to perform form parsing on the target association result table, and identify data attribute fields included in the target association result table according to the form parsing result as expected data attribute fields; A comparison and determination module, configured to compare the expected data attribute fields with the actual data attribute fields, and determine target data streams respectively corresponding to all the expected data attribute fields according to the comparison result; A service data acquisition module, configured to acquire service data required finally for the target association service from a target data resource library according to the target data streams respectively corresponding to all the expected data attribute fields and source difference information of the multi-source data streams; A multi-source data stream association module, configured to map the service data into the association result table according to the expected data attribute fields, and complete the association of the multi-source data streams.
9. A computer device, characterized in that, It includes a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the data stream association method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the data stream association method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and equipment for modeling on heterogeneous data source based on unified grammar
CN116450609A
Bayesian Sleep Fusion
US20140297600A1