Data source switching method, device, equipment and storage medium

By performing multiple comparisons and characteristic parameter analyses during the data source switching process and dividing the pilot areas into groups for batch switching, the problems of low accuracy and security in data source switching were solved, local memory usage was reduced, and data processing efficiency was improved.

CN119311675BActive Publication Date: 2025-09-12INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410573556.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2025-09-12
Estimated Expiration
2044-05-10

AI Technical Summary

Technical Problem

The data source switching process has problems such as low accuracy and security, as well as large local memory usage.

Method used

When the host source table data and the platform source table data are consistent, the data is stored in the data lake, an entry file is generated, and a consistency check is performed; the predicted influence value is determined based on the characteristic parameters of the region, and the region is divided into k pilot switching area groups, and the data source is switched in batches.

Benefits of technology

It improves the accuracy and security of data source switching, reduces local memory usage, and reduces workload and the risk of data omission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311675B_ABST
    Figure CN119311675B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data source switching method, which can be applied to the fields of computer technology, big data technology, or financial technology. The method includes: storing host source table data in a data lake to obtain a host entry file, and storing platform source table data in a data lake to obtain a platform entry file; determining the predicted influence value generated when the data source of the area is switched based on the characteristic parameters of the area; dividing the area into k pilot switching area groups based on the predicted influence value; switching the data sources corresponding to the k pilot switching area groups in batches based on the influence level obtained using the predicted influence value, switching the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switching the data source for generating the host entry file from the host source table data to the platform source table data. The present disclosure also provides a data source switching device, equipment, storage medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of computer technology, big data technology, or financial technology, and specifically to a data source switching method, apparatus, device, storage medium, and program product. Background Art

[0002] With the development of innovation in information technology applications, more and more data sources need to be gradually switched from host systems to platform systems, and the data in the data sources need to be processed using applications on the platform systems.

[0003] In the process of realizing the inventive concept of the present disclosure, the inventors discovered that the following problems generally exist in the related art: when switching data sources, the data source switching system generally switches the data source directly, making it difficult to ensure the accuracy and security of the switched data source. In addition, the data source switching system generally processes the data source locally, which not only makes the local memory occupied by the data source switching system larger, but also brings certain waste to local resources. Overall, the data source switching process of the related art has the problems of low accuracy and security, as well as large local memory occupation. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a data source switching method, apparatus, device, storage medium and program product that improve the accuracy and security of the data source switching process and reduce local memory usage.

[0005] One aspect of the present disclosure provides a data source switching method, the method comprising: when the host source table data and the platform source table data are consistent, storing the host source table data in the data lake to obtain a host entry file, and storing the platform source table data in the data lake to obtain a platform entry file, wherein both the host source table data and the platform source table data include a region identifier; when the host entry file and the platform entry file are consistent, and the processing result file is consistent with the reference processing result file, determining the predicted influence value generated when the data source of the region is switched according to the characteristic parameters of the region, wherein the region includes the above The area represented by the regional identifier, the above-mentioned processing result file is obtained by logically processing the above-mentioned host entry lake file, and the above-mentioned reference processing result file is obtained by calling the above-mentioned host source table data; according to the above-mentioned predicted influence value, the above-mentioned area is divided into k pilot switching area groups to realize k batches of switching data sources, k is a positive integer; according to the influence level obtained by using the above-mentioned predicted influence value, the data sources corresponding to the above-mentioned k pilot switching area groups are switched batch by batch, so as to switch the data source for generating the anti-corrosion file from the above-mentioned reference processing result file to the above-mentioned processing result file, and switch the data source for generating the above-mentioned host entry lake file from the above-mentioned host source table data to the above-mentioned platform source table data.

[0006] Another aspect of the present disclosure further provides a data source switching device, the device comprising: a storage module for storing the host source table data in a data lake to obtain a host entry file, and storing the platform source table data in the data lake to obtain a platform entry file, when the host source table data and the platform source table data are consistent, wherein both the host source table data and the platform source table data include a region identifier; a determination module for determining, when the host entry file and the platform entry file are consistent, and the processing result file is consistent with the reference processing result file, the predicted influence value generated when the data source of the above region is switched according to the characteristic parameters of the region, wherein the above region includes the above The area represented by the area identifier, the above-mentioned processing result file is obtained by logically processing the above-mentioned host entry lake file, and the above-mentioned reference processing result file is obtained by calling the above-mentioned host source table data; the division module is used to divide the above-mentioned area into k pilot switching area groups according to the above-mentioned predicted influence value, so as to realize k batch switching of data sources, k is a positive integer; the switching module is used to switch the data sources corresponding to the above-mentioned k pilot switching area groups in batches according to the influence level obtained by using the above-mentioned predicted influence value, so as to switch the data source for generating the anti-corrosion file from the above-mentioned reference processing result file to the above-mentioned processing result file, and switch the data source for generating the above-mentioned host entry lake file from the above-mentioned host source table data to the above-mentioned platform source table data.

[0007] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the data source switching method.

[0008] Another aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instruction stored thereon, which implements the steps of the above-mentioned data source switching method when the computer program or instruction is executed by a processor.

[0009] Another aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implements the steps of the above-mentioned data source switching method when executed by a processor.

[0010] According to the data source switching method, apparatus, device, storage medium and program product provided by the embodiments of the present disclosure, a host entry-lake file and a platform entry-lake file are obtained when the host source table data and the platform source table data are consistent; when the host entry-lake file and the platform entry-lake file are consistent, and the processing result file and the reference processing result file are consistent, the predicted influence value of switching the data source is determined according to the characteristic parameters of the region; the region is divided into k pilot switching region groups according to the predicted influence value; and the data sources of the k pilot switching region groups are switched batch by batch according to the influence level. Since the data and files are compared multiple times during the data source switching process, and when the comparisons are consistent, the predicted influence value of the switching data source is determined based on the characteristic parameters of the area where the data source is to be switched, and the switching is performed batch by batch based on the influence level determined by the predicted influence value, this at least partially overcomes the problem of low data switching accuracy and security caused by direct switching of data sources in related technologies. Moreover, the processing result file of the present application is obtained by processing the host entry file after the host source table data is stored in the data lake, and there is no need to process the host source table data locally. This at least partially overcomes the problem of large local memory occupied by the data source switching system, thereby achieving the technical effect of improving the accuracy and security of the data source switching process and reducing the local memory occupation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0012] Figure 1 Schematically illustrates an application scenario diagram of a data source switching method, apparatus, device, storage medium, and program product according to an embodiment of the present disclosure;

[0013] Figure 2The following schematically shows a flow chart of a data source switching method according to an embodiment of the present disclosure;

[0014] Figure 3 A system architecture diagram of a data source switching method according to an embodiment of the present disclosure is schematically shown;

[0015] Figure 4 A schematic diagram schematically illustrates data comparison during data source switching according to an embodiment of the present disclosure;

[0016] Figure 5 Schematically shows an architecture diagram for generating a predicted influence value according to an embodiment of the present disclosure;

[0017] Figure 6 Schematically shows a flow chart of data source switching according to another embodiment of the present disclosure;

[0018] Figure 7 A structural block diagram of a data source switching device according to an embodiment of the present disclosure is schematically shown; and

[0019] Figure 8 A block diagram of an electronic device suitable for implementing a data source switching method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0020] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0021] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0023] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0024] In the technical solution disclosed herein, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0025] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide users with corresponding operation portals for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge, and skills, and have reached a certain level of professionalism.

[0026] Currently, data source switching systems typically occupy local resources, processing host source table data locally. Once the processed results are obtained, they are stored in the data lake. This not only results in a large amount of local memory occupied by the data source switching system, but also results in a certain waste of local system resources. Furthermore, data source switching systems typically directly replace the platform data source with the host data source and independently generate two lines for storage in the data lake. When downstream callers need to access data, they generally need to retrieve and organize all files on the line to facilitate the transformation. This not only results in a large workload and a high risk of data omissions, but also leads to low accuracy and low security in the data switching process.

[0027] In view of this, embodiments of the present disclosure provide a data source switching method, apparatus, device, storage medium, and program product for improving the accuracy and security of data source switching, as well as reducing local memory usage, workload, and the risk of data omission. Specifically, the method includes: when the host source table data and the platform source table data are consistent, storing the host source table data in the data lake to obtain a host entry file, and storing the platform source table data in the data lake to obtain a platform entry file, wherein both the host source table data and the platform source table data include a region identifier; when the host entry file and the platform entry file are consistent, and the processing result file and the reference processing result file are consistent, determining, according to the characteristic parameters of the region, a predicted influence value generated when the data source of the region is switched, wherein the region includes the region represented by the region identifier, the processing result file is obtained by logically processing the host entry file, and the reference processing result file is obtained by calling based on the host source table data; dividing the region into k pilot switching region groups according to the predicted influence value to realize k batch switching of data sources, where k is a positive integer; according to the influence level obtained by using the predicted influence value, switching the data sources corresponding to the k pilot switching region groups in batches, so as to switch the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switch the data source for generating the host entry file from the host source table data to the platform source table data.

[0028] It should be noted that the data source switching method and device determined in the embodiments of the present disclosure can be used in the fields of computer technology, big data technology or financial technology, and can also be used in any field other than the fields of computer technology, big data technology or financial technology. The embodiments of the present disclosure do not limit the application fields of the determined data source switching method and device.

[0029] Figure 1 The application scenario diagram of the data source switching method, apparatus, device, storage medium and program product according to the embodiments of the present disclosure is schematically shown.

[0030] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0031] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, such as sending a data source switching request or receiving a data source switching result. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as data source switching applications, host system applications, platform system applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0032] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0033] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports data source switching requests sent by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received requests and other data, and feed back processing results (e.g., web pages, information, or data obtained or generated according to the requests) to the terminal devices.

[0034] It should be noted that the data source switching method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the data source switching device provided in the embodiment of the present disclosure can generally be set in the server 105. The data source switching method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the data switching device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0036] The following will be based on Figure 1 The scene described by Figures 2 to 5 The data source switching method of the disclosed embodiment is described in detail.

[0037] Figure 2The flowchart of the data source switching method according to an embodiment of the present disclosure is schematically shown.

[0038] like Figure 2 As shown, the data source switching method of this embodiment includes operations S210 to S240.

[0039] In operation S210, if the host source table data and the platform source table data are consistent, the host source table data is stored in the data lake to obtain a host entry file, and the platform source table data is stored in the data lake to obtain a platform entry file, wherein both the host source table data and the platform source table data include region identifiers.

[0040] In operation S220, when the host entry file and the platform entry file are consistent, and the processing result file is consistent with the reference processing result file, the predicted influence value generated when the data source of the area is switched is determined based on the characteristic parameters of the area, where the area includes the area represented by the area identifier, the processing result file is obtained by logically processing the host entry file, and the reference processing result file is obtained by calling based on the host source table data.

[0041] In operation S230 , the region is divided into k pilot switching region groups according to the predicted influence values ​​to implement k batches of switching data sources, where k is a positive integer.

[0042] In operation S240, according to the influence level obtained by using the predicted influence value, the data sources corresponding to the k pilot switching area groups are switched batch by batch, so as to switch the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switch the data source for generating the host entry file from the host source table data to the platform source table data.

[0043] Optionally, a host system can run on the host, and a platform system can run on the platform. The platform system can be an autonomous and controllable system. With the development of innovative information technology applications, more and more industries require the use of autonomous and controllable platform systems for related business processing. Therefore, it is necessary to switch the system providing the data source from the host to the platform.

[0044] Optionally, the host source table data may be data from a data source provided by the host system, and the platform source table data may be data from a data source provided by the platform system. In one embodiment, when a transaction needs to be stored, a dual-write system of related art can be used to simultaneously store the information in the transaction registration table (e.g., the region identifier of the transaction area, the transaction account, the transaction amount, the transaction currency, and other information) in both the host system's data source and the platform system's data source.

[0045] It is understood that the process of storing information in the transaction registration form can be conducted with the permission of the information provider. For example, before storing the information in the transaction registration form, a request to retrieve and store the information can be sent to the information provider, and a corresponding operation interface can be provided for the user to authorize or deny. If the information provider agrees or authorizes the retrieval and storage of the information, the storage operation can proceed. If the information provider denies the retrieval and storage operation, the expert decision-making process can be entered. The entire process complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, and does not violate public order and good morals.

[0046] Optionally, when the host source table data and the platform source table data are consistent, it can be considered that the data source of the host system is preliminarily consistent with the data source of the platform system. At this time, the host source table data can be stored in the data lake in the form of a host source table entry file to obtain a host entry file (for example, a host source table entry table), and the platform source table data can be stored in the data lake in the form of a platform source table entry file to obtain a platform entry file (for example, a platform source table entry table). It can be understood that the contents of the host entry file and the host source table entry file, the platform entry file and the platform source table entry file can be the same. The difference is that the host entry file and the platform entry file are the file names stored after the host source table entry file and the platform source table entry file are stored in the data lake. In one embodiment, the data lake can use a distributed storage method to store host entry files and platform entry files.

[0047] Optionally, a comparison can be performed between the host and platform files entering the data lake. The comparison results can be used to determine whether there are storage errors when storing the host source table data and the platform source table data in the data lake, and it can also be considered that the data source of the host system is further consistent with the data source of the platform system.

[0048] Optionally, the processing result file (such as a processing result table) is obtained by logically processing the host entry file in the data lake, and the reference processing result file (such as a reference processing result table) is obtained by calling based on the host source table data. For example, the relevant technology will process the host source table data in the local system, and after obtaining the processing result (such as the host processing result entry file), the processing result is stored in the data lake, thereby obtaining a reference processing result file. The reference processing result file can be an existing file, and the embodiment of the present disclosure can be directly called according to the identifier of the host source table data when used. By comparing the processing result file and the reference processing result file, not only can it be determined whether there is a storage error problem in the processing process in the data lake based on the comparison results, but it can also be considered that the data source of the host system is consistent with the data source of the platform system, thereby providing support for the subsequent data source switching process.

[0049] Optionally, if all of the above comparisons are consistent, a predicted impact value of switching the data source can be determined based on characteristic parameters of the region, such as the total number of user groups in the sub-regions of the region, the types of users in the user groups, and the traffic volume of the application groups in the sub-regions. It will be appreciated that the predicted impact value can be used to represent the expected impact of switching the data source in the region.

[0050] Optionally, the region can be divided into k pilot switching region groups according to the predicted influence value, and the data sources corresponding to the k pilot switching region groups are switched in batches according to the influence levels obtained using the predicted influence value until all data sources in the region are switched.

[0051] Optionally, the specific operation of data source switching may include switching the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switching the data source for generating the host entry file from the host source table data to the platform source table data.

[0052] Optionally, the anti-corrosion file can be a file for receiving data from upstream (such as a data source) and providing calls to downstream (such as a data caller). The meaning of the word "anti-corrosion" can indicate that the data format in the file is fixed, that is, when the format of the received upstream data is arbitrarily changed, the data called by the downstream caller can still be the data format in the anti-corrosion file, and will not be affected by the change in the upstream data format. The data format in the anti-corrosion file can also shield the modification of the downstream caller. The content in the anti-corrosion file can be determined according to the needs of the downstream caller, and the required content is transmitted to the anti-corrosion file according to the required content of the downstream caller. Each time the downstream caller requires different content, it is only necessary to implement file replacement, and then there is no need to call out all the files on the line and sort them out to cooperate with the downstream caller, thereby reducing the workload and the risk of data omissions, and improving the use efficiency of the downstream caller.

[0053] According to the data source switching method, apparatus, device, storage medium and program product provided by the embodiments of the present disclosure, a host entry-lake file and a platform entry-lake file are obtained when the host source table data and the platform source table data are consistent; when the host entry-lake file and the platform entry-lake file are consistent, and the processing result file and the reference processing result file are consistent, the predicted influence value of switching the data source is determined according to the characteristic parameters of the region; the region is divided into k pilot switching region groups according to the predicted influence value; and the data sources of the k pilot switching region groups are switched batch by batch according to the influence level. Since the data and files are compared multiple times during the data source switching process, and when the comparisons are consistent, the predicted influence value of the switching data source is determined based on the characteristic parameters of the area where the data source is to be switched, and the switching is performed batch by batch based on the influence level determined by the predicted influence value, this at least partially overcomes the problem of low data switching accuracy and security caused by direct switching of data sources in related technologies. Moreover, the processing result file of the present application is obtained by processing the host entry file after the host source table data is stored in the data lake, and there is no need to process the host source table data locally. This at least partially overcomes the problem of large local memory occupied by the data source switching system, thereby achieving the technical effect of improving the accuracy and security of the data source switching process and reducing the local memory occupation.

[0054] Figure 3 The system architecture diagram of the data source switching method according to an embodiment of the present disclosure is schematically shown.

[0055] like Figure 3 As shown, the data lake of the embodiment of the present disclosure may include a first sub-data lake 301 and a second sub-data lake 302. The first sub-data lake 301 may be used to store data, and the second sub-data lake 302 may be used to process the data in the first sub-data lake 301.

[0056] The reference processing result file used in the embodiments of this disclosure can be obtained as follows: Related technologies process host source table data in a local system to obtain host processing result entry file 303, which is then stored in the first sub-data lake 301, thereby obtaining reference processing result file 304. In related technologies, the data source of anti-corrosion file 305 can be this reference processing result file 304.

[0057] In the disclosed embodiments, the process of logically processing a host-entered file to obtain a processed result file may include the following operations: In the second sub-data lake, logically processing the host-entered file with multi-table joint values ​​to obtain a processed result file. Multi-table joint value extraction can be a process of merging the value results of multiple tables.

[0058] According to the embodiments of the present disclosure, the process of processing the host source table data to obtain the processing result file is transferred from the original processing in the local system occupying local resources to the data lake, and the high efficiency advantage of the distributed high scalability of the data lake is fully utilized. This not only improves the data processing efficiency, but also releases the resources of the local system, saves the memory usage of the data source switching system, and improves the utilization rate of local system resources.

[0059] like Figure 3 As shown, the host source table data 306 is stored in the first sub-data lake 301 to obtain a host entry file 307 , and the host entry file 307 is logically processed in the second sub-data lake 302 to obtain a processing result file 308 .

[0060] Optionally, the platform source table data 309 is stored in the first sub-data lake 301 , and a platform entry lake file 310 can be obtained.

[0061] Optionally, continue with reference to Figure 3 The data source switching process of the embodiment of the present disclosure can be divided into two parts. The first part can be to replace the reference processing result file 304 with the processing result file 308; the second part can be to replace the host source table data 306 with the platform source table data 309.

[0062] Optionally, the second switching process may include the following operations: calling the platform source table data; and regenerating the host entry lake file based on the platform source table data to switch the host source table data to the platform source table data. For example, the platform source table data may be restored to the first sub-data lake to load the platform source table data into the host entry lake file.

[0063] Figure 4 The following schematically illustrates data comparison during data source switching according to an embodiment of the present disclosure.

[0064] like Figure 4 As shown, the embodiment of the present disclosure can perform comparisons between the host source table data 306 and the platform source table data 309, between the host entry-lake file 307 and the platform entry-lake file 310, and between the processing result file 308 and the reference processing result file 304, and perform data switching when the host source table data 306 and the platform source table data 309 are consistent, the host entry-lake file 307 and the platform entry-lake file 310 are consistent, and the processing result file 308 and the reference processing result file 304 are consistent, so as to effectively ensure the accuracy of the data source switching process.

[0065] In one embodiment, the comparison process of host source table data and platform source table data, host entry-lake files and platform entry-lake files, and processing result files and reference processing result files may include at least one of the following: performing consistency verification between host source table data and platform source table data, between host entry-lake files and platform entry-lake files, and between processing result files and reference processing result files, wherein the consistency verification includes at least one of the following: verifying the total transaction amount, verifying the field format of transaction elements, and verifying the name of transaction elements; selecting the field value of the primary key field for accuracy verification between host source table data and platform source table data, between host entry-lake files and platform entry-lake files, and between processing result files and reference processing result files; querying the occurrence frequency of the primary key field between host source table data and platform source table data, between host entry-lake files and platform entry-lake files, and between processing result files and reference processing result files to perform uniqueness verification.

[0066] Optionally, when any of the consistency check, accuracy check, and uniqueness check fails, a comparison error file is generated based on the file that fails the check; and the comparison error file is pushed to the target object.

[0067] Optionally, the consistency check can be, for example, comparing the total transaction volume, the field format of the transaction elements, and the names of the transaction elements between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file to see if they are consistent. For example, if the total transaction volume, the field format of the transaction elements, and the names of the transaction elements between the host source table data and the platform source table data are completely consistent, it can be considered that the host source table data and the platform source table data are consistent; if there is an incomplete consistency, a comparison error file can be generated based on the incompletely consistent elements (such as the total transaction volume, the field format of the transaction elements, and the names of the transaction elements, etc.) to push it to the target object (such as a person or automated system or automated robot that can process the comparison error file), and the target object processes the comparison error file. The comparison between the host entry file and the platform entry file, and between the processing result file and the reference processing result file can refer to the above process and will not be repeated here.

[0068] Optionally, the accuracy check is, for example, to compare whether the information of the primary key fields between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file is consistent. For example, if the user number, user account, transaction currency, and transaction amount and other information between the host source table data and the platform source table data are completely consistent. If the user number, user account, transaction currency, and transaction amount and other information between the host source table data and the platform source table data are completely consistent, it can be considered that the host source table data and the platform source table data are consistent; if they are not completely consistent, a comparison error file can be generated based on the incompletely consistent elements (such as user number, user account, transaction currency, and transaction amount and other information) to push it to the target object, and the target object processes the comparison error file. The comparison between the host entry file and the platform entry file, and between the processing result file and the reference processing result file can refer to the above process and will not be repeated here.

[0069] Optionally, the uniqueness check is, for example, to compare whether the frequency of occurrence of the primary key fields between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file is consistent, so as to ensure that there is no duplication or omission between the data or files. For example, if the frequency of occurrence of fields such as user number, user account, transaction currency, and transaction amount or the number of primary key fields between the host source table data and the platform source table data is consistent, it can be considered that the host source table data and the platform source table data are consistent; if the frequency of occurrence of the fields or the number of primary key fields are inconsistent, a comparison error file can be generated based on the incompletely consistent primary key fields (such as user number, user account, transaction currency, and transaction amount, etc.) to push to the target object, and the target object processes the comparison error file. The comparison between the host entry file and the platform entry file, and between the processing result file and the reference processing result file can refer to the above process and will not be repeated here.

[0070] It is understood that the comparison of information such as user ID, user account number, transaction currency, and transaction amount can be conducted with the permission of the information provider. For example, before the information is compared, a request to obtain and compare the information can be sent to the information provider, with a corresponding entry point provided for the user to authorize or deny. If the information provider agrees or authorizes the acquisition and comparison of the information, the comparison will proceed; if the information provider denies the acquisition and comparison, the expert decision-making process will proceed. The entire process complies with relevant laws, regulations, and standards, employs necessary confidentiality measures, and does not violate public order and good morals.

[0071] According to an embodiment of the present disclosure, by comparing the host source table data with the platform source table data, the host lake-in file with the platform lake-in file, and the processed result file with the reference processed result file, and when they are consistent, performing the operation of switching the data source, which can improve the accuracy rate of the data source switching process.

[0072] Optionally, when the above-mentioned consistency check, accuracy check, and uniqueness check are all performed between the host source table data and the platform source table data, the host lake-in file and the platform lake-in file, and the processed result file and the reference processed result file, the operation of determining the prediction influence value can be performed.

[0073] In one embodiment, the area may include sub-areas, and the characteristic parameters may include the total number of user groups in the sub-areas, the types of users in the user groups, and the business volume of the application groups in the sub-areas; the process of determining the prediction influence value generated when switching the data source of the area may include the following operations: generating a first influence value according to the total number of user groups in the sub-areas, and screening out the target sub-areas from the sub-areas according to the first influence value; within the target sub-areas, generating a second influence value according to the types of users; within the target sub-areas, generating a third influence value according to the business volume of the application groups in the sub-areas per unit time period; generating a prediction influence value according to the first influence value, the second influence value, and the third influence value.

[0074] Optionally, the calculation method of the total number of user groups N may be the total number of users in a single sub-area. In one embodiment, the value-taking logic of the first influence value n of the total number of user groups N may be as follows: n(0≤N≤1 million)=1; n(1 million<N≤5 million)=2; n(5 million<N≤10 million)=3; n(N>10 million)=4.

[0075] Optionally, the process of screening out the target sub-areas from the sub-areas according to the first influence value may include the following operations: generating the priority of the sub-areas according to the first influence value; sorting the sub-areas according to the priority to obtain the first sorting result; screening out the sub-areas that meet the predetermined conditions from the first sorting result to obtain the target sub-areas.

[0076] Optionally, in order to avoid the impact brought by the data source switching process, the area with the lowest influence value can be used as the priority area. Exemplarily, if there are 5 sub-areas in the area, namely sub-area A, sub-area B, sub-area C, sub-area D, and sub-area E, after matching the above-mentioned value-taking logic of the first influence value, the first influence values of sub-area A, sub-area B, sub-area C, sub-area D, and sub-area E can be: n = 1, n = 2, n = 3, n = 4, n = 1. The priorities of the sub-areas generated according to the first influence value can be: first priority, second priority, third priority, fourth priority, and first priority (the priority degree from the first priority to the fourth priority can gradually decrease).

[0077] The sorting result obtained by sorting the sub-areas according to this priority can be sub-area A and area E, sub-area B, sub-area C, sub-area D.

[0078] Sub-areas A and area E that meet the predetermined conditions (for example, the highest priority and also the lowest first influence value) can be selected from the sorting result as the target sub-areas for piloting from the target sub-areas.

[0079] Optionally, within the target sub-area, a second influence value g can be generated according to the type G of the user. In one embodiment, the value-taking logic of the second influence value g of the user type G can be as follows: g (small enterprise with enterprise scale less than the predetermined scale threshold) = 1; g (large enterprise with enterprise scale greater than or equal to the predetermined scale threshold) = 2; g (enterprise meeting the predetermined attributes among small enterprises) = 3; g (institutional enterprise) = 4; g (others) = 5. The predetermined scale threshold can be adaptively modified according to actual needs; the predetermined attributes (such as real estate, etc.) can be adaptively adjusted according to actual needs; others can include the remaining enterprises that do not belong to g = 1, 2, 3, and 4.

[0080] Optionally, within the target sub-area, a third influence value a can also be generated according to the application group business volume A of the sub-area within a unit time period (such as within a single day). In one embodiment, the value-taking logic of the third influence value a of the application group business volume A can be as follows: a (0 ≤ A ≤ 10,000) = 1; a (10,000 < A ≤ 50,000) = 2; a (50,000 < A ≤ 100,000) = 3; a (A > 100,000) = 4.

[0081] Figure 5 Schematically shows an architecture diagram for generating a predicted influence value according to an embodiment of the present disclosure.

[0082] As Figure 5As shown, when generating the predicted influence value, the total number of user groups N in each sub-area (for example, N1, N2, N3, N4, ...) can be calculated first, and then all user types G under the total number of user groups in each sub-area (for example, G1, G2, G3, G4, G5, ...) can be determined, and then the sub-area application group business volume A under each user type can be determined (for example, A1, A2, A3, A4, ...), and the first influence value, the second influence value and the third influence value can be determined in turn according to the above-mentioned value logic.

[0083] Optionally, the embodiment of the present disclosure may configure a data source switch to control the range of each switch. Specifically, the switching range of the data source switch may be Q = {Z, C, B}, where sub-area Z = {z1, z2, z3, ..., z r}; User type number C = {c1, c2, c3, ..., c m}; Apply B = {b1, b2, b3, ..., b p}. Where r, m and p are all positive integers.

[0084] The process of generating a predicted influence value according to the first influence value, the second influence value, and the third influence value can be shown as formula (1).

[0085] Predicted influence value = n(N(z i ))×0.5+g(G(c j ))×0.2+a(A(b x ))×0.3 (1)

[0086] Among them, i belongs to 1,...,r; j belongs to 1,...,m, and x belongs to 1,...,p.

[0087] According to the embodiments of the present disclosure, by jointly generating a predicted influence value based on the first influence value, the second influence value, the third influence value and their respective weights, the impact of the total number of user groups, user types and the total amount of application group business on switching data sources can be personalized considered, thereby improving the reference value of the predicted influence value, thereby providing convenience for improving the accuracy and security of the data source switching process.

[0088] Optionally, once the predicted influence values ​​are obtained, the region can be divided into k pilot switching region groups using the predicted influence values ​​to implement k-batch data source switching. Specifically, the process may include the following operations: sorting the predicted influence values ​​to obtain a second sorting result; performing k-quantile quantification on the second sorting result to obtain k-1 k-quantile nodes; and dividing the region into k pilot switching region groups based on the k-1 k-quantile nodes to implement k-batch data source switching.

[0089] Optionally, the above formula (1) can be used to obtain the predicted influence value for the total number of user groups in each sub-region, for each user type, and for each business group in the region. By arranging all the predicted influence values ​​in the region in ascending order, a second ranking result can be obtained.

[0090] In the second sorting result, k quantile values ​​are taken, resulting in k-1 k-point nodes. Taking quartile values ​​as an example, in the second sorting result, quartile values ​​are taken, resulting in three quartile nodes, which are then arranged into four equal parts from smallest to largest. The predicted influence value of each part can correspond to one or more sub-regions, and the pilot switching region group can be composed of these one or more sub-regions. Each time, the data source of the pilot switching region group with the lowest influence is switched, achieving k data source switches, thereby achieving a smooth transition of the data source switching process.

[0091] Optionally, when switching the data sources of the pilot switching area groups, the switching can be performed batch by batch based on the influence levels obtained using the predicted influence values. Specifically, the process may include the following operations: determining k influence levels for the k pilot switching area groups based on the positions of the k-quantile nodes in the second sorting result; determining the switching order for switching the data sources of the pilot switching area groups based on the k influence levels; and sequentially switching the data sources of the pilot switching area groups with the lowest influence levels among the k pilot switching area groups according to the switching order, until the data sources of all the k pilot switching area groups are switched.

[0092] Optionally, k-1 nodes can be divided into k equal parts from small to large. Taking three quartile nodes (i.e., the first quartile node, the second quartile node, and the third quartile node) as an example, the three quartile nodes can divide the predicted influence value from small to large into four equal parts (i.e., the first pilot switching area group before the first quartile node, the second pilot switching area group between the first quartile node and the second quartile node, the third pilot switching area group between the second quartile node and the third quartile node, and the fourth pilot switching area group after the third quartile node). Taking the order of the predicted influence values ​​from small to large from left to right as an example, the positions of the first quartile node, the second quartile node, and the third quartile node in the second sorting result can also be from left to right. Accordingly, the influence levels of the first pilot switching area group, the second pilot switching area group, the third pilot switching area group, and the fourth pilot switching area group can be the first influence level, the second influence level, the third influence level, and the fourth influence level, respectively (the degree of influence from the first influence level to the fourth influence level can be gradually increased).

[0093] The order of switching the data sources of the pilot switching area groups determined based on the influence level may be: switching the first pilot switching area group, the second pilot switching area group, the third pilot switching area group, and the fourth pilot switching area group in sequence. The switching of the data sources can be completed according to this order.

[0094] Optionally, when switching data sources, the data sources of the pilot switching region group with the lowest influence level among the k pilot switching region groups may be switched sequentially. Specifically, this process may include the following operations: determining a switching range for the pilot switching region group with the lowest influence level based on the influence level; gradually expanding the switching range at a predetermined frequency; and within the switching range, switching the data source for generating anti-corrosion files from reference processing result files to processing result files, and switching the data source for generating host input files from host source table data to platform source table data, until all data sources of the pilot switching region groups are switched.

[0095] Alternatively, for example, among the first, second, third, and fourth pilot switching area groups, the first pilot switching area group, which has the lowest level of influence, may be determined to be the first pilot switching area group to be switched first. Because this pilot switching area group has a low influence, the scope of the first switch may be 15%. 15% of the host source table data may be switched to platform source table data, and 15% of the reference result processing file data may be switched to the processing result file. After reading according to the above strategy, the read files are merged. The scope of the second switch may be expanded by 5% at a predetermined frequency, for example, to 20% of the remaining data, and the switching process is the same as above. The scope of the third switch may be expanded by 10% at a predetermined frequency, for example, to 30% of the remaining data, and the switching process is the same as above. This process continues in this manner until all data sources of the first pilot switching area group are switched from host source table data to platform source table data, and the reference result processing file is switched to the processing result file.

[0096] After all data sources in the first pilot switching region group have been switched, the second pilot switching region group with the lowest influence level can be selected from the second, third, and fourth pilot switching region groups based on their influence levels for data source switching. When switching data sources, because the influence of the second pilot switching region group is higher than that of the already switched first pilot switching region group, the first data source switching range in the second pilot switching region group can be 10%, expanding at a predetermined frequency of 3%. The second switching range can be 13%, the third switching range can be 16%, and so on. The switching process continues as above, and so on, until the data source of the second pilot switching region group is completed.

[0097] After all data sources in the second pilot switching region group have been switched, the third pilot switching region group with the lowest influence level can be selected from the third and fourth pilot switching region groups based on their influence level for data source switching. This switching range can start at 8% and expand at a predetermined frequency of 2% until all data sources in the third pilot switching region group are completed.

[0098] After all the data sources of the third pilot switching area group are switched, the fourth pilot switching area group can be switched for data source. Since the influence level of this pilot switching area is higher than the influence levels of the first three pilot switching area groups, the switching range can start from 5% and expand at a predetermined frequency of 1% until the data source of the fourth pilot switching area group is completed. It should be noted that the switching range and predetermined frequency of the above process can be adaptively adjusted according to actual needs. For example, when it is not the first time to switch the data source, the predetermined frequency can be increased or decreased according to the result of the previous switching of the data source. For example, if there is no error in the result of the previous or previous multiple switching of the data source, the predetermined frequency can be appropriately increased; if there is an error in the result of the previous or previous multiple switching of the data source, the predetermined frequency can be appropriately reduced or not used.

[0099] According to an embodiment of the present disclosure, by switching the data source according to the predicted influence value from low to high and the switching range from small to large, the smooth switching of the data source can be effectively guaranteed, thereby improving the security of the data source switching.

[0100] Figure 6 The flowchart of data source switching according to another embodiment of the present disclosure is schematically shown.

[0101] like Figure 6 As shown, the process of switching the data source in this embodiment may include operations S610 to S630.

[0102] In operation S610 , the file is stored in the data lake.

[0103] For example, data from host source tables is imported into the lake and stored in the host inbound file. Data from platform source tables is also imported into the lake and stored in the platform inbound file. In the second sub-data lake, logical processing (such as multi-table joins) is performed based on the host inbound file. The inbound and processing processes are described above and will not be repeated here.

[0104] In operation S620 , a data comparison process is performed.

[0105] For example, we compare host source table data with platform source table data, host inbound files with platform inbound files, and processed result files with reference processed result files. The comparison process is similar to the one described above and will not be repeated here.

[0106] In operation S630, the switching of the data source is completed.

[0107] The data source switching switch can be used for control. By using influence, priorities can be divided according to the total number of regional user groups N. The target area can be determined from small to large. Within the target area, a small-scale control can be switched according to the customer classification G and the total amount of application group business A. The degree of influence can be divided into four levels from low to high, and the flow can be switched in batches. Within the pilot range, new data sources are taken (such as the platform source table data or processing result files after switching), and outside the range, old data sources are taken (such as the host source table data or reference processing result files before replacement). After reading them separately, the data files are merged, and the flow switching range is gradually expanded until a smooth transition to full-line flow switching is achieved. Optionally, the data source switching switch can also have a one-button switchback function to achieve an emergency switchback guarantee function during the data source switching process. The specific process can refer to the data source switching process described above, which will not be repeated here.

[0108] The data source switching method provided by the disclosed embodiments can switch reference processing result files to processing result files, and host source table data to platform source table data. Using an anti-corrosion table model, the method allows for gradual, pilot-based switching. Data comparison tasks are set up at key nodes to monitor data consistency and accuracy at each stage, ensuring a safe and smooth transition. It also provides data source switching control for emergency fallback.

[0109] According to the embodiments of the present disclosure, by migrating the data processing process to the data lake, the advantages of the distributed, highly scalable and efficient data lake are utilized to improve data processing efficiency and release local system resources.

[0110] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0111] Based on the above data source switching method, the present disclosure also provides a data source switching device. Figure 7 The device is described in detail.

[0112] Figure 7 The structural block diagram of the data source switching device according to an embodiment of the present disclosure is schematically shown.

[0113] like Figure 7 As shown, the data source switching device 700 of this embodiment includes a storage module 710 , a determination module 720 , a division module 730 and a switching module 740 .

[0114] Storage module 710 is used to store the host source table data in the data lake to obtain a host entry file, and to store the platform source table data in the data lake to obtain a platform entry file, when the host source table data and the platform source table data are consistent. Both the host source table data and the platform source table data include region identifiers.

[0115] Determination module 720 is used to determine the predicted influence value generated when the data source of the area is switched according to the characteristic parameters of the area when the host entry file and the platform entry file are consistent, and the processing result file is consistent with the reference processing result file. The area includes the area represented by the area identifier, the processing result file is obtained by logically processing the host entry file, and the reference processing result file is obtained by calling based on the host source table data.

[0116] The division module 730 is used to divide the region into k pilot switching region groups according to the predicted influence value to implement k batches of switching data sources, where k is a positive integer.

[0117] The switching module 740 is used to switch the data sources corresponding to the k pilot switching area groups in batches according to the influence level obtained by using the predicted influence value, so as to switch the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switch the data source for generating the host entry file from the host source table data to the platform source table data.

[0118] According to the data source switching method, apparatus, device, storage medium and program product provided by the embodiments of the present disclosure, a host entry-lake file and a platform entry-lake file are obtained when the host source table data and the platform source table data are consistent; when the host entry-lake file and the platform entry-lake file are consistent, and the processing result file and the reference processing result file are consistent, the predicted influence value of switching the data source is determined according to the characteristic parameters of the region; the region is divided into k pilot switching region groups according to the predicted influence value; and the data sources of the k pilot switching region groups are switched batch by batch according to the influence level. Since the data and files are compared multiple times during the data source switching process, and when the comparisons are consistent, the predicted influence value of the switching data source is determined based on the characteristic parameters of the area where the data source is to be switched, and the switching is performed batch by batch based on the influence level determined by the predicted influence value, this at least partially overcomes the problem of low data switching accuracy and security caused by direct switching of data sources in related technologies. Moreover, the processing result file of the present application is obtained by processing the host entry file after the host source table data is stored in the data lake, and there is no need to process the host source table data locally. This at least partially overcomes the problem of large local memory occupied by the data source switching system, thereby achieving the technical effect of improving the accuracy and security of the data source switching process and reducing the local memory occupation.

[0119] According to an embodiment of the present disclosure, the determination module may include a first generation submodule, a second generation submodule, a third generation submodule, and a fourth generation submodule.

[0120] The first generating submodule is configured to generate a first influence value according to the total number of user groups in the sub-region, and to screen out a target sub-region from the sub-region according to the first influence value.

[0121] The second generating submodule is configured to generate a second influence value according to the type of user in the target sub-area.

[0122] The third generating submodule is configured to generate a third influence value within the target sub-area according to the application group traffic volume of the sub-area within a unit time period.

[0123] The fourth generating submodule is used to generate a predicted influence value according to the first influence value, the second influence value and the third influence value.

[0124] According to an embodiment of the present disclosure, the first generating submodule may include a generating unit, a sorting unit, and a screening unit.

[0125] A generating unit is configured to generate a priority of the sub-region according to the first influence value.

[0126] The sorting unit is configured to sort the sub-regions according to the priority level to obtain a first sorting result.

[0127] The screening unit is configured to screen out sub-regions that meet a predetermined condition from the first sorting result to obtain a target sub-region.

[0128] According to an embodiment of the present disclosure, the partitioning module may include a sorting submodule, a value obtaining submodule and a partitioning submodule.

[0129] The sorting submodule is used to sort the predicted influence values ​​to obtain a second sorting result.

[0130] The value taking submodule is used to take k-quantile values ​​in the second sorting result and obtain k-1 k-quantile nodes.

[0131] The division submodule is used to divide the area into k pilot switching area groups according to k-1 k-quantile nodes to achieve k batch switching of data sources.

[0132] According to an embodiment of the present disclosure, the switching module may include a first determining submodule, a second determining submodule, and a switching submodule.

[0133] The first determination submodule is configured to determine k influence levels of the k pilot switching area groups according to positions of the k quantile nodes in the second sorting result.

[0134] The second determining submodule is configured to determine a switching order of data sources of the switching pilot switching area group according to the k influence levels.

[0135] The switching submodule is used to switch the data source of the pilot switching area group with the lowest influence level among the k pilot switching area groups in sequence according to the switching order until the data sources of the k pilot switching area groups are all switched.

[0136] According to an embodiment of the present disclosure, the switching submodule may include a determining unit and a switching unit.

[0137] The determination unit is used to determine the switching range according to the influence level.

[0138] The switching unit is used to gradually expand the switching range according to the predetermined frequency, and switch the data source for generating anti-corrosion files from the reference processing result file to the processing result file within the switching range, and switch the data source for generating host entry files from the host source table data to the platform source table data, until the data sources of the pilot switching area group are all switched.

[0139] According to an embodiment of the present disclosure, the switching module may further include a calling submodule and a fifth generating submodule.

[0140] The calling submodule is used to call the platform source table data.

[0141] The fifth generation submodule is used to regenerate the host entry file based on the platform source table data to switch the host source table data to the platform source table data.

[0142] According to an embodiment of the present disclosure, the determination module may include a processing submodule.

[0143] The processing sub-module is used to perform logical processing on the host input files in the second sub-data lake by combining multiple tables to obtain the processed result files.

[0144] According to an embodiment of the present disclosure, the data source switching device may further include a first verification module, a second verification module, and a third verification module.

[0145] The first verification module is used to perform consistency verification between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file, wherein the consistency verification includes at least one of the following: verifying the total transaction amount, verifying the field format of the transaction element, and verifying the name of the transaction element.

[0146] The second verification module is used to select the field value of the primary key field for accuracy verification between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file.

[0147] The third verification module is used to query the frequency of occurrence of the primary key field between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file to perform uniqueness verification.

[0148] According to an embodiment of the present disclosure, the data source switching device may further include a generating module and a pushing module.

[0149] The generation module is used to generate a comparison error file based on the file that fails the verification when the consistency verification, accuracy verification and uniqueness verification fail.

[0150] The push module is used to push the comparison error file to the target object.

[0151] According to an embodiment of the present disclosure, any multiple modules of the storage module 710, the determination module 720, the partitioning module 730, and the switching module 740 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the storage module 710, the determination module 720, the partitioning module 730, and the switching module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the storage module 710, the determination module 720, the partitioning module 730, and the switching module 740 can be at least partially implemented as a computer program module, which can perform the corresponding function when the computer program module is executed.

[0152] It should be noted that the data source switching device part in the embodiment of the present disclosure corresponds to the data source switching method part in the embodiment of the present disclosure. The description of the data source switching device part specifically refers to the data source switching method part and will not be repeated here.

[0153] Figure 8A block diagram of an electronic device suitable for implementing a data source switching method according to an embodiment of the present disclosure is schematically shown.

[0154] like Figure 8 As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0155] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0156] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage portion 808 including a hard disk; and a communication portion 809 including a network interface card such as a LAN card or a modem. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 810 as needed, so that a computer program read therefrom can be installed into the storage portion 808 as needed.

[0157] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0158] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.

[0159] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the data source switching method provided in the embodiments of the present disclosure.

[0160] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0161] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0162] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0163] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0165] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.

[0166] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A data source switching method, characterized in that: The method comprises: If the host source table data and the platform source table data are consistent, the host source table data is stored in the data lake to obtain a host entry file, and the platform source table data is stored in the data lake to obtain a platform entry file, wherein both the host source table data and the platform source table data include a region identifier; When the host entry file and the platform entry file are consistent, and the processed result file is consistent with the reference processed result file, a predicted influence value generated when the data source of the region is switched is determined based on characteristic parameters of the region, wherein the region includes the region represented by the region identifier, the processed result file is obtained by logically processing the host entry file, and the reference processed result file is obtained by calling the host source table data; Dividing the region into k pilot switching region groups according to the predicted influence value to implement k batches of switching data sources, where k is a positive integer; According to the influence level obtained using the predicted influence value, the data sources corresponding to the k pilot switching area groups are switched batch by batch, so as to switch the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and switch the data source for generating the host entry file from the host source table data to the platform source table data.

2. The method according to claim 1, characterized in that The area includes sub-areas, and the characteristic parameters include the total number of user groups in the sub-areas, the types of users in the user groups, and the application group traffic volume in the sub-areas; The determining, based on the characteristic parameters of the region, a predicted influence value generated when the data source of the region is switched, includes: generating a first influence value according to the total number of user groups in the sub-region, and screening a target sub-region from the sub-region according to the first influence value; In the target sub-area, generating a second influence value according to the type of the user; generating a third influence value within the target sub-area according to the application group traffic volume of the sub-area within a unit time period; The predicted influence value is generated according to the first influence value, the second influence value, and the third influence value.

3. The method according to claim 2, characterized in that The step of selecting a target sub-region from the sub-regions according to the first influence value includes: generating a priority of the sub-region according to the first influence value; sorting the sub-regions according to the priority to obtain a first sorting result; The sub-regions that meet a predetermined condition are screened out from the first sorting results to obtain the target sub-region.

4. The method according to claim 1, wherein The step of dividing the region into k pilot switching region groups according to the predicted influence value to implement k batches of switching data sources includes: Sorting the predicted influence values ​​to obtain a second sorting result; In the second sorting result, k-quantile values ​​are taken to obtain k-1 k-quantile nodes; According to the k-1 k-quantile nodes, the area is divided into k pilot switching area groups to implement k batches of switching data sources.

5. The method according to claim 4, characterized in that Switching the data sources corresponding to the k pilot switching area groups in batches according to the influence levels obtained by using the predicted influence values ​​includes: determining k influence levels of the k pilot switching area groups according to positions of the k quantile nodes in the second sorting result; Determining a switching order of switching data sources of the pilot switching area group according to the k influence levels; According to the switching order, the data sources of the pilot switching area group with the lowest influence level among the k pilot switching area groups are switched in sequence until the data sources of the k pilot switching area groups are all switched.

6. The method according to claim 5, characterized in that The step of sequentially switching the data source of the pilot switching area group with the lowest influence level among the k pilot switching area groups includes: For the pilot switching area group with the lowest influence level, perform the following switching operations: Determining a switching range based on the influence level; The switching range is gradually expanded according to a predetermined frequency, and within the switching range, the data source for generating the anti-corrosion file is switched from the reference processing result file to the processing result file, and the data source for generating the host entry file is switched from the host source table data to the platform source table data, until the data sources of the pilot switching area group are all switched.

7. The method according to claim 1, characterized in that Switching the data source for generating the host entry file from the host source table data to the platform source table data includes: Calling the platform source table data; The host entry file is regenerated according to the platform source table data to switch the host source table data to the platform source table data.

8. The method according to claim 1, characterized in that The data lake includes a first sub-data lake and a second sub-data lake, and the host-entered data file is stored in the first sub-data lake; The processing result file is obtained by logically processing the host input file, including: In the second sub-data lake, the host input file is subjected to logical processing of multi-table joint value extraction to obtain the processing result file.

9. The method according to claim 1, characterized in that The comparison of the host source table data with the platform source table data, the comparison of the host entry file with the platform entry file, and the comparison of the processing result file with the reference processing result file include at least one of the following: Performing consistency checks between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file, wherein the consistency check includes at least one of the following: verifying the total transaction amount, verifying the field format of the transaction element, and verifying the name of the transaction element; Between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file, the field value of the primary key field is selected for accuracy verification; The occurrence frequency of the primary key field is queried between the host source table data and the platform source table data, between the host entry file and the platform entry file, and between the processing result file and the reference processing result file to perform uniqueness verification.

10. The method according to claim 9, characterized in that The method further comprises: If the existence check among the consistency check, the accuracy check, and the uniqueness check fails, generating a comparison error file based on the file that fails the check; Push the comparison error file to the target object.

11. A data source switching device, characterized in that: The device comprises: a storage module configured to, when the host source table data and the platform source table data are compared and consistent, store the host source table data in the data lake to obtain a host entry-lake file, and store the platform source table data in the data lake to obtain a platform entry-lake file, wherein both the host source table data and the platform source table data include a region identifier; a determination module configured to determine, based on characteristic parameters of a region, a predicted influence value generated when the data source of the region is switched, if the host entry file and the platform entry file are consistent, and the processed result file and the reference processed result file are consistent, wherein the region includes the region represented by the region identifier, the processed result file is obtained by logically processing the host entry file, and the reference processed result file is obtained by calling the host source table data; a division module, configured to divide the region into k pilot switching region groups according to the predicted influence value, so as to implement k batches of switching data sources, where k is a positive integer; A switching module is used to switch the data sources corresponding to the k pilot switching area groups in batches according to the influence level obtained by using the predicted influence value, so as to switch the data source for generating the anti-corrosion file from the reference processing result file to the processing result file, and to switch the data source for generating the host entry file from the host source table data to the platform source table data.

12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, The method is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Switching between mediator services for a storage system

    CN112470142A

  • System and method for transmitting data from data warehouse to data lake

    CN116361402A