Data matching method and device, storage medium and electronic equipment
By merging and processing subtask configuration information to generate group configuration information, the problem of multiple data matching is solved, and an efficient data matching process is achieved, which is suitable for data matching needs of multiple tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DUXIAOMAN TECH (BEIJING) CO LTD
- Filing Date
- 2023-05-31
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, different models require multiple data matching processes, which makes the data matching process cumbersome and time-consuming.
By acquiring the configuration information of each subtask in the subtask set, performing fusion processing to generate group configuration information, and using the group configuration information to identify the indicator features, the dataset required by each subtask can be quickly obtained, thereby achieving data matching for multiple tasks.
It simplifies the data matching process, improves data matching efficiency, reduces the time spent on repeated matching, and is suitable for different types of models.
Smart Images

Figure CN116860796B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a data matching method, apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, data matching has been applied in various fields. For example, modeling work in deep learning and machine learning is mainly based on the matched sample sets for model training. Various model training methods have been widely applied in various fields, and data mining exploration based on these methods can provide highly analytical insights into enterprise data. However, existing technologies typically require multiple data matching operations for different model-related tasks, and the result of a single data matching operation can only be used for model training for the current task. This makes the data matching process cumbersome and time-consuming. Therefore, how to conveniently perform data matching for multiple tasks, thereby improving the efficiency of data matching, has become a research hotspot. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data matching method, apparatus, storage medium, and electronic device to solve the problem that different models need to perform multiple data matchings and the matching time is long. In other words, embodiments of the present invention can conveniently perform data matching for multiple tasks, thereby improving the efficiency of data matching.
[0004] According to one aspect of the present invention, a data matching method is provided, the method comprising:
[0005] The configuration information of each subtask in the subtask set is obtained, and the configuration information of each subtask is fused to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0006] The first dataset for each feature indicated by a feature identifier in the group configuration information is determined respectively. The first dataset for a feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask.
[0007] Based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset is assigned to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0008] According to another aspect of the present invention, a data matching apparatus is provided, the apparatus comprising:
[0009] The acquisition unit is used to acquire the configuration information of each subtask in the subtask set;
[0010] The processing unit is used to perform fusion processing on the configuration information of each subtask to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0011] The processing unit is further configured to determine the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of the feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask;
[0012] The processing unit is further configured to allocate a second dataset for each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device including a processor and a memory storing a program, wherein the program includes instructions; the instructions, when executed by the processor, cause the processor to perform the following steps:
[0014] The configuration information of each subtask in the subtask set is obtained, and the configuration information of each subtask is fused to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0015] The first dataset for each feature indicated by a feature identifier in the group configuration information is determined respectively. The first dataset for a feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask.
[0016] Based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset is assigned to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0017] According to another aspect of the present invention, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the following steps:
[0018] The configuration information of each subtask in the subtask set is obtained, and the configuration information of each subtask is fused to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0019] The first dataset for each feature indicated by a feature identifier in the group configuration information is determined respectively. The first dataset for a feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask.
[0020] Based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset is assigned to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0021] In this embodiment of the invention, after obtaining the configuration information of each subtask in the subtask set, the configuration information of each subtask can be fused to obtain the group configuration information of the group tasks corresponding to the subtask set. Each configuration information includes a feature identifier for each feature in at least one feature. The feature identifier in the configuration information indicates the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask, so as to quickly obtain the matching data required by each subtask through the group configuration information. Then, a first dataset for the feature indicated by each feature identifier in the group configuration information can be determined. The first dataset for the feature indicated by one feature identifier in the group configuration information includes: the data required to match the corresponding feature in each subtask. Based on this, a second dataset can be allocated to each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset for the feature indicated by each feature identifier in the group configuration information. As can be seen, the embodiments of the present invention can conveniently perform data matching for multiple tasks through group configuration information, thereby improving the efficiency of data matching; based on this, for different types of models, only one data matching is required to achieve the data filtering of all required associated feature vectors, which can effectively save the time of multiple repeated matching operations, and the operation is simple and easy to implement. Attached Figure Description
[0022] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 A flowchart illustrating a data matching method according to an exemplary embodiment of the present invention is shown;
[0024] Figure 2 A schematic diagram of a configuration interface according to an exemplary embodiment of the present invention is shown;
[0025] Figure 3 A flowchart illustrating another data matching method according to an exemplary embodiment of the present invention is shown;
[0026] Figure 4 A schematic diagram of a data fragmentation process according to an exemplary embodiment of the present invention is shown;
[0027] Figure 5 A flowchart illustrating yet another data matching method according to an exemplary embodiment of the present invention is shown;
[0028] Figure 6 A schematic diagram of a data matching platform according to an exemplary embodiment of the present invention is shown;
[0029] Figure 7A schematic block diagram of a data matching apparatus according to an exemplary embodiment of the present invention is shown;
[0030] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0031] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0032] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0033] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0034] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0036] It should be noted that the execution subject of the data matching method provided in this embodiment of the invention can be one or more electronic devices, and this invention does not limit this; wherein, the electronic device can be a terminal (i.e., a client) or a server. Therefore, when the execution subject includes multiple electronic devices, and among the multiple electronic devices includes at least one terminal and at least one server, the data matching method provided in this embodiment of the invention can be jointly executed by the terminal and the server. Accordingly, the terminal mentioned herein may include, but is not limited to: smartphones, tablets, laptops, desktop computers, smartwatches, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server mentioned herein can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0037] Based on the above description, embodiments of the present invention propose a data matching method, which can be executed by the aforementioned electronic device (terminal or server); or, the data matching method can be executed jointly by the terminal and the server. For ease of explanation, the following description will use the execution of the data matching method by an electronic device as an example; such as Figure 1 As shown, the data matching method may include the following steps S101-S103:
[0038] S101, obtain the configuration information of each subtask in the subtask set, and perform fusion processing on the configuration information of each subtask to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0039] The feature identifier can be a numerical identifier (such as a feature number) or a text identifier (such as a feature name), etc.; this invention does not limit this. In the embodiments of this invention, the methods for obtaining the configuration information of each subtask in the above-mentioned subtask set include, but are not limited to, the following:
[0040] The first method of acquisition: The electronic device can obtain the task download link of the subtask set and download the task according to the task download link, thereby obtaining the configuration information of each subtask in the subtask set.
[0041] The second method of acquisition: If an electronic device stores multiple subtasks, then the electronic device can select at least one subtask from the multiple subtasks. The electronic device can then use the selected subtask as a subtask in the subtask set, thereby obtaining the configuration information of each subtask in the subtask set, that is, the configuration information of each subtask in the selected at least one subtask.
[0042] The third method of acquisition: The electronic device may have a configuration interface. Based on this, when a configuration instruction for a data matching task targeting the configuration interface is detected, the electronic device can acquire the configuration information indicated by the configuration instruction, such as... Figure 2 As shown; then, the electronic device can generate the data matching task, use the configuration information indicated by the configuration instruction as the configuration information of the data matching task, and add the data matching task to the subtask set, thereby making the data matching task a subtask in the subtask set. It should be noted that the user can perform configuration operations on the configuration interface for the data matching task. When the electronic device detects the configuration operation performed by the user, it can detect the configuration instruction for the data matching task on the configuration interface. The configuration operation can be a swipe operation performed according to a preset gesture, or a click operation on a specific component of the configuration interface, etc.; this invention does not limit this. It should be understood that... Figure 2 The configuration interface is shown only as an example, and the present invention does not limit the specific content of the configuration interface; for example, the configuration interface may not include a matching time range; or the task name in the configuration interface can only be configured by selection, etc.
[0043] In this embodiment of the invention, since the feature identifier in the configuration information is used to indicate the features required for the corresponding task, for any subtask in the subtask set, the electronic device can determine the configuration information of any subtask based on the features required for that subtask, that is, determine the feature identifier included in the configuration information of that subtask. In this case, the user (i.e., the modeler or task matching personnel, etc.) can establish the features required for the current task (i.e., the features that the current task needs to match), thereby realizing the task configuration. Wherein, when the current task is a modeling task, the user can configure the model according to the different feature vectors to be matched as needed.
[0044] For example, suppose that the features required by model 1 are features 2 and 3, the features required by model 2 are features 1 and 3, and the features required by model 3 are features 2, 3 and 4. Then the user can configure the sub-tasks corresponding to each model as a sub-task. In this case, the electronic device can obtain a group task that includes data matching of model 1, model 2 and model 3, and the group task is composed of sub-task 1 corresponding to model 1, sub-task 2 corresponding to model 2 and sub-task 3 corresponding to model 3. Furthermore, assuming that feature 1 is identified by feature a, feature 2 by feature b, feature 3 by feature c, and feature 4 by feature d, then the configuration information of subtask 1 may include feature b and feature c, the configuration information of subtask 2 may include feature a and feature c, and the configuration information of subtask 3 may include feature b, feature c, and feature d. Based on this, the electronic device can perform fusion processing on the configuration information of each subtask, that is, it can summarize the configuration information of each subtask to obtain the group configuration information of the group task, and the group configuration information may include feature a, feature b, feature c, and feature d. At this time, the number of features required by the group task is less than the sum of the number of features required by each subtask.
[0045] Optionally, the features mentioned in the embodiments of the present invention can be the result of classifying features according to different characteristics. That is, the features mentioned in the embodiments of the present invention can be a specific feature or multiple features with the same characteristics. The present invention does not limit this. For example, assuming that the features "velocity" and "acceleration" have the same characteristics, after classifying according to characteristics, these two features can be used as features in feature 1, so that feature 1 includes the two features "velocity" and "acceleration". Correspondingly, the feature identifier of feature 1 can be used to indicate these two features, and when obtaining sample data of feature 1, sample data of each of these two features can be obtained separately.
[0046] S102, determine the first dataset of the feature indicated by each feature identifier in the group configuration information. The first dataset of the feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask.
[0047] It should be noted that, for any feature identifier in the group configuration information, if the feature identifier does not exist in the configuration information of a certain subtask in the subtask set, then the data that the feature indicated by the feature identifier needs to match in the subtask corresponding to the configuration information that does not include the feature identifier is empty; accordingly, the first dataset of the feature indicated by the feature identifier includes: the data that the corresponding feature needs to match in each target subtask, and a target subtask refers to the subtask that matches the feature identifier (or the feature indicated by the feature identifier), that is, the subtask that includes the feature identifier in the configuration information.
[0048] For example, assuming the group task includes subtask 1, subtask 2, and subtask 3, when the group configuration information includes feature identifier 'a' for feature 1, the first dataset for feature 1 includes: the data that feature 1 needs to match in subtask 1, the data that feature 1 needs to match in subtask 2, and the data that feature 1 needs to match in subtask 3. Further, assuming the configuration information for subtask 2 and subtask 3 includes feature identifier 'a', while the configuration information for subtask 1 does not include feature identifier 'a', then the data that feature 1 needs to match in subtask 1 is empty, that is, there is no need to match data for feature 1 for subtask 1. Based on this, the first dataset for feature 1 may include: the data of feature 1 that subtask 2 needs to match, and the data of feature 1 that subtask 3 needs to match. At this time, the target subtasks can be subtask 2 and subtask 3.
[0049] S103, based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, assign a second dataset to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0050] In this embodiment of the invention, after the group task is completed (i.e., after the first dataset of the features indicated by each feature identifier in the group configuration information is determined respectively), for different sub-tasks, the electronic device can perform the required feature retrieval for each sub-task, thereby obtaining the second dataset of each feature required by each sub-task.
[0051] For example, assuming that the features required for subtask 1 include feature 1 and feature 2, then after the group task is completed, the electronic device can pull a second dataset of the required feature 1 for subtask 1, and a second dataset of the required feature 2 for subtask 1, and so on.
[0052] In this embodiment of the invention, after obtaining the configuration information of each subtask in the subtask set, the configuration information of each subtask can be fused to obtain the group configuration information of the group tasks corresponding to the subtask set. Each configuration information includes a feature identifier for each feature in at least one feature. The feature identifier in the configuration information indicates the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask, so as to quickly obtain the matching data required by each subtask through the group configuration information. Then, a first dataset for the feature indicated by each feature identifier in the group configuration information can be determined. The first dataset for the feature indicated by one feature identifier in the group configuration information includes: the data required to match the corresponding feature in each subtask. Based on this, a second dataset can be allocated to each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset for the feature indicated by each feature identifier in the group configuration information. As can be seen, the embodiments of the present invention can conveniently perform data matching for multiple tasks through group configuration information, thereby improving the efficiency of data matching; based on this, for different types of models, only one data matching is required to achieve the data filtering of all required associated feature vectors, which can effectively save the time of multiple repeated matching operations, and the operation is simple and easy to implement.
[0053] Based on the above description, this embodiment of the invention also proposes a more specific data matching method, wherein the first dataset of a feature is composed of a first data fragment set of the corresponding feature, and a data fragment includes data stored within a corresponding time range (i.e., a data fragment refers to all data within a certain time range); configuration information includes the matching level and matching time range of the feature indicated by the corresponding feature identifier, and the matching level of a feature is used to indicate the splitting method of the sample data of the corresponding feature. Accordingly, this data matching method can be executed by the aforementioned electronic device (terminal or server); or, the data matching method can be executed jointly by the terminal and the server. For ease of explanation, the following description will use the execution of this data matching method by an electronic device as an example; please refer to [link to relevant documentation]. Figure 3 The data matching method may include the following steps S301-S304:
[0054] S301, obtain the configuration information of each subtask in the subtask set, and perform fusion processing on the configuration information of each subtask to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0055] It should be noted that before model training, further feature vector matching of the required sample data is usually necessary. Then, the feature vectors derived from the sample associations are used for data analysis. Finally, a subset of these feature vectors are input into the model to generate the final model file. The existing feature vector matching mechanism involves: first, obtaining a first set of sample data; then, generating a related second set of sample data based on this data; and finally, selecting the required data for model training. In this approach, historical data matching can be challenging due to the time span involved, making it difficult to manage historical data effectively and resulting in lengthy and highly accurate matching. For example, it requires calculating the correlation between each sample data point in the historical data and each sample data point in the first set of sample data to obtain the second set of sample data, leading to a large computational load and making implementation difficult.
[0056] In this embodiment of the invention, the electronic device can manage and match historical data by adding date versions (i.e., time ranges) in a fragmented manner. That is, the electronic device can split any historical data into at least one data fragment, and manage and match any historical data according to the at least one fragmented data, thereby achieving fast and efficient data matching while also managing historical data in a more orderly manner. Here, data fragmentation can also be called data slicing. In one implementation, the electronic device can store data generated before the current moment within the current partition time (i.e., the current time range). That is, the data stored within the current time range can include data stored within the current time range and one or more time ranges preceding the current time range, so that the data within the current time range includes all data before the current moment. In this case, the characteristic of data fragmentation is that the data generated based on the current time range is the most accurate feature vector data at the current moment. In another implementation, the electronic device can store only the data generated within the current time range, and so on.
[0057] S302, for any feature identifier in the group configuration information, obtain the target sample data of the feature indicated by the feature identifier, and perform data splitting processing on the target sample data according to the matching level of the feature indicated by the feature identifier to obtain the data fragmentation result of the feature indicated by the feature identifier.
[0058] It should be noted that before obtaining the target sample data of any feature indicated by the feature identifier, the user can select the initial sample data of the feature to be matched to build the current model. That is, the user can select the initial sample data for data matching of the group task. The initial sample data can be stored in the storage space of the electronic device itself or in other storage devices outside the electronic device. This invention does not limit this.
[0059] Correspondingly, configuration information may include filtering instruction information for the feature indicated by the corresponding feature identifier. When the electronic device acquires target sample data for the feature indicated by any feature identifier, it can acquire initial sample data for that feature. Then, based on this initial sample data, it determines the sample data to be filtered and uses the filtering instruction information for that feature to select the target sample data for the feature indicated by that feature identifier from the sample data to be filtered. The filtering instruction information may be a device identifier that meets the selection criteria (such as a device ID, or a selection time range, etc.); this invention does not limit this. It should be noted that the process of selecting target sample data can also be called a sample processing process.
[0060] Specifically, when determining the sample data to be screened based on the initial sample data, if the storage indication information of each sample data in the initial sample data does not match the screening indication information, the electronic device can perform data conversion on the initial sample data to obtain the sample data to be screened, so that the storage indication information of each sample data in the sample data to be screened matches the screening indication information; if the storage indication information of each sample data in the initial sample data matches the screening indication information, then the initial sample data can be used as the sample data to be screened. Here, the storage indication information can refer to the device identifier, or it can refer to the data generation time, etc., and this invention does not limit it; correspondingly, the storage indication information of any sample data in the target sample data matches the screening indication information.
[0061] For example, suppose the storage indication information for each sample in the initial sample data is the device name, and the filtering indication information is the device ID that corresponds to the filtering condition. In this case, the storage indication information and the filtering indication information for each sample in the initial sample data do not match. The electronic device can then perform data conversion on the initial sample data to obtain the sample data to be filtered. In this case, the storage indication information for each sample in the sample data to be filtered is the corresponding device ID. Alternatively, suppose that both the storage indication information and the filtering indication information for each sample in the initial sample data are device IDs. In this case, the storage indication information and the filtering indication information for each sample in the initial sample data match.
[0062] In this embodiment of the invention, because the updates and fluctuations of sample data for some features are not particularly significant, but the updates of sample data for other features are relatively fast and fluctuate considerably, this embodiment of the invention can provide precise matching at the daily level, relative fuzzy matching at the weekly and monthly levels, etc.; that is, the above matching levels can be daily, weekly, monthly, etc., and this invention does not limit them. Based on this, the electronic device can perform time-slicing processing on the target sample data, and the matching level can also be called the splitting time; correspondingly, the data slicing result can include at least one data slice, and the duration formed by the time range corresponding to a data slice is the same as the matching level, and the target sample data includes the sample data in each data slice of at least one data slice. It should be understood that the number of samples corresponding to any data slice in the data slicing result is greater than the preset number of samples, which can be set according to experience or according to actual needs, and this invention does not limit it.
[0063] For example, such as Figure 4 As shown, taking the node with a time range marked below it as an example to illustrate the data sharding results, let's assume that the feature indicated by any of the above feature identifiers is feature 1, and the matching level of the feature indicated by any of the feature identifiers is monthly. Then, the electronic device can split the target sample data by month to obtain the data sharding results of feature 1. At this time, the target sample data is the sample data of feature 1 during the historical period from January 2019 to December 2020. Assuming that the target sample data is mainly distributed in 6 months, that is, the sample data of feature 1 stored in these 6 months meets the sharding conditions. If the number of corresponding samples is greater than the preset number of samples, then the data sharding results of feature 1 can include 6 data shards, namely January 2019, June 2019, November 2019, January 2020, May 2020, and September 2020.
[0064] It should be noted that the time range corresponding to the data shards for each feature identifier in the group configuration information is not necessarily completely consistent. The selection of data shards is determined based on the specific characteristics of each feature. For example, such as... Figure 4 As shown, the data shards of features 2 and 3 are not completely consistent with those of feature 1; specifically, feature 2 has a data shard for February 2019, which corresponds to a different time range than the data shards of feature 1. This is because the two features have richer data at different times.
[0065] S303, based on the matching time range of the feature indicated by any feature identifier, match the first data fragment set of the feature indicated by any feature identifier from the data fragment splitting result of the feature indicated by any feature identifier, so as to obtain the first dataset of the feature indicated by any feature identifier, and the time range corresponding to any data fragment in the first data fragment set of a feature matches the matching time range of the corresponding feature.
[0066] In this system, the number of matching time ranges for a single feature is at least one. For any matching time range of the feature indicated by any of the aforementioned feature identifiers, based on that matching time range and the time ranges corresponding to each data fragment in the data fragmentation results of the feature indicated by that feature identifier, a target time range matching that matching time range is determined. The target time range is not located after any matching time range, and the distance between the target time range and any matching time range is less than the distance between the time ranges corresponding to other data fragments in the corresponding data fragmentation results and any matching time range. That is, the target time range is located before or equal to any matching time range. Here, "other data fragments" refers to any data fragment in the corresponding data fragmentation results other than the data fragments within the target time range. Based on this, the electronic device can match the data fragments corresponding to the target time range from the data fragmentation results of the feature indicated by any feature identifier and add the matched data fragments to the first data fragment set of the feature indicated by any feature identifier. It should be understood that the electronic device can search for the most recently stored data fragment for matching within any matching time range. It should be noted that when the number of matching time ranges for a feature is at least one, the time range corresponding to any data segment in the first data segment set of a feature being matched with the matching time range of the corresponding feature means that the time range corresponding to any data segment in the first data segment set of a feature is matched with a certain time range in at least one matching time range of the corresponding feature. In other words, the distance between the time range and a certain time range in at least one matching time range of the corresponding feature is less than the distance between the time ranges corresponding to other data segments in the data segmentation results of the corresponding feature and the aforementioned time range.
[0067] For example, such as Figure 4As shown, assuming any of the above matching time ranges is February 2019, when the feature indicated by any of the above feature identifiers is feature 1, feature 2, or feature 3, the electronic device can match the data fragment of January 2019 in the data fragmentation result of feature 1, match the data fragment of February 2019 in the data fragmentation result of feature 2, and match the data fragment of February 2019 in the data fragmentation result of feature 3. It should be understood that when any of the above matching time ranges is December 2019, the electronic device can match the data fragment of November 2019 in the data fragmentation results of Feature 1, the data fragment of September 2019 in the data fragmentation results of Feature 2, and the data fragment of September 2019 in the data fragmentation results of Feature 3. This is because although the distance between the time range corresponding to the data fragment of January 2020 in the data fragmentation results of Feature 2 and any matching time range is shorter, it is impossible to obtain relevant data for 2020 for the time range of December 2019. Such matching is meaningless and inaccurate. Therefore, the matching principle of this embodiment of the invention is to find the nearest slice in any of the above matching time ranges or forward.
[0068] S304. Based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset is assigned to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0069] In this embodiment of the invention, the second dataset of a feature is composed of the second data fragment set of the corresponding feature. Based on this, for any subtask in the subtask set, the electronic device can traverse each feature identifier in the configuration information of any subtask and take the feature indicated by the currently traversed feature identifier as the current feature. Then, the target matching time range of the current feature can be determined from the configuration information of any subtask, and the matching time range of the current feature in the group configuration information includes the target matching time range. Correspondingly, the electronic device can select the data fragment corresponding to the target matching time range from the first data fragment set of the current feature and add the selected data fragment to the second data fragment set of the current feature under any subtask to obtain the second data fragment set of the current feature. After traversing each feature identifier in the configuration information of any subtask, the second data fragment set of each feature required by any subtask is obtained to obtain the second dataset of each feature required by any subtask, such as... Figure 5 As shown. It should be noted that, Figure 5This is merely an illustrative example illustrating the process of data matching for each subtask, and the present invention does not limit this; for example, model 1 may also include feature 3, or model 2 may only include feature 1, etc.; wherein, one model corresponds to one subtask, that is, the features required for subtask 1 corresponding to model 1 include feature 1 and feature 2, the features required for subtask 2 corresponding to model 2 include feature 1 and feature 3, the features required for subtask 3 corresponding to model 3 include feature 2, feature 3 and feature 4, and the process of matching the first data fragment set of features indicated by each feature identifier in the group configuration information for the group task can be called feature slicing method matching, thereby providing the first data fragment set of features indicated by each feature identifier in the group configuration information for subtask allocation.
[0070] It should be noted that the number of target matching time ranges for the current feature under any of the above sub-tasks can be at least one. Therefore, when selecting a data segment corresponding to the target matching time range from the first data segment set of the current feature and adding the selected data segment to the second data segment set of the current feature under any sub-task, the electronic device can traverse each target matching time range in at least one target matching time range of the current feature and use the currently traversed target matching time range as the current target matching time range. Furthermore, the electronic device can select a data segment corresponding to the current target matching time range from the first data segment set of the current feature and add the selected data segment to the second data segment set of the current feature under any sub-task. After traversing each target matching time range in at least one target matching time range, the second data segment set of the current feature under that sub-task is obtained.
[0071] It should be noted that the configuration information of a subtask may include the data processing rules for the corresponding subtask, and a data processing rule is used to indicate the method of generating a matching feature vector result based on data sharding of each feature in at least one feature. Based on this, for the i-th subtask in the subtask set, the electronic device may determine the set of data shards to be processed for each feature required by the i-th subtask based on the second set of data shards for each feature required by the i-th subtask, where i is a positive integer and i is less than or equal to the number of subtasks in the subtask set. Then, the electronic device may process the set of data shards to be processed for each feature required by the i-th subtask based on the data processing rules of the i-th subtask to obtain the matching feature vector result of the i-th subtask, so that the matching feature vector result of the i-th subtask can be used as the training data for the model corresponding to the i-th subtask.
[0072] In the specific implementation, when determining the data fragment set to be processed for each feature required by the i-th sub-task based on the second data fragment set of each feature required by the i-th sub-task, for the j-th feature required by the i-th sub-task, if there is a data fragment to be processed in the second data fragment set of the j-th feature, the electronic device can process the data fragment to be processed to obtain the data fragment set to be processed for the j-th feature, such that there is no data fragment to be processed in the data fragment set to be processed for the j-th feature, where j is a positive integer and j is less than or equal to the number of features required by the i-th sub-task; correspondingly, if there is no data fragment to be processed in the second data fragment set of the j-th feature, then the second data fragment set of the j-th feature can be used as the data fragment set to be processed for the j-th feature. Here, a data fragment to be processed refers to a data fragment that needs to be processed according to the feature information, that is, a data fragment that does not meet the feature information.
[0073] It should be noted that the aforementioned feature information includes, but is not limited to, data processing logic and data processing methods, etc.; this invention does not limit these. Specifically, data processing logic can refer to a specified data representation method. When the feature information includes data processing logic, the data fragment to be processed refers to a data fragment whose data representation method differs from the specified data representation method. For example, the data fragment to be processed may use a binary representation method, while the specified data representation method is a decimal representation method, etc. In this case, the data representation method of each data fragment in the set of data fragments to be processed is the specified data representation method. Correspondingly, data processing method can refer to a specified data storage format. When the feature information includes data processing method, the data fragment to be processed refers to a data fragment whose data storage format differs from the specified data storage format, etc. In this case, the data storage format of each data fragment in the set of data fragments to be processed is the specified data storage format, and this invention does not limit either the specified data representation method or the specified data storage format. Furthermore, when the feature information includes both data processing logic and data processing method, the data representation method of each data fragment in the set of data fragments to be processed is the specified data representation method, and the data storage format of each data fragment is the specified data storage format, etc.
[0074] In this embodiment of the invention, the above-mentioned data processing rules include, but are not limited to, addition, subtraction, and logarithmic operations, etc., and the invention does not limit these operations. For example, assuming that the data processing rule for the i-th subtask is addition, and the required features for the i-th subtask include feature 1 and feature 2, then the electronic device can perform addition on the data fragment set to be processed for feature 1 and the data fragment set to be processed for feature 2 to obtain the matching feature vector result of the i-th subtask, that is, the corresponding addition result, such that any matching feature vector in the matching feature vector result of the i-th subtask is: the addition result between a sample data of feature 1 in the data fragment set to be processed under the i-th subtask and a sample data of feature 2 in the data fragment set to be processed under the i-th subtask.
[0075] Furthermore, the electronic device can also record data of the data fragment set of each feature required by each sub-task, or record data of the second data fragment set of each feature required by each sub-task, etc., to obtain data recording results, and perform order-of-magnitude statistics or matching data statistics on the data recording results to obtain statistical results, thereby selecting features or sorting features based on the statistical results.
[0076] In this embodiment of the invention, after obtaining the configuration information of each subtask in the subtask set, the configuration information of each subtask is fused to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier for each feature in at least one feature, so as to facilitate rapid matching of data shards through the group configuration information. In this case, for any feature identifier in the group configuration information, the target sample data of the feature indicated by the feature identifier can be obtained, and the target sample data can be split according to the matching level of the feature indicated by the feature identifier to obtain the data sharding result of the feature indicated by the feature identifier; then, according to the matching time range of the feature indicated by the feature identifier, the first data shard set of the feature indicated by the feature identifier is matched from the data sharding result of the feature indicated by the feature identifier to obtain the first dataset of the feature indicated by the feature identifier, thereby pulling the first data shard set of the feature indicated by each feature identifier in the group configuration information at one time, so as to avoid multiple feature pulling for different subtasks and avoid repeated feature pulling, thereby saving the time of repeated feature pulling. Furthermore, based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset can be allocated to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask. It is evident that this embodiment of the invention can conveniently perform data matching on multiple subtasks through group tasks, thereby improving the efficiency of data matching. Moreover, this embodiment of the invention can manage and match historical data (such as target sample data) by splitting it into data shards and adding date versions, thereby enabling fast and efficient data matching while also allowing for more orderly management of historical data.
[0077] It should be noted that existing technologies require hard-coding of the current logic to support the unique logic of each model when processing data for different models. However, if the data matching conditions and processing logic change, developers need to develop the corresponding logic and then redeploy, which is time-consuming and makes it difficult to quickly respond to corresponding logic adjustments. Therefore, if... Figure 6As shown, the electronic device mentioned in the embodiments of the present invention may include a data matching platform (a low-code development platform device), and the data matching method proposed in the embodiments of the present invention can be applied to the data matching platform in the electronic device, thereby conveniently realizing data matching for multiple sub-tasks through the data matching platform. The data matching platform is a high-efficiency, high-performance visual application development platform. The data matching platform can modularize various functions and decouple them from each other. At the same time, it abstracts the cumbersome underlying architecture and infrastructure into a graphical interface, providing users with a convenient operation page, thereby reducing hard-coding development and shortening the deployment process and time through configuration.
[0078] Accordingly, the data matching platform includes, but is not limited to: a task management module, a task scheduling module (i.e., a scheduling engine), a task feature configuration module, and a task execution module (including a task executor), etc.; this invention does not limit these components. The task management module is primarily responsible for scheduling and executing all tasks, such as adjusting priorities, starting tasks, stopping tasks, and notifying and alerting about task execution status. The task feature configuration module can configure relevant task feature information (such as time range) and required features. During task execution, all used configuration information is read from the task feature configuration module; that is, certain parameters and components must be matched from the task feature configuration module before task execution begins. The task execution module can monitor the current task execution status, system resource usage, and task health checks. The scheduling engine provides functions for setting, managing, and executing scheduled tasks, can create and maintain multiple scheduled tasks, listen for scheduled task execution events, and execute tasks that meet the execution conditions.
[0079] In this scenario, after a task is submitted to the data matching platform in the electronic device via an API (Application Programming Interface), the data matching platform generates at least one set of tasks and corresponding sub-task sets for each set of tasks. These tasks are added to the job pool of the data matching platform and await invocation. Here, the API consists of predefined functions designed to provide applications and developers (i.e., users) with the ability to access a set of routines based on certain software or hardware without needing to access the source code or understand the details of the internal working mechanism.
[0080] In this scenario, the aforementioned group tasks can obtain task execution rights through the scheduling engine, thereby parsing the group configuration information to obtain each feature identifier in the group configuration information. Correspondingly, the task execution module can compile the group tasks (e.g., combining task blocks or decrypting data based on resource distribution). A task can include at least one task block (e.g., data matching for a feature). Then, the group tasks can use the task executor to perform data transformation (i.e., determining the sample data to be filtered) and sample processing (i.e., selecting target sample data from the sample data to be filtered). Then, the features to be matched by different sub-tasks are fused (i.e., integrated and summarized) and configuration read to obtain the group configuration information of the group tasks. Furthermore, the task executor can perform data interaction with different data sources (i.e., initial sample data) for the group tasks, thereby completing the feature retrieval task of the group tasks to obtain the first data fragment set of the feature indicated by each feature identifier in the group configuration information. When the group tasks are completed, the first data fragment set of all matched features can be placed into the feature pool, so that the feature pool includes the first data fragment sets retrieved by all group tasks.
[0081] Correspondingly, when the scheduling engine allocates resources to a subtask, the subtask gains execution rights. In this case, the task executor can directly pull the features required by the subtask from the feature pool (i.e., pull the second data shard set of each feature required by the corresponding subtask). Then, it can process and parse the pulled second data shard set to obtain the data shard set of each feature required by the corresponding subtask. The configured data processing rules (i.e., the processing rules implemented by the rule engine) can then be used to process the data shard set of each feature required by the corresponding subtask (i.e., result processing) to obtain the matching feature vector result for the corresponding task. This rule engine, evolved from the inference engine, is a component embedded in the application. It separates business decisions from application code, uses predefined semantic modules to write business decisions, accepts data input, interprets business rules, and makes business decisions based on those rules. It should be noted that users can operate the data matching platform through a graphical interface (i.e., the configuration interface) and submit data matching tasks, allowing even non-technical personnel to operate the platform through this interface.
[0082] It should be understood that, based on the data matching platform proposed in this embodiment of the invention, data matching can be completed with only simple configuration for different sub-tasks; at the same time, for sub-tasks corresponding to some complex models (i.e., models with special logic), the rule engine can also process them to obtain the matching feature vector results for the corresponding tasks. Therefore, the data matching platform proposed in this embodiment of the invention not only provides users with a convenient operation interface, reducing the online process and time through configuration, but also supports data slicing matching and simultaneous data matching for multiple sub-tasks, and can quickly respond to corresponding needs after configuration information is adjusted, providing convenience for users.
[0083] Based on the description of the relevant embodiments of the data matching method above, this invention also proposes a data matching device, which can be a computer program (including program code) running in an electronic device; such as Figure 7 As shown, the data matching device may include an acquisition unit 701 and a processing unit 702. The data matching device can perform... Figure 1 or Figure 3 The data matching method shown, i.e., the data matching device can operate the above-mentioned unit:
[0084] Acquisition unit 701 is used to acquire the configuration information of each subtask in the subtask set;
[0085] The processing unit 702 is used to perform fusion processing on the configuration information of each subtask to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask.
[0086] The processing unit 702 is further configured to determine the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of the feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask;
[0087] The processing unit 702 is further configured to allocate a second dataset for each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
[0088] In one implementation, the first dataset of a feature is composed of a first data shard set of the corresponding feature, and a data shard includes data stored within a corresponding time range; a configuration information includes the matching level and matching time range of the feature indicated by the corresponding feature identifier, and the matching level of a feature is used to indicate the splitting method of the sample data of the corresponding feature. When the processing unit 702 determines the first dataset of the feature indicated by each feature identifier in the group configuration information, it can specifically be used to:
[0089] For any feature identifier in the group configuration information, obtain the target sample data of the feature indicated by the feature identifier, and perform data splitting processing on the target sample data according to the matching level of the feature indicated by the feature identifier to obtain the data fragmentation result of the feature indicated by the feature identifier.
[0090] Based on the matching time range of the feature indicated by any feature identifier, the first data fragment set of the feature indicated by any feature identifier is matched from the data fragment splitting result of the feature indicated by any feature identifier to obtain the first dataset of the feature indicated by any feature identifier, and the time range corresponding to any data fragment in the first data fragment set of a feature matches the matching time range of the corresponding feature.
[0091] In another implementation, the number of matching time ranges for a feature is at least one; when the processing unit 702 matches the first data fragment set of the feature indicated by the feature identifier from the data fragment splitting result of the feature indicated by the feature identifier according to the matching time range of the feature indicated by the feature identifier, it can be specifically used for:
[0092] For any matching time range of the feature indicated by any feature identifier, based on the any matching time range and the time range corresponding to each data segment in the data segmentation result of the feature indicated by any feature identifier, a target time range that matches the any matching time range is determined. The target time range is not located after the any matching time range, and the distance between the target time range and the any matching time range is less than the distance between the time range corresponding to other data segments in the corresponding data segmentation result and the any matching time range.
[0093] From the data fragmentation results of the feature indicated by any feature identifier, the data fragment corresponding to the target time range is matched, and the matched data fragment is added to the first data fragment set of the feature indicated by any feature identifier.
[0094] In another implementation, configuration information includes filtering indication information for the feature indicated by the corresponding feature identifier; when the processing unit 702 acquires target sample data for the feature indicated by any of the feature identifiers, it may specifically be used to:
[0095] Obtain initial sample data for the feature indicated by any of the feature identifiers;
[0096] Based on the initial sample data, the sample data to be screened is determined, and the target sample data indicating the feature indicated by the feature identifier is selected from the sample data to be screened using the screening indication information of the feature identifier indicated by the feature identifier.
[0097] In another implementation, the second dataset for a feature is composed of a second data fragment set of the corresponding feature. When the processing unit 702 allocates the second dataset for each feature required by each sub-task according to the feature identifier in the configuration information of each sub-task and the first dataset of the feature indicated by each feature identifier in the group configuration information, it can specifically be used to:
[0098] For any subtask in the subtask set, traverse each feature identifier in the configuration information of any subtask, and take the feature indicated by the currently traversed feature identifier as the current feature.
[0099] From the configuration information of any of the subtasks, determine the target matching time range of the current feature, wherein the matching time range of the current feature in the group configuration information includes the target matching time range;
[0100] From the first data fragment set of the current feature, select the data fragment corresponding to the target matching time range, and add the selected data fragment to the second data fragment set of the current feature under any subtask to obtain the second data fragment set of the current feature;
[0101] After traversing through all feature identifiers in the configuration information of any subtask, a second data fragment set of each feature required by any subtask is obtained, so as to obtain a second dataset of each feature required by any subtask.
[0102] In another implementation, the second dataset for a feature is composed of a second data fragment set for the corresponding feature. The configuration information for a subtask includes the data processing rules for the corresponding subtask, and a data processing rule is used to indicate the manner in which matching feature vector results are generated based on the data fragments of each feature in at least one feature. The processing unit 702 can also be used for:
[0103] For the i-th subtask in the subtask set, based on the second data shard set of each feature required by the i-th subtask, determine the data shard set to be processed for each feature required by the i-th subtask, where i is a positive integer and i is less than or equal to the number of subtasks in the subtask set.
[0104] Based on the data processing rules of the i-th sub-task, the data fragments of each feature required by the i-th sub-task are processed to obtain the matching feature vector result of the i-th sub-task, so that the matching feature vector result of the i-th sub-task can be used as the training data of the model corresponding to the i-th sub-task.
[0105] In another implementation, the configuration information of each subtask is configured through a configuration interface, and the acquisition unit 701 can also be used for:
[0106] When a configuration instruction for a data matching task targeting the configuration interface is detected, the configuration information indicated by the configuration instruction is obtained;
[0107] Processing unit 702 can also be used for:
[0108] Generate the data matching task, use the configuration information indicated by the configuration instruction as the configuration information of the data matching task, and add the data matching task to the subtask set.
[0109] According to one embodiment of the present invention, Figure 1 or Figure 3 Each step involved in the method shown can be derived from... Figure 7 This is performed by the individual units in the data matching device shown. For example, Figure 1 Step S101 shown can be performed by Figure 7 The acquisition unit 701 and processing unit 702 shown are executed together, and steps S102 and S103 can both be performed by... Figure 7 The processing unit 702 shown executes this. For example, Figure 3 Step S301 shown can be performed by Figure 7 The acquisition unit 701 and processing unit 702 shown are executed together, and steps S302-S304 can all be performed by [the relevant unit / organization]. Figure 7 The processing unit 702 shown executes, etc.
[0110] According to another embodiment of the present invention, Figure 7Each unit in the data matching device shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effect of the embodiments of the present invention. The above units are based on logical function division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of the present invention, any data matching device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0111] According to another embodiment of the present invention, it is possible to perform operations such as those described above by running on a general-purpose electronic device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 1 or Figure 3 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 7 The data matching apparatus shown herein, and the data matching method for implementing embodiments of the present invention, are described. The computer program may be recorded on, for example, a computer storage medium, loaded onto the aforementioned electronic device via the computer storage medium, and run therein.
[0112] In this embodiment of the invention, after obtaining the configuration information of each subtask in the subtask set, the configuration information of each subtask can be fused to obtain the group configuration information of the group tasks corresponding to the subtask set. Each configuration information includes a feature identifier for each feature in at least one feature. The feature identifier in the configuration information indicates the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask, so as to quickly obtain the matching data required by each subtask through the group configuration information. Then, a first dataset for the feature indicated by each feature identifier in the group configuration information can be determined. The first dataset for the feature indicated by one feature identifier in the group configuration information includes: the data required to match the corresponding feature in each subtask. Based on this, a second dataset can be allocated to each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset for the feature indicated by each feature identifier in the group configuration information. As can be seen, the embodiments of the present invention can conveniently perform data matching for multiple tasks through group configuration information, thereby improving the efficiency of data matching; based on this, for different types of models, only one data matching is required to achieve the data filtering of all required associated feature vectors, which can effectively save the time of multiple repeated matching operations, and the operation is simple and easy to implement.
[0113] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.
[0114] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0115] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0116] refer to Figure 8 The present invention will now be described in the form of a structural block diagram of an electronic device 800 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0117] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0118] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 808 may include, but is not limited to, disks and optical discs. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0119] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the data matching method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the data matching method by any other suitable means (e.g., by means of firmware).
[0120] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0125] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0126] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A data matching method, characterized in that, include: The configuration information of each subtask in the subtask set is obtained, and the configuration information of each subtask is fused to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask. The first dataset for each feature indicated by a feature identifier in the group configuration information is determined respectively. The first dataset for a feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask. Based on the feature identifiers in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, a second dataset is assigned to each feature required by each subtask. The first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
2. The method according to claim 1, characterized in that, A first dataset for a feature is composed of a first data shard set for the corresponding feature, and a data shard includes data stored within a corresponding time range; a configuration information includes the matching level and matching time range of the feature indicated by the corresponding feature identifier, and the matching level of a feature is used to indicate the splitting method of the sample data of the corresponding feature. The step of determining the first dataset for each feature indicated by each feature identifier in the configuration information includes: For any feature identifier in the group configuration information, obtain the target sample data of the feature indicated by the feature identifier, and perform data splitting processing on the target sample data according to the matching level of the feature indicated by the feature identifier to obtain the data fragmentation result of the feature indicated by the feature identifier. Based on the matching time range of the feature indicated by any feature identifier, the first data fragment set of the feature indicated by any feature identifier is matched from the data fragment splitting result of the feature indicated by any feature identifier to obtain the first dataset of the feature indicated by any feature identifier, and the time range corresponding to any data fragment in the first data fragment set of a feature matches the matching time range of the corresponding feature.
3. The method according to claim 2, characterized in that, The number of matching time ranges for a feature is at least one; the step of matching the first data fragment set of the feature indicated by the feature indicated by the feature identifier from the data fragment splitting results of the feature indicated by the feature identifier based on the matching time range of the feature indicated by the feature identifier includes: For any matching time range of the feature indicated by any feature identifier, based on the any matching time range and the time range corresponding to each data segment in the data segmentation result of the feature indicated by any feature identifier, a target time range that matches the any matching time range is determined. The target time range is not located after the any matching time range, and the distance between the target time range and the any matching time range is less than the distance between the time range corresponding to other data segments in the corresponding data segmentation result and the any matching time range. From the data fragmentation results of the feature indicated by any feature identifier, the data fragment corresponding to the target time range is matched, and the matched data fragment is added to the first data fragment set of the feature indicated by any feature identifier.
4. The method according to claim 2, characterized in that, A configuration information includes filtering indication information for the feature indicated by the corresponding feature identifier; obtaining the target sample data for the feature indicated by any of the feature identifiers includes: Obtain initial sample data for the feature indicated by any of the feature identifiers; Based on the initial sample data, the sample data to be screened is determined, and the target sample data indicating the feature indicated by the feature identifier is selected from the sample data to be screened using the screening indication information of the feature identifier indicated by the feature identifier.
5. The method according to claim 2, characterized in that, The second dataset for a feature is composed of the second data fragment set of the corresponding feature. The process of allocating the second dataset for each feature required by each sub-task, based on the feature identifiers in the configuration information of each sub-task and the first dataset indicating the feature for each feature in the group configuration information, includes: For any subtask in the subtask set, traverse each feature identifier in the configuration information of any subtask, and take the feature indicated by the currently traversed feature identifier as the current feature. From the configuration information of any of the subtasks, determine the target matching time range of the current feature, wherein the matching time range of the current feature in the group configuration information includes the target matching time range; From the first data fragment set of the current feature, select the data fragment corresponding to the target matching time range, and add the selected data fragment to the second data fragment set of the current feature under any subtask to obtain the second data fragment set of the current feature; After traversing through all feature identifiers in the configuration information of any subtask, a second data fragment set of each feature required by any subtask is obtained, so as to obtain a second dataset of each feature required by any subtask.
6. The method according to any one of claims 1-5, characterized in that, The second dataset for a feature is composed of a second data shard set for the corresponding feature. The configuration information for a subtask includes the data processing rules for the corresponding subtask, and a data processing rule indicates the method for generating matching feature vector results based on data shards of each feature in at least one feature. The method further includes: For the i-th subtask in the subtask set, based on the second data shard set of each feature required by the i-th subtask, determine the data shard set to be processed for each feature required by the i-th subtask, where i is a positive integer and i is less than or equal to the number of subtasks in the subtask set. Based on the data processing rules of the i-th sub-task, the data fragments of each feature required by the i-th sub-task are processed to obtain the matching feature vector result of the i-th sub-task, so that the matching feature vector result of the i-th sub-task can be used as the training data of the model corresponding to the i-th sub-task.
7. The method according to any one of claims 1-5, characterized in that, The configuration information for each subtask is configured through a configuration interface. The method also includes: When a configuration instruction for a data matching task targeting the configuration interface is detected, the configuration information indicated by the configuration instruction is obtained; Generate the data matching task, use the configuration information indicated by the configuration instruction as the configuration information of the data matching task, and add the data matching task to the subtask set.
8. A data matching device, characterized in that, The device includes: The acquisition unit is used to acquire the configuration information of each subtask in the subtask set; The processing unit is used to perform fusion processing on the configuration information of each subtask to obtain the group configuration information of the group task corresponding to the subtask set. Each configuration information includes a feature identifier of each feature in at least one feature. The feature identifier in the configuration information is used to indicate the features required by the corresponding task, and the number of features required by the group task is less than or equal to the sum of the number of features required by each subtask. The processing unit is further configured to determine the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of the feature indicated by a feature identifier in the group configuration information includes: the data that the corresponding feature needs to match in each subtask; The processing unit is further configured to allocate a second dataset for each feature required by each subtask according to the feature identifier in the configuration information of each subtask and the first dataset of the feature indicated by each feature identifier in the group configuration information, wherein the first dataset of a feature includes the second dataset of the corresponding feature under each subtask.
9. An electronic device, characterized in that, include: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and equipment for detecting continuous time signal data
CN106681991A
Data processing method and device, electronic equipment and computer readable storage medium
CN109597826A