A data integration method, electronic device, and storage medium

By acquiring and matching user feature data and using a random forest model to determine identity identifiers, the problem of integrating user data across devices and time periods was solved, achieving highly accurate data integration.

CN120610981BActive Publication Date: 2026-07-24ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
Filing Date
2025-04-02
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve continuous user tracking and behavior analysis across devices and time periods, and the complexity of user identification makes data integration difficult.

Method used

By obtaining the initial user ID and feature data list from the target database, the feature data of the specified user is obtained, and matching and insertion are performed when the conditions are met, or key feature data is collected and integrated within the target time period. The random forest model is used to determine the identity identifier, thereby improving the accuracy of data integration.

Benefits of technology

It enables continuous user tracking and behavior analysis across devices and time periods, improving the accuracy of data integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610981B_ABST
    Figure CN120610981B_ABST
Patent Text Reader

Abstract

The application provides a data integration method, an electronic device and a storage medium, and comprises the following steps: obtaining a target database, obtaining specified feature data corresponding to a specified user, obtaining a unique identifier corresponding to the specified user when the specified feature data meets a preset specified condition, and inserting the specified feature data into a corresponding initial feature data list to realize data integration, obtaining a target time period when the specified feature data does not meet the preset specified condition, obtaining key feature data corresponding to the specified user, matching the key feature data with data in the target database, and inserting the key feature data into the target database to realize data integration. The application adopts different ways to integrate data based on whether the user behavior characteristics can be matched with the data in the database, reasonably sets the time period for obtaining the behavior characteristics of the user, and improves the accuracy of data integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data integration method, electronic device, and storage medium. Background Technology

[0002] With the rapid development and popularization of mobile internet technology, it has become commonplace for users to use various service platforms on different devices in their daily lives. However, the diversity of data ports and devices corresponding to various service platforms has led to the complexity of user identification. Users may use the same service on different devices or use the same service platform on the same device under different network environments. This causes the user's behavioral characteristics to change with changes in network environment and other factors, making it difficult to achieve continuous tracking and behavioral analysis of the same user across devices and time. Therefore, it is particularly important to develop a method that can effectively integrate the behavioral data generated by users at different times, and achieve continuous tracking and behavioral analysis of users across devices and time. Summary of the Invention

[0003] To address the aforementioned technical problems, the present invention adopts the following technical solution: a data integration method, comprising the following steps:

[0004] S10, Obtain the target database, wherein the target database stores several initial user IDs and a list of initial feature data corresponding to each initial user ID. The initial user ID is a unique identifier representing the identity of the initial user. The list of initial feature data includes several initial feature data, which are behavioral data generated by the initial user connecting to several target service platforms through data ports within a historical time period.

[0005] S20, obtain specified feature data corresponding to a specified user, wherein the specified feature data is behavioral data generated when the specified user connects with a target service platform at the current time.

[0006] S30, when the specified feature data meets the preset specified conditions, obtain the unique identifier corresponding to the specified user and insert the specified feature data into the corresponding initial feature data list to achieve data integration, wherein the unique identifier corresponding to the specified user is the initial user ID corresponding to the initial feature data that matches the specified feature data obtained from the target database.

[0007] S40, when the specified feature data does not meet the preset specified conditions, obtain the target time period, wherein the target time period is the time period that can confirm the identity of the specified user and requires the collection of specified user behavior feature data based on the initial feature data list corresponding to several initial user IDs.

[0008] S50, acquire key feature data corresponding to a specified user, wherein the key feature data is behavioral feature data generated by the specified user connecting with several target service platforms within a target time period.

[0009] S60: After matching the key feature data with the data in the target database, the key feature data is inserted into the target database to achieve data integration.

[0010] The present invention protects a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method of data integration.

[0011] The present invention protects a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data integration method.

[0012] This invention has at least the following beneficial effects: A data integration method, the method comprising the following steps: acquiring a target database, acquiring specified feature data corresponding to a specified user, when the specified feature data meets preset specified conditions, acquiring a unique identifier corresponding to the specified user and inserting the specified feature data into the corresponding initial feature data list to achieve data integration, when the specified feature data does not meet the preset specified conditions, acquiring a target time period, acquiring key feature data corresponding to the specified user, matching the key feature data with data in the target database, and then inserting the key feature data into the target database to achieve data integration. This invention, based on whether user behavior features can be matched with data in the database, adopts different methods for data integration, reasonably sets the time period for acquiring user-corresponding behavior features, and improves the accuracy of data integration. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating a data integration method provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example

[0017] This embodiment provides a data integration method, which includes the following steps: Figure 1 As shown:

[0018] S10, Obtain the target database, wherein the target database stores several initial user IDs and a list of initial feature data corresponding to each initial user ID. The initial user ID is a unique identifier representing the identity of the initial user. The list of initial feature data includes several initial feature data, which are behavioral data generated by the initial user connecting to several target service platforms through data ports within a historical time period.

[0019] Specifically, the historical time period refers to the time period preceding the current moment.

[0020] Specifically, the target service platform is an internet application developed based on internet technology that can perform several functions.

[0021] Specifically, the target database is a database constructed based on the associated users determined by the above method and the target associated data corresponding to the associated users.

[0022] Specifically, in S10, the target database is obtained through the following steps:

[0023] S100, obtain the target associated data list corresponding to the target service platform, wherein the target associated data list includes several target associated data, and the target associated data is data uploaded through multiple data ports corresponding to the target service platform.

[0024] Specifically, the target service platform is an internet application developed based on internet technology that can perform several functions.

[0025] Specifically, each target associated data corresponds to a data port, which is a port for user login and use, such as the data port of an APP, mini program, or web page.

[0026] S200, based on the target associated data list, obtain a first target data list, wherein the first target data list includes a plurality of first target data, and the first target data is target associated data with a specified identifier obtained from the target associated data list.

[0027] Specifically, the designated identifier is a unique identifier that represents the identity of the target user, such as a login account or other designated identifier.

[0028] Furthermore, the target user is the user who connects to the target service platform using the data port corresponding to the target associated data.

[0029] S300, based on the first target data list, associate all target users corresponding to the first target data with the same specified identifier in the first target data list to determine the associated users. The associated users are used to indicate that the associated users are the same user.

[0030] Specifically, when the similarity between specified identifiers is 1, the specified identifiers are determined to be consistent. As those skilled in the art know, any method for obtaining identifier similarity in the prior art falls within the protection scope of this invention, and will not be elaborated here.

[0031] S400, based on the target associated data list, obtain a second target data list, wherein the second target data list includes a plurality of second target data, and the second target data is target associated data obtained from the target associated data list that does not have a specified identifier.

[0032] S500, based on the first target data list and the second target data list, if several feature data included in the second target data and several corresponding feature data included in the first target data meet the preset comparison rules, then the user corresponding to the second target data and the user corresponding to the first target data are determined as associated users.

[0033] Specifically, the S500 also includes the following steps:

[0034] S501, obtain the target feature vector list corresponding to the target associated data list, wherein the target feature vector list includes several target feature vectors, each target associated data corresponds to one target feature vector, and the target feature vector is a vector generated based on the data corresponding to several target features in each target associated data.

[0035] Specifically, in S501, the target feature vector is obtained through the following steps:

[0036] S1. Obtain a target feature list based on the target associated data list, wherein the target feature list includes several target features, and the target features are initial features selected from the initial feature list corresponding to the target data list.

[0037] Specifically, S1 also includes the following steps:

[0038] S11, Obtain the initial feature data list set B = {B1, ..., B1} corresponding to the target associated data list. i , ..., B n}, B i ={B i1 , ..., B ij , ..., B im}, B ij Let j be the data under the j-th initial feature in the initial feature data corresponding to the i-th target associated data, where j ranges from 1 to m, m is the number of initial features, i ranges from 1 to n, and n is the number of target associated data in the target associated data list.

[0039] Specifically, the initial features are the network features represented when the target user connects to the target service platform using the data port corresponding to the target associated data, such as device model, operating system, IP data, max address, network card, and other initial features.

[0040] Furthermore, if the target associated data does not include data under a certain initial feature, then the data under that initial feature is 0.

[0041] S12, when B is in B ij When all values ​​are not zero, the j-th initial feature is selected as the target feature.

[0042] S2, obtain the number of target features, wherein the number of target features is the number of target features in the target feature list.

[0043] S3, based on the number of target features, obtain the target feature vector. In S3, the target feature vector is obtained through the following steps:

[0044] S31, when the number of target features is less than 2 or the number of target features is not less than 20, obtain the target feature vector, wherein the target feature vector is a vector generated by concatenating the vectors generated based on the associated data of each target.

[0045] Specifically, for example: when the number of target features is 1, the target feature vector is a vector generated based on the associated data of a single target; when the number of target features is not less than 20, the target feature vector is obtained as L = (L1, ..., L...). v , ..., L b ), L vThe vector generated for the associated data of the v-th target, L v =(L v1 , ..., L vx , ..., L vp ), L vx For L v The value of the x-th bit in the vector is given, where x ranges from 1 to p, p is the dimension of the vector generated for each target-related data, and v ranges from 1 to b, where b is the number of target features.

[0046] Specifically, as those skilled in the art will know, any method in the prior art for generating vectors from text falls within the protection scope of this invention, and will not be elaborated further here.

[0047] S32, when the number of target features is not less than 2 and the number of target features is less than 20, obtain a specified priority, wherein the specified priority is the weight corresponding to the specified feature, and the specified feature is any target feature other than IP data obtained from the target feature list.

[0048] Specifically, the specified priority value ranges from 0 to 0.5. As those skilled in the art will know, the specified priority can be selected according to actual needs, and all such selections fall within the protection scope of this invention, which will not be elaborated further here.

[0049] S33, based on the number of target features and the specified priority, confirm the target feature vector. The target feature vector is obtained in S33 through the following steps:

[0050] S331, when the number of target features is not less than 2 and the number of target features is less than the target value, obtain the target feature vector F = (F1, ..., F2). v , ..., F b ), F v =η 1 v ×β v ,β v η is a vector generated based on the target association data corresponding to the v-th target feature. 1 v For β v The corresponding first weight, η 1 v The ratio of 2 to v is subtracted from the specified priority, where the target value is the reciprocal of the specified priority, the b-th target feature is IP data, v takes values ​​from 1 to b, and b is the number of target features.

[0051] S332, when the number of target features is not less than the target value and the number of target features is less than 20, obtain the target feature vector E = (E1, ..., E2). v , ..., E b Ev =η 2 v ×β v ,β v η is a vector generated based on the target association data corresponding to the v-th target feature. 2 v For β v The corresponding second weight, η 2 v The target value is 10 times the ratio of the specified priority value to v, where the target value is the reciprocal of the specified priority, the b-th target feature is IP data, v takes values ​​from 1 to b, and b is the number of target features.

[0052] S502, obtain a target similarity list corresponding to each target feature vector, wherein the target similarity list includes several target similarities, the target similarity is the similarity between the target feature vector and the candidate feature vector, and the candidate feature vector is any target feature vector other than the target feature vector in the target feature vector list.

[0053] S503, when the target similarity meets the preset conditions, determine the target users corresponding to the target feature vectors corresponding to this target similarity as associated users.

[0054] Specifically, the preset condition is that the similarity is not less than a preset similarity threshold. The preset similarity threshold ranges from 0.7 to 0.8. As those skilled in the art know, the preset similarity threshold can be selected according to actual needs, and all of these fall within the protection scope of this invention. Therefore, it will not be elaborated further here.

[0055] S600 constructs a target database based on the identified associated users and the target associated data corresponding to those users.

[0056] As mentioned above, when judging the similarity between users, different weights are assigned to the data of several features corresponding to users. At the same time, based on the different number of features corresponding to users, different methods are used to obtain the weight corresponding to each feature, which improves the accuracy of obtaining the feature vector corresponding to users on the data end, and makes the accuracy of obtaining the same user higher.

[0057] S20, obtain specified feature data corresponding to a specified user, wherein the specified feature data is behavioral data generated when the specified user connects with a target service platform at the current time.

[0058] S30, when the specified feature data meets the preset specified conditions, obtain the unique identifier corresponding to the specified user and insert the specified feature data into the corresponding initial feature data list to achieve data integration, wherein the unique identifier corresponding to the specified user is the initial user ID corresponding to the initial feature data that matches the specified feature data obtained from the target database.

[0059] Specifically, the preset specified condition is that there is initial feature data in the target database that can match the specified feature data, wherein the matching means that the similarity is greater than the preset specified similarity threshold.

[0060] Furthermore, the preset specified similarity threshold ranges from 0.7 to 0.8. As those skilled in the art know, the preset specified similarity threshold can be selected according to actual needs, and all such selections fall within the protection scope of this invention, which will not be elaborated further here.

[0061] S40, when the specified feature data does not meet the preset specified conditions, obtain the target time period, wherein the target time period is the time period that can confirm the identity of the specified user and requires the collection of specified user behavior feature data based on the initial feature data list corresponding to several initial user IDs.

[0062] Specifically, the starting time of the target time period is the current moment, and the time span corresponding to the target time period is the target time span.

[0063] Furthermore, the target time span is obtained through a target random forest.

[0064] Furthermore, S40 also includes the following steps:

[0065] S501, Obtain several intermediate features, wherein the intermediate features are features of the initial user behavior, including the number of times the APP is opened, the time interval between APP usage, and the duration of APP opening.

[0066] S502, for any intermediate feature, combine several preset time intervals corresponding to the intermediate feature into a target sample to obtain several target samples.

[0067] Specifically, the preset time interval is a preset target time span, wherein those skilled in the art set the preset time interval according to actual needs.

[0068] S503 obtains several historical sample datasets from several target samples through sampling with replacement.

[0069] Specifically, as those skilled in the art will know, the number of sampling times can be selected according to actual needs, all of which fall within the protection scope of this invention, and will not be elaborated further here.

[0070] S504: For any historical sample dataset, split the root node of the constructed initial decision tree at a preset time interval as a feature subset until a preset stopping condition is met to obtain the target decision tree.

[0071] Specifically, the preset stopping condition is that during the process of constructing the initial decision tree into the target decision tree, the number of leaf node samples is less than a preset threshold or the depth of the initial decision tree reaches a preset value.

[0072] S505, based on the target decision tree corresponding to each intermediate feature, obtain the target random forest model.

[0073] S50, acquire key feature data corresponding to a specified user, wherein the key feature data is behavioral feature data generated by the specified user connecting with several target service platforms within a target time period.

[0074] S60: After matching the key feature data with the data in the target database, the key feature data is inserted into the target database to achieve data integration.

[0075] This embodiment provides a data integration method, which includes the following steps: acquiring a target database, acquiring specified feature data corresponding to a specified user, when the specified feature data meets preset specified conditions, acquiring a unique identifier corresponding to the specified user and inserting the specified feature data into the corresponding initial feature data list to achieve data integration, when the specified feature data does not meet the preset specified conditions, acquiring a target time period, acquiring key feature data corresponding to the specified user, matching the key feature data with the data in the target database, and then inserting the key feature data into the target database to achieve data integration. This invention uses different methods to integrate data based on whether user behavior features can be matched with data in the database, and reasonably sets the time period for acquiring user-corresponding behavior features, thereby improving the accuracy of data integration.

[0076] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.

[0077] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0078] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.

Claims

1. A method for data integration, characterized in that, The method includes the following steps: S10, Obtain the target database, wherein the target database stores several initial user IDs and a list of initial feature data corresponding to each initial user ID. The initial user ID is a unique identifier representing the identity of the initial user. The list of initial feature data includes several initial feature data, which are behavioral data generated by the initial user connecting to several target service platforms through data ports within a historical time period. The target database is obtained in S10 through the following steps: S100, Obtain the target associated data list corresponding to the target service platform, wherein the target associated data list includes several target associated data, and the target associated data is data uploaded through multiple data ports corresponding to the target service platform; S200, based on the target associated data list, obtain a first target data list, wherein the first target data list includes a plurality of first target data, and the first target data is target associated data with a specified identifier obtained from the target associated data list; S300, based on the first target data list, associate all target users corresponding to the first target data with the same specified identifier in the first target data list to determine the associated users. The associated users are used to indicate that the associated users are the same user. S400, based on the target associated data list, obtain a second target data list, wherein the second target data list includes a plurality of second target data, and the second target data is target associated data obtained from the target associated data list that does not have a specified identifier; S500, based on the first target data list and the second target data list, if several feature data included in the second target data and several corresponding feature data included in the first target data meet the preset comparison rules, then the user corresponding to the second target data and the user corresponding to the first target data are determined as associated users; S600: Based on the identified associated users and the target associated data corresponding to the associated users, construct the target database; S20, obtain specified feature data corresponding to a specified user, wherein the specified feature data is behavioral data generated when the specified user connects with a target service platform at the current time; S30, when the specified feature data meets the preset specified conditions, obtain the unique identifier corresponding to the specified user and insert the specified feature data into the corresponding initial feature data list to achieve data integration, wherein the unique identifier corresponding to the specified user is the initial user ID corresponding to the initial feature data that matches the specified feature data obtained from the target database; S40, when the specified feature data does not meet the preset specified conditions, obtain the target time period, wherein the target time period is the time period that can confirm the identity of the specified user and requires the collection of the specified user behavior feature data based on the initial feature data list corresponding to several initial user IDs; S50, acquire key feature data corresponding to a specified user, wherein the key feature data is behavioral feature data generated by the specified user connecting with several target service platforms within a target time period; S60: After matching the key feature data with the data in the target database, the key feature data is inserted into the target database to achieve data integration.

2. The data integration method according to claim 1, characterized in that, The preset specified condition is that there is initial feature data in the target database that can match the specified feature data, wherein the matching means that the similarity is greater than the preset specified similarity threshold.

3. The data integration method according to claim 2, characterized in that, The preset similarity threshold ranges from 0.7 to 0.

8.

4. The data integration method according to claim 1, characterized in that, The starting time of the target time period is the current moment, and the time span corresponding to the target time period is the target time span.

5. The data integration method according to claim 1, characterized in that, S40 also includes the following steps: S501, Obtain several intermediate features, wherein the intermediate features are features of the initial user behavior that include the number of times the APP is opened, the time interval between APP usage, and the duration of APP opening; S502, for any intermediate feature, combine several preset time intervals corresponding to the intermediate feature into a target sample to obtain several target samples; S503: Obtain several historical sample datasets from several target samples through sampling with replacement; S504: For any historical sample dataset, split the root node of the constructed initial decision tree at a preset time interval as a feature subset until a preset stopping condition is met to obtain the target decision tree. S505, based on the target decision tree corresponding to each intermediate feature, obtain the target random forest model.

6. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the method as described in any one of claims 1-5.

7. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 6.