Cross-platform data association method and apparatus, electronic device, and computer storage medium
Through cross-platform behavior matching and association probability analysis, and the use of multi-dimensional behavior data instead of precise identifiers, the problem of cross-platform user behavior association is solved, effective user behavior association is achieved in a privacy-protected environment, and the accuracy of market evaluation and advertising delivery is improved.
Patent Information
- Application Number
- PCT/CN2025/078376
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-02-21
- Publication Date
- 2025-10-09
AI Technical Summary
With the improvement of privacy protection laws and regulations, it is difficult for brands to achieve correlation analysis of cross-platform user behavior through accurate user identifiers, making traditional methods difficult to implement.
By collecting user behavior data of multiple user accounts on multiple platforms, cross-platform behavior matching is performed to generate a set of behavior data pairs, and the behavior data pairs of the same user account are determined based on the association probability value, and multiple types and multi-dimensional behavior data are used to replace precise user identifiers for association.
In the absence of accurate user identifiers, effective association of cross-platform user behavior can be achieved, reducing trial and error costs and improving the accuracy of market evaluation and advertising delivery.
Smart Images

Figure CN2025078376_09102025_PF_FP_ABST
Abstract
Description
Cross-platform data association method, device, electronic device and computer storage medium Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a cross-platform data association method, device, electronic device, and computer storage medium. Background Art
[0002] In an increasingly complex digital environment, brands, faced with a multitude of websites, applications (APPs), and mini-programs developed by different developers, have a practical need to correlate user behaviors across different software (or platforms). For example, a user browses an ad on platform A and then completes a purchase on platform B, or reads an article on platform C and then transacts on platform D. This correlation is crucial for brands to evaluate marketing effectiveness, and the market demand is significant.
[0003] Traditional methods rely on precise user identifiers, such as ID numbers and mobile phone numbers, to analyze cross-platform user behavior. However, with the evolution of privacy laws and regulations, and the tightening of user privacy data controls by various platforms, it is becoming increasingly difficult for brands to obtain these private user fields from the platforms, making traditional correlation analysis based on precise ID matching difficult to implement.
[0004] Currently, no effective solution has been found for the above-mentioned problems existing in the related technologies. Summary of the Invention
[0005] The present application provides a cross-platform data association method, device, electronic device and computer storage medium to solve the above-mentioned technical problems existing in the related art.
[0006] According to one embodiment of the present application, a cross-platform data association method is provided, including: collecting multiple user behavior data generated by multiple user accounts on multiple platforms; performing cross-platform behavior matching on the multiple user behavior data to generate a behavior data pair set, and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms, respectively, and the association probability value is used to characterize the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; and determining, based on the association probability value, a number of behavior data pairs in which the first behavior data and the second behavior data in the behavior data pair set belong to the same user account.
[0007] According to another embodiment of the present application, a cross-platform data matching device is provided, including: an acquisition module for acquiring multiple user behavior data generated by multiple user accounts on multiple platforms; a matching module for performing cross-platform behavior matching on the multiple user behavior data, generating a behavior data pair set, and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms, respectively, and the association probability value is used to characterize the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; a determination module for determining, based on the association probability value, several behavior data pairs in which the first behavior data and the second behavior data in the behavior data pair set belong to the same user account.
[0008] Optionally, the matching module includes an extraction unit for extracting multiple feature fields of each user behavior data; a data processing unit for standardizing the feature values of the multiple feature fields respectively to obtain multiple intermediate features corresponding to the multiple feature fields; a matching unit for determining the weight corresponding to each intermediate feature, inputting the intermediate features and the weights into a pre-built matching model, and using the matching model to perform feature similarity matching on user behavior data across platforms to generate a set of behavior data pairs and an association probability value for each behavior data pair.
[0009] Optionally, the data processing unit includes an extraction subunit, which is used to extract multiple feature values of each feature field for each feature field, wherein each of the multiple feature values comes from a user behavior data, and each user behavior data corresponds to a platform; and a conversion unit, which is used to convert the multiple feature values into intermediate features in a unified format using a preset standard format.
[0010] Optionally, the cross-platform data association device also includes a first acquisition module for acquiring positive samples and negative samples, wherein the positive samples are behavioral data pairs whose association probability values are greater than or equal to a preset threshold, and the negative samples are behavioral data pairs whose association probability values are less than the preset threshold; and a training module for training the initial model using the positive samples and the negative samples to obtain a pre-constructed matching model.
[0011] Optionally, the training module also includes an acquisition subunit for obtaining the value range corresponding to each hyperparameter in the initial model; a listing unit for listing multiple hyperparameter combinations within the value range; a verification unit for inputting the positive sample and the negative sample into the initial model for verification for each hyperparameter combination in the multiple hyperparameter combinations to obtain multiple verification scores; and a selection unit for selecting the target hyperparameter combination with the highest verification score, and setting the hyperparameters of the initial model to the values of the target hyperparameter combination as a pre-constructed matching model.
[0012] Optionally, the cross-platform data association device also includes a first identification module, used to identify a first timestamp corresponding to the transaction behavior, and a second timestamp corresponding to the browsing behavior, wherein the first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior; an analysis module, used to analyze the order of the transaction behavior and the browsing behavior based on the first timestamp and the second timestamp; a first adjustment module, used to increase the association probability value of the behavior data pair if the transaction behavior is after the browsing behavior.
[0013] Optionally, the cross-platform data association device also includes a second identification module for identifying the browsing mode of the browsing behavior and / or the type of operation on the browsing content, wherein the first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior; a second adjustment module for adjusting the association probability value based on the association strength value between the browsing mode and the transaction behavior, and / or adjusting the association probability value based on the association strength value between the operation type and the transaction behavior.
[0014] According to another embodiment of the present application, a computer storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned device embodiments when running.
[0015] According to another embodiment of the present application, an electronic device is also provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein: the memory is used to store computer programs; the processor is used to execute the steps in the above method by running the program stored in the memory.
[0016] According to another embodiment of the present application, a computer program product comprising instructions is provided. When the computer program product is run on a computer, the computer is caused to execute the steps in the above method.
[0017] Through the embodiments of the present application, multiple user behavior data generated by multiple user accounts on multiple platforms are collected; cross-platform behavior matching is performed on the multiple user behavior data to generate a behavior data pair set and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms respectively, and the association probability value is used to characterize the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; according to the association probability value, several behavior data pairs in which the first behavior data and the second behavior data in the behavior data pair set belong to the same user account are determined. When it is impossible to obtain or use an accurate user identifier to associate user behaviors across platforms, behavior data is used instead of the accurate user identifier to achieve cross-platform user behavior association, thereby overcoming the limitations of traditional user matching that relies on a single accurate identifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] FIG1 is a block diagram of the hardware structure of a computer according to an embodiment of the present application;
[0020] FIG2 is a flow chart of a cross-platform data association method according to an embodiment of the present application;
[0021] FIG3 is a schematic diagram of an implementation flow of a cross-platform data association method according to an embodiment of the present application;
[0022] FIG4 is a structural block diagram of a cross-platform data matching device according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of this application. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] Example 1
[0026] The method embodiment provided in the first embodiment of the present application can be executed in a mobile phone, a computer, a tablet or a similar computing device. Taking running on a computer as an example, Figure 1 is a hardware structure block diagram of a computer in an embodiment of the present application. As shown in Figure 1, the computer may include one or more (only one is shown in Figure 1) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned computer may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that the structure shown in Figure 1 is only for illustration and does not limit the structure of the above-mentioned computer. For example, the computer may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0027] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as a computer program corresponding to a cross-platform data association method in an embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0028] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0029] In this embodiment, a cross-platform data association method is provided. FIG2 is a flow chart of a cross-platform data association method according to an embodiment of the present application. As shown in FIG2 , the process includes the following steps:
[0030] Step S10, collecting multiple user behavior data generated by multiple user accounts on multiple platforms;
[0031] In this embodiment, an application capable of compliant data collection can be registered on multiple platforms that need to be associated, thereby collecting multiple user behavior data generated by multiple user accounts on the current platform. In this embodiment, two platforms (Platform A and Platform B) that need to be behaviorally associated are used as an example. For example, Platform A is mainly used for product recommendations, and users can browse product recommendation notes, articles, videos, etc. on Platform A. Platform B is mainly used for product transactions, and users can purchase products on Platform B.
[0032] Step S20: performing cross-platform behavior matching on the multiple user behavior data to generate a set of behavior data pairs and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms, respectively, and the association probability value is used to represent the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account;
[0033] Optionally, cross-platform behavior matching is performed on multiple user behavior data to generate a set of behavior data pairs. The method can be to combine any user behavior data in a certain platform with any user behavior data in another platform in pairs to obtain multiple behavior data pairs. The multiple behavior data of each behavior data pair come from different platforms. In this embodiment, each behavior data pair includes first behavior data and second behavior data, and the first behavior data and the second behavior data correspond to different platforms. In this embodiment, the first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior. For example, the transaction behavior of platform A can be associated with the browsing behavior of platform B. Other behavior data associations can also be achieved, such as the association of reading behavior with search behavior, etc., which are not limited in this embodiment. Each behavior data pair corresponds to an association probability value to characterize the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account.
[0034] Step S30, determining, based on the association probability value, a number of behavior data pairs in the behavior data pair set, in which the first behavior data and the second behavior data belong to the same user account;
[0035] Obtain a preset threshold and determine whether the association probability value of the current behavior data pair is greater than the preset threshold. If the association probability value is greater than the preset threshold, determine that the first behavior data and the second behavior data in the current behavior data pair belong to the same user account; if the association probability value is less than the preset threshold, determine that the first behavior data and the second behavior data in the current behavior data pair do not belong to the same user account. Extract several behavior data pairs in which the first behavior data and the second behavior data belong to the same user account from the behavior data pair set. This can be used by brands for market evaluation, advertising placement and other businesses, to understand the behavior and decision-making paths of cross-platform buyers, and allow decision makers to find the optimal solution from multiple placement plans, reducing trial and error costs.
[0036] Through the above steps, multiple user behavior data generated by multiple user accounts on multiple platforms are collected; cross-platform behavior matching is performed on the multiple user behavior data to generate a behavior data pair set and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms respectively, and the association probability value is used to characterize the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; according to the association probability value, several behavior data pairs in which the first behavior data and the second behavior data in the behavior data pair set belong to the same user account are determined. When it is impossible to obtain or use accurate user identifiers to associate cross-platform user behaviors, behavior data is used instead of accurate user identifiers to achieve cross-platform user behavior association, thereby overcoming the limitations of traditional reliance on a single accurate identifier for user matching.
[0037] In one implementation of this embodiment, cross-platform behavior matching is performed on the plurality of user behavior data to generate a set of behavior data pairs, and the association probability value of each behavior data pair includes:
[0038] A21, extracts multiple feature fields of each user behavior data;
[0039] A22, performing standardization processing on the feature values of the plurality of feature fields to obtain a plurality of intermediate features corresponding to the plurality of feature fields;
[0040] A23, determine the weight corresponding to each intermediate feature, input the intermediate feature and the weight into a pre-built matching model, use the matching model to perform feature similarity matching on user behavior data across platforms, and generate a set of behavior data pairs and an association probability value for each behavior data pair.
[0041] In this embodiment, the characteristic fields of different user behavior data may be different. For example, the characteristic fields that can be collected by each platform include but are not limited to behavioral characteristic fields such as network IP address, geographic location, browsing time, browsing content, device information used for browsing (device brand, device model, operating system and version), browser information, browsing time, browsing source (search, information flow, etc.).
[0042] In addition to the above behavioral feature fields, user feature fields and browsing content feature fields can also be obtained. Browsing content feature fields include content categories and keywords, while user feature fields include gender, age group, marital status, hobbies, etc. In one example, a piece of browsing behavior data includes but is not limited to the following feature fields: browsing time, browsing duration, merchants of related products in the browsing content, network IP address and device information at the time of browsing, etc.; a piece of transaction behavior data includes but is not limited to the following feature fields: user unique ID, merchant corresponding to the transaction, product ID of the transaction, transaction time, network IP address and device information, etc.
[0043] To facilitate subsequent feature similarity matching, the feature values of multiple feature fields are standardized separately to obtain multiple intermediate features corresponding to the multiple feature fields. The weight corresponding to each intermediate feature is determined, and the intermediate features and weights are input into a pre-built matching model. The matching model is used to perform feature similarity matching on user behavior data across platforms to generate a set of behavior data pairs and an association probability value for each behavior data pair.
[0044] Optionally, an initial weight value can be pre-configured for each intermediate feature based on comprehensive factors such as the feature's data recognition, data quality, and data coverage, and this initial weight value is used as the initial input to the matching model. In one example, the data coverage is inversely proportional to the weight. For example, each user behavior data set includes three fields: province, city, and network IP address. The data coverage of the city is smaller than that of the province, and the accuracy of the city's association with user behavior is higher than that of the province. Therefore, the weight of the city is higher than that of the province. The network IP address is more accurate than the city, so the weight of the network IP address is higher than that of the city. In another example, data recognition is directly proportional to the weight. For example, each user behavior data set includes two fields: gender and age group. Gender may include male, female, and unknown, and has a lower recognition of the user, so its value is relatively small. The age group may be divided into multiple areas such as 0-10 years old and 10-20 years old, and its value is greater than that of gender. Therefore, the weight of the age group field is higher than that of the gender field.
[0045] In this embodiment, a matching model is used to perform feature similarity matching on various types and dimensions of features in user behavior data. Behavioral data generally carries other fields, such as the network IP address, geographic location, and device information where the behavior occurs. Such behavioral features cannot directly determine whether two behavioral data belong to the same user account, but in the absence of an ID Mapping (mapping between precise user IDs) field, these fields will greatly increase the success rate of behavioral fuzzy matching. For example, the probability of two cross-platform behaviors being associated with the same device information under the same network IP address or geographic location is relatively high.
[0046] In this embodiment, the characteristic values of the plurality of characteristic fields are respectively standardized to obtain a plurality of intermediate features corresponding to the plurality of characteristic fields, including:
[0047] a1. For each feature field, extract multiple feature values of the feature field, wherein each feature value in the multiple feature values comes from a piece of user behavior data, and each piece of user behavior data corresponds to a platform;
[0048] a2, uses a preset standard format to convert multiple feature values into intermediate features in a unified format.
[0049] Because data from various data sources may not come from a single API and have different data formats, non-standard data needs to be converted to a standard format. Therefore, in this implementation, multiple feature values are extracted for each feature field. Each of these multiple feature values comes from a piece of user behavior data, and each piece of user behavior data corresponds to a platform. These multiple feature values are converted into intermediate features in a unified format using a preset standard format. For example, a detailed address string accurate to a neighborhood or building needs to be converted into a unified format of longitude and latitude.
[0050] Optionally, in this embodiment, multiple feature fields can also be cleaned and filtered, and invalid fields in a piece of user behavior data can be discarded. For example, if the address information in the geographic location field cannot be parsed into longitude and latitude, or the network IP address format is chaotic or garbled, the invalid fields in the user behavior data will be discarded, and other fields will be retained.
[0051] This implementation method is based on the analysis and identification of the full amount of user data collected by the two platforms, processed through database data (or big data technology), and cleaned out irrelevant data.
[0052] In this embodiment, before inputting the intermediate features and the weights into the pre-constructed matching model, the method further includes: obtaining positive samples and negative samples, wherein the positive samples are behavioral data pairs whose associated probability values are greater than or equal to a preset threshold, and the negative samples are behavioral data pairs whose associated probability values are less than the preset threshold; and using the positive samples and the negative samples to train the initial model to obtain a pre-constructed matching model.
[0053] In this embodiment, the positive samples can be pairs of behavioral data that have been verified to be accurate and belong to the same user account, and the negative samples can be pairs of behavioral data that have been verified to be accurate and do not belong to the same user account. Collect sample data and divide the data set into training samples and test samples according to a preset ratio (for example, 6:4). Standardize the features of the training set and the test set, for example, unify the data into a standard table format. Input the training feature set and the test feature set into the initial model for training to obtain a pre-built matching model.
[0054] The positive samples and the negative samples are used to train the initial model to obtain a pre-constructed matching model, which includes: obtaining the value range corresponding to each hyperparameter in the initial model; listing multiple hyperparameter combinations within the value range; for each hyperparameter combination in the multiple hyperparameter combinations, inputting the positive sample and the negative sample into the initial model for verification to obtain multiple verification scores; selecting the target hyperparameter combination with the highest verification score, and setting the hyperparameters of the initial model to the values of the target hyperparameter combination as the pre-constructed matching model.
[0055] In this embodiment, the initial model can adopt the K-Nearest Neighbors (KNN) model. The hyperparameters in the initial model include but are not limited to K value (number of neighbors), p (distance metric), and weight (weighting method). According to the ratio and distribution of positive and negative samples, the hyperparameters such as the number of neighbors, distance metric, and weighting method in the initial model are set to a wide range, and GridSearchCV is used to select the optimal parameters.
[0056] This implementation traverses preset hyperparameter combinations and uses cross-validation to evaluate the model performance under each set of hyperparameters to find the optimal hyperparameter settings.
[0057] This embodiment realizes the use of various types and dimensions of behavioral data to replace user privacy fields when it is impossible to obtain or use accurate user identifiers to correlate cross-platform user behaviors, and performs AI algorithm modeling to achieve cross-platform user behavior correlation. It also uses known positive and negative samples to tune the program and continuously improve the accuracy of the cross-platform user behavior correlation effect.
[0058] In another implementation of this embodiment, the first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior. After generating a set of behavior data pairs and an association probability value for each behavior data pair, the method further includes:
[0059] B1, identifying a first timestamp corresponding to the transaction behavior and a second timestamp corresponding to the browsing behavior;
[0060] B2, analyzing the sequence of the transaction behavior and the browsing behavior based on the first timestamp and the second timestamp;
[0061] B3. If the transaction behavior occurs after the browsing behavior, the association probability value of the behavior data pair is increased.
[0062] In this embodiment, when associating cross-platform transaction behavior with browsing behavior, the contribution of browsing behavior to transaction behavior is considered. The first timestamp corresponding to the transaction behavior in the behavior data pair and the second timestamp corresponding to the browsing behavior in the behavior data pair are identified. Based on the first and second timestamps, the order of the transaction behavior and the browsing behavior is analyzed. If the transaction behavior occurs after the browsing behavior, it indicates a typical user conversion behavior from browsing to purchase, that is, the browsing behavior contributes to the transaction conversion, and the association probability value of the behavior data pair is increased. If the transaction behavior occurs before the browsing behavior, it indicates that the browsing behavior does not contribute to the transaction conversion, and the association probability value of the behavior data pair is decreased, thereby improving the overall association accuracy.
[0063] In another implementation of this embodiment, the first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior. After generating a set of behavior data pairs and an association probability value for each behavior data pair, the method further includes:
[0064] C1, identifying the browsing mode of the browsing behavior and / or the type of operation on the browsing content;
[0065] C3: adjusting the association probability value based on the association strength value between the browsing mode and the transaction behavior, and / or adjusting the association probability value based on the association strength value between the operation type and the transaction behavior.
[0066] In this embodiment, the browsing mode refers to the way a user browses a piece of content, including but not limited to information flow recommendations or searches. User-initiated searches have a higher association strength than information flow recommendations. The types of operations performed on the browsed content include but are not limited to browsing time, sharing, and saving, each representing a different level of association strength. The association probability value is adjusted based on the association strength between the browsing mode and the transaction behavior, and / or based on the association strength between the operation type and the transaction behavior, where the association strength value is proportional to the association probability value.
[0067] In this embodiment, for a specific user or a specific product, the transaction decision generally has certain characteristics. For example, the more expensive the product, the longer the user will spend to understand it or the more frequently the user will browse information to understand it. Therefore, if the same transaction behavior is associated with multiple browsing behaviors, the association probability value is adjusted according to the time difference between the browsing behavior and the transaction behavior. Among them, the time of the browsing behavior is before the time of the transaction behavior. The smaller the time difference, the larger the corresponding association probability value.
[0068] The solution of this embodiment provides a cross-platform data association method. Figure 3 is an implementation flow chart of an embodiment of the present application, including: users visit platform A and platform B, collect user behavior data of platform A and platform B respectively, and perform feature engineering on the user behavior data of platform A and platform B, that is, data standardization processing, input the standardized features into a fuzzy matching model, and obtain the matching behavior and corresponding probability value output by the fuzzy matching model. This embodiment realizes the use of non-privacy fields to achieve cross-platform user behavior fuzzy matching when the user's privacy fields cannot be obtained, and continuously optimizes the feature weights and positive and negative sample data to meet the brand's needs for cross-platform user behavior association analysis.
[0069] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0070] Example 2
[0071] In this embodiment, a cross-platform data matching device is also provided for implementing the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0072] FIG4 is a structural block diagram of a cross-platform data matching device according to an embodiment of the present application. As shown in FIG4 , the device includes:
[0073] A collection module 40 is used to collect multiple user behavior data generated by multiple user accounts on multiple platforms;
[0074] a matching module 42 configured to perform cross-platform behavior matching on the plurality of user behavior data to generate a set of behavior data pairs and an association probability value for each behavior data pair, wherein the behavior data pair set includes a plurality of behavior data pairs, each behavior data pair includes browsing behavior and transaction behavior, and the association probability value is used to represent the probability that the browsing behavior and transaction behavior in the same behavior data pair belong to the same user account;
[0075] The determination module 44 is configured to determine the user to which each behavior data pair in the behavior data pair set belongs according to the association probability value.
[0076] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0077] Example 3
[0078] An embodiment of the present application further provides a computer storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when running.
[0079] Optionally, in this embodiment, the computer storage medium may be configured to store a computer program for performing the following steps:
[0080] S1, collects multiple user behavior data generated by multiple user accounts on multiple platforms;
[0081] S2. Perform cross-platform behavior matching on the multiple user behavior data to generate a set of behavior data pairs and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms, respectively, and the association probability value is used to represent the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account;
[0082] S3, determining, according to the association probability value, a number of behavior data pairs in the behavior data pair set, in which the first behavior data and the second behavior data belong to the same user account.
[0083] Optionally, in this embodiment, the above-mentioned computer storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0084] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0085] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0086] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0087] S1, collects multiple user behavior data generated by multiple user accounts on multiple platforms;
[0088] S2. Perform cross-platform behavior matching on the multiple user behavior data to generate a set of behavior data pairs and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data correspond to different platforms, respectively, and the association probability value is used to represent the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account;
[0089] S3, determining, according to the association probability value, a number of behavior data pairs in the behavior data pair set, in which the first behavior data and the second behavior data belong to the same user account.
[0090] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0091] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0092] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0094] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0095] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a computer storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned computer storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0097] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A cross-platform data association method, characterized in that: The method comprises: Collect multiple user behavior data generated by multiple user accounts on multiple platforms; Performing cross-platform behavior matching on the multiple user behavior data to generate a set of behavior data pairs and an association probability value for each behavior data pair, wherein the behavior data pair set includes multiple behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data respectively corresponding to different platforms, and the association probability value is used to represent the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; Determine, according to the association probability value, a number of behavior data pairs in the behavior data pair set, in which the first behavior data and the second behavior data belong to the same user account.
2. The method according to claim 1, characterized in that Performing cross-platform behavior matching on the plurality of user behavior data to generate a set of behavior data pairs, and the association probability value of each behavior data pair includes: Extract multiple feature fields from each user's behavior data; Normalizing the characteristic values of the plurality of characteristic fields to obtain a plurality of intermediate characteristics corresponding to the plurality of characteristic fields; Determine the weight corresponding to each intermediate feature, input the intermediate feature and the weight into a pre-built matching model, use the matching model to perform feature similarity matching on user behavior data across platforms, and generate a set of behavior data pairs and an association probability value for each behavior data pair.
3. The method according to claim 2, characterized in that Normalizing the characteristic values of the plurality of characteristic fields to obtain a plurality of intermediate characteristics corresponding to the plurality of characteristic fields, including: For each feature field, extract multiple feature values of the feature field, wherein each feature value in the multiple feature values comes from a piece of user behavior data, and each piece of user behavior data corresponds to a platform; A preset standard format is used to convert multiple feature values into intermediate features in a unified format.
4. The method according to claim 2, characterized in that Before inputting the intermediate features and the weights into a pre-built matching model, the method further includes: Acquire positive samples and negative samples, wherein the positive samples are behavior data pairs whose association probability values are greater than or equal to a preset threshold, and the negative samples are behavior data pairs whose association probability values are less than the preset threshold; The positive samples and the negative samples are used to train the initial model to obtain a pre-built matching model.
5. The method according to claim 4, characterized in that The initial model is trained using the positive samples and the negative samples to obtain a pre-built matching model including: Obtaining the value range corresponding to each hyperparameter in the initial model; List multiple hyperparameter combinations within the value range; For each hyperparameter combination in the multiple hyperparameter combinations, input the positive sample and the negative sample into the initial model for verification to obtain multiple verification scores; The target hyperparameter combination with the highest validation score is selected, and the hyperparameters of the initial model are set to the values of the target hyperparameter combination as the pre-built matching model.
6. The method according to claim 1, characterized in that The first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior. After generating a set of behavior data pairs and an association probability value for each behavior data pair, the method further includes: Identifying a first timestamp corresponding to the transaction behavior and a second timestamp corresponding to the browsing behavior; analyzing a sequence of the transaction behavior and the browsing behavior according to the first timestamp and the second timestamp; If the transaction behavior occurs after the browsing behavior, the association probability value of the behavior data pair is increased.
7. The method according to claim 1, characterized in that The first behavior data in the behavior data pair is a transaction behavior, and the second behavior data is a browsing behavior. After generating a set of behavior data pairs and an association probability value for each behavior data pair, the method further includes: Identify the browsing behavior and / or the type of operation on the browsing content; The association probability value is adjusted based on the association strength value between the browsing mode and the transaction behavior, and / or the association probability value is adjusted based on the association strength value between the operation type and the transaction behavior.
8. A cross-platform data association device, characterized in that: include: The collection module is used to collect multiple user behavior data generated by multiple user accounts on multiple platforms; a matching module, configured to perform cross-platform behavior matching on the plurality of user behavior data, generate a set of behavior data pairs, and an association probability value for each behavior data pair, wherein the behavior data pair set includes a plurality of behavior data pairs, each behavior data pair includes first behavior data and second behavior data, the first behavior data and the second behavior data corresponding to different platforms, respectively, and the association probability value is used to represent the probability that the first behavior data and the second behavior data in the same behavior data pair belong to the same user account; A determination module is used to determine, based on the association probability value, a number of behavior data pairs in the behavior data pair set, in which the first behavior data and the second behavior data belong to the same user account.
9. An electronic device, characterized in that: The system comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; wherein: Memory for storing computer programs; A processor, configured to execute the cross-platform data association method according to any one of claims 1 to 7 by running a program stored in a memory.
10. A computer storage medium, characterized in that The computer storage medium includes a stored program, wherein the cross-platform data association method according to any one of claims 1 to 7 is executed when the program is run.
Citation Information
Patent Citations
User recognition method and device
CN108197190A
Cross-platform social network user identity recognition method
CN108897789A
Cross-device user identification method and device, electronic equipment and storage medium
CN111241502A
Cross-platform data information matching management method and system
CN117575717A
Cross-platform data association method and device, electronic equipment and computer storage medium
CN118246953A