A data processing method, device and electronic equipment
By analyzing common data and calculating the strength of associations among users, this technology solves the problem of analyzing the strength of associations among users who do not interact, enabling in-depth exploration of the deep structure of complex social networks and precise decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CETC NEW SMART CITY RES INST CO LTD
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to reveal the deep structures and implicit connections behind complex social networks, especially the strength of user associations between users who do not interact.
By identifying users participating in multiple target events, we conduct correlation analysis to obtain common data among the second users. Based on the common data and the associated first users, we determine the user correlation strength, extract common data across different data types, configure business scenario weights, and calculate the target correlation strength by combining the user and community correlation strength.
It expands the scope of association analysis, avoids missing valuable associations, uncovers hidden needs, supports accurate decision-making, enables association analysis of users without interactive behavior, and improves the ability to explore user relationships.
Smart Images

Figure CN121278672B_ABST
Abstract
Description
A data processing method, apparatus and electronic device Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a data processing method, apparatus and electronic device. Background Technology
[0002] Currently, in the practice of social relationship analysis, most methods are still limited to identifying the surface relationships between users based on direct interactions (such as comments, likes, private messages, etc.), and it is difficult to further reveal the deep structure and non-explicit connections hidden behind complex social networks.
[0003] For example, there are users A, B, and C. Users A and B have interactive behaviors. Current social analysis methods can analyze the strength of the social relationship between users A and B, but this analysis method has limitations. Summary of the Invention
[0004] This application provides a data processing method, apparatus, and electronic device that can analyze the strength of user associations between users who do not interact.
[0005] In a first aspect, embodiments of this application provide a data processing method, comprising: identifying multiple users participating in multiple target events; performing correlation analysis on different target events to identify a first user and a second user from the multiple users; the first user being a user who participates in multiple target events simultaneously, and the second user being a user who participates in some of the target events; the second users being associated with each other through the first user; obtaining common data among the second users; the common data being identical user data generated when the second users do not interact with each other; and determining the user association strength among the second users based on the common data and the first user associated with each second user.
[0006] Optionally, obtaining common data among the second users includes: obtaining first user data of different second users; classifying the first user data according to data type to obtain second user data corresponding to different data types; extracting the same second user data from multiple second user data as the common data; the same second user data includes the same user characteristics and / or the same behavioral characteristics of the second users.
[0007] Optionally, determining the user association strength between the second users based on the common data and the first users associated with the second users includes: obtaining the user association strength based on the number of first users associated with the second users, a first weight corresponding to the number of users, the common data, and a second weight corresponding to the common data.
[0008] Optionally, the common data includes first common data and second common data. The first common data is common data associated with a business scenario, and the second common data is common data not associated with the business scenario. The second weight corresponding to the common data includes a third weight corresponding to the first common data and a fourth weight corresponding to the second common data. The method further includes: obtaining the business scenario; different business scenarios are associated with different common data; configuring the third weight for the first common data associated with the business scenario, and configuring the fourth weight for the second common data not associated with the business scenario; the third weight is greater than the fourth weight.
[0009] Optionally, the method further includes: acquiring multiple events in which the second user participates; performing correlation analysis on every two events to obtain multiple event correlation strengths; each event correlation strength is used to characterize the closeness of the correlation between every two events; obtaining a community correlation strength based on the multiple event correlation strengths; the community correlation strength is used to characterize the probability that different second users are in the same community; and obtaining a target correlation strength between the second users based on the user correlation strength and the community correlation strength.
[0010] Optionally, the step of performing correlation analysis on every two events among the multiple events to obtain the correlation strength of multiple events includes: obtaining the correlation strength of the events based on a first quantity and a second quantity; the first quantity is the number of the same second user appearing in every two events, and the second quantity is the number of different second users in every two events.
[0011] Optionally, obtaining the target association strength between the second users based on the user association strength and the community association strength includes: obtaining a first association strength between the second users based on the user association strength, a fifth weight corresponding to the user association strength, the community association strength, and a sixth weight corresponding to the community association strength; obtaining a second association strength based on the user association strength and the community association strength; and obtaining the target association strength based on the first association strength and the second association strength.
[0012] Optionally, identical second user data among the second users is used to characterize that the similarity between the second user data of the second users is greater than a preset similarity.
[0013] Secondly, embodiments of this application provide a data processing apparatus, comprising: a user determination module, configured to determine multiple users participating in multiple target events; a correlation analysis module, configured to perform correlation analysis on different target events, and identify a first user and a second user from the multiple users; the first user is a user who participates in multiple target events simultaneously, and the second user is a user who participates in some of the target events; the second users are associated with each other through the first user; an acquisition module, configured to acquire common data among the second users; the common data is identical user data generated when the second users do not interact with each other; and a correlation strength determination module, configured to determine the user correlation strength among the second users based on the common data and the first user associated with the second users.
[0014] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data processing method provided in the first aspect.
[0015] The beneficial effects of the embodiments in this application compared with the prior art are:
[0016] It can identify the first users associated with the second users and the common data between the second users, and then determine the strength of the user association between the second users based on the associated first users and the common data. It is not limited to analyzing the strength of user association between users who have interactive behavior, but also realizes the analysis of user association between second users who do not have interactive behavior, thereby exploring the hidden user relationships between the second users. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 is a flowchart of the steps of a data processing method provided in an embodiment of this application;
[0019] Figure 2 is a flowchart of the steps of a data processing method provided in an embodiment of this application;
[0020] Figure 3 is a schematic diagram of dividing the first user data of a second user into different second user data according to an embodiment of this application;
[0021] Figure 4 is a flowchart of the steps of a data processing method provided in an embodiment of this application;
[0022] Figure 5 is a flowchart of the steps of a data processing method provided in an embodiment of this application;
[0023] Figure 6 is a block diagram of a data processing apparatus provided in an embodiment of this application. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0028] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] Figure 1 shows a schematic flowchart of the data processing method provided in this application. It is an example and not a limitation. The data processing method can be applied to electronic devices such as mobile phones, computers, and tablets. The data processing method includes the following steps:
[0031] S101, identify multiple users who participate in multiple target events.
[0032] The target event is the interactive event of current interest, which includes the event content, the platform on which the event is participated, the location of the event, the time of the event, and the users participating in the event. The event publishing platform is the social platform used by users to participate in the target event.
[0033] For each target event, users participating in that target event can come from multiple event participation platforms, and different users under the same target event are interested in the same event content.
[0034] For example, the following interactive behaviors could occur for the target event "AI Technology Summit":
[0035] (1) User A posted a post about the AI Technology Summit on Platform X. The content of the event is "AI Technology Summit", the platform involved in the event is "Platform X", and the user involved in the event is "User A".
[0036] (2) User B sent an invitation to the AI Technology Summit to User C through Platform Y. The content of the event is "AI Technology Summit", the platform involved in the event is "Platform Y", and the users involved in the event are "User B and User C".
[0037] (3) User C had a 15-minute phone call with User A through Platform Z to discuss the details of attending the AI Technology Summit. The content of the event is "AI Technology Summit", the platform involved in the event is "Platform Z", and the users involved in the event are "User A and User C".
[0038] (4) User B forwarded the AI Technology Summit post by User A through Platform W and actively participated in the comments. Therefore, the content of the event is "AI Technology Summit", the platform involved in the event is "Platform W", and the users involved in the event are "User A and User B".
[0039] For example, the following interactive behaviors can occur for the target event "discussion on technological trends":
[0040] (5) User B posted his opinion on “Technology Trend Discussion” through platform H. The content of the event is “Technology Trend Discussion”, the platform participating in the event is “Platform H”, and the user participating in the event is “User B”.
[0041] (6) User D sent a technology trend discussion email to User B through Platform O. The content of the event is "technology trend discussion", the platform involved in the event is "Platform O", and the users involved in the event are "User B and User D".
[0042] (7) User A and User D discussed technology trends on the same platform G. The content of the event is “Technology Trend Discussion”, the platform involved in the event is “Platform G”, and the users involved in the event are “User A and User D”.
[0043] For example, regarding the target event "investment evaluation", the following interactive behaviors may occur:
[0044] (8) User A initiated an investment assessment to User B through Platform Y and Platform Z. The content of the event is "investment assessment", the participating platforms are "Platform Y and Platform Z", and the participating users are "User A and User B".
[0045] S102, perform correlation analysis on different target events to identify the first user and the second user from multiple users.
[0046] The first user is the user who participates in multiple target events simultaneously.
[0047] For example, regarding the aforementioned target event "AI Technology Summit," the participating users include "User A, User B, and User C"; regarding the aforementioned target event "Technology Trend Discussion," the participating users include "User A, User B, and User D"; and regarding the aforementioned target event "Investment Evaluation," the participating users include "User A and User B." It is evident that "User A and User B" are the first users simultaneously participating in all three target events.
[0048] The second user is a user who participates in some of the target events among multiple target events. It can also be understood as the second user being a user other than the first user among the multiple users who participate in the target events. The second users are associated with each other through the first user.
[0049] For example, in the three target events mentioned above, namely "AI Technology Summit", "Technology Trend Discussion" and "Investment Evaluation", users C and D are the second users.
[0050] Optionally, different target events can be constructed into different hyperedges, with each hyperedge representing a target event. Then, hyperedge correlation analysis can be performed on the different hyperedges to obtain the first user who participates in different hyperedges simultaneously.
[0051] For example, in the above example, the "AI Technology Summit" can be constructed as hyperedge e1 (AI Technology Summit): {A, B, C}; the "Technology Trend Discussion" can be constructed as hyperedge e2 (Technology Trend Discussion): {A, B, D}; and the "Investment Assessment" can be constructed as hyperedge e3: {A, B}. In all three hyperedges e1, e2, and e3, the nodes {A, B} are shared. Therefore, users A and B are the first users shared across all three target events. Users C and D each participate in a portion of the target events, so users C and D are the second users.
[0052] S103, Obtain common data among the second users.
[0053] Common data refers to identical user data generated when the second user does not interact with each other. This user data includes the same geographical location information, active time periods, user characteristics, and shared interests. User characteristics include the same age group and the same gender.
[0054] For example, if the second user is either user C or user D, and both user C and user D visited city A on November 19, 2025, although user C and user D do not have direct interaction, they share the same user data as "city A"; or if user C and user D are both frequently active on the same social media platform between 10 pm and 11 pm, this also constitutes the same user data.
[0055] Optionally, common data between second users can be determined by calculating the cosine similarity or cosine distance between the second user data. If the cosine similarity between the second user data is greater than a preset similarity or the cosine distance is less than a preset distance, then the two second user data are considered to be common data.
[0056] For example, if user C has been to city A, and user D has not been to city A, but the distance between the geographical locations of the two users is less than a preset distance, it is also considered that the two users have been to the same geographical location and have the common data of the same geographical location.
[0057] S104. Based on common data and the first users associated with the second users, determine the user association strength between the second users.
[0058] Among them, the user association strength between the second users is used to characterize the tightness of the relationship between the second users. The greater the user association strength, the tighter the relationship between the second users.
[0059] In this context, the first user associated between the second users is the user who associates the second users who do not have interactive behavior. For example, in the example above, user A and user B are the connection bridge between user C and user D.
[0060] Optionally, the user association strength can be obtained based on the number of first users associated with the second users, the first weight corresponding to the number of users, common data, and the second weight corresponding to the common data. The first weight corresponding to each common data point can be different, and the first and second weights can also be different; the first and second weights can be set according to actual needs.
[0061] For example, the formula for calculating user association strength is as follows:
[0062] (1)
[0063] In formula (1), It is the strength of the user association between the second user. It is the intercept of the J-th hidden neuron / hidden node; It represents the importance of the i-th feature to the j-th hidden neuron; It is the i-th input feature, and each input feature represents common data among users, for example... It is the shared geographical location information between the second user. It is the shared active time period among the second group of users. It is the number of first users associated with the second users; It is an activation function.
[0064] Optionally, after obtaining the user association strength, it can be input into an unsupervised deep relational reasoning network to further obtain the user association strength between the second users, and the calculation formula is as follows:
[0065] (2)
[0066] In formula (2), This is the score for the improved user association strength; the higher the value, the higher the confidence in predicting a implicit relationship between the second users. kThe intercept of the k-th response is scaled, resulting in the final bias term; d k The scaling factor for the k-th response is c, and the scaling factor for the output layer is c. k and d k Together, we adjust the range and distribution of the final output; k The intercept of the k-th response is the bias term of the output layer; is the coefficient between the j-th hidden node and the k-th response, representing the contribution of the j-th pattern to the final relation judgment; N H The number of hidden nodes determines the number of complex relational patterns that the model can learn and represent, and is a key hyperparameter for model capacity.
[0067] The input features in formula (1) above can be obtained through the following formula:
[0068] Query vector:
[0069] Key vector:
[0070] Value vector:
[0071] Attention weights:
[0072] Enhanced features: (3)
[0073] In formula (3), , , A trainable weight matrix for query, key, and value; , , For the corresponding bias term; The dimension of the key vector, used for scaling; A represents the input features enhanced by the attention mechanism, while A represents the input features before enhancement.
[0074] In related technologies, in low-density network environments, direct connections (such as phone calls and direct transactions) between nodes (e.g., users) are relatively rare. Most nodes exist in a sparse state with no direct connections, resulting in limited interaction and making it difficult to discover implicit relationships between users. Traditional solutions filter out common data without direct connections, such as mutual friends or similar transaction characteristics, as noise, leading to missed detections of this data. Furthermore, in coefficient data, a few occasional direct connections, such as accidental transfers between users, are treated as strong correlations, leading to false positives for strong relationship information.
[0075] The above technical solution can identify the common data between the first users and the second users, thereby determining the strength of the user association between the second users. It is not limited to analyzing the user association strength between users who have interacted, but also enables the analysis of user associations between second users who do not have interacting behavior. This allows for the exploration of hidden social relationships, which has the following advantages:
[0076] First, expand the scope of association analysis to avoid "missing" valuable associations. For example, the "first user" is a brand merchant, and the "second user" is a consumer who has never communicated directly but has purchased the same high-priced product from the merchant. These second users are essentially "potential colleagues or target customers" but have no direct interaction.
[0077] Secondly, uncovering "implicit needs or behavioral logic" can support precise decision-making. Social software has discovered that some users frequently use the resume template download function but lack any substantial interactive behavior. Modules such as "job referrals" and "interview communication" can be developed to address these implicit needs.
[0078] Figure 2 is an exemplary embodiment involved in step S103 above, which is used to interpret an exemplary scheme for obtaining common data between the second users, including the following steps:
[0079] S103-1, Obtain first user data from different second users.
[0080] Optionally, the system can acquire the first user data of the second user within a preset time period, such as the first user data generated by the second user within the past month, as shown in Figure 3. This first user data includes the second user's basic attributes, social behaviors, active time points, and geographical location information. The second user's basic attributes include the second user's gender, age group, height, and appearance. The first user data generated by the second user can originate from different social media platforms.
[0081] S103-2, the first user data is classified according to data type to obtain second user data corresponding to different data types.
[0082] As shown in Figure 3, the data types include text, images, videos, instructions, transaction information, etc. The first user data of the second user can be classified into the corresponding data type according to the type of the first user data. There is at least one second user data under the same data type. At least one second user data can be second user data from the same second user or different second users.
[0083] For example, referring to Figure 3, taking the second users, including users C, D, and E, as an example, if user C posts a text message on a social media platform, mentioning city A; user D posts an image showing a landmark building in city A; and user E posts a video with tags mentioning city A, then user C's post will be classified as a text data type, user D's image will be classified as an image data type, and user E's video will be classified as a video data type, thus obtaining second user data for different users under different data types.
[0084] S103-3, Extract the common second user data from multiple second user data as the common data.
[0085] The data types of the same second user data can be different or the same. The same second user data includes the same user characteristics and / or the same behavioral characteristics of the second user. User characteristics include basic information such as the user's age, height, and appearance, while behavioral characteristics include geographical location information and active time periods of the user.
[0086] Continuing with the example above, referring to Figure 3, we can combine three different types of second user data—text posted by user C, images posted by user D, and videos posted by user E—into a three-data set. This set collectively represents that all three different users have been to City A. The common data among the three users is City A, which is the same user data among the three users.
[0087] It is understandable that, in addition to geographic location information, other common data among multiple second users can also be considered, such as interaction intensity (e.g., number of likes, comments, activity times), communication characteristics (call content, call duration, call time, communication platform), social influence, and device type. This disclosure does not impose any restrictions on these. For example, the common behavior of multiple users liking a post can be considered common data, as can the time period during which multiple users are active together, or the same device type used by multiple users.
[0088] In related technologies, common data of the same type is obtained from different second users. However, this can lead to the omission of common data across other data types, making it impossible to accurately analyze the strength of user associations between different second users. For example, comparing only the text or images of three users may miss key clues that different types of data point to the same behavior. For instance, user C mentions city A in text (city A can be a character mentioned in the text), user D takes a picture of a landmark in city A (the landmark in city A can be the pixel coordinates in the image), and user E uses a video with the city A tag (the city A tag can be the latitude and longitude of city A). If compared by the same type, the three have no direct commonality and would be judged as "unrelated". However, after cross-type extraction, the commonality of "city A" is clearly identified, directly providing a core anchor for association analysis.
[0089] Through the above technical solutions, on the one hand, second user data across different data types such as text, images, videos, and signals can provide richer second user data. The complementarity between multimodal second user data can reduce the omission of common data and uncover potential relationships between users. On the other hand, the common data extracted from second user data of different data types is not common data under a single data type, but common data under multimodal data. This can prove from multiple dimensions that the common data between second users is real data, rather than common data from accidental collisions, making the obtained common data more accurate. Therefore, the user association strength obtained based on the common data will also be more accurate.
[0090] Figure 4 is an exemplary embodiment involved in this disclosure, which is used to illustrate an exemplary scheme for configuring corresponding weights for different common data based on different business scenarios, including the following steps:
[0091] S105, Obtain the business scenario.
[0092] Different business scenarios are used to represent business scenarios in different industries. For example, business scenarios include community discovery, financial scenarios, and traffic anomaly trajectory detection scenarios. Different business scenarios are associated with different common data, which includes first common data and second common data. The first common data is common data associated with the business scenario, while the second common data is common data not associated with the business scenario. The second weight corresponding to the common data includes the third weight corresponding to the first common data and the fourth weight corresponding to the second common data.
[0093] For example, taking community discovery as a business scenario, which is a specific scenario within the field of network analysis, the focus is more on users' likes, comments, topic follows, and active time periods, rather than on users' movement patterns, device types, and geographical locations. Therefore, the first common data associated with community discovery includes likes, comments, topic follows, and active time periods shared by multiple users, and this first common data has a higher third weight. On the other hand, the second common data not associated with community discovery includes movement patterns, device types, and geographical locations shared by multiple users, and this second common data has a lower fourth weight.
[0094] For example, taking a financial scenario as an example, the financial scenario pays more attention to users' transaction records, login devices, account associations, etc. Therefore, the first common data associated with the financial scenario includes transaction records, login devices, and associated accounts that are common to multiple users, and the third weight of this part of the first common data will be greater; while the second common data that is not associated with the financial scenario includes likes, comments, topic follows, and active time periods that are common to multiple users, and the fourth weight of this part of the second common data will be smaller.
[0095] S106, configure the third weight for the first common data associated with the business scenario, and configure the fourth weight for the second common data not associated with the business scenario.
[0096] It is understandable that multiple common data points with commonalities can be extracted from the second user data of multiple users. However, based on the different business scenarios that are of interest, the first common data points that are related to the business scenario will be selected from the extracted common data points and assigned a higher third weight, while the second common data points that are not related to the business scenario will be selected and assigned a lower fourth weight. This allows the final user association strength to adapt to changes in the business scenario.
[0097] Optionally, the user association strength can be obtained based on the first common data, the third weight corresponding to the first common data, the second common data, the fourth weight corresponding to the second common data, the number of users of the first user, and the first weight corresponding to the number of users.
[0098] The above technical solution allows for the configuration of different weights for common data among second users in different business scenarios. This weights, based on the common data weights across different business scenarios, yield the user association strength appropriate for that specific scenario. Firstly, this ensures the calculated user association strength more closely matches the actual business scenario, achieving adaptive and precise association analysis. Secondly, it eliminates the need to develop a separate association analysis model for each business scenario—for example, a separate model for community discovery or financial scenarios. A single solution can be reused, allowing for the accurate determination of user association strength within a given business scenario by adjusting the weights of common data associated with each scenario. This weight adjustment approach replaces model redevelopment, saving development costs. Thirdly, this method of reducing multi-dimensional common data to common data associated with business scenarios reduces computational load. The overall solution focuses solely on the common data associated with each business scenario, rather than individual data points, significantly reducing computational complexity.
[0099] Figure 5 is an exemplary embodiment of this disclosure, which provides a further scheme for interpreting the user association strength, including the following steps:
[0100] S107, retrieve multiple events involving the second user.
[0101] Here, the multiple events in which the second user participates are obtained with the second user as the central node, which means obtaining all the events the second user has participated in or the events that the user has been watching over in the historical period; while the users who participate in the target event are obtained with the target event as the central node.
[0102] For example, it is possible to obtain multiple events in which user C participated, or multiple events in which user D participated.
[0103] S108 performs correlation analysis on every two events among multiple events to obtain the correlation strength of multiple events.
[0104] The association strength of each event is used to characterize the degree of association between any two events, where each pair of events involves two different second users. For example, multiple events involving user C can be paired with multiple events involving user D to form multiple sets of events, each set including events from both user C and user D.
[0105] Optionally, the event association strength can be obtained based on the first quantity and the second quantity.
[0106] The first quantity is the number of times the same second user appears in every two events, and the second quantity is the number of times a different second user appears in every two events. For example, taking events e1 and e2 as an example, the formula for calculating the event association strength based on the first and second quantities is as follows:
[0107] (4)
[0108] In formula (4), It represents the strength of the event association between event e1 and event e2 among multiple events; This is the number of common nodes between events e1 and e2. For example, if the second user in event e1 is e1={A,B,C}, and the user set in event e2 is e2={A,B,D}, then the users who appear in both events e1 and e2 are user A and user B, and the number of common nodes is 2. The number of nodes that are not repeated in events e1 and e2 is 2. The different second users that appear in events e1 and e2 are user C and user D, respectively.
[0109] It is understandable that the association strength between any two events in a set of events can be calculated using formula (4). First, identify the events in which the second user participates without direct interaction. Then, perform association analysis on each pair of events to obtain the association strength between each pair of events, and thus obtain multiple sets of association strengths.
[0110] S109, based on the correlation strength of multiple events, the community correlation strength is obtained.
[0111] The community association strength is used to characterize the probability that different second users belong to the same community. The sum of the association strengths of multiple events can be used as the community association strength. When the community association strength is greater than a preset strength, the two different second users are considered to belong to the same community. The calculation formula is as follows:
[0112] (5)
[0113] In formula (5), It refers to the strength of the community association between second users who do not have direct interaction behavior, such as the strength of the community association between user C and user D; It is the set of events where user C is located; It is any event in the event set where user C is located; It is the set of events where user D is located; It is any event in the event set where user D is located; It represents the strength of the event association between events in which user C participates and events in which user D participates.
[0114] As can be seen from formula (5), the sum of the event association strengths between any events in which two different second users participate can be used as the community association strength between the two second users, thereby determining the possibility that two different users belong to the same community.
[0115] S110, based on the user association strength and community association strength, obtain the target association strength between the second users.
[0116] Optionally, the first association strength between the second users can be obtained based on the user association strength, the fifth weight corresponding to the user association strength, the community association strength, and the sixth weight corresponding to the community association strength; the second association strength can be obtained based on the user association strength and the community association strength; and the target association strength can be obtained based on the first association strength and the second association strength. For example, the calculation formula is as follows:
[0117] R_final = α×R_GELU +β×T(C,D)+γ×(R_GELU×T(C,D))(6)
[0118] In formula (6), R_final is the final target association strength; R_GELU is the user association strength, α is the fifth weight corresponding to the user association strength; T(C,D) is the community association strength, β is the sixth weight corresponding to the community association strength; R_GELU×T(C,D) is the second association strength obtained based on the user association strength and the community association strength, and γ is the seventh weight corresponding to the second association strength.
[0119] As can be seen from the above formula (6), user association strength and community association strength have their own weights. If more attention is paid to the micro-level overlap of behaviors between the second users, such as browsing the same content together, using similar login devices, or having similar active time periods, the fifth weight of user association strength can be configured higher. If more attention is paid to the macro-level overlap of behaviors between the second users, such as whether the two second users are in the same community, whether they have common interests, or whether they follow the same celebrity, the sixth weight of community association strength can be set higher. If attention is paid to the synergistic effect of user association strength and community association strength, the seventh weight corresponding to the second association strength can be set higher.
[0120] Understandably, in a scheme that only considers the micro-level user association strength as the target association strength, there may be coincidences in the micro-level user association strength. For example, two second users may both like to shop at 8 a.m., but they do not know each other. If only the micro-level user association strength is considered, the coincidental user association strength may be misjudged as the real user association strength, thus causing unrelated second users to be misjudged as related second users.
[0121] In a scheme that only considers the macro-level community association strength as the target association strength, the macro-level community association strength may be passively bound. For example, two second users may both be in a large group of 500 people, but they never interact or share any common behaviors. This will also deeply bind ordinary users in the same community, causing unrelated second users to be mistakenly identified as related second users.
[0122] In a scheme that only considers the first association strength after the fusion of user association strength and community association strength as the target association strength, the fusion result after the user association strength and community association strength mutually corroborate each other is ignored. For example, the credibility of two second users having common data and being in the same community is ignored, which is higher than the credibility of having only common data or being in the same community. The dual strength after the combination of user association strength and community association strength is missed.
[0123] Through the above technical solutions, the first aspect of obtaining the target association strength can avoid one-sidedness. It covers both micro-level user association strength and macro-level community association strength, solving the coincidence phenomenon brought about by micro-level user association strength and the passive binding phenomenon brought about by macro-level community association strength, making the obtained target association strength more comprehensive. The second aspect of obtaining the target association strength has higher credibility. It uses the second association strength to synergistically assist user association strength and community association strength. It can make the credibility of two second users having common data and belonging to the same community higher than the credibility of single data. This leads to higher user strength of real association and lower user strength of false association, thus increasing the credibility of the obtained target association strength.
[0124] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0125] Figure 6 shows a structural block diagram of the data processing apparatus provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0126] Referring to Figure 6, the data processing device 600 includes: a user determination module 610, a correlation analysis module 620, an acquisition module 630, and a correlation strength determination module 640.
[0127] User identification module 610 is used to identify multiple users who participate in multiple target events;
[0128] The correlation analysis module 620 is used to perform correlation analysis on different target events and identify a first user and a second user from multiple users; the first user is a user who participates in multiple target events simultaneously, and the second user is a user who participates in some of the target events; the second users are associated with each other through the first user;
[0129] The acquisition module 630 is used to acquire common data among the second users; the common data is the same user data generated when the second users do not interact.
[0130] The association strength determination module 640 is used to determine the user association strength between the second users based on the common data and the first users associated with the second users.
[0131] Optionally, the acquisition module 630 is further configured to acquire first user data of different second users; classify the first user data according to data type to obtain second user data corresponding to different data types; extract the same second user data from multiple second user data as the common data; the same second user data includes the same user characteristics and / or the same behavioral characteristics of the second users.
[0132] Optionally, the association strength determination module 640 is further configured to obtain the user association strength based on the number of users of the first user associated with the second user, the first weight corresponding to the number of users, the common data, and the second weight corresponding to the common data.
[0133] Optionally, the common data includes first common data and second common data. The first common data is common data associated with the business scenario, and the second common data is common data not associated with the business scenario. The second weight corresponding to the common data includes a third weight corresponding to the first common data and a fourth weight corresponding to the second common data. The data processing device 600 includes:
[0134] The scenario acquisition module is used to acquire the business scenario; different business scenarios are associated with different common data.
[0135] The configuration module is used to configure the third weight for the first common data associated with the business scenario, and to configure the fourth weight for the second common data not associated with the business scenario; the third weight is greater than the fourth weight.
[0136] Optionally, the data processing apparatus 600 includes:
[0137] The event acquisition module is used to acquire multiple events in which the second user participates;
[0138] The analysis module is used to perform correlation analysis on every two events among the multiple events to obtain the correlation strength of multiple events; the correlation strength of each event is used to characterize the closeness of the correlation between every two events;
[0139] The community module is used to obtain the community association strength based on the association strength of multiple events; the community association strength is used to characterize the probability that different second users are in the same community;
[0140] The calculation module is used to obtain the target association strength between the second user based on the user association strength and the community association strength.
[0141] Optionally, the analysis module is further configured to obtain the event association strength based on a first quantity and a second quantity; the first quantity is the number of identical second users appearing in every two events, and the second quantity is the number of different second users in every two events.
[0142] Optionally, the calculation module is further configured to obtain a first association strength between the second users based on the user association strength, the fifth weight corresponding to the user association strength, the community association strength, and the sixth weight corresponding to the community association strength; obtain a second association strength based on the user association strength and the community association strength; and obtain the target association strength based on the first association strength and the second association strength.
[0143] Optionally, identical second user data among the second users is used to characterize that the similarity between the second user data of the second users is greater than a preset similarity.
[0144] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0145] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0146] This application also provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0147] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0148] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.
[0149] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0150] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0151] Computer program code for performing the operations of the embodiments of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages such as Python, Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0152] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0153] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0154] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A data processing method, characterized in that, include: Identify multiple users who participate in multiple target events; A correlation analysis is performed on the different target events to identify the first user and the second user from among the multiple users; The first user is a user who participates in multiple target events simultaneously, and the second user is a user who participates in some of the target events among the multiple target events; The second users are associated with each other through the first user; Obtain common data among the second users; The common data refers to the same user data generated when the second users do not interact with each other. Based on the common data and the first users associated with the second users, the user association strength between the second users is determined; the method further includes: acquiring multiple events in which the second users participate; performing correlation analysis on every two events to obtain multiple event association strengths; each event association strength is used to characterize the closeness of association between every two events, where every two events are events in which two different second users participate; obtaining community association strength based on the multiple event association strengths; the community association strength is used to characterize the probability that different second users are in the same community; obtaining the target association strength between the second users based on the user association strength and the community association strength; the correlation analysis is performed on every two events in the multiple events. The analysis yields multiple event association strengths, including: obtaining the event association strength based on a first quantity and a second quantity; the first quantity is the number of identical second users appearing in every two events, and the second quantity is the number of different second users in every two events; obtaining the target association strength between the second users based on the user association strength and the community association strength includes: obtaining a first association strength between the second users based on the user association strength, a fifth weight corresponding to the user association strength, and a sixth weight corresponding to the community association strength; obtaining a second association strength based on the product of the user association strength and the community association strength; and obtaining the target association strength based on the first association strength and the second association strength.
2. The method as described in claim 1, characterized in that, The step of obtaining common data among the second users includes: obtaining first user data of different second users; classifying the first user data according to data type to obtain second user data corresponding to different data types; extracting the same second user data from multiple second user data as the common data; the same second user data includes the same user characteristics and / or the same behavioral characteristics of the second users.
3. The method as described in claim 1, characterized in that, The step of determining the user association strength between the second users based on the common data and the first users associated with the second users includes: obtaining the user association strength based on the number of first users associated with the second users, a first weight corresponding to the number of users, the common data, and a second weight corresponding to the common data.
4. The method as described in claim 3, characterized in that, The common data includes first common data and second common data. The first common data is common data associated with the business scenario, and the second common data is common data not associated with the business scenario. The second weight corresponding to the common data includes a third weight corresponding to the first common data and a fourth weight corresponding to the second common data. The method further includes: obtaining the business scenario; different business scenarios are associated with different common data; configuring the third weight for the first common data associated with the business scenario, and configuring the fourth weight for the second common data not associated with the business scenario; the third weight is greater than the fourth weight.
5. The method as described in claim 2, characterized in that, The second user data that is the same among the second users is used to characterize the similarity between the second user data of the second users being greater than a preset similarity.
6. A data processing apparatus, characterized in that, include: The user identification module is used to identify multiple users who participate in multiple target events; The correlation analysis module is used to perform correlation analysis on different target events and identify the first user and the second user from multiple users; The first user is a user who participates in multiple target events simultaneously, and the second user is a user who participates in some of the target events among the multiple target events; The second users are associated with each other through the first user; The acquisition module is used to acquire common data among the second users; The common data refers to the same user data generated when the second users do not interact with each other. The association strength determination module is used to determine the user association strength between the second users based on the common data and the first users associated with the second users; The event acquisition module is used to acquire multiple events in which the second user participates; The analysis module is used to perform correlation analysis on every two events among the multiple events to obtain the correlation strength of multiple events; the correlation strength of each event is used to characterize the closeness of the correlation between every two events, and every two events are events in which two different second users participate; The community module is used to obtain the community association strength based on the association strength of multiple events; the community association strength is used to characterize the probability that different second users are in the same community; The calculation module is used to obtain the target association strength between the second user based on the user association strength and the community association strength; The analysis module is further configured to obtain the event association strength based on a first quantity and a second quantity; the first quantity is the number of identical second users appearing in every two events, and the second quantity is the number of different second users in every two events; the calculation module is further configured to obtain a first association strength between the second users based on the user association strength, a fifth weight corresponding to the user association strength, a community association strength, and a sixth weight corresponding to the community association strength; obtain a second association strength based on the product of the user association strength and the community association strength; and obtain the target association strength based on the first association strength and the second association strength.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
User recommendation method, device and equipment and computer readable storage medium
CN114401242A
Method and system for generating intimacy between users and social network display method
CN117726473A