Vehicle-related user group analysis method and device, electronic equipment and storage medium
By combining user profiles and spatiotemporal activity data, related user sets are segmented and clustered, solving the problem of insufficient accuracy in traditional data mining schemes. This achieves high-accuracy and high-correlation-value user group identification, which is suitable for real-time analysis of intelligent vehicle systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional data mining solutions lack accuracy and offline correlation value when identifying user groups, making it difficult to guide subsequent business operations.
By acquiring user profile data and spatiotemporal activity data, we can divide related user sets based on user correlation and use spatiotemporal activity data for clustering to identify similar user groups.
It improves the accuracy of user group segmentation and offline correlation value, is suitable for parallel computing, and meets real-time requirements.
Smart Images

Figure CN122489846A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data mining technology, and in particular to methods, devices, electronic equipment and storage media for analyzing vehicle-related user groups. Background Technology
[0002] Traditional data mining solutions typically focus on analyzing users' social interactions, likes, or text tags on the internet. While these solutions can identify similar users, their accuracy needs improvement, and the identified similar user groups often lack offline connections, making it difficult to guide subsequent business operations. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this specification provides methods, devices, electronic devices and storage media for analyzing vehicle-related user groups.
[0004] According to a first aspect of the embodiments of this specification, a method for analyzing vehicle-related user groups is provided, the method comprising: For at least two users to be segmented, user profile data and spatiotemporal activity data generated by the user's vehicle usage are obtained for each user. The user profile data is determined based on the user's personal attributes and / or social relationships.
[0005] Based on the user profile data of each user, the user correlation degree corresponding to each basic user pair consisting of the at least two users is determined. Based on the user correlation degree, related user pairs are determined from the basic user pairs, wherein each basic user pair contains two different users among the at least two users. Based on each related user pair, the at least two users are divided into at least one related user set, wherein the related user pairs to which the same user belongs are divided into the same related user set.
[0006] For each set of associated users, clustering is performed on each user in the set based on the spatiotemporal activity data of each user in the set to obtain at least one similar user group.
[0007] According to a second aspect of the embodiments of this specification, a vehicle-related user group analysis apparatus is provided, the apparatus comprising: The data acquisition module is used to acquire user profile data and spatiotemporal activity data generated by the user's use of the vehicle for at least two users to be segmented. The user profile data is determined based on the user's personal attributes and / or social relationships.
[0008] The set generation module is used to determine the user correlation degree corresponding to each basic user pair consisting of at least two users based on the user profile data of each user, determine the associated user pairs from the basic user pairs based on the user correlation degree, wherein each basic user pair contains two different users among the at least two users; and, based on each associated user pair, divide the at least two users into at least one associated user set, wherein each associated user pair to which the same user belongs is divided into the same associated user set.
[0009] The group segmentation module is used to perform clustering processing on each user in each associated user set based on the spatiotemporal activity data of each user in the set, so as to obtain at least one similar user group.
[0010] According to a third aspect of the embodiments of this specification, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described in the first aspect.
[0011] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method as described in the first aspect.
[0012] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: The user group analysis process related to vehicles described in this specification includes a first stage of dividing related user sets based on user profiles and a second stage of dividing similar user groups based on spatiotemporal activity data. In the first stage, the user correlation degree corresponding to each basic user pair consisting of at least two users to be divided is calculated using user profile data. Related user pairs are determined from the basic user pairs based on the user correlation degree, and then multiple users are divided into at least one related user set based on each related user pair. The calculation of user correlation degree based on user profile data and the preliminary division of related user sets ensures that users within the same related user set have highly similar user profiles. Further, in the second stage, users in each related user set are clustered using spatiotemporal activity data generated when users use the vehicle, thereby further dividing the related user sets into at least one similar user group.
[0013] As can be seen, this solution not only identifies users based on user profile data, thus initially segmenting users with different user profiles, but more importantly, it also introduces spatiotemporal activity data highly correlated with vehicle usage to identify the activity characteristics of each user in time, space, and other dimensions. Based on this, users in each associated user set are further segmented into different similar user groups. Clearly, users in the different similar user groups obtained through this method not only have different personal attributes and / or social relationships, but also different spatiotemporal activity characteristics. Therefore, the segmentation results are not only highly accurate but also possess significant offline correlation value, which can be used to guide subsequent business operations.
[0014] In addition, in the second stage, the clustering processes for different sets of related users are independent of each other. Therefore, this solution is naturally suitable for parallel / distributed computing logic. This stage only requires a very short processing time, which can meet the real-time requirements of the vehicle terminal for discovering similar user groups.
[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.
[0017] Figure 1 This is a flowchart illustrating a vehicle-related user group analysis method according to an exemplary embodiment of this specification.
[0018] Figure 2 This is a schematic diagram illustrating a vehicle-related user group analysis process according to an exemplary embodiment of this specification.
[0019] Figure 3 This is a schematic diagram illustrating an active location clustering based on spatiotemporal activity data according to an exemplary embodiment of this specification.
[0020] Figure 4 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of this specification.
[0021] Figure 5 This is a block diagram illustrating a vehicle-related user group analysis device according to an exemplary embodiment of this specification. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0023] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0025] With the widespread adoption of intelligent connected vehicles, in-vehicle systems have become the second largest mobile intelligent terminal after smartphones. During daily driving and vehicle use, users generate massive amounts of high-frequency spatiotemporal activity data through onboard sensors, GPS positioning modules, and vehicle-to-everything (V2X) platforms. This data not only includes precise geographic coordinates but also implicitly contains in-depth features such as time distribution, driving trajectories, dwell time, and travel frequency, recording the user's complete lifestyle from the physical world to the digital space.
[0026] Grouping vehicle-related users with similar behavioral patterns or social attributes into the same group has broad application value in the automotive aftermarket, social media promotion, and urban services. For example, for identified car enthusiast groups with similar interests and overlapping trajectories, "commuter carpooling groups," "weekend camping circles," or "brand enthusiast clubs" can be automatically created to increase user stickiness to the in-vehicle infotainment system. For example, for identified groups with similar charging / refueling habits in terms of time and space, targeted push notifications of charging station availability, refueling discounts, or route guidance can be provided. For example, identifying travel groups with common time and space activities can also be used to optimize urban parking space allocation and ride-sharing matching.
[0027] The present solution will now be described in detail with reference to the accompanying drawings and corresponding embodiments.
[0028] Figure 1 This is a flowchart illustrating a method for analyzing vehicle-related user groups according to an exemplary embodiment. The method is applied to a data mining server. Figure 1 As shown, the method includes steps 101-103: Step 101: For at least two users to be segmented, obtain user profile data and spatiotemporal activity data generated by the user's vehicle usage for each user. The user profile data is determined based on the user's personal attributes and / or social relationships.
[0029] Step 102: Determine the user correlation degree corresponding to each basic user pair consisting of the at least two users based on the user profile data of each user; determine the associated user pairs from the basic user pairs based on the user correlation degree, wherein each basic user pair contains two different users among the at least two users; and, classify the at least two users into at least one associated user set based on each associated user pair, wherein each associated user pair to which the same user belongs is classified into the same associated user set.
[0030] Step 103: For each set of associated users, cluster each user in the set based on the spatiotemporal activity data of each user in the set to obtain at least one similar user group.
[0031] Each of the at least two users to be segmented can refer to an account used to log in to the vehicle's infotainment system. The account on the vehicle's infotainment system can be linked to the vehicle. The spatiotemporal activity data generated by the user using the vehicle can include the vehicle's location information during its journey, collected in real-time by the vehicle's positioning system while the user is logged into the infotainment system, and the corresponding timestamp for that location information.
[0032] The acquired location information can be the coordinates of the vehicle's parking location, which can be represented by longitude and latitude. The vehicle parking location may include, for example, the location where the vehicle was powered off / parked, or the location where the vehicle remained for more than a preset time (e.g., 30 minutes). Of course, each of the above locations can refer to a specific spatial point (e.g., a parking space) or a spatial area (e.g., a parking lot, a park, a city, etc.).
[0033] Because vehicles generate a large amount of GPS (Global Positioning System) coordinate information during operation, statistical interference caused by unintentional stops such as traffic lights and congestion is filtered out using time thresholds to retain the locations representing the user's real-life goals. This not only effectively eliminates geographical noise but also compresses / refines the user's spatiotemporal activity data, reducing the computational load of clustering based on spatiotemporal activity data.
[0034] From the perspective of data sources, the user profile data of any user described in this specification can be generated by the user during the use of at least one third-party APP (Application). This data can be collected by the client or server of the APP and then sent to the aforementioned data mining server. The client of any APP can be deployed and run on the electronic device used by the user (such as a mobile phone, computer, wearable device, etc.); or it can be deployed and run in the vehicle used by the user, such as being deployed locally in the vehicle and managed by the vehicle's infotainment system. The user can interact with the client through the in-vehicle screen and / or voice system to achieve preset application functions.
[0035] Similarly, the spatiotemporal activity data of any user described in this specification is generated by the vehicle used by that user. This data can be collected by the vehicle and uploaded to the vehicle management service or directly uploaded to the aforementioned data mining service, which will not be elaborated further. Specifically, this data can be collected by a third-party APP managed by the vehicle's in-vehicle system (i.e., the vehicle-side APP), or it can be collected by the in-vehicle system itself. This specification does not limit the embodiments in this way.
[0036] For example, any of the aforementioned third-party apps can have functions such as navigation, music, social networking, and / or shopping. For instance, data such as a user's posts, interactions, and friend lists in a car infotainment app, or friend lists and following lists generated when using a social networking app, can serve as the user's user profile data. Similarly, data such as saved locations and generated navigation records while using a navigation app can serve as the user's spatiotemporal activity data, which will not be elaborated further.
[0037] In addition, user profile data can also be generated from the vehicles used by users, such as vehicle operation data, driving behavior characteristics, and in-vehicle environment preferences generated by the user's driving. For example, historical average vehicle speed, preferred driving modes, in-vehicle environment settings preferences, and vehicle maintenance records, etc.
[0038] User profile data can be determined based on a user's personal attributes and / or social relationships. From a feature classification perspective, for example, user profile data includes personal attribute data determined based on a user's personal attributes; and / or, social relationship data determined based on a user's social relationships.
[0039] Personal attributes represent the characteristics of a user as an independent individual, reflecting their personal traits and daily preferences. For example, personal attribute data may include age, gender, occupation, and interests.
[0040] Social relationships are used to characterize the social connections between users. For example, social relationship data may include the friendship between at least two users to be segmented, records of jointly participated events, and the frequency of interactions (such as likes, comments, and shares).
[0041] A basic user pair is a pair of users that make up every two users in the set of at least two users to be partitioned. For example, if the set of at least two users to be partitioned is {A, B, C, D, E}, then the basic user pairs are AB, AC, AD, AE, BC, BD, BE, CD, CE, and DE. Each basic user pair contains two distinct users from the set of at least two users to be partitioned.
[0042] User relevance is a quantitative indicator used in this scheme to measure the similarity of personal attributes and / or the closeness of the relationship between two users.
[0043] If the user correlation between any two users is not lower than the correlation threshold, it indicates that there is a "correlation" between the two users. This "correlation" can be reflected in the similarity of the two users' lifestyles and / or social circles. If the user correlation between any two users is lower than the correlation threshold, it indicates that the two users are completely unrelated, meaning that they will not have any interaction in their lives, for example, their lifestyles and social circles are different.
[0044] Therefore, when identifying related user pairs from basic user pairs based on user relevance, basic user pairs with a corresponding user relevance of not less than the relevance threshold can be identified as related user pairs. The higher the relevance threshold, the stronger the "relevance" between the two users in the selected related user pairs.
[0045] Of course, some users may have different personal attributes but close social relationships. Their user correlation may still be higher than the correlation threshold, and they can be regarded as related user pairs that may have a "correlation relationship".
[0046] Although some users may not have social relationships, their personal attributes are highly consistent. Their user correlation may still be higher than the correlation threshold, and they can be regarded as related user pairs that may have a "correlation relationship".
[0047] When determining the degree of user association between two users based on their user profile data, the association can be determined solely based on the users' individual attribute data, solely based on their individual social relationship data, or simultaneously based on both their individual attribute data and social relationship data.
[0048] like Figure 2 As shown, a user to be partitioned can be assigned to at least one set of associated users. In the partitioning result, if all users to be partitioned have direct or indirect relationships with each other, the user to be partitioned may be assigned to a unique set of associated users, meaning the partitioning result has only one set of associated users. Of course, the partitioning result may also have multiple sets of associated users, where there is no direct or indirect relationship between users in any two sets. In other words, all associated user pairs belonging to the same user are assigned to the same set of associated users.
[0049] In an exemplary implementation of partitioning users into at least one set of associated users based on each associated user pair, an empty set of associated users can be initialized, and a user can be randomly selected from the users to be partitioned. Other users who form an associated user pair with the randomly selected user can be added to this set of associated users. For each user in the current set of associated users, iterate through the set to find other users who form an associated user pair with that user. If other users who could form an associated user pair are not in this set, then add that user to the set of associated users. Users already iterated through in this set are no longer iterated through; continue searching for other users who have not been iterated through, and then add them to the set of associated users until no new associated user pair can be found. After excluding users from the first set of associated user pairs from the users to be partitioned, continue executing the above logic to determine a new set of associated users.
[0050] For example, such as Figure 2 As shown, the user to be partitioned is assigned to one of the associated user sets 201, in which each of the associated user pairs belonging to any user in the associated user set 201 is partitioned to the associated user set 201.
[0051] Users within the same set of related users can have direct or indirect relationships. For example, although users A and C do not form a pair of related users (i.e., there is no direct relationship between them), users A and E are related, and users E and C are also related. Therefore, users A and C have an indirect relationship.
[0052] As for users G and H among the users to be divided, since they have no association with any user in the associated user set 201, they were not divided into the associated user set 201.
[0053] User association can be used to group users with similar personal attributes and / or similar social relationships into a set of related users.
[0054] However, while users within the same associated user set may have similar profiles, their spatiotemporal activity patterns may not be identical. Therefore, relying solely on the similarity of user profiles while ignoring the similarity of spatiotemporal activity patterns may result in invalid user grouping results, making them unsuitable for various subsequent vehicle-related application scenarios. For example, in a car enthusiast recommendation scenario, the system might recommend a user in Beijing to a user in Guangzhou with similar interests.
[0055] Therefore, based on this, for each set of related users, clustering is performed on each user in the set based on their spatiotemporal activity data to obtain at least one similar user group. For example, clustering related user set 201 can yield similar user group 202 and similar user group 203.
[0056] In this embodiment, the spatiotemporal activity data generated by the vehicle system is high-frequency and dense, and the computational power requirement for direct clustering would explode exponentially. However, this solution has already initially divided a large number of users into multiple smaller sets of related users through user profile data. When performing spatiotemporal clustering operations on each set of related users, it can significantly alleviate the computational complexity of spatiotemporal clustering operations and reduce the computational requirements of the hardware.
[0057] The user group analysis process related to vehicles described in this embodiment includes a first stage of dividing related user sets based on user profiles and a second stage of dividing similar user groups based on spatiotemporal activity data. In the first stage, the user correlation degree corresponding to each basic user pair consisting of at least two users to be divided is calculated using user profile data. Related user pairs are determined from the basic user pairs based on the user correlation degree, and then multiple users are divided into at least one related user set based on each related user pair. The calculation of user correlation degree based on user profile data and the preliminary division of related user sets ensures that users within the same related user set have highly similar user profiles. Further, in the second stage, users in each related user set are clustered using spatiotemporal activity data generated when users use the vehicle, thereby further dividing the related user sets into at least one similar user group.
[0058] As can be seen, this solution not only identifies users based on user profile data, thus initially segmenting users with different user profiles, but more importantly, it also introduces spatiotemporal activity data highly correlated with vehicle usage to identify the activity characteristics of each user in time, space, and other dimensions. Based on this, users in each associated user set are further segmented into different similar user groups. Clearly, users in the different similar user groups obtained through this method not only have different personal attributes and / or social relationships, but also different spatiotemporal activity characteristics. Therefore, the segmentation results are not only highly accurate but also possess significant offline correlation value, which can be used to guide subsequent business operations.
[0059] In addition, in the second stage, the clustering processes for different sets of related users are independent of each other. Therefore, this solution is naturally suitable for parallel / distributed computing logic. This stage only requires a very short processing time, which can meet the real-time requirements of the vehicle terminal for discovering similar user groups.
[0060] An exemplary implementation of determining the user association degree of each basic user pair consisting of at least two users to be segmented based on user profile data of each user: For example, a deep learning model can be trained by inputting user profile data of each user into the deep learning model, and the output is the user association degree of each basic user.
[0061] For example, user profile data for each user can be converted into embedded vectors, and then the correlation between each base user and its corresponding user can be calculated using vector similarity.
[0062] For example, a global association graph corresponding to at least two users to be divided can be constructed based on the user profile data of each user, and the user association degree of each basic user can be calculated based on the topological relationship between each user represented by the global association graph.
[0063] When constructing a global relationship graph based on user profile data, to represent the topological relationships between users in the global relationship graph, vertices can be identified, with each vertex representing a user. Based on the user's personal attribute data, the attribute characteristics of that user's vertices are determined. Based on the user's social relationship data, the structural characteristics of the edges between that user's vertex and other vertices indicated by the social relationship data are determined. The weights of the edges can be determined by the strength of the association, for example, based on data such as the frequency of interaction between two users, the number of shared events, and the level of friendship between them.
[0064] In this embodiment, by constructing a global association graph, the spatial topology information of users' social relationships can be effectively preserved, thereby more accurately calculating the user association degree between two users.
[0065] An exemplary implementation of calculating the association degree between each basic user and its corresponding user, based on the topological relationships between users represented by a global association graph: For example, the following is executed for each basic user pair: The similarity between the vertices of two users in a basic user pair in the global association graph can be calculated based on a random walk algorithm, which can be used to characterize the user association between the two users in the basic user pair.
[0066] For example, graph neural networks, such as Node2Vec (Node to Vector) or GraphSAGE (Graph Sample and Aggregate), can be used to mine the similarity between the vertices of two users in a basic user pair in the global association graph, which can be used to characterize the user association degree of the basic user pair.
[0067] For example, attribute similarity can be calculated based on the individual attribute data of the first and second users included in the base user pair.
[0068] The structural similarity between the first and second users is determined based on the connections between the vertices corresponding to the first and second users in the global association graph and other vertices. Each vertex in the global association graph identifies a user.
[0069] The degree of user association between the first user and the second user is determined based on attribute similarity and structural similarity.
[0070] Attribute similarity can reflect the similarity of users in terms of personal identity, behavioral preferences, and value orientations. For example, user personal attribute data can be converted into embedded vectors, and attribute similarity can be calculated using vector similarity calculation methods.
[0071] Structural similarity can reflect whether two users share similar social circles. For example, for any given user, structural features such as social influence and the number of mutual friends can be determined based on the connections between that user's vertices and those of other users. If a user's vertices have numerous connections to other users' vertices, it indicates that the user has many friends and significant social influence. Similarly, if a user shares many vertices with another user, it reflects the similarity of their social circles. Of course, other structural features can also be used to determine the structural similarity between two users.
[0072] When determining the user association degree between the first user and the second user based on attribute similarity and structural similarity, the results of attribute similarity calculation and structural similarity calculation can be weighted and summed to obtain the user association degree of the basic user pair including the first user and the second user.
[0073] In this embodiment, by comprehensively determining user association based on attribute similarity and structural similarity, deeply associated user pairs can be discovered, a higher quality set of associated users can be constructed, and the number of users in the set of associated users that need to be processed in subsequent clustering steps can be further reduced, thereby alleviating the computational complexity of the clustering algorithm and reducing the computational requirements of the hardware.
[0074] An exemplary implementation of determining the structural similarity between the first user and the second user based on the connection relationships between the vertices corresponding to the first user and the second user in the global association graph and other vertices: The structural index values of the vertices corresponding to the first user and the second user can be determined separately, and the structural similarity between the first user and the second user can be calculated based on the structural index values of the two vertices.
[0075] Structural metrics may include degree centrality metrics, clustering coefficient metrics, and / or neighbor similarity metrics.
[0076] Degree centrality is used to calculate the number of connections a vertex has to other vertices. For example, if a vertex is connected to 5 other vertices, its degree centrality value is 5. If a vertex is not connected to any other vertex, its degree centrality value is 0. In a global graph, degree centrality can measure a user's social influence.
[0077] Compared to degree centrality, the clustering coefficient metric measures not only the number of connections between a vertex and its neighbors, but also the number of connections between those neighbors. A high clustering coefficient indicates that the user is in a closed, small circle where people know each other. Conversely, a low clustering coefficient indicates that the user's social connections are more dispersed, with acquaintances belonging to different circles.
[0078] The neighbor similarity metric is used to calculate the ratio between the set of common neighbor vertices between two vertices and the set of all neighbors of those two vertices. It can be used to measure the degree of overlap in social relationships between two users.
[0079] Of course, the structural indices used in this embodiment are not limited to the three mentioned above; other structural indices such as betweenness centrality and compact centrality can also be used. When determining the structural index value for each vertex based on the global correlation graph, one or more structural index values can be determined. This scheme does not limit the combination of multiple structural index values.
[0080] If only one structural index value is used, for any two vertices using the same structural index value, the two values can be subtracted, and the result is used to measure the structural similarity between the two vertices. If multiple structural index values are used, for any two vertices using the same structural index value, the corresponding structural index values can be subtracted, and then the results of the subtractions for each structural index can be weighted and summed to obtain the structural similarity between the two vertices.
[0081] For each associated user set, clustering is performed on each user in the set based on their spatiotemporal activity data to obtain at least one similar user group. An exemplary implementation is as follows: For each set of associated users, the spatial range and temporal range of each user in the set can be determined based on their spatiotemporal activity data.
[0082] Determine the similarity of the activity space range between any two users in the set, and determine the similarity of the activity time range between any two users in the set.
[0083] Cluster every two users whose similarity in both activity space and activity time is not less than the corresponding threshold into the same similar user group.
[0084] For example, a user can be randomly selected from the users in the set as the core user, and it can be determined whether the similarity of the activity range and the activity time range between the other users and the core user are not less than the corresponding threshold.
[0085] If so, other eligible users are added to the cluster containing the core user. If the number of users in a cluster exceeds a certain threshold, the cluster can be considered a similar user group. Otherwise, the cluster is not valid.
[0086] For the remaining users who have not joined any similar user group, a user can be randomly selected from these users as a new core user. Then, it is determined whether the similarity of the activity range and the activity time range between the other users not joined in any similar user group and the newly selected core user is not less than the corresponding thresholds. If so, these users who meet the conditions can be added to the new cluster where the new core user belongs. If the number of users in the new cluster is greater than the threshold, the new cluster can be considered a newly added similar user group. If the number of users in the new cluster is not greater than the threshold, the cluster is not considered a similar user group.
[0087] This process continues until all users have been clustered.
[0088] In this embodiment, when clustering each user in the associated user set, the similarity of the user's activity space range and activity time range is considered at the same time. By using the consistency of user spatiotemporal patterns as a constraint, the objectivity and accuracy of similar user group division are further enhanced.
[0089] In one embodiment, the map can be divided into equally sized grid regions, and the locations of all dwelling points within a user's activity space over a period of time can be assigned to the grid region closest to that dwelling point. When calculating the similarity of the activity space ranges between any two users, it can be measured by dividing the number of grids where the two users co-occur by the total number of grids covered by the two users.
[0090] The above methods can only measure the similarity of two users appearing in the same location, but cannot measure whether two users appear in the same location at the same time. For example, user A has been charging their vehicle at that location every evening for the past month, while user B has been charging their vehicle at that location every midday for the past month. Although there is spatial overlap, the time range of their activities does not overlap, and the requirement of spatiotemporal consistency is still not met.
[0091] Based on this, the average overlap of the activity time ranges of two users appearing in the same grid can be used as the similarity of their activity time ranges. For example, for any grid that appears in the same grid, if both appear in the morning, the overlap of the activity time ranges of that grid is high; if one appears in the morning and the other appears in the evening, the overlap of the activity time ranges of that grid is low.
[0092] However, the above method is based on statistics of all the locations a user appears in, and cannot accurately distinguish between a user's "routine daily routine" and "occasional occurrence." For example, user A regularly appears at the company every day, while user B only visits the company once. The spatiotemporal similarity of the locations is very high, but their actual daily routines are completely different.
[0093] In response, this application provides another improved embodiment: When determining the spatial and temporal activity ranges of each user in a given set of associated users based on their spatiotemporal activity data, the following can be executed for any user in each set of associated users: Based on any user's spatiotemporal activity data, determine at least one active location of the vehicle used by the user and the time of the vehicle's appearance at at least one active location, wherein the at least one active location is used to characterize the user's activity spatial range and the appearance time is used to characterize the user's activity temporal range.
[0094] or, Based on any user's spatiotemporal activity data, determine multiple active locations of the vehicle used by the user or the vehicle's driving trajectory between multiple active locations, and determine the time of the vehicle's appearance at multiple active locations. The multiple active locations or driving trajectory are used to characterize the user's activity spatial range, and the time of appearance is used to characterize the user's activity temporal range.
[0095] For example, when determining at least one active location of a user's vehicle based on any user's spatiotemporal activity data, such as Figure 3 As shown, the full spatiotemporal activity data of users can be clustered to obtain at least one cluster center, with each cluster center serving as an active location.
[0096] Clustering simplifies a user's massive location coordinates into a few active locations. These active locations represent anchor points in a user's life and offer a better measure of their lifestyle compared to a chaotic mass of location coordinates. For example, clustering a user's full spatiotemporal activity data into two active locations could semantically represent home and work, respectively.
[0097] When characterizing the user's activity space, it can be characterized by at least one active location, and / or by the vehicle's travel trajectory between multiple active locations.
[0098] For example, the occurrence time of an active location can be represented by the set of occurrence times of all locations within the cluster to which the active location belongs. For instance, if there are 5 locations in the cluster, 3 of which are visited in the morning and 2 of which are visited in the evening, the occurrence time of this active location can be characterized as {morning: 3; noon: 0; evening: 2}, quantized as a time feature vector of [3, 0, 2].
[0099] When determining the similarity of the activity space range between any two users in the set, the distance similarity between at least one active location corresponding to each of the two users can be determined as the similarity of the activity space range between the two users, provided that at least one active location is used to characterize the activity space range of the user.
[0100] When multiple active locations or driving trajectories are used to characterize the user's activity space range, the distance similarity between the multiple active locations corresponding to each of the two users is determined, the trajectory similarity between the driving trajectories corresponding to each of the two users is determined, and the similarity of the activity space range of each of the two users is determined based on the distance similarity and trajectory similarity.
[0101] For example, when there is only one active location, the distance similarity between the active locations of any two users can be calculated.
[0102] For example, when there are multiple active locations, multiple pairs of active locations can be identified separately, and then the distance similarity between each pair of active locations can be calculated separately. The average of all distance similarities can be used as the distance similarity between the multiple active locations corresponding to each pair of users.
[0103] Suppose user 1 has locations A, B, and C; user 2 has locations D, E, and F. If location A is closest to location D, location B is closest to location F, and location C is closest to location E, then we can calculate the distance similarity between location A and location D, the distance similarity between location B and location F, and the distance similarity between location C and location E, and then average these three distance similarities.
[0104] For example, when there are multiple active locations, the driving trajectories of two users can be determined separately, and then the trajectory similarity between the driving trajectories of these two users can be determined.
[0105] In this scheme, distance similarity can be used to identify whether two users share the same social circles, while trajectory similarity can be used to identify whether two users have similar travel routes. Both can reflect the spatial similarity of users.
[0106] Therefore, this scheme can determine the similarity of the activity space ranges of any two users based solely on distance similarity. Alternatively, it can determine the similarity of the activity space ranges of any two users based solely on trajectory similarity. Of course, it can also determine the similarity of the activity space ranges of any two users based on both distance similarity and trajectory similarity simultaneously.
[0107] When determining the similarity of the activity time range between any two users in the set, the similarity between the occurrence times of each two users can be determined and used as the similarity of the activity time range between any two users.
[0108] For example, referring to the aforementioned embodiments, after obtaining the time feature vectors corresponding to the occurrence times of the two users, the vector similarity between the two time feature vectors can be calculated as the similarity between the two occurrence times.
[0109] To further eliminate the influence of noise in the clustering results, for each of the at least one similar user groups obtained through clustering, the cluster density index and / or spatiotemporal consistency index value for that group are determined. The cluster density index characterizes the average distribution density of users in that group within the corresponding activity space, while the spatiotemporal consistency index characterizes the degree of overlap in the activity time ranges of the users in that group.
[0110] If the cluster density index value is not greater than the density threshold and / or the spatiotemporal consistency index value is not greater than the consistency threshold, then the group is removed.
[0111] For example, cluster density can be measured by the proportion of the total number of users in a group to the smallest coverage area of all users' activity space. A higher cluster density value indicates more users per unit area, a higher degree of spatial overlap in the group, and a more concentrated activity range for similar user groups. This indicator can be used to eliminate spatially sparse similar user groups.
[0112] Cluster density can also be measured by the average distance from the active locations of all users in the group to the center point of the smallest coverage area.
[0113] Spatiotemporal consistency metrics can be measured by the average similarity of the activity time ranges between each pair of users within a group. Spatiotemporal consistency metrics can measure the overall regularity of the group's occurrence.
[0114] In this embodiment, filtering by cluster density and / or spatiotemporal consistency metrics can further eliminate clustering noise, ensuring that the final segmented similar user groups have a higher degree of correlation, thereby improving actual business value. For example, when making product recommendations or engaging in social interactions, it is easier to find truly like-minded users.
[0115] Corresponding to the embodiments of the foregoing methods, this specification also provides embodiments of the apparatus and the terminal to which it is applied.
[0116] Figure 4 This is a schematic diagram illustrating the structure of an electronic device according to an exemplary embodiment. Figure 4 As shown, at the hardware level, the electronic device 400 includes a processor 402, an internal bus 404, a network interface 406, memory 408, and non-volatile memory 410, and may also include other hardware required for business operations. One or more embodiments of this specification can be implemented in software, for example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into memory 408 and then runs it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic module, but can also be hardware or logic devices.
[0117] Figure 5 This is a block diagram illustrating a vehicle-related user group analysis device according to an exemplary embodiment of this specification. Figure 5 As shown, this device can be applied to, for example Figure 4 The electronic device 400 shown implements the technical solution of this specification. The device includes: The data acquisition module 502 is used to acquire user profile data and spatiotemporal activity data generated by the user using the vehicle for at least two users to be segmented. The user profile data is determined based on the user's personal attributes and / or social relationships.
[0118] The set generation module 504 is used to determine the user correlation degree corresponding to each basic user pair consisting of the at least two users based on the user profile data of each user, determine the associated user pairs from the basic user pairs based on the user correlation degree, wherein each basic user pair contains two different users among the at least two users; and, based on each associated user pair, divide the at least two users into at least one associated user set, wherein each associated user pair to which the same user belongs is divided into the same associated user set.
[0119] The group segmentation module 506 is used to perform clustering processing on each user in each associated user set based on the spatiotemporal activity data of each user in the set, so as to obtain at least one similar user group.
[0120] Optionally, the set generation module 504 is specifically used to construct a global association graph corresponding to the at least two users based on the user profile data of each user, and to calculate the user association degree of each basic user to the corresponding user based on the topological relationship between each user represented by the global association graph.
[0121] Optionally, the user profile data includes personal attribute data. The set generation module 504 is specifically used to perform the following for each basic user pair: calculate attribute similarity based on the personal attribute data of the first user and the second user included in the basic user pair; determine the structural similarity between the first user and the second user according to the connection relationship between the vertices corresponding to the first user and the second user in the global association graph and other vertices; each vertex in the global association graph is used to identify a user; and determine the user association degree between the first user and the second user according to the attribute similarity and the structural similarity.
[0122] Optionally, the set generation module 504 is specifically used to determine the structural index values of the vertices corresponding to the first user and the vertices corresponding to the second user respectively; and to calculate the structural similarity between the first user and the second user based on the structural index values of the two vertices respectively.
[0123] Optionally, the set generation module 504 is specifically used to, for each associated user set, determine the activity spatial range and activity time range of each user in the set based on the spatiotemporal activity data of each user in the set; determine the similarity of the activity spatial range between every two users in the set, and determine the similarity of the activity time range between every two users in the set; and cluster every two users whose similarity of both activity spatial range and activity time range is not less than the corresponding threshold into the same similar user group.
[0124] Optionally, the set generation module 504 is specifically configured to perform the following for any user in each associated user set: determining at least one active location of the vehicle used by the user and the appearance time of the vehicle at the at least one active location based on the spatiotemporal activity data of the user, wherein the at least one active location is used to characterize the user's activity spatial range and the appearance time is used to characterize the user's activity time range; or, determining multiple active locations of the vehicle used by the user or the driving trajectory of the vehicle between multiple active locations based on the spatiotemporal activity data of the user, and determining the appearance time of the vehicle at the multiple active locations, wherein the multiple active locations or the driving trajectory are used to characterize the user's activity spatial range and the appearance time is used to characterize the user's activity time range.
[0125] Optionally, the set generation module 504 is specifically configured to: when the at least one active location is used to characterize the user's activity space range, determine the distance similarity between at least one active location corresponding to each of the two users, and use it as the similarity of the activity space range of each of the two users; when multiple active locations or the driving trajectory are used to characterize the user's activity space range, determine the distance similarity between multiple active locations corresponding to each of the two users, and determine the trajectory similarity between the driving trajectories corresponding to each of the two users, and determine the similarity of the activity space range of each of the two users based on the distance similarity and the trajectory similarity; determine the similarity of the activity time range between each of the two users in the set, including: determining the similarity between the occurrence times corresponding to each of the two users, and using it as the similarity of the activity time range between each of the two users.
[0126] Optionally, the device further includes a result optimization module 508, used to determine the cluster density index value and / or spatiotemporal consistency index value for each group in at least one similar user group obtained by clustering processing; the cluster density index value is used to characterize the average distribution density of users in the group within the corresponding activity space range, and the spatiotemporal consistency index value is used to characterize the degree of overlap of the activity time ranges of users in the group; if the cluster density index value is not greater than a density threshold and / or the spatiotemporal consistency index value is not greater than a consistency threshold, then the group is removed.
[0127] Optionally, the set generation module is specifically used to determine basic user pairs whose corresponding user association degree is not lower than the association degree threshold as associated user pairs.
[0128] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0129] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0130] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned vehicle-related user group analysis methods provided in this application.
[0131] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0132] This specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the aforementioned vehicle-related user group analysis methods.
[0133] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with relevant laws, regulations and standards, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.
Claims
1. A method for analyzing vehicle-related user groups, characterized in that, The method includes: For at least two users to be segmented, obtain user profile data and spatiotemporal activity data generated by the user's vehicle usage for each user, wherein the user profile data is determined based on the user's personal attributes and / or social relationships; Based on the user profile data of each user, the user correlation degree corresponding to each basic user pair consisting of the at least two users is determined. Based on the user correlation degree, related user pairs are determined from the basic user pairs, wherein each basic user pair contains two different users among the at least two users. Based on each related user pair, the at least two users are divided into at least one related user set, wherein the related user pairs to which the same user belongs are divided into the same related user set. For each set of associated users, clustering is performed on each user in the set based on the spatiotemporal activity data of each user in the set to obtain at least one similar user group.
2. The method according to claim 1, characterized in that, The determination of user association degree for each basic user pair consisting of at least two users based on user profile data of each user includes: Based on the user profile data of each user, a global association graph corresponding to at least two users is constructed, and based on the topological relationship between each user represented by the global association graph, the association degree between each basic user and the corresponding user is calculated.
3. The method according to claim 2, characterized in that, The user profile data includes personal attribute data. The calculation of the association degree between each basic user and its corresponding user, based on the topological relationships represented by the global association graph, includes: Execution for each basic user: Calculate attribute similarity based on the personal attribute data of the first and second users included in the basic user group; The structural similarity between the first user and the second user is determined based on the connection relationships between the vertices corresponding to the first user and the second user in the global association graph and other vertices; each vertex in the global association graph is used to identify a user; The user association degree between the first user and the second user is determined based on the attribute similarity and the structural similarity.
4. The method according to claim 3, characterized in that, The step of determining the structural similarity between the first user and the second user based on the connection relationships between the vertices corresponding to the first user and the second user respectively and other vertices in the global association graph includes: Determine the structural index values of the vertices corresponding to the first user and the vertices corresponding to the second user, respectively. The structural similarity between the first user and the second user is calculated based on the structural index values of the two vertices.
5. The method according to claim 4, characterized in that, The structural index values include degree centrality index value, clustering coefficient and / or neighbor similarity.
6. The method according to claim 1, characterized in that, For each associated user set, clustering is performed on each user in the set based on the spatiotemporal activity data of each user in the set to obtain at least one similar user group, including: For each set of associated users, the spatial range and temporal range of each user in the set are determined based on the spatiotemporal activity data of each user in the set. Determine the similarity of the activity space range between every two users in the set, and determine the similarity of the activity time range between every two users in the set; Cluster every two users whose similarity in both activity space and activity time is not less than the corresponding threshold into the same similar user group.
7. The method according to claim 6, characterized in that, For each associated user set, the process of determining the activity spatial range and activity time range of each user in the set based on the spatiotemporal activity data of each user in the set includes: Execute for any user in each associated user set: Based on the spatiotemporal activity data of any user, determine at least one active location of the vehicle used by that user and the time of appearance of the vehicle at that at least one active location, wherein the at least one active location is used to characterize the user's activity spatial range, and the time of appearance is used to characterize the user's activity temporal range; or... Based on the spatiotemporal activity data of any user, determine multiple active locations of the vehicle used by the user or the driving trajectory of the vehicle between multiple active locations, and determine the appearance time of the vehicle at the multiple active locations, wherein the multiple active locations or the driving trajectory are used to characterize the user's activity spatial range, and the appearance time is used to characterize the user's activity time range.
8. The method according to claim 7, characterized in that, Determining the similarity of the activity space range between any two users in the set includes: In the case where the at least one active location is used to characterize the user's activity space range, the distance similarity between the at least one active location corresponding to each of the two users is determined and used as the similarity of the activity space range of the two users; When the multiple active locations or the driving trajectory are used to characterize the user's activity space range, the distance similarity between the multiple active locations corresponding to each of the two users is determined, and the trajectory similarity between the driving trajectories corresponding to each of the two users is determined. The similarity of the activity space range of each of the two users is determined based on the distance similarity and the trajectory similarity. Determine the similarity of the activity time ranges between every two users in this set, including: Determine the similarity between the occurrence times of each pair of users, and use this similarity as the similarity of the activity time range between each pair of users.
9. The method according to claim 1, characterized in that, The method further includes: For each of the at least one similar user groups obtained by clustering, determine the cluster density index value and / or spatiotemporal consistency index value of the group; the cluster density index value is used to characterize the average distribution density of users in the group within the corresponding activity space range, and the spatiotemporal consistency index value is used to characterize the degree of overlap of the activity time ranges of users in the group. If the cluster density index value is not greater than the density threshold and / or the spatiotemporal consistency index value is not greater than the consistency threshold, then the group is removed.
10. The method according to claim 1, characterized in that, The process of determining associated user pairs from basic user pairs based on the user association degree includes: Basic user pairs whose corresponding user correlation degree is not lower than the correlation degree threshold are identified as related user pairs.
11. The method according to any one of claims 1-10, wherein the user profile data comprises: Personal attribute data determined based on the user's personal attributes; And / or, Social relationship data is determined based on the user's social relationships.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-11.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1-11.