Cross-data-table efficient personnel circling method
By decomposing the primary population package into multi-dimensional population package, and combining bitmap technology, layer compression and graph convolution network, the problem of inefficiency in traditional circle selection methods when processing cross-data table related information and massive data is solved, and efficient and accurate user group division is achieved.
Patent Information
- Application Number
- CN202510223889.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional personnel selection method cannot effectively process the correlation information among multiple data tables, and is less efficient when processing massive data, and cannot fully utilize the intrinsic relationships between various types of data, making it difficult to achieve accurate user group division.
By decomposing the primary population package into attribute classes, behavior classes and dynamic asset population packages, and generating corresponding bitmap vectors, combining layer compression strategies and graph convolution networks for feature learning, performing merging and deduplication strategies, and finally generating the target user group through data analysis strategies.
It significantly improves the performance of the circle selection process, reduces the computational complexity, improves processing speed, and achieves more accurate user group division, which is suitable for processing complex and multi-dimensional user data.
Smart Images

Figure CN120067165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data table circle selection, and specifically to an efficient method for circle selecting personnel across data tables. Background Art
[0002] With the continuous growth of data volume and the increasing demand for data processing, how to efficiently perform personnel circle selection and data analysis has become an important issue in various data analysis and marketing decision-making systems.
[0003] Traditional personnel circle selection methods usually rely on simple database query, screening, and filtering operations, and often cannot handle the associated information across multiple data tables, and have low processing efficiency for massive data. In addition, the circle selection requirements based on multi-dimensional data such as user attributes, behaviors, and dynamic assets are also becoming increasingly apparent. How to integrate different types of data and accurately divide and identify target user groups while ensuring high efficiency and accuracy is a major challenge in the current technology. In the process of circle selecting personnel across data tables, the common practice is to screen based on user attribute data, behavior data, and dynamic asset data. However, this approach cannot fully utilize the internal associations between various types of data and has low efficiency in processing complex data structures.
[0004] This solution proposes an efficient method for circle selecting personnel across data tables, which can effectively improve the performance of the circle selection process when processing complex user data across data tables. By decomposing the primary population package into population packages of attribute type, behavior type, and dynamic asset type, and combining the bitmap vector and layer compression strategy, the computational complexity can be significantly reduced while ensuring data integrity, and the processing speed can be improved. Summary of the Invention
[0005] The present invention provides an efficient method for circle selecting personnel across data tables, which helps to solve the problems mentioned in the above background art.
[0006] In a first aspect, the present application provides an efficient method for circle selecting personnel across data tables, adopting the following technical solution: The efficient method for circle selecting personnel across data tables includes: S1. Obtain the data tables storing user data, where the user data includes user profile data, user attribute data, and user behavior data; Obtain the query conditions for querying user data, and input the query conditions on the UI interface to obtain the records in the data table, and form a primary population package; S2. Execute a circle selection parsing algorithm on the primary population package to decompose the primary population package into three types of population packages, namely, population packages of attribute type, behavior type, and dynamic asset type; S3. Generate bitmap vectors for the population packages of attribute type, behavior type, and dynamic asset type respectively; S4. Execute the layer compression strategy based on the generated bitmap vector to obtain a primary embedding feature matrix; S5. Execute a merging strategy on the primary embedding features of each user to generate an in-group packet result; S6. Execute a data parsing strategy to parse the in-group packet result and generate a target user group.
[0007] Preferably, perform a selection and parsing algorithm on the primary population packet to decompose the primary population packet into three types of population packets, including: The primary population packet is {g 1 , g 2 ,..., g n}, where g i is the i-th record and n is the number of records in the primary population packet; The record g i includes an attribute field set A i , a behavior field set B i and a dynamic asset field set C i ; Denote the condition for the attribute type population packet as D 1 ; Denote the condition for the behavior type population packet as D 2 ; Denote the condition for the dynamic asset type population packet as D 3 ; Then the attribute type population packet is expressed as E 1 ={g i |A i ∈D 1}, the behavior type population packet is expressed as E 2 ={g i |B i ∈D 2}, and the dynamic asset type population packet is expressed as E 3 ={g i |C i ∈D 3}; By performing a selection and analysis algorithm on the primary population package and decomposing it into three types of population packages: attribute-based, behavior-based, and dynamic asset-based, it is possible to effectively process and analyze data from different dimensions. The advantage of this strategy is that the diversity of data is decomposed into more targeted and hierarchical groups, facilitating subsequent precise screening and analysis. For example, the attribute-based population package focuses on the basic information of users, such as age, gender, and region; the behavior-based population package focuses on the dynamic behavior characteristics of users, such as consumption behavior, access frequency, and activity records; the dynamic asset-based population package combines the asset changes of users, such as purchase history and asset evaluation. The separation of these data levels helps to avoid interference between different data categories and improves the accuracy of analysis.
[0008] Preferably, the generation of bitmap vectors for the attribute-based population package, behavior-based population package, and dynamic asset-based population package respectively includes: Generate bitmaps for the attribute-based population package, behavior-based population package, and dynamic asset-based population package: Bitmap(E i )={b i1 ,b i2 ,…,b in},i = 1, 2, 3; Wherein, b ij = 1 indicates that user j belongs to group E i , b ij = 0 indicates that user j does not belong to group E i .
[0009] By generating bitmap vectors for the attribute-based population package, behavior-based population package, and dynamic asset-based population package respectively, the processing of each type of population package can be transformed into binary operations, thereby greatly improving the speed and space utilization efficiency of data processing. As an efficient data compression and storage method, the Bitmap technology can represent a large amount of data in a very small space, reducing storage requirements while accelerating data retrieval and analysis. In this method, the generation of bitmap vectors allows us to quickly mark and query whether a user belongs to a specific group, thus greatly improving the efficiency of group selection, especially when facing a large amount of data, its advantages are particularly obvious.
[0010] Preferably, the execution of the layer compression strategy based on the generated bitmap vectors to obtain the primary embedding feature matrix includes: Extract the structural features of the user data graph based on social network analysis; Construct the user data graph (V, U), where V represents the user set, including all users, V = {g 1 , g 2 ,…, g n}, where $U$ represents the edge set, including user relationships; Calculate the weight $\omega$ of user relationships ij , representing the strength of the user relationship between user $i$ and user $j$: where, $|I$ i $\cap I$ j $|$ is the number of common interests of user $i$ and user $j$, $|I$ i $\cup I$ j $|$ is the total number of interests of user $i$ and user $j$, $I$ i and $I$ j respectively represent the interest sets of user $i$ and user $j$, $f$ ij represents the interaction frequency between user $i$ and user $j$.
[0011] By executing the layer compression strategy based on the generated bitmap vector, a primary embedded feature matrix can be obtained, which can further improve the efficiency of feature extraction in the data hierarchy. The core advantage of the layer compression strategy lies in its ability to effectively reduce the dimension of the feature space while maintaining the structural information of the data. Through compression processing, the data can be represented in a more concise and compact form, making the subsequent feature learning and pattern recognition processes more efficient. In addition, the compressed feature matrix can provide more representative data inputs for models such as graph convolutional networks, helping the models better capture the implicit relationships in the data.
[0012] Preferably, the step of executing the layer compression strategy based on the generated bitmap vector to obtain the primary embedded feature matrix further includes: Feature learning based on a graph convolutional network: Establish a user feature matrix Execute graph convolution operations: $H$ (l+1) $ = \sigma\times(D$ -1 / 2 $\times A\times D$ -1 / 2 $\times H$ (l) $\times W$ (l) ), where $D$ is the degree matrix, representing the degree of each user in the user data graph; $A$ is the adjacency matrix, containing the structural information of the user data graph, $D$ -1 / 2 $\times A\times D$ -1 / 2 represents the normalized adjacency matrix; $H$ (l) is the user feature matrix of the $l$-th layer, representing the feature representation of each user in the $l$-th layer, $l\geq0$; $\sigma$ is the activation function; $W$ (l) is the weight matrix of the $l$-th layer, used to learn the way of information fusion between users; At the l+1 layer, the feature of each user is the weighted sum of its current feature and the features of all neighboring users: where N(i) represents the set of neighboring users of user i, and w ij represents the weight of the user relationship between user i and user j.
[0013] By performing feature learning based on the Graph Convolutional Network (GCN), more profound structural features can be extracted from the user data graph, improving the accuracy of group division. The graph convolutional network is a deep learning model suitable for graph data, which can effectively capture the relationships between nodes and perform feature aggregation based on the connection structure between nodes. In this solution, the user data is regarded as a graph, where each user is a node, and the relationships between users (such as interactions, common interests, etc.) form the edges of the graph. Through graph convolutional operations, GCN can learn high-level feature representations for each user from the graph structure, thus more accurately capturing the implicit relationships between users during group division.
[0014] By adopting GCN, feature learning can not only consider the single-attribute data of users but also combine the relationships between users and other users, comprehensively reflecting the behavioral characteristics and dynamic changes of users. This comprehensive analysis method can help the system identify user groups with similar behaviors or interests, while traditional methods often can only be divided based on single-dimensional data and are difficult to comprehensively reflect the diversity of users. GCN aggregates the features of neighboring nodes layer by layer, making the feature representation of each user not only depend on its own data but also fuse the features of surrounding users, thus improving the accuracy of group division. This method is particularly suitable for user data with complex relationships, such as social networks and consumer behavior analysis.
[0015] Preferably, the implementation of the merging strategy on the primary embedding features of each user to generate the in-group packet result includes: Deduplication based on the graph convolutional network: Calculate the similarity between user i and user j where H i and H j are the user feature matrices of user i and user j respectively, and r is the dimension of the user feature vector; Set the similarity threshold; Compare the similarity between user i and user j with the similarity threshold; If the similarity between user i and user j is less than the similarity threshold; Then it is determined that user i and user j are redundant, and any one of user i and user j is randomly selected for deduplication; Record the matrix after the deduplication operation as the primary deduplication result.
[0016] By implementing a duplicate removal strategy on the primary embedding features of each user, redundant users can be effectively eliminated, thereby improving the quality and efficiency of group division. During the actual process of selecting personnel, there are a certain number of duplicate or similar users, and the duplication of these users can lead to deviations in the analysis results and even affect the accuracy of decision-making. Therefore, adopting a duplicate removal strategy can help remove these redundant users, making the final user group more accurate. By calculating the similarity between users, determining whether they are redundant users, and performing duplicate removal based on a set threshold, it is possible to effectively avoid the repeated appearance of users with the same features in the same group, improving the uniqueness and diversity of the selection results.
[0017] Preferably, the implementation of the merging strategy on the primary embedding features of each user to generate the in-group package result includes: Performing bitwise operations on the primary duplicate removal result: The primary duplicate removal result is represented as a column vector as bitmap 1 , bitmap 2 …bitmap z , where z is the number of elements in the primary duplicate removal result; Obtaining the number of elements x of each element in bitmap 1 , bitmap 2 …bitmap z ; Calculating the values of 16 - x, 10 - x, 8 - x, and 2 - x respectively, and determining whether there are positive numbers. If there are positive numbers, obtain the smallest positive number and denote it as the target base; After bitmap 1 , bitmap 2 …bitmap z , supplement 0 to make the number of bits of each column vector equal to the target base; Representing the primary duplicate removal result using the target base respectively to obtain the secondary duplicate removal result; If it is determined that there are no positive numbers, after bitmap 1 , bitmap 2 …bitmap z , supplement 0 to make the number of bits of each column vector equal to a multiple of 16; Representing the primary duplicate removal result using hexadecimal respectively to obtain the secondary duplicate removal result; The secondary duplicate removal result is represented as bitmap' 1 , bitmap' 2 …bitmap' z ; Performing the bitwise operation bitmap' 1 ∧bitmap' 2 ∧…∧bitmap'z Get the result of the group entry package.
[0018] By performing bitwise operations on the primary deduplication result, the process of group division can be further optimized, significantly improving the calculation efficiency and accuracy. As a low-level binary operation, bitwise operations can process data at extremely high speeds. Especially when dealing with large-scale data sets, they can effectively shorten the calculation time and reduce the system burden. In this solution, the primary deduplication result is processed through bitwise operations, enabling the feature information of each user to be stored and calculated in the most compact form, thus achieving efficient data screening and group identification. Through bitwise operations, the system can complete batch processing of a large amount of data within a fixed time. The efficiency of bitwise operations is extremely high and is particularly suitable for scenarios that require high-frequency and real-time data processing. In this solution, by first converting user data into bit vectors and then mapping them to the target groups through bitwise operations, it can ensure that the group division process is more concise and efficient, avoiding repeated calculations and complex operations. At the same time, bitwise operations can reduce memory occupancy and improve the space efficiency of data processing. Especially in application scenarios with a large amount of data, it can significantly save computing resources.
[0019] Preferably, the execution of the data parsing strategy to parse the result of the group entry package and generate the target user group includes: Obtain the result of the group entry package {b' 1 ,b' 2 ,…,b' t} where t is the number of elements in the result of the group entry package; Determine whether each element in the result of the group entry package is equal to 1; Obtain the users corresponding to all elements equal to 1 to form the target user group.
[0020] The present invention has the following beneficial effects: 1. This efficient method for selecting personnel across data tables can accurately divide the primary population package into three types of population packages: attribute type, behavior type, and dynamic asset type through the application of the selection and parsing algorithm. Each type of population package represents different characteristic dimensions of users, making the analysis process more accurate and personalized. Through this method, it is possible to focus on the basic attributes, behavior patterns, and asset dynamics of users respectively, providing a more detailed user group division. This not only helps to accurately identify the true needs of each user but also provides basic data support for subsequent personalized marketing and recommendation systems. Through the multi-dimensional data analysis after decomposition, not only the accuracy of group division is improved, but also deeper data mining and decision-making are made possible.
[0021] 2. The efficient personnel selection method across data tables can efficiently mark and query population packages of attribute classes, behavior classes, and dynamic asset classes by generating bitmap vectors. The application of bitmap vectors can convert the status of whether each user belongs to a certain group into binary flags, making data storage and processing more efficient. Through bit operations, the query operation becomes faster and can obtain results in constant time without scanning data row by row. This method greatly reduces the time consumption of query operations in traditional methods and improves the system response speed, especially in scenarios that require real-time analysis and quick response. Through this efficient query method, not only computing resources are saved, but also the processing ability in the big data environment is improved.
[0022] 3. The efficient personnel selection method across data tables can effectively reduce the dimension of the feature matrix and improve the computing efficiency through the layer compression strategy. Layer compression accelerates the group division process by reducing redundant features and focusing the calculation on the most informative features. When dealing with large-scale user data, this strategy can greatly shorten the calculation time and reduce memory occupancy, ensuring that the system can process more data in a short time. On datasets with a higher dimension, layer compression can not only reduce the system burden but also help the model better focus on the most important data features, avoid unnecessary computational complexity, and ensure the efficiency and accuracy of group division.
[0023] 4. The efficient personnel selection method across data tables can improve the accuracy of group division by performing feature learning through a graph convolutional network (GCN). GCN can capture the complex relationships between users, especially social relationships and interaction patterns, which are crucial for group division. By performing convolutional operations on the user data graph, GCN can further integrate the relationships between a user and other users based on considering the user's independent features, thus providing a more accurate representation of user features. This method is particularly suitable for complex user data, can effectively capture potential information that traditional methods may ignore, make group division more in line with the needs of actual scenarios, and improve the effects of personalized recommendation and precision marketing.
[0024] 5. The efficient personnel selection method across data tables can significantly reduce redundant users through a deduplication strategy, ensuring the purity and quality of group division. The deduplication operation calculates the similarity between users, identifies highly similar or duplicate users, and removes redundant data. This process ensures that the members within each user group are independent and representative, thus avoiding information redundancy and duplicate calculations in group division. This deduplication mechanism not only improves the accuracy of group division but also reduces the waste of computing resources, making data processing more efficient and precise. The deduplicated dataset can better support subsequent analysis and decision-making, providing a more reliable user profile and behavior prediction for enterprises.
[0025] 6. The efficient personnel selection method across data tables can further optimize group division and improve calculation accuracy by performing bitwise operations on the primary deduplication result. Bitwise operations can operate on user groups with extremely high efficiency and improve the representation accuracy of features through strategies such as the least common multiple and base conversion. In large-scale data processing, bitwise operations can transform data operations from traditional complex calculations into simple binary operations, greatly enhancing the calculation speed. This optimization strategy can ensure more accurate group division and avoid errors caused by redundant information and duplicate calculations in traditional methods, providing more accurate and efficient data processing capabilities.
[0026] 7. The efficient personnel selection method across data tables can further improve the quality of group division through bitwise operations after deduplication. The deduplicated data is converted into a more compact bit vector, and the efficiency of bitwise operations makes group division faster and more accurate. During the deduplication process, redundant users are removed to ensure that the members of each group are unique, and through further bitwise operation processing, group division can be quickly completed. This method compresses data into bit vectors and performs bitwise operations, avoiding the one-by-one comparison of a large amount of redundant data in traditional processing methods, improving calculation efficiency and accuracy. Finally, through this process, the division of user groups is more precise and can better support subsequent analysis and decision-making.
[0027] 8. The efficient personnel selection method across data tables can further optimize the in-group package result and accurately identify the target user group by implementing a data parsing strategy. The data parsing strategy can identify users who meet specific conditions by analyzing the in-group package result in detail, and then generate a high-quality target user group. This process ensures the accuracy and representativeness of the final user group by checking each item of the group division result. In practical applications, this data parsing strategy can be flexibly adjusted according to different business requirements, thus improving the applicability of group division. Through this refined parsing method, not only can the quality of group division be ensured, but also more reliable user data support can be provided for applications such as precision marketing and personalized recommendation. Brief Description of the Drawings
[0028] Figure 1 This is a schematic flowchart of the method of the present invention.
[0029] Figure 2 This is a schematic flowchart of the duplicate removal process of the method of the present invention. Detailed Embodiments
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] Embodiment 1. Refer to Figure 1 , an efficient method for selecting personnel across data tables, including: S1. Obtain the data tables storing user data, where the user data includes user profile data, user attribute data, and user behavior data; Obtain the query conditions for querying user data, and input the query conditions on the UI interface to obtain the records in the data table, forming a primary population package; S2. Execute a selection parsing algorithm on the primary population package to decompose the primary population package into three types of population packages, namely attribute-based population packages, behavior-based population packages, and dynamic asset-based population packages; S3. Generate bitmap vectors for the attribute-based population package, behavior-based population package, and dynamic asset-based population package respectively; S4. According to the generated bitmap vectors, execute a layer compression strategy to obtain a primary embedding feature matrix; S5. Execute a merging strategy on the primary embedding features of each user to generate an in-group package result; S6. Execute a data parsing strategy to parse the in-group package result and generate a target user group.
[0032] Execute a selection parsing algorithm on the primary population package to decompose the primary population package into three types of population packages, including: The primary population package is {g 1 , g 2 ,..., g n}, where g i is the i-th record, and n is the number of records in the primary population package; Record g i includes an attribute field set A i , a behavior field set B i and a dynamic asset field set C i; Denote the condition for meeting the attribute - based population package as D 1 ; Denote the condition for meeting the behavior - based population package as D 2 ; Denote the condition for meeting the dynamic - asset - based population package as D 3 ; Then the attribute - based population package is represented as E 1 ={g i |A i ∈D 1}, the behavior - based population package is represented as E 2 ={g i |B i ∈D 2}, the dynamic - asset - based population package is represented as E 3 ={g i |C i ∈D 3}; Through the circle - selection parsing algorithm, the primary population package can be subdivided into three sub - populations: attribute - based, behavior - based, and dynamic - asset - based, thus improving the accuracy and flexibility of population division. Each type of sub - population represents different dimensions in user data, focusing on the basic attributes, behavior characteristics, and asset dynamics of users respectively, and can provide more detailed and accurate analysis when processing different data dimensions. This subdivision enables different types of users to be processed separately according to their specific characteristics, avoiding processing all data at once, thereby optimizing the quality of population division. In addition, this method provides a clearer user portrait for subsequent precision marketing, personalized recommendation, etc., helping to better meet the needs of different user groups. Through the decomposed population package, the system can more flexibly adjust and optimize the data analysis strategy, enhancing the controllability of population division, ensuring that each sub - population can meet different business requirements, and maximizing the efficiency and value of data utilization.
[0033] Generate bitmap vectors for the attribute - based population package, behavior - based population package, and dynamic - asset - based population package respectively, including: Generate bitmaps for the attribute - based population package, behavior - based population package, and dynamic - asset - based population package: Bitmap(E i )={b i1 ,b i2 ,…,b in}, i = 1, 2, 3; Among them, b ij = 1 indicates that user j belongs to population E i , b ij = 0 indicates that user j does not belong to population E i .
[0034] By generating a bitmap vector, the processing speed of group division and query can be greatly improved. Bitmap is an efficient binary data representation method that can convert whether each user belongs to a specific group into a binary flag, thus enabling faster storage and query. When performing data screening, the bitmap vector can quickly determine whether a user meets certain conditions through bit operations, which greatly reduces the time consumption in traditional data processing methods. In the application scenarios of large-scale data analysis and high-frequency queries, the advantages of the bitmap technology are particularly prominent. It can convert complex query operations into simple bit operations, greatly improving the computing efficiency. In this way, the system can process a large amount of data in a very short time and provide real-time and fast query results. This efficient query method is particularly suitable for fields such as real-time data analysis, dynamic group division, and precise recommendation. It can save system resources and reduce the computing burden while ensuring high efficiency, thus providing stronger support for business applications.
[0035] According to the generated bitmap vector, execute the layer compression strategy to obtain the primary embedded feature matrix, including: Extract the structural features of the user data graph based on social network analysis; Construct the user data graph (V, U), where V represents the user set, containing all users, V = {g 1 , g 2 , …, g n}, and U represents the edge set, containing user relationships; Calculate the weight ω ij of the user relationship, representing the strength of the user relationship between user i and user j: Among them, | I i ∩I j | is the number of common interests of user i and user j, | I i ∪I j | is the total number of interests of user i and user j, I i and I j respectively represent the interest sets of user i and user j, and f ij represents the interaction frequency between user i and user j.
[0036] Optimizing the feature matrix through the layer compression strategy can significantly improve the efficiency and accuracy of data processing. Layer compression effectively removes redundant information by reducing the dimension of the feature matrix, enabling the system to be more efficient when processing large amounts of data. Different from traditional high-dimensional data processing methods, layer compression can compress the computational amount, thereby reducing memory occupancy and consumption of computing resources. Especially in the case of extremely large data volume or limited computing power, the compressed feature matrix can ensure faster and more energy-efficient data processing. In addition, the compressed feature matrix retains the most valuable information in the data, enabling the system to focus on key features when processing high-dimensional data, avoiding interference from excessive irrelevant data in the calculation process, and further improving the accuracy of group division. Layer compression can not only improve computational efficiency but also help maintain higher analysis accuracy when dealing with complex data sets, providing more accurate data support for subsequent intelligent analysis and decision-making, especially suitable for applications in big data environments.
[0037] According to the generated bitmap vector, execute the layer compression strategy to obtain the primary embedded feature matrix, further including: Perform feature learning based on the graph convolutional network: Establish the user feature matrix Execute graph convolution operation: H (l+1) =σ×(D -1 / 2 ×A×D -1 / 2 ×H (l) ×W (l) ), where D is the degree matrix, representing the degree of each user in the user data graph; A is the adjacency matrix, containing the structural information of the user data graph, and D -1 / 2 ×A×D -1 / 2 represents the normalized adjacency matrix; H (l) is the user feature matrix of the l-th layer, representing the feature representation of each user at the l-th layer, l≥0; σ is the activation function; W (l) is the weight matrix of the l-th layer, used to learn the way of information fusion between users; At the (l + 1)-th layer, the feature of each user is the weighted sum of its current feature and the features of all neighbor users: where N(i) represents the set of neighbor users of user i, and w ij represents the weight of the user relationship between user i and user j.
[0038] Feature learning through the Graph Convolutional Network (GCN) can effectively improve the depth and accuracy of group division. By considering the relationships between users, GCN can not only process the individual characteristics of users but also capture the potential interaction patterns and social network information among users through the graph structure. The graph convolutional network can automatically learn and fuse the similarities and relationships between different users based on the user data graph, thereby generating more accurate user feature representations. In traditional group division methods, only the independent characteristics of users are usually relied on for analysis, while the interaction relationships between users are ignored, which may lead to incomplete or distorted group division. The graph convolutional network can effectively capture the relationship information between users, making the group division more in line with actual behaviors and interests, thus improving the accuracy of group division. In addition, the non-linear learning ability of GCN enables it to process complex data patterns, ensuring that the group division does not stay at the analysis of surface features but can delve into the deep relationships in the data. This makes the group division based on the graph convolutional network more intelligent and automated, improving the efficiency and effectiveness of large-scale user data analysis.
[0039] Execute a merging strategy on the primary embedding features of each user to generate the in-group package result, including: Deduplication based on the graph convolutional network: Calculate the similarity between user i and user j where, H i and H j are the user feature matrices of user i and user j respectively, and r is the dimension of the user feature vector; Set the similarity threshold; Compare the similarity between user i and user j with the similarity threshold; If the similarity between user i and user j is less than the similarity threshold; Then it is determined that user i and user j are redundant, and arbitrarily select one of user i and user j for deduplication; Denote the matrix after the deduplication operation as the primary deduplication result.
[0040] Through the deduplication strategy, redundant users can be effectively eliminated, ensuring the accuracy and representativeness of group division. In practical applications, redundant data not only increases the computational burden but also may lead to distortion of the group division results. The deduplication strategy calculates the similarity between users and identifies duplicate or highly similar users, avoiding the impact of such redundant data on group division. Through this operation, it is ensured that the members within each group are unique and representative, thus improving the accuracy of group division. The deduplicated dataset can better reflect the characteristics of real users, making subsequent group analysis and behavior prediction more in line with the actual situation. At the same time, this deduplication process also reduces the waste of computing resources and optimizes the efficiency of data processing. In large-scale data analysis scenarios, the deduplication strategy not only improves the quality of group division but also provides more reliable data support for applications such as precision marketing, targeted advertising, and personalized recommendations.
[0041] Execute the merging strategy on the primary embedding features of each user to generate the result of the group entry package, including: Perform bitwise operations on the primary deduplication result: The primary deduplication result is represented as a column vector in bitmap 1 , bitmap 2 …bitmap z , where z is the number of elements in the primary deduplication result, z = 3; Obtain the number x of each element in bitmap 1 , bitmap 2 …bitmap z ; Calculate the values of 16 - x, 10 - x, 8 - x, and 2 - x respectively, and determine whether there are positive numbers. If there are positive numbers, obtain the smallest positive number and denote it as the target base; After bitmap 1 , bitmap 2 …bitmap z , supplement 0 to make the number of bits of each column vector equal to the target base; Represent the primary deduplication result using the target base respectively to obtain the secondary deduplication result; If it is judged that there are no positive numbers, after bitmap 1 , bitmap 2 …bitmap z , supplement 0 to make the number of bits of each column vector equal to a multiple of 16; Represent the primary deduplication result using hexadecimal respectively to obtain the secondary deduplication result; The secondary deduplication result is represented as bitmap' 1 , bitmap' 2 …bitmap' z; Perform bitwise operation on bitmap' 1 ∧bitmap' 2 ∧…∧bitmap' z Obtain the result of the group entry package.
[0042] In this embodiment, refer to Figure 2 .
[0043] By performing bitwise operations on the deduplicated data, it is possible to significantly improve the data processing speed while ensuring the accuracy of group division. Bitwise operations, with their high efficiency and low resource consumption, have become a common optimization method in big data processing. After the deduplication operation is completed, the system converts the user data into bit vectors and uses bitwise operations to quickly perform group division and feature matching. Through bitwise operations, instead of comparing data one by one, group screening is achieved through efficient binary operations, greatly improving the processing speed. In addition, bitwise operations can accurately handle the intersection, union, and difference set operations of multiple groups, ensuring that the accuracy of group division is not affected. Through this optimization strategy, the process of group division becomes more concise and efficient, especially suitable for scenarios of real-time data analysis and dynamic group updates. This method can achieve efficient operations on large-scale data sets, avoid the computational bottlenecks in traditional methods, and ensure that the system can still maintain good performance and stability in complex data processing.
[0044] By performing bitwise operations on the deduplicated data, the process of group division can be further optimized and redundant operations can be reduced. By converting the data into a compact binary form, bitwise operations can quickly and accurately determine whether a user belongs to a specific group when performing group division. Different from traditional methods that require complex calculations and a large number of queries, bitwise operations can complete group screening in constant time, greatly improving the computational efficiency. In addition, bitwise operations make the data processing process more concise, avoiding repeated calculations and invalid operations, thus saving computational resources. During the process of group division, bitwise operations not only ensure the accuracy of the results but also reduce unnecessary calculation steps, ensuring more efficient use of resources. In this way, the quality of group division is improved and the system performance is optimized. Especially when processing large-scale user data, it can significantly improve the processing speed and meet the real-time computing requirements.
[0045] Execute the data parsing strategy, parse the result of the group entry package, and generate the target user group, including: Obtain the result of the group entry package {b' 1 ,b' 2 ,…,b' t}, where t is the number of elements in the result of the group entry package; Determine whether each element in the result of the group entry package is equal to 1; Retrieve the users corresponding to all elements equal to 1 to form the target user group.
[0046] Processing the in-group package results through a data parsing strategy can further optimize the group division and ensure the accuracy of the target user group. The data parsing strategy analyzes the results of each in-group package meticulously to ensure that each group member meets the predetermined conditions, thereby improving the quality of group division. This strategy helps the system verify and optimize the group division results, further eliminating incorrect data and redundant information, making the finally generated target user group more accurate and reliable. Through this strategy, the system can flexibly adjust the group division criteria according to different business requirements to ensure that the division results highly match the actual needs. The data parsing strategy not only improves the accuracy of group division but also provides more accurate user data for applications such as precision marketing and user customization services, enhancing the decision-making support ability. Through this optimization process, the system can more intelligently meet the business requirements, improving the efficiency and reliability of the entire group division process.
[0047] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0048] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An efficient personnel selection method across data tables, characterized in that: include: S1. Obtain a data table storing user data, wherein the user data includes user profile data, user attribute data, and user behavior data; Obtain the query conditions for querying user data, and enter the query conditions on the UI interface to obtain records in the data table to form a primary population package; S2. Execute the circle selection and parsing algorithm on the primary crowd package to decompose the primary crowd package into three types of crowd packages, namely, attribute crowd package, behavior crowd package and dynamic asset crowd package; S3, generating bitmap vectors of attribute crowd packages, behavior crowd packages and dynamic asset crowd packages respectively; S4. Execute the layer compression strategy according to the generated bitmap vector to obtain the primary embedding feature matrix; S5, executing a merging strategy on each user's primary embedding features to generate a group entry package result; S6. Execute the data analysis strategy, analyze the group entry package results, and generate the target user group.
2. The method for efficiently selecting personnel across data tables according to claim 1, characterized in that: The primary crowd package is subjected to the circle selection and parsing algorithm to decompose the primary crowd package into three types of crowd packages, including: The primary population package is {g1,g2,...,g n }, where g i is the i-th record, n is the number of records in the primary population package; Record i Includes attribute field set A i , behavior field set B i and dynamic asset field set C i ; The condition that satisfies the attribute-based crowd package is recorded as D1, and the condition that satisfies the behavior-based crowd package is recorded as D2; The condition that satisfies the dynamic asset class population package is recorded as D3; Then the attribute class population package is expressed as E1 = {g i |A i ∈D1}, the behavior group package is represented by E2 = {g i |B i ∈D2}, the dynamic asset class group package is represented by E3={g i |C i ∈D3}; 3. The method for efficiently selecting personnel across data tables according to claim 1 is characterized in that: The bitmap vectors of the attribute crowd package, the behavior crowd package and the dynamic asset crowd package are generated respectively, including: Generate bitmaps of attribute crowd packages, behavior crowd packages, and dynamic asset crowd packages: Bitmap(E i )={b i1 ,b i2 ,…,b in },i=1,2,3; Among them, b ij =1 means user j belongs to group E i , b ij =0 means user j does not belong to group E i .
4. The method for efficiently selecting personnel across data tables according to claim 2, characterized in that: According to the generated bitmap vector, the layer compression strategy is executed to obtain the primary embedding feature matrix, including: Extract structural features of user data graph based on social network analysis; Construct the user data graph (V, U), where V represents the user set, including all users, V = {g1, g2, ..., g n }, U represents the edge set, including user relations; Calculate the weight ω of the user relationship ij , represents the user relationship strength between user i and user j: Among them, |I i ∩I j | is the number of common interests between user i and user j, |I i ∪I j | is the total number of interests of user i and user j, I i and I j denote the interest sets of user i and user j respectively, and f ij Represents the interaction frequency between user i and user j.
5. The method for efficiently selecting personnel across data tables according to claim 4 is characterized in that: The method further comprises: executing a layer compression strategy according to the generated bitmap vector to obtain a primary embedding feature matrix; Feature learning based on graph convolutional networks: Build a user feature matrix Perform graph convolution operation: H (l+1) =σ×(D -1 / 2 ×A×D -1 / 2 ×H (l) ×W (l) ), where D is the degree matrix, representing the degree of each user in the user data graph; A is the adjacency matrix, which contains the structural information of the user data graph, and D -1 / 2 ×A×D -1 / 2 represents the normalized adjacency matrix, H (l) is the user feature matrix of the lth layer, which represents the feature representation of each user at the lth layer, l ≥ 0, and σ is the activation function; W (l) is the weight matrix of the lth layer, which is used to learn the way of information fusion between users; At the l+1th layer, each user’s feature is the weighted sum of its current feature and the features of all neighboring users: Among them, N(i) represents the set of neighbor users of user i, w ij Represents the weight of the user relationship between user i and user j.
6. The method for efficiently selecting personnel across data tables according to claim 5, characterized in that: The merging strategy is executed on the primary embedding features of each user to generate a group entry package result, including: Deduplication based on graph convolutional network: Calculate the similarity between user i and user j Among them, H i and H j are the user feature matrices of user i and user j respectively, and r is the dimension of the user feature vector; Set similarity threshold; Compare the similarity between user i and user j with the similarity threshold; If the similarity between user i and user j is less than the similarity threshold; Then it is determined that user i and user j are redundant, and one of user i and user j is randomly selected for deduplication; The matrix after the deduplication operation is performed is recorded as the primary deduplication result.
7. The method for efficiently selecting personnel across data tables according to claim 1, characterized in that: The merging strategy is executed on the primary embedding features of each user to generate a group entry package result, including: Perform bitwise operations on the primary deduplication results: The primary deduplication result is represented by a column vector as bitmap1, bitmap2…bitmap z , z is the number of elements in the primary deduplication result, z = 3; Get bitmap1, bitmap2...bitmap z The number of each element in x; Calculate the values of 16-x, 10-x, 8-x, and 2-x respectively, and determine whether they contain positive numbers. If there are positive numbers, obtain the smallest positive number and record it as the target base. In bitmap1, bitmap2...bitmap z Add 0 at the end so that the number of bits in each column vector is equal to the target base; The primary deduplication results are respectively expressed in the target base to obtain secondary deduplication results; If it is determined that there is no positive number, then in bitmap1, bitmap2...bitmap z Add 0 at the end so that the number of bits in each column vector is a multiple of 16; The primary deduplication results are expressed in hexadecimal to obtain secondary deduplication results; The secondary deduplication result is represented by bitmap'1, bitmap'2...bitmap' z ; Perform bit operations bitmap'1∧bitmap'2∧…∧bitmap' z Get the group entry package result.
8. The method for efficiently selecting personnel across data tables according to claim 1, characterized in that: The execution of the data analysis strategy, parsing the group entry package results, and generating the target user group includes: Get the group entry package result {b'1,b'2,…,b' t }, t is the number of elements in the group entry packet result; Check whether each element in the group packet result is equal to 1; Get all users corresponding to the elements that are equal to 1 to form the target user group.