Data labeling method and device based on graph structure and community discovery and storage medium
By constructing heterogeneous graph network and community discovery algorithms, the problem of data silos and flattening labeling systems in traditional customer data annotation methods is solved, unified understanding of customer behavior and dynamic value mining is achieved, and a hierarchical customer cognitive framework is established.
Patent Information
- Application Number
- CN202510779395.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Traditional customer data annotation methods are caused by data silos, single data dimension restricts dynamic value mining, and flattening and restricting hierarchical cognitive construction of labeling systems.
By building a heterogeneous graph network that includes the basic physical layer, the behavioral semantic layer and the environment association layer, combined with the community discovery algorithm, the customer group is divided into multiple customer communities, and a hierarchical tag tree is generated to realize data fusion and hierarchical cognition across business departments.
It realizes a unified understanding of customer behavior, in-depth exploration of multi-dimensional customer relationships, dynamic evolution capture of customer value, and hierarchical customer cognition construction, solves the limitations of data silos and labeling systems, and supports cross-departmental applications.
Smart Images

Figure CN120298083A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data annotation method, device and storage medium based on graph structure and community discovery. Background Art
[0002] In the process of digital transformation of the retail industry, customer data labeling, as a key link in data preprocessing, is of great significance for building refined customer portraits and supporting intelligent marketing decisions. Traditional customer data labeling methods in retail scenarios mainly revolve around specific business needs, such as product recommendation systems, member management systems, etc., and static labels are given to customers (such as "high-value customers" and "preferors of maternal and child products") through manual or rule engines. For example, the patent application with publication number CN118535775A and invention name "A retail shopping guide recommendation method and system based on user portrait labels" discloses a method of combining real-time consumer needs and historical consumer preferences to construct portrait labels and make personalized product recommendations to users. For another example, the patent application with publication number CN115619454A and invention name "A member management system and method based on member middle platform" discloses a method of adding user labels to members through multi-channel data classification, and generating personalized marketing activities based on labels. Although this type of method meets the customer classification needs in basic business scenarios to a certain extent, as retail scenarios evolve towards omni-channel and personalized directions, its limitations are becoming increasingly prominent: (1) Excessive task orientation leads to data silos: Retail business naturally involves the collaboration of multiple departments (such as procurement, inventory, marketing, and after-sales), and the problem of data fragmentation between departments is particularly prominent. The traditional labeling process is highly dependent on specific business goals (such as promotional activity response prediction), and different business departments often independently build labeling systems to meet their respective specific business needs. For example, the marketing department may only focus on transaction data and label "high-value customers" based on transaction data; while the risk control department may only focus on behavioral data and label "Double 11 high-spending users" based on behavioral data. This approach of independently building a labeling system based on departmental KPIs has exacerbated the phenomenon of data silos, causing the same customer to be labeled with fragmented labels in different retail departments, making it difficult to form a unified customer perception across departments, which in turn affects the overall operational efficiency and customer experience of retail companies; (2) Single data dimension restricts value mining: Traditional methods overly rely on structured data such as transaction amount and purchase frequency, ignoring the associated value of unstructured information such as customers’ behavior trajectory in physical stores (such as heat maps of traffic flow), online browsing paths (such as product detail page jump logic), and environmental factors (such as holiday promotion cycles and weather changes). As a result, the annotation results can only reflect the static cross-section of customer value and cannot capture dynamic evolution characteristics such as “adjusting shopping categories due to seasonal changes”. (3) The flat label system limits the depth of cognition: The existing technologies usually adopt a flat structure of "single-layer label + weight value" (such as "Preference for mother and baby products: 0.8"). It can neither establish a hierarchical cognitive framework such as "High-value customers → Preference for mother and baby products → Associated purchase of milk powder", nor can it quantify temporal features such as "Volatility of purchase cycle" and "Decay rate of activity response", resulting in a lack of data support for precision marketing in complex scenarios (such as life cycle value prediction).
[0003] In view of the technical problems in the above-mentioned existing technologies, that is, the traditional customer data annotation method has strong task orientation leading to data islands, the single data dimension restricts the mining of dynamic value, and the flat label system limits the construction of hierarchical cognition, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present disclosure provide a data annotation method, device, and storage medium based on graph structure and community discovery, so as to at least solve the technical problems in the existing technologies, that is, the traditional customer data annotation method has strong task orientation leading to data islands, the single data dimension restricts the mining of dynamic value, and the flat label system limits the construction of hierarchical cognition.
[0005] According to one aspect of the embodiments of the present disclosure, a data annotation method based on graph structure and community discovery is provided, which is applied to a data annotation platform in a retail scenario. The data annotation method includes: obtaining multi-dimensional data of a customer group to be labeled with tags; wherein, the multi-dimensional data includes customer transaction data, commodity attribute data, customer behavior trajectory data, and environmental data; constructing a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavior semantics layer, and an environmental association layer. The basic physical layer is used to represent the transaction association relationship between customers and commodities and the physical attributes of commodities. The behavior semantics layer is used to represent the potential behavior patterns of customers and the semantic association features between commodities. The environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between commodities and customers; and based on the heterogeneous graph network, dividing the customer group into multiple customer communities through a community discovery algorithm, and generating a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes a first-level label, a second-level label, and a third-level label. The first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer and serves as the top-level node of the hierarchical label tree. The second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavior semantics layer and serves as the middle-level node of the hierarchical label tree. The third-level label is constructed based on the dynamic trajectory temporal features of customers in the customer community in the heterogeneous graph network and serves as the bottom-level node of the hierarchical label tree.
[0006] According to another aspect of the embodiments of the present disclosure, there is also provided a storage medium, which includes a stored program, wherein the method described in any one of the above is executed by a processor when the program runs.
[0007] According to another aspect of the embodiments of the present disclosure, there is also provided a data annotation device based on graph structure and community discovery, including: a data acquisition module, configured to acquire multi-dimensional data of a customer group to be labeled with tags; wherein the multi-dimensional data includes customer transaction data, commodity attribute data, customer behavior trajectory data, and environmental data; a graph construction module, configured to construct a heterogeneous graph network according to the multi-dimensional data; wherein the heterogeneous graph network includes a basic physical layer, a behavior semantics layer, and an environmental association layer, the basic physical layer is used to represent the transaction association relationship between customers and commodities and the physical attributes of commodities, the behavior semantics layer is used to represent the potential behavior patterns of customers and the semantic association features between commodities, and the environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between commodities and customers; and an annotation module, configured to divide the customer group into multiple customer communities based on the heterogeneous graph network through a community discovery algorithm, and generate a hierarchical label tree for each customer community; wherein the hierarchical label tree includes a first-level label, a second-level label, and a third-level label, the first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer and serves as the top-level node of the hierarchical label tree; the second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavior semantics layer and serves as the middle-level node of the hierarchical label tree; the third-level label is constructed based on the dynamic trajectory time series features of customers in the customer community in the heterogeneous graph network and serves as the bottom-level node of the hierarchical label tree.
[0008] According to another aspect of the embodiments of the present disclosure, there is also provided a data annotation device based on graph structure and community discovery, including: a processor; and a memory, connected to the processor, for providing instructions for the processor to process the following steps: obtaining multi-dimensional data of a customer group to be labeled with tags; wherein, the multi-dimensional data includes customer transaction data, commodity attribute data, customer behavior trajectory data, and environmental data; constructing a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavioral semantics layer, and an environmental association layer, the basic physical layer is used to represent the transaction association relationship between customers and commodities and the physical attributes of commodities, the behavioral semantics layer is used to represent the potential behavior patterns of customers and the semantic association features between commodities, and the environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between commodities and customers; and based on the heterogeneous graph network, dividing the customer group into multiple customer communities through a community discovery algorithm, and generating a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes a first-level label, a second-level label, and a third-level label, the first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer and serves as the top node of the hierarchical label tree; the second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavioral semantics layer and serves as the middle-layer node of the hierarchical label tree; the third-level label is constructed based on the dynamic trajectory time series characteristics of customers in the customer community in the heterogeneous graph network and serves as the bottom node of the hierarchical label tree.
[0009] This application first obtains multi-dimensional data such as customer transaction data, product attribute data, customer behavior trajectory data, and environmental data in the retail scenario, breaking the data fragmentation caused by departmental task differences in the traditional annotation process. Then, by constructing a heterogeneous graph network including a basic physical layer, a behavioral semantics layer, and an environmental association layer, these data from different sources and different modalities are uniformly mapped into a graph structure to express the complex relationships among customers, products, and the environment in the form of a graph structure, thereby realizing the deep integration of multi-dimensional data and multi-dimensional relationship modeling in the retail scenario, and providing a unified data foundation for comprehensively understanding customer behavior. Secondly, through the community discovery algorithm, customer groups with cross-business department business value are identified in the heterogeneous graph network. These groups are not defined based on the needs of a single department, but through the cross-layer connection characteristics in the graph network, the common behavior patterns and value characteristics in different business scenarios are naturally mapped, thus forming a unified cognitive framework for customers. This framework supports cross-business department applications and solves the problem of fragmented labels being attached to the same customer in the traditional annotation system. Finally, based on the above cross-business department customer data, a hierarchical label tree is constructed to establish a hierarchical cognitive framework. This hierarchical label tree not only reflects the logical relationship between different labels, but also finely represents multiple dimensions and levels of customer value through the construction of multi-level labels. Thus, the technical effects of unified understanding of customer behavior, in-depth mining of multi-dimensional customer relationships, capture of the dynamic evolution of customer value, and construction of hierarchical customer cognition are achieved. Furthermore, the technical problems existing in the prior art, such as data islands caused by overly task-oriented traditional customer data annotation methods, single data dimension restricting dynamic value mining, and flat label systems restricting hierarchical cognitive construction, are solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present disclosure, and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings: Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure; Figure 2 is a schematic diagram of the hardware architecture of a data annotation system based on a graph structure and community discovery according to Embodiment 1 of the present disclosure; Figure 3 is a flowchart of a data annotation method based on a graph structure and community discovery according to Embodiment 1 of the present application; Figure 4 is a schematic diagram of a data annotation device based on a graph structure and community discovery according to Embodiment 2 of the present application; Figure 5 is a schematic diagram of a data annotation device based on a graph structure and community discovery according to Embodiment 3 of the present application. Detailed implementation manners
[0011] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0012] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0013] Embodiment 1 According to this embodiment, a method embodiment of a data annotation method based on graph structure and community discovery is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0014] The method embodiment provided in this embodiment can be executed in a server or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing a data annotation method based on graph structure and community discovery is shown. As Figure 1 shown, the computing device may include one or more processors (the processors may include, but are not limited to, processing devices such as microprocessor MCUs or programmable logic devices FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. Among them, the memory, the transmission device, and the input / output interface are connected to the processor through a bus. In addition, it may further include: a display, a keyboard, and a cursor control device connected to the input / output interface. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include moreFigure 1 more or fewer components shown, or having a configuration different from that Figure 1 shown.
[0015] It should be noted that one or more of the above-mentioned processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0016] The memory can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data annotation method based on graph structure and community discovery in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizes the data annotation method based on graph structure and community discovery of the above application program. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include memories remotely located relative to the processor, and these remote memories can be connected to the computing device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0017] The transmission device is used to receive or send data via a network. Specific examples of the above network can include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0018] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computing device.
[0019] It should be noted here that in some alternative embodiments, the above Figure 1 shown computing device can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1This is just an example of a specific concrete instance and is intended to illustrate the types of components that may exist in the above computing device.
[0020] Figure 2 is a schematic diagram of a data annotation system based on graph structure and community discovery according to the present embodiment. Refer to Figure 2 As shown, the system includes: a server 300 of a data annotation platform. Among them, each terminal device (such as terminal device 200) running a client of the data annotation platform can be communicatively connected to the server 300. The server 300 can be communicatively connected to a database 400 in the data annotation platform, supporting reading or storing corresponding data from the database 400.
[0021] A user or business party 100 with data annotation requirements can send a data annotation request with multi-dimensional data (including customer transaction data, product attribute data, customer behavior trajectory data, and environmental data) to the server 300 through a terminal device (such as terminal device 200) running a client of the data annotation platform. The server 300 can clean the multi-dimensional data in the request and store it in the database 400 for subsequent reading when processing. In addition, after the server 300 completes the annotation of the customer group, it can store the intermediate data and the final annotation result generated during the annotation process in the database 400 for query by the terminal device, so that the user of the terminal device can subsequently perform downstream task analysis or model training based on the final annotation result obtained from the server 300.
[0022] It should be noted that the server 300 of the data annotation platform can be applicable to the above-mentioned hardware structure.
[0023] Under the above operating environment, according to the first aspect of the present embodiment, a data annotation method based on graph structure and community discovery is provided. This method is implemented by Figure 2 the server 300 shown in Figure 3 shows a schematic flow diagram of this method. Refer to Figure 3 As shown, this method includes: S302: Obtain multi-dimensional data of the customer group to be labeled with tags; wherein, the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; S304: Construct a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavioral semantics layer, and an environmental association layer. The basic physical layer is used to represent the transaction association relationship between customers and products and the physical attributes of products. The behavioral semantics layer is used to represent the potential behavior patterns of customers and the semantic association characteristics between products. The environmental association layer is used to represent the dynamic impact of external environmental factors on the interaction behavior between products and customers; and S306: Based on the heterogeneous graph network, divide the customer group into multiple customer communities through a community discovery algorithm, and generate a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes first-level labels, second-level labels, and third-level labels. The first-level labels are constructed based on the interaction patterns between the basic physical layer and the environmental association layer, and serve as the top-level nodes of the hierarchical label tree; the second-level labels are constructed based on the potential behavior patterns of customers within the customer community in the behavioral semantics layer, and serve as the middle-level nodes of the hierarchical label tree; the third-level labels are constructed based on the dynamic trajectory time series characteristics of customers within the customer community in the heterogeneous graph network, and serve as the bottom-level nodes of the hierarchical label tree.
[0024] In an embodiment of the present invention, when users or business parties in any business department (such as procurement, inventory, marketing, after-sales, etc.) have a need to label a certain customer group, they can send a data annotation request to the server 300 through a terminal device. The data annotation request includes multi-dimensional data of the customer group, providing data support for subsequent annotation processing. The multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data. Specifically, customer transaction data such as order ID, transaction amount, transaction timestamp, etc., product attribute data refers to information used to describe product characteristics, performance, functions, appearance, specifications, etc., customer behavior trajectory data such as the stay duration on the product detail page, add-to-cart events, etc., and environmental data such as temperature, passenger flow density, holidays, etc.
[0025] After the server 300 obtains the multi-dimensional data of the customer group, it is necessary to construct a heterogeneous graph network including a basic physical layer, a behavioral semantics layer, and an environmental association layer based on these data. This heterogeneous graph network aims to comprehensively depict the complex relationships among customers, products, and the environment. The specific construction steps are as follows: (1) The server 300 first defines the node types in the heterogeneous graph network, such as customer nodes, product nodes, environmental nodes, etc., and determines the attribute information of each node type. For example, customer nodes may include attributes such as customer ID, age, gender, etc.; product nodes may include attributes such as product ID, category, price, etc.
[0026] (2) Starting from the dimension of the transaction association relationship between customers and products and the mapping relationship of product physical attributes, the server 300 constructs the basic physical layer. Specifically, the server 300 constructs a basic physical layer for representing the transaction association relationship between customers and products and product physical attributes based on the node attributes of customer nodes, product nodes, and area nodes and the corresponding edge relationships.
[0027] (3) Starting from the dimension of the potential behavior patterns of customers and the semantic association features among products, the server 300 constructs a behavior semantics layer. Specifically, the server 300 mines the potential behavior patterns of customers, such as browsing habits and purchase preferences, through the analysis of customer behavior trajectory data, and uses these patterns as the node attributes of the behavior semantics layer. In addition, the server 300 also calculates the semantic association features among products, such as extracting keywords from product descriptions through text analysis and constructing semantic association edges among products. After that, the server 300 constructs the behavior semantics layer based on the node attributes of customer behavior nodes and product combination nodes, the corresponding edge relationships, the potential behavior patterns of customers, the semantic association features, and the semantic association edges among products.
[0028] (4) Starting from the dimension of the dynamic impact of external environmental factors on the interaction behavior between products and customers, the server 300 constructs an environment association layer. Specifically, the server 300 integrates environmental data, such as weather, holidays, promotional activities, etc., and analyzes how these factors affect the interaction behavior between customers and products. For example, the server 300 can analyze the impact of rainy days on the customer arrival rate and product sales volume, and use this impact as the edge weight of the environment association layer. After that, the server 300 constructs the environment association layer based on the node attributes of environmental nodes, the corresponding edge relationships, and the edge weights.
[0029] (5) The server 300 connects the basic physical layer, the behavior semantics layer, and the environment association layer into an integrated heterogeneous graph network.
[0030] Through the above method, the server 300 successfully constructs a heterogeneous graph network that can comprehensively reflect the complex relationships among customers, products, and the environment, achieving the goal of uniformly mapping multi-dimensional data into the graph structure. This network not only expresses the complex relationships among customers, products, and the environment, but also supports cross-business department data fusion and multi-dimensional relationship modeling, providing strong support for subsequent data analysis and mining.
[0031] Immediately afterwards, based on the heterogeneous graph network, the server 300 divides the customer groups into multiple customer communities through the community discovery algorithm and generates a hierarchical label tree for each customer community. Specifically, the server 300 performs the following operations: (1) The server 300 selects a community discovery algorithm suitable for heterogeneous graph networks, such as existing deep learning-based community discovery algorithms (such as GraphSAGE, GAT) or traditional community discovery algorithms (such as Louvain, Leiden algorithms). These algorithms can consider the characteristics of different node types and edge types in the heterogeneous graph network, thereby more accurately discovering customer communities. The server 300 adjusts algorithm parameters according to the characteristics of the heterogeneous graph network, such as the number of communities, the number of iterations, etc., to optimize the community discovery effect. Through the community discovery algorithm, the server 300 divides the customer group into multiple customer communities with internal similarity and external difference.
[0032] (2) For each discovered customer community, the server 300 generates a hierarchical label tree to comprehensively and deeply describe the characteristics of the customers within the community. Specifically, the server 300 extracts the basic value attributes and value fluctuation coefficients of the customer community based on the interaction patterns of the basic physical layer and the environmental association layer. These attributes are quantified as first-level labels, serving as the top-level nodes of the hierarchical label tree, reflecting the overall value characteristics of the customer community. The server 300 analyzes the interaction behavior patterns of the customers within the customer community with the products in the behavioral semantics layer, such as purchase preferences, browsing habits, etc. Through clustering analysis or pattern mining, the server 300 generates second-level labels, serving as the middle-level nodes of the hierarchical label tree, deeply depicting the behavioral characteristics of the customer community. The server 300 captures the dynamic trajectory time series characteristics of the customers within the customer community in the heterogeneous graph network, such as the change trend of purchase frequency, the evolution of behavior patterns, etc. These time series characteristics are transformed into third-level labels, serving as the bottom-level nodes of the hierarchical label tree, dynamically reflecting the change characteristics of the customer community.
[0033] Through the above steps, the server 300 not only realizes the community division of the customer group, but also generates a comprehensive and in-depth hierarchical label tree for each customer community as the final annotation result. Subsequently, downstream task analysis or model training can be performed based on this annotation result.
[0034] As described in the background technology, the existing customer data annotation methods have defects such as data islands caused by overly strong task orientation, dynamic value mining restricted by single data dimension, and hierarchical cognitive construction restricted by flat label systems.
[0035] In view of this, the present application first obtains multi-dimensional data such as customer transaction data, product attribute data, customer behavior trajectory data, and environmental data in the retail scenario, breaking the data fragmentation caused by department task differences in the traditional annotation process. Then, by constructing a heterogeneous graph network including a basic physical layer, a behavior semantics layer, and an environmental association layer, these data from different sources and different modalities are uniformly mapped into the graph structure to express the complex relationships among customers, products, and the environment in the form of a graph structure, so as to achieve the deep integration of multi-dimensional data and multi-dimensional relationship modeling in the retail scenario, providing a unified data foundation for comprehensively understanding customer behavior. Secondly, through the community discovery algorithm, customer groups with cross-business department business value are identified in the heterogeneous graph network. These groups are not defined based on the needs of a single department, but through the cross-layer connection characteristics in the graph network, naturally mapping out the common behavior patterns and value characteristics in different business scenarios, thus forming a unified cognitive framework for customers. This framework supports cross-business department applications and solves the problem that the same customer is labeled with fragmented labels in the traditional annotation system. Finally, based on the above cross-business department customer data, a hierarchical label tree is constructed to establish a hierarchical cognitive framework. This hierarchical label tree not only reflects the logical relationship between different labels, but also finely represents multiple dimensions and levels of customer value through the construction of multi-level labels. Thus, the technical effects of unified understanding of customer behavior, in-depth mining of multi-dimensional customer relationships, capture of the dynamic evolution of customer value, and construction of hierarchical customer cognition are achieved. Furthermore, the technical problems existing in the prior art, such as data islands caused by overly task-oriented traditional customer data annotation methods, dynamic value mining restricted by single data dimension, and hierarchical cognitive construction limited by flat label systems, are solved.
[0036] Optionally, the operation of constructing a heterogeneous graph network according to the multi-dimensional data includes: determining the node attributes of six predefined types of nodes based on the multi-dimensional data; where the six types of nodes include customer nodes, product nodes, area nodes, environmental nodes, customer behavior nodes, and product combination nodes; analyzing the multi-dimensional data to construct the edge relationships between the six types of nodes; based on the node attributes and corresponding edge relationships of the customer nodes, the product nodes, and the area nodes, constructing a basic physical layer for characterizing the transaction association relationship between customers and products and the physical attributes of products; on top of the basic physical layer, based on the node attributes and corresponding edge relationships of the customer behavior nodes and the product combination nodes, constructing a behavior semantics layer for characterizing the potential behavior patterns of customers and the semantic association features between products; on top of the behavior semantics layer, based on the node attributes and corresponding edge relationships of the environmental nodes, constructing an environmental association layer for characterizing the dynamic impact of external environmental factors on the interaction behavior between products and customers, obtaining a preliminary heterogeneous graph network; and establishing a cross-layer interaction mechanism for the preliminary heterogeneous graph network and performing multi-granularity attribute fusion processing to generate a final heterogeneous graph network.
[0037] In an embodiment of the present invention, the server 300 may pre-define six types of nodes, namely, customer nodes, commodity nodes, area nodes, environment nodes, customer behavior nodes, and commodity combination nodes. Then, the server 300 analyzes the multi-dimensional data to determine the node attributes of these six types of nodes. Specifically, the server 300 analyzes the customer transaction data and extracts customer basic information (such as ID, age, gender, membership, etc.) as the node attributes of the customer nodes. The server 300 analyzes the commodity attribute data and extracts information such as commodity ID, name, category, price, size, weight, etc. as the node attributes of the commodity nodes. The server 300 analyzes the customer behavior trajectory data and extracts information such as the ID, location, area, and area sales of the area where the customer is located as the node attributes of the area nodes. The server 300 analyzes the environment data and extracts environmental characteristics (such as temperature, humidity, passenger flow density, promotion activity ID, time, discount rate, etc.) as the node attributes of the environment nodes. The server 300 analyzes the customer behavior trajectory data and extracts customer behavior characteristics (such as click, browse, add-to-cart event, behavior trajectory movement line, stay duration, etc.) as the node attributes of the customer behavior nodes. The server 300 analyzes the commodity attribute data and extracts the characteristics of frequently co-occurring commodity combinations (such as commodity co-occurrence frequency, complementarity coefficient, and combination recommendation weight) as the node attributes of the commodity combination nodes.
[0038] Next, the server 300 also needs to further analyze the multi-dimensional data and construct the edge relationships between the above six types of nodes. Specifically, the server 300 further analyzes the customer transaction data to establish a transaction association edge between the customer and the product, and the edge weight is related to, for example, the transaction amount and frequency. The server 300 further analyzes the product attribute data to establish an attribute association edge between the products, and the edge weight is related to, for example, the product similarity. The server 300 further analyzes the customer behavior trajectory data to establish a behavior sequence edge between the customer behavior node and the customer node, and the edge weight is related to, for example, the frequency of occurrence or conversion rate of the behavior. The server 300 further analyzes the customer transaction data and the product attribute data, extracts the product combination sales data therefrom, and then analyzes the product combination sales data to establish a combination association edge between the product combination node and the product node, and the edge weight is related to, for example, the contribution degree of the product participating in the combined sales. The server 300 jointly analyzes the environmental data, the customer transaction data, and the product attribute data, extracts the environmental impact information among the three, and constructs an environmental impact edge between the environmental node and the customer node, and an environmental impact edge between the product nodes according to the environmental impact information. The edge weight is related to, for example, the degree of influence of environmental factors on customer behavior or product sales. The server 300 jointly analyzes the environmental data, the customer transaction data, and the product attribute data, extracts the seasonal change information among the three, and constructs an environmental adjustment edge between the environmental node and the product combination node according to the seasonal change information. The edge weight is related to, for example, the influence of the season factor on the combined sales of products.
[0039] Subsequently, based on the node attributes and the corresponding edge relationships of the customer node, the product node, and the regional node, the server 300 constructs a basic physical layer for characterizing the transaction association relationship between the customer and the product and the physical attributes of the product. Specifically, the server 300 constructs a transaction association network between the customer and the product according to the transaction edge relationship between the customer node and the product node as the preliminary basic physical layer. The weight of the transaction edge can reflect key indicators such as the transaction amount and frequency to quantify the transaction intensity between the customer and the product. The server 300 maps the physical attributes of the product node (such as size, color, material, etc.) into the basic physical layer, and these attributes can be connected to the product node through the attribute edge to form a mapping network of the product physical attributes. For example, the size attribute of the product can be represented by a virtual node and connected to the specific product node through the attribute edge. The server 300 integrates the sales characteristics of the regional node (such as regional sales volume, inventory, etc.) into the basic physical layer according to the sales edge relationship between the regional node and the product node. The weight of the sales edge can reflect the sales situation of the product in different regions. Thus, a basic physical layer that can accurately characterize the transaction association relationship between the customer and the product and the physical attributes of the product is constructed.
[0040] Subsequently, based on the node attributes and corresponding edge relationships of the customer behavior nodes and the product combination nodes, the server 300 constructs a behavior semantics layer on top of the basic physical layer to represent the potential customer behavior patterns and the semantic association features between products. Specifically, according to the behavior sequence edge relationships between the customer behavior nodes and the product combination nodes, the server 300 constructs a potential customer behavior pattern network on top of the basic physical layer as a preliminary behavior semantics layer. The weight of the behavior sequence edge can reflect the frequency of occurrence, conversion rate, or temporal importance of the behavior to quantify the association strength between the customer behavior and the product combination. The server 300 maps the semantic association features of the product combination nodes (such as product co-occurrence frequency, combined recommendation weight, etc.) into the behavior semantics layer, and these features can be connected to the product combination nodes through semantic association edges to form a mapping network of semantic associations between products. For example, a product combination that is frequently purchased together can be represented by a virtual node and connected to the specific product combination node through a semantic association edge to reflect the semantic association features between products. Thus, a behavior semantics layer that can accurately represent the potential customer behavior patterns and the semantic association features between products is constructed.
[0041] Subsequently, based on the node attributes and corresponding edge relationships of the environment nodes, the server 300 constructs an environment association layer on top of the behavior semantics layer to represent the dynamic impact of external environmental factors on the interaction behavior between products and customers, obtaining a preliminary heterogeneous graph network. Specifically, according to the known environment node attributes (such as temperature, humidity, promotion activity identifier, etc.) and the preset edge relationship rules (such as the association strength between the "high-temperature environment" and "cold drink products"), the server 300 maps the impact of environmental factors on the customer behavior pattern (such as the change rate of purchase frequency) into an environment impact edge through weight configuration, and at the same time maps the regulatory effect of environmental factors on the combined product sales (such as the promotion coefficient of combined products during holidays) into an environment adjustment edge. Then, on top of the behavior semantics layer, according to the node attributes of the environment nodes and the corresponding environment impact edges and environment adjustment edges, the server 300 constructs the environment association layer.
[0042] Finally, the server 300 establishes a cross-layer interaction mechanism for the preliminary heterogeneous graph network and performs multi-granularity attribute fusion processing to generate the final heterogeneous graph network. Specifically, the server 300 realizes the information interaction between the basic physical layer, the behavior semantics layer, and the environment association layer through the preset cross-layer connection rules (such as the indirect association path between the customer node and the product combination node), and adopts a dynamic weight adjustment strategy (such as real-time adjustment of the behavior sequence edge weight according to environmental changes) to enhance the adaptability of the network to complex scenarios. At the attribute fusion level, the server 300 performs weighted fusion on the node attributes (such as transaction amount, behavior frequency), and finally generates a heterogeneous graph network with cross-layer interaction and multi-granularity attribute fusion capabilities.
[0043] In the above manner, a heterogeneous graph network that can comprehensively reflect the complex relationships among customers, products, and the environment is constructed, achieving the goal of uniformly mapping multi-dimensional data into a graph structure. This network not only expresses the complex relationships among customers, products, and the environment but also supports cross-business department data fusion and multi-dimensional relationship modeling, providing strong support for subsequent data analysis and mining.
[0044] Optionally, the operation of determining the node attributes of six predefined types of nodes based on the multi-dimensional data includes: extracting customer static features from the customer transaction data as the node attributes of the customer nodes; wherein, the customer static features include customer ID, membership level, and life cycle stage; extracting product static features from the product attribute data as the node attributes of the product nodes; wherein, the product static features include product ID, physical attributes, and price range; extracting regional features from the customer behavior trajectory data as the node attributes of the regional nodes; wherein, the regional features include sales area coordinates and delivery scope; extracting environmental features from the environmental data as the node attributes of the environmental nodes; wherein, the environmental features include time dimension, weather condition, and season factor; extracting customer behavior features from the customer behavior trajectory data as the node attributes of the customer behavior nodes; wherein, the customer behavior features include behavior type, behavior timestamp, and stay duration; and extracting product association features from the product attribute data as the node attributes of the product combination nodes; wherein, the product association features include product co-occurrence frequency, complementarity coefficient, and combination recommendation weight.
[0045] In an embodiment of the present invention, the server 300 first extracts customer static features from the customer transaction data as the node attributes of the customer nodes. Specifically, the server 300 extracts the static features of the customer from the customer transaction data, such as customer ID, membership level, and life cycle stage, and directly associates these features with the identity and status of the customer as the node attributes of the customer nodes.
[0046] Next, the server 300 extracts product static features from the product attribute data as the node attributes of the product nodes. Specifically, the server 300 extracts the static features of the product from the product attribute data, including product ID, physical attributes, and price range, etc., as the node attributes of the product nodes. These features describe the basic attributes and value range of the product and serve as the attribute information of the product nodes.
[0047] Then, the server 300 extracts regional features from the customer behavior trajectory data as the node attributes of the regional nodes. Specifically, the server 300 analyzes the customer behavior trajectory data and extracts regional features such as sales area coordinates and delivery ranges as the node attributes of the regional nodes. These features define the spatial attributes and sales coverage capabilities of the regions, providing basic attribute information for the regional nodes.
[0048] Immediately afterwards, the server 300 extracts environmental features from the environmental data as the node attributes of the environmental nodes. Specifically, the server 300 extracts features such as time dimension, weather conditions, and seasonal factors from the environmental data as the node attributes of the environmental nodes. These features reflect the changing trends of the external environment, providing key attribute information that affects customer behavior and commodity sales for the environmental nodes.
[0049] In addition, the server 300 also extracts customer behavior features from the customer behavior trajectory data as the node attributes of the customer behavior nodes. Specifically, the server 300 analyzes the customer behavior trajectories and extracts features such as behavior types, behavior timestamps, and stay durations as the node attributes of the customer behavior nodes. These features record the behavior patterns and preferences of the customers, providing detailed attribute information for the customer behavior nodes.
[0050] Finally, the server 300 extracts commodity association features from the commodity attribute data as the node attributes of the commodity combination nodes. Specifically, the server 300 analyzes the commodity attribute data and extracts association features such as commodity co-occurrence frequencies, complementarity coefficients, and combination recommendation weights as the node attributes of the commodity combination nodes. These features reveal the internal relationships and combined sales potential among the commodities, providing important attribute information for the commodity combination nodes.
[0051] In the above manner, the server 300 accurately characterizes the six types of node attributes based on multi-dimensional data, achieving the goal of extracting structured features from transaction behaviors, commodity attributes, spatial regions, environmental dynamics, user behavior trajectories, and commodity association relationships. This process not only assigns multi-dimensional attribute labels to customer nodes, commodity nodes, regional nodes, environmental nodes, customer behavior nodes, and commodity combination nodes, but also transforms the scattered data into graph-structured attribute information through feature engineering, laying a data foundation for constructing a heterogeneous graph network that integrates spatio-temporal semantics and behavior patterns.
[0052] Optionally, the operation of analyzing the multi-dimensional data and constructing the edge relationships between the six types of nodes includes: analyzing the transaction frequency and amount in the customer transaction data, and constructing a transaction edge between the customer node and the product node; wherein, the transaction edge represents the weighted sum of the transaction frequency and the amount; analyzing the physical attributes and category information in the product attribute data, and constructing an attribute association edge between the product nodes; wherein, the attribute association edge represents the similarity of product attributes; analyzing the behavior sequence information in the customer behavior trajectory data, and constructing a behavior sequence edge between the customer behavior node and the customer node; wherein, the behavior sequence edge represents the frequency or conversion rate of the behavior occurrence; analyzing the product combination sales data between the customer transaction data and the product attribute data, and constructing a combination association edge between the product combination node and the product node; wherein, the combination association edge represents the contribution degree of the product participating in the combination sales; analyzing the environmental impact information among the environmental data, the customer transaction data, and the product attribute data, and constructing an environmental impact edge between the environment node and the customer node and the product node; wherein, the environmental impact edge represents the degree of influence of environmental factors on customer behavior or product sales; analyzing the seasonal change information among the environmental data, the customer transaction data, and the product attribute data, and constructing an environmental adjustment edge between the environment node and the product combination node; wherein, the environmental adjustment edge represents the influence of the season factor on the product combination sales.
[0053] In an embodiment of the present invention, the server 300 first analyzes the transaction frequency and amount in the customer transaction data, and constructs a transaction edge between the customer node and the product node. Specifically, the server 300 deeply analyzes the customer transaction data, accurately captures the two core indicators of transaction frequency and amount, and combines them in a weighted sum manner to construct a transaction edge between the customer node and the product node. This edge not only intuitively reflects the direct transaction connection between the customer and the product, but also reveals the transaction activity and value magnitude through the quantified weighted sum, laying a solid foundation for subsequent transaction behavior analysis.
[0054] Next, the server 300 analyzes the physical attributes and category information in the product attribute data, and constructs an attribute association edge between the product nodes. Specifically, the server 300 extracts the physical attributes and category information from the product attribute data, and constructs an attribute association edge between the product nodes by calculating the attribute similarity between products. This edge effectively reveals the potential internal connections between products, and provides strong support for application scenarios such as product recommendation and inventory optimization based on the similarity of products in physical attributes and categories.
[0055] Then, the server 300 analyzes the behavioral sequence information in the customer behavior trajectory data and constructs a behavioral sequence edge between the customer behavior node and the customer node. Specifically, the server 300 deeply mines the customer behavior trajectory data, accurately extracts the behavioral sequence information, and constructs a behavioral sequence edge between the customer behavior node and the customer node by quantifying the frequency or conversion rate of behavior occurrence. This edge details the customer's behavior pattern and provides key clues for understanding customer preferences and predicting purchase intentions.
[0056] Immediately afterwards, the server 300 analyzes the product combination sales data between the customer transaction data and the product attribute data and constructs a combined association edge between the product combination node and the product node. Specifically, the server 300 combines the customer transaction data and the product attribute data, deeply analyzes the product combination sales data, and constructs a combined association edge between the product combination node and the product node by calculating the contribution degree of the product's participation in the combined sales. This edge effectively reveals the combined sales potential between products and provides an important basis for formulating strategies such as product combination recommendations and bundling sales.
[0057] In addition, the server 300 analyzes the environmental impact information among the environmental data, the customer transaction data, and the product attribute data and constructs an environmental impact edge between the environmental node and the customer node and the product node. Specifically, the server 300 comprehensively considers the environmental data, the customer transaction data, and the product attribute data, extracts the environmental impact information, and constructs an environmental impact edge between the environmental node and the customer node and the product node by quantifying the impact degree of environmental factors on customer behavior or product sales. This edge deeply reflects the regulatory role of the external environment on the relationship between customers and products and provides strong support for formulating environmentally adaptive marketing strategies.
[0058] Finally, the server 300 analyzes the seasonal change information among the environmental data, the customer transaction data, and the product attribute data and constructs an environmental adjustment edge between the environmental node and the product combination node. Specifically, the server 300 particularly focuses on the seasonal change information in the environmental data, the customer transaction data, and the product attribute data and constructs an environmental adjustment edge between the environmental node and the product combination node by analyzing the impact of seasonal factors on product combination sales. This edge effectively reveals the regulatory role of seasonal changes on the product combination sales model and provides an important basis for formulating seasonal marketing strategies.
[0059] In the above manner, the server 300 constructs complex edge relationships between six types of nodes based on multi-dimensional data. It not only realizes the in-depth characterization of the customer-product relationship from multiple perspectives such as transaction behavior, product attributes, and environmental dynamics, but also injects rich business logic interpretability into the heterogeneous graph network through semantic indicators such as weighted sum, similarity, frequency, and influence degree. The establishment of these edge relationships enables the heterogeneous graph network to more comprehensively reflect the complex associations in the real world and provides strong support for subsequent data analysis and mining.
[0060] Optionally, the operation of establishing a cross-layer interaction mechanism for the preliminary heterogeneous graph network and performing multi-granularity attribute fusion processing to generate a final heterogeneous graph network includes: establishing a cross-layer interaction mechanism between the basic physical layer, the behavior semantic layer, and the environmental association layer in the preliminary heterogeneous graph network through hyperedges or meta-paths; fusing the static attributes of the customer nodes and the dynamic attributes of the customer behavior nodes in the preliminary heterogeneous graph network to generate a comprehensive customer feature attribute, and adding the comprehensive customer feature attribute to the node attributes of the customer nodes; determining the semantic association feature attributes between products according to the node attributes of the product nodes and the node attributes of the product combination nodes, and adding the semantic association feature attributes to the node attributes of the product nodes; and dynamically adjusting the edge weights in the preliminary heterogeneous graph network according to the node attributes of the environmental nodes to obtain the final heterogeneous graph network.
[0061] In an embodiment of the present invention, the server 300 first establishes a cross-layer interaction mechanism between the basic physical layer, the behavioral semantics layer, and the environmental association layer in the preliminary heterogeneous graph network through hyperedges or meta-paths. Specifically, the server 300 uses hyperedges or meta-paths as bridges to cleverly connect the basic physical layer, the behavioral semantics layer, and the environmental association layer, realizing cross-layer interaction. Among them, a hyperedge, as a special edge that can connect multiple nodes, the server 300 connects entities in the basic physical layer (such as customers, goods), behavioral patterns in the behavioral semantics layer (such as purchase, browsing), and environmental factors in the environmental association layer through hyperedges, breaking through the limitation that an edge in a traditional graph network only connects two nodes, and more realizing the organic integration of multi-layer information. For example, in a retail scenario, when customer A purchases product X on an e-commerce platform in rainy weather, the server 300 can construct a cross-layer hyperedge: one end of this hyperedge connects the customer node in the basic physical layer (attributes include the ID and membership level of customer A) and the product node (attributes include the category and price of product X), and the other end is associated with the customer behavior node in the behavioral semantics layer (recording the "purchase" behavior type and timestamp), and at the same time penetrates to the environmental association layer to connect the environmental node (marking the "rainy" weather condition and the seasonal factor of the day). Furthermore, this hyperedge can also be attached with dynamic attributes, such as the transaction amount, the influence coefficient of the delivery time limit, and the stimulation index of the demand for the category to which product X belongs caused by the rain. Through this design, a single hyperedge realizes the complete modeling of the complex event of "a specific customer generates a specific transaction behavior under the drive of a specific environment". Compared with the traditional two-node edge, the hyperedge encapsulates the three elements of physical entities, behavioral semantics, and environmental impacts into a computable relationship unit through multi-dimensional attribute aggregation and cross-layer node connection, providing a structured context carrier for subsequent customer community mining.
[0062] The server 300 can also define the connection mode between nodes through specific types of edges and node sequences by introducing meta-paths, thereby explicitly encoding cross-layer semantic associations. For example, by constructing a meta-path of "customer - product - environment - product", the server 300 can capture the deep behavioral pattern of customers purchasing products under specific environmental conditions. This path not only spans the physical layer (customers, products) and the environmental layer (environmental nodes), but also establishes a logical association through the semantic layer (purchase behavior). The introduction of meta-paths enables the server 300 to accurately depict the structured semantic relationships in the network. For example, through the meta-path of "product - combined association edge - product combination node - environmental adjustment edge - environmental node", the influence weight of seasonal factors on the combined sales of products can be quantified, providing an interpretable path basis for cross-layer knowledge reasoning. This mechanism not only breaks the single connection mode between nodes in a traditional graph network, but also enables the heterogeneous graph network to more comprehensively capture the complex relationships among customers, products, and the environment by introducing higher-level abstract connections, laying a solid foundation for subsequent multi-granularity attribute fusion processing.
[0063] Next, the server 300 performs a fusion process on the static attributes of the customer nodes and the dynamic attributes of the customer behavior nodes in the preliminary heterogeneous graph network to generate customer comprehensive feature attributes, and adds the customer comprehensive feature attributes to the node attributes of the customer nodes. Specifically, the server 300 first extracts the static attributes of the customer nodes from the basic physical layer, such as identity identification information like customer ID, membership level, life cycle stage, etc. At the same time, it parses the dynamic attributes of the customer behavior nodes from the behavioral semantics layer, including behavioral types such as browsing, adding to cart, and payment, as well as their corresponding time stamps, residence durations and other temporal features. To achieve cross-layer feature fusion, the server 300 uses an existing graph neural network model based on the attention mechanism, which can dynamically calculate the semantic association weights between static attributes and dynamic behaviors. For example, for the customer node 'CUST_001', the model can identify the strong correlation between its 'platinum member' level and the 'high-frequency night payment' behavior, and generate a comprehensive feature vector containing the label of 'high-value night active user' through weighted summation. Furthermore, this fusion process also introduces the real-time context information of the environment association layer. For example, during the 'Double Eleven' promotion period, the model will increase the weight of the 'adding to cart behavior' to reflect the promotion sensitivity. The finally generated customer comprehensive feature attributes include both a multi-dimensional label system of'membership level - category preference - environment sensitivity', and quantify the implicit association between customer value and behavior tendency through vector embedding technology. The server 300 injects this high-order feature into the attribute matrix of the customer nodes, upgrading the node representation from a single identity identification to a three-dimensional portrait including consumption ability, behavior pattern, and environment response. For example, the updated customer node attributes can support the graph neural network in the link prediction task, simultaneously considering whether a certain customer belongs to a high-value group (static attribute), whether their recent browsing behavior points to the target product (dynamic attribute), and whether it is currently in the seasonal promotion period (environmental attribute), thereby improving the recommendation accuracy. This attribute fusion mechanism not only enhances the semantic richness of the node representation, but also provides a structured feature foundation for subsequent precision marketing and relationship reasoning based on the heterogeneous graph network.
[0064] Then, the server 300 determines the semantic association feature attributes between products based on the node attributes of the product nodes and the node attributes of the product combination nodes, and adds the semantic association feature attributes to the node attributes of the product nodes. Specifically, the server 300 first extracts the inherent attributes of the product nodes from the basic physical layer, including structured information such as product ID, category label, physical attributes (such as size, color), and price range. At the same time, it analyzes the association attributes of the product combination nodes from the behavioral semantics layer. For example, by analyzing the co-occurrence frequency of products in customer transaction data, the complementarity coefficient (such as the combined purchase ratio of 'toothpaste' and 'toothbrush'), and the combined recommendation weight (the potential for cross-selling calculated based on the collaborative filtering algorithm). To achieve cross-layer semantic association, the server 300 uses an existing algorithm model that can quantify the semantic association strength between product nodes and product combination nodes. For example, for the product node 'PROD_001', the model calculates its co-occurrence frequency with other products in the'summer promotion combination' (such as the combined purchase rate of'sunscreen' and'sunglasses' increases by 30%), and combines the seasonal factor in the environmental association layer (such as the increased recommendation weight for cooling products in summer) to generate a semantic association feature vector containing labels such as 'high cross-selling potential - seasonally sensitive products'. Furthermore, the model also incorporates real-time feedback information from customer behavior trajectory data. For example, in a certain marketing campaign, if the click-through rate of the combination of 'PROD_001' and 'PROD_002' increases by 25%, the semantic association weight is dynamically adjusted to reflect the market response. The finally generated semantic association feature attributes include both a multi-dimensional label system of 'category similarity - complementarity coefficient - environmental adaptability' and quantify the implicit association rules between products through a graph attention network (GAT). The server 300 injects this high-order feature into the attribute matrix of the product nodes, upgrading the node representation from a single physical attribute to a three-dimensional portrait including cross-selling potential, seasonal trends, and market response. For example, the updated product node attributes can support the heterogeneous graph network in link prediction tasks, simultaneously considering whether a product belongs to a high-repurchase category (static attribute), its historical combined sales performance (dynamic attribute), and whether it is currently in a seasonal clearance period (environmental attribute), thereby increasing the accuracy of combined recommendations by more than 18%. This fusion mechanism of semantic association feature attributes not only enhances the semantic richness of product node representations but also provides a structured knowledge base for subsequent accurate product selection, dynamic pricing, and associated marketing based on heterogeneous graph networks.
[0065] Finally, the server 300 dynamically adjusts the edge weights in the preliminary heterogeneous graph network according to the node attributes of the environmental nodes to obtain the final heterogeneous graph network. Specifically, the server 300 first extracts multi-dimensional environmental signals from the environmental association layer, including structured features such as the time dimension (e.g., weekday / weekend, holiday identifier), weather status (e.g., heavy rain, high temperature warning), seasonal factor (e.g., spring / autumn mode), and market dynamics (e.g., promotion activity cycle). To achieve weight adjustment for environmental perception, the server 300 deploys a dynamic weight adjustment algorithm, which realizes fine-grained regulation through a three-layer mechanism: The first layer is the rule engine layer, which preliminarily adjusts the edge weights based on preset threshold rules. For example, when the 'heavy rain' weather status is detected, the algorithm automatically reduces the timeliness weight of the edge between the delivery area node and the commodity node by 30%, and at the same time increases the recommendation weight of the 'flood control materials' commodity combination edge within the area node by 50%. The second layer is the machine learning layer, which uses the existing spatio-temporal graph convolutional network (ST-GCN) to model historical environment-behavior data and capture complex environmental impact patterns. For example, by analyzing the transaction data during the '618' promotion period in the past three years, the model learns that the attention of customers to the commodity combination node during the promotion period shows spatio-temporal characteristics of'slow climb during the preheating period - exponential growth during the explosion period', and accordingly dynamically adjusts the weight decay coefficient of the promotion-related edges. The third layer is the real-time feedback layer, which continuously calibrates the weight parameters through an online learning mechanism. For example, when a 'high temperature warning' suddenly occurs in a certain area node, the system monitors the change in the click-through rate of the 'cold drink' commodity in the customer behavior nodes in this area in real time. If the click-through rate increases by 20% within 10 minutes, the weight enhancement mechanism for the edge between the commodity node and the area node is immediately triggered. This multi-level adjustment architecture not only realizes the direct impact modeling of environmental factors on customer behavior (e.g., 'heavy rain weather inhibits offline store traffic') and commodity sales (e.g.,'summer promotion increases the weight of the air conditioner category combination'), but also reconstructs the distribution of the association strength between nodes in the heterogeneous graph network through the spatio-temporal dynamic change of the edge weights. For example, after dynamic adjustment, the transaction edge weight between the customer node 'CUST_001' and the commodity node 'PROD_001' may increase by 40% due to the 'Double 11' promotion environmental factor, while the delivery edge weight between this commodity node and the area node 'AREA_001' decreases by 25% due to the 'logistics congestion warning'. Finally, through this weight optimization for environmental perception, the server 300 constructs a heterogeneous graph network with adaptive capabilities. Its edge weight matrix not only contains static business rules, but also embeds dynamic environmental semantics, resulting in an 18% increase in the accuracy of the link prediction task and a 22% increase in the conversion rate of the combined recommendation system. This dynamic adjustment mechanism provides a more interpretable network foundation for subsequent advanced graph analysis tasks such as causal reasoning and anomaly detection based on spatio-temporal context.
[0066] In the above manner, the server 300 deeply optimizes and upgrades the preliminary heterogeneous graph network. By establishing a cross-layer interaction mechanism, integrating multi-granularity attributes, and dynamically adjusting edge weights, etc., the server 300 successfully generates a heterogeneous graph network that can comprehensively reflect the complex relationships among customers, commodities, and the environment. This network not only enhances the expression ability of node attributes and the semantic richness of edge relationships, but also adapts to changes in the external environment through a dynamic adjustment mechanism, providing more accurate and comprehensive graph-structured data support for subsequent data analysis and mining.
[0067] Optionally, the operation of dividing the customer groups into multiple customer communities based on the heterogeneous graph network includes: performing community division based on transaction association at the basic physical layer of the heterogeneous graph network to form an initial community centered on commodity physical attributes and transaction frequencies; performing community optimization based on behavior patterns at the behavioral semantic layer of the heterogeneous graph network, and adjusting the community boundaries of the initial community by combining customer behavior sequences and commodity combination association characteristics; and introducing dynamic environmental factors at the environmental association layer of the heterogeneous graph network to perform temporal sensitivity correction on the adjusted initial community to obtain the corresponding customer communities.
[0068] In the embodiment of the present invention, the server 300 performs coarse-grained community division based on transaction association at the basic physical layer of the heterogeneous graph network to form an initial community centered on commodity physical attributes and transaction frequencies. Specifically, the server 300 can adopt the existing Louvain algorithm based on modularity optimization to process customer nodes, commodity nodes, and transaction edges in the basic physical layer. For example, in the maternal and child products scenario, the server 300 extracts the physical attributes of commodity nodes (such as the category code is "infant milk powder" and the price range is marked as "500 - 800 yuan") and transaction frequencies (such as "customer A purchased a certain milk powder 5 times in the past 30 days"), and quantifies the community cohesion through the existing modularity calculation formula. This algorithm iteratively optimizes the community structure, aggregates node pairs with close transaction associations into an initial community, such as aggregating "customers who frequently purchase the same brand of milk powder" and "milk powder commodities with a high repurchase rate", and excluding low-frequency transaction nodes at the same time. The initial community generated in this step is oriented towards commodity physical attributes, but does not consider customer behavior patterns and environmental dynamics.
[0069] At the behavioral semantics layer of the heterogeneous graph network, the server 300 performs community optimization based on behavioral patterns, and adjusts the community boundaries of the initial community by combining customer behavior sequences and commodity combination association features. Specifically, the server 300 deploys an existing graph attention network combined with LSTM sequence encoding to perform temporal modeling on the customer behavior node sequence (such as the "click - add to cart - payment" path), and generates a behavioral pattern hidden vector. For example, if some customers in a community show a behavioral pattern of "comparing prices frequently and then buying low - price goods" (the cosine distance of the behavioral hidden vector encoded by LSTM exceeds 0.75), while another part of the customers show the characteristic of "making quick decisions to buy high - end goods", the algorithm will adjust the community boundary through the graph attention mechanism and split the original community into two sub - communities. At the same time, the server 300 uses the existing Apriori algorithm to mine high - frequency commodity combinations (such as the combination purchase frequency of "diapers + wet wipes" exceeds the threshold), and quantifies the combination association strength as the semantic edge weight between commodity nodes to further optimize the community structure. This step enables the community division to distinguish customer groups with similar physical attributes but very different behavioral patterns by introducing behavioral semantic information.
[0070] At the environmental association layer of the heterogeneous graph network, the server 300 introduces dynamic environmental factors to perform temporal sensitivity correction on the adjusted initial community to obtain the final customer community. Specifically, the server 300 uses an existing dynamic graph attention network (DGAT) to convert environmental node attributes (such as holiday identifiers, weather conditions, promotion intensity) into temporal feature vectors, and dynamically adjusts the influence of historical environmental factors through a time decay function (such as , where, is the weight at time t, is the weight size at the initial moment (t = 0), γ is the decay coefficient, which is a constant between 0 and 1, is the time interval from the initial moment to the current moment). For example, in heavy rain weather, the algorithm may increase the edge weight between the "flood control supplies" commodity node and the regional distribution node, and at the same time reduce the cohesion within the "outdoor supplies" commodity node community. In addition, the server 300 also uses an online learning mechanism to monitor environmental changes in real - time (such as when a "high - temperature warning" suddenly occurs in a certain area, the system monitors the change in the click - through rate of "cold drink" commodities in that area in real - time). If the click - through rate increases by 20% within 10 minutes, the edge weight enhancement mechanism between the commodity node and the regional node will be immediately triggered. This step enables the community structure to adapt to external changes. For example, during the summer promotion period, customers originally belonging to the "daily necessities community" may temporarily join the "cool household appliances community" due to the drive of environmental factors.
[0071] In the above manner, the server 300 realizes the transformation from the coarse-grained division driven by physical attributes, to the fine-grained optimization driven by behavioral semantics, and then to the temporal correction of environmental perception, finally generating a customer community with static features, dynamic behaviors, and environmental adaptability. This hierarchical processing mechanism not only improves the accuracy of community division but also enhances the response ability of the community structure to changes in business scenarios, providing high-quality group portrait support for downstream tasks such as precision marketing and personalized recommendation.
[0072] Optionally, the operation of generating the hierarchical label tree for each customer community includes: determining the basic value attributes of each customer community based on the basic physical layer; determining the value fluctuation coefficient of each customer community based on the environmental association layer; determining the first-level labels of each customer community according to the basic value attributes and the value fluctuation coefficient, as the top nodes of the hierarchical label tree; determining the behavioral sequence information of customers within each customer community based on the behavioral semantics layer; determining the product purchase situation of customers within each customer community based on the product nodes and product combination nodes in the heterogeneous graph network; determining the second-level labels of each customer community according to the behavioral sequence information and the product purchase situation, as the middle-layer nodes of the hierarchical label tree; and calculating the temporal feature indicators of customers within each customer community based on the basic physical layer, the behavioral semantics layer, and the environmental association layer, and determining the third-level labels of each customer community according to the temporal feature indicators, as the bottom nodes of the hierarchical label tree; where the temporal feature indicators include the volatility of the purchase cycle and the change rate of the activity participation frequency.
[0073] In an embodiment of the present invention, the server 300 takes a customer community of maternal and child products as an example to describe in detail the generation of the hierarchical label tree for this customer community of maternal and child products: (1) The server 300 extracts the basic value attributes of a customer community for a certain mother and baby product from the basic physical layer. For example, for the "milk powder purchase community", the server 300 uses the existing RFM model to quantify the "number of days since the last milk powder purchase (R value = 15 days)", "frequency of diaper purchases in the past 90 days (F value = 8 times)", and "average customer unit price (M value = 680 yuan)" of the customers in this community, and generates a basic value score through weighted summation (such as the score of this community = 0.4×15 + 0.3×8 + 0.3×680 = 216.2). At the same time, the server 300 imports dynamic environmental factors from the environmental correlation layer, such as the "summer promotion period" identifier, and uses the existing GARCH model to calculate the value volatility coefficient. For example, on the first day of the promotion period, due to the sharp increase in the purchase volume of heatstroke prevention supplies (such as baby cooling mats) in this community, the value volatility coefficient soars to 1.8 (>1.5 threshold), triggering the top-level label correction mechanism. Finally, the server 300 fuses the basic value attributes with the value volatility coefficient (such as the top-level label = basic value score × volatility coefficient) to generate the first-level label. For example, this community is marked as the "high-value environment-sensitive mother and baby core group" due to its high value attributes (score 216.2) and high volatility coefficient (1.8).
[0074] Then, the server 300 extracts the behavioral sequence information of the customers in this mother and baby community from the behavioral semantics layer. For example, through the existing Markov chain model, the server 300 monitors that 60% of the customers follow the behavioral chain of "browsing the milk powder details page → adding to the shopping cart → abandoning payment" (transition probability 0.6), while 40% of the customers complete the conversion path of "browsing complementary foods → receiving a full reduction coupon → combined payment" (transition probability 0.4). At the same time, the server 300 uses the existing Apriori algorithm to mine the combined purchase patterns of products and finds that the support degree of the combination of "milk powder + complementary foods" reaches 45% and the confidence level is 0.9, indicating that this combination has a strong correlation. Finally, the server 300 fuses the behavioral patterns with the combined product characteristics to generate the second-level label. For example, this community is marked as the "decision-making hesitant milk powder consumer group" due to the high proportion of payment interruption behaviors, and at the same time is marked as the "complementary food-related loyal customer" due to the high-frequency combined purchases.
[0075] Finally, the server 300 calculates the time-series feature metrics of the mother and baby community across the basic physical layer, behavioral semantics layer, and environmental association layer. For example, an existing ARIMA model is used to quantify the volatility of the purchase cycle, and it is detected that the variance of the purchase cycle in the community drops from 12 days to 3 days (a 75% decrease in volatility) 30 days before the "618 Big Promotion"; an existing exponential smoothing method is used to predict the change rate of activity participation frequency, and it is found that the click-through rate of the promotion page soars from 2% per day to 18% (a change rate of 800%). Finally, the server 300 encodes the time-series features into three-level labels. For example, this community is labeled as a "promotion-driven explosive mother and baby group" due to its low cycle volatility (σ = 3) and high activity response rate (Δ = 800%).
[0076] In the above manner, the server 300 constructs a hierarchical label tree containing three-level labels. For example, for the "high-value environment-sensitive mother and baby core group" (first-level label), the behavioral characteristics of its "decision-making hesitant milk powder consumer group + complementary food-related loyal customers" (second-level label) can be traced downwards, and finally the time-series response pattern of the "promotion-driven explosive mother and baby group" (third-level label) can be located. This structure not only realizes the multi-dimensional characterization from static value assessment to dynamic behavior analysis and then to time-series environmental response, but also provides an interpretable decision-making path for precision marketing. For example, during the summer promotion period, the system can quickly locate this community group, and based on the characteristics of its underlying labels, automatically trigger the push of a "milk powder buy-one-get-one-free + complementary food full reduction" combination coupon to improve the accuracy of recommendations. This hierarchical label system significantly enhances the business operability of customer segmentation and provides a standardized output framework for value mining in heterogeneous graph networks.
[0077] It should be added that in the above process, this embodiment only takes the mother and baby scenario as an example to illustrate the construction logic of the hierarchical label tree. The generation process of the customer community label system in other business fields follows the same technical framework, and its implementation path and core algorithm mechanism are universal, so it will not be elaborated in detail.
[0078] In addition, referring to Figure 1 As shown, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program runs, the method described in any one of the above is executed by a processor.
[0079] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0081] Embodiment 2 Figure 4 It shows a data annotation device based on graph structure and community discovery according to this embodiment, and this device corresponds to the method according to Embodiment 1. Refer to Figure 4 As shown, the device includes: a data acquisition module 410, configured to acquire multi-dimensional data of a customer group to be labeled with tags; wherein, the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; a graph construction module 420, configured to construct a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavior semantics layer, and an environmental association layer. The basic physical layer is used to represent the transaction association relationship between customers and products and the physical attributes of products. The behavior semantics layer is used to represent the potential behavior patterns of customers and the semantic association features between products. The environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between products and customers; and an annotation module 430, configured to divide the customer group into multiple customer communities based on the heterogeneous graph network through a community discovery algorithm, and generate a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes a first-level label, a second-level label, and a third-level label. The first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer, and serves as the top-level node of the hierarchical label tree. The second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavior semantics layer, and serves as the middle-level node of the hierarchical label tree. The third-level label is constructed based on the dynamic trajectory time series features of customers in the customer community in the heterogeneous graph network, and serves as the bottom-level node of the hierarchical label tree.
[0082] Thus, according to this embodiment, first, multi-dimensional data such as customer transaction data, product attribute data, customer behavior trajectory data, and environmental data in the retail scenario are obtained, breaking the data fragmentation caused by departmental task differences in the traditional annotation process. Then, by constructing a heterogeneous graph network including a basic physical layer, a behavior semantics layer, and an environmental association layer, these data from different sources and different modalities are uniformly mapped into the graph structure to express the complex relationships among customers, products, and the environment in the form of a graph structure, thereby realizing the deep integration of multi-dimensional data and multi-dimensional relationship modeling in the retail scenario and providing a unified data foundation for comprehensively understanding customer behavior. Secondly, through the community discovery algorithm, customer groups with cross-business department business value are identified in the heterogeneous graph network. These groups are not defined based on the needs of a single department, but through the cross-layer connection characteristics in the graph network, the common behavior patterns and value characteristics in different business scenarios are naturally mapped, thus forming a unified cognitive framework for customers. This framework supports cross-business department applications and solves the problem that the same customer is labeled with fragmented labels in the traditional annotation system. Finally, based on the above cross-business department customer data, a hierarchical label tree is constructed to establish a hierarchical cognitive framework. This hierarchical label tree not only reflects the logical relationship between different labels, but also finely characterizes multiple dimensions and levels of customer value through the construction of multi-level labels. Thus, the technical effects of unified understanding of customer behavior, deep mining of multi-dimensional customer relationships, capturing of the dynamic evolution of customer value, and construction of hierarchical customer cognition are achieved. Furthermore, the technical problems existing in the prior art, such as data islands caused by overly strong task orientation in the traditional customer data annotation method, dynamic value mining restricted by single data dimension, and hierarchical cognitive construction restricted by the flat label system, are solved.
[0083] Embodiment 3 Figure 5 shows a data annotation device based on a graph structure and community discovery according to this embodiment. This device corresponds to the method described in Embodiment 1. Refer to Figure 5As shown, the device includes: a processor 510; and a memory 520, connected to the processor 510, for providing instructions for the processor 510 to process the following processing steps: obtaining multi-dimensional data of a customer group for which labels are to be marked; wherein the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; constructing a heterogeneous graph network based on the multi-dimensional data; wherein the heterogeneous graph network includes a basic physical layer, a behavior semantics layer, and an environmental association layer, the basic physical layer is used to represent the transaction association relationship between customers and products and the physical attributes of products, the behavior semantics layer is used to represent the potential behavior patterns of customers and the semantic association features between products, and the environmental association layer is used to represent the dynamic impact of external environmental factors on the interaction behavior between products and customers; and based on the heterogeneous graph network, dividing the customer group into multiple customer communities through a community discovery algorithm, and generating a hierarchical label tree for each customer community; wherein the hierarchical label tree includes a first-level label, a second-level label, and a third-level label, the first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer and serves as the top-level node of the hierarchical label tree; the second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavior semantics layer and serves as the middle-level node of the hierarchical label tree; the third-level label is constructed based on the dynamic trajectory time sequence features of customers in the customer community in the heterogeneous graph network and serves as the bottom-level node of the hierarchical label tree.
[0084] Thus, according to this embodiment, first, multi-dimensional data such as customer transaction data, product attribute data, customer behavior trajectory data, and environmental data in the retail scenario are obtained, breaking the data fragmentation caused by differences in department tasks in the traditional annotation process. Then, by constructing a heterogeneous graph network including a basic physical layer, a behavioral semantics layer, and an environmental association layer, these data from different sources and different modalities are uniformly mapped into a graph structure to express the complex relationships among customers, products, and the environment in the form of a graph structure, thereby realizing the deep integration of multi-dimensional data and multi-dimensional relationship modeling in the retail scenario and providing a unified data foundation for comprehensively understanding customer behavior. Secondly, through the community discovery algorithm, customer groups with cross-business department business value are identified in the heterogeneous graph network. These groups are not defined based on the needs of a single department, but through the cross-layer connection characteristics in the graph network, the common behavior patterns and value characteristics in different business scenarios are naturally mapped, thus forming a unified cognitive framework for customers. This framework supports cross-business department applications and solves the problem of fragmented labels being attached to the same customer in the traditional annotation system. Finally, based on the above cross-business department customer data, a hierarchical label tree is constructed to establish a hierarchical cognitive framework. This hierarchical label tree not only reflects the logical relationship between different labels, but also through the construction of multi-level labels, finely characterizes multiple dimensions and levels of customer value. Thus, the technical effects of unified understanding of customer behavior, in-depth mining of multi-dimensional customer relationships, capture of the dynamic evolution of customer value, and construction of hierarchical customer cognition are achieved. Furthermore, the technical problems existing in the prior art, such as data islands caused by overly task-oriented traditional customer data annotation methods, single data dimension restricting dynamic value mining, and flat label systems restricting hierarchical cognition construction, are solved.
[0085] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0086] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0087] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0088] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0089] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0090] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0091] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A data annotation method based on graph structure and community discovery, characterized in that, A data annotation platform applied to the retail scenario, and the data annotation method includes: Obtain multi-dimensional data of the customer group to be labeled with tags; wherein, the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; Construct a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavior semantics layer, and an environmental association layer. The basic physical layer is used to represent the transaction association relationship between customers and products and the physical attributes of products. The behavior semantics layer is used to represent the potential behavior patterns of customers and the semantic association characteristics between products. The environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between products and customers; and Based on the heterogeneous graph network, divide the customer group into multiple customer communities through a community discovery algorithm, and generate a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes a first-level label, a second-level label, and a third-level label. The first-level label is constructed based on the interaction mode of the basic physical layer and the environmental association layer and serves as the top-level node of the hierarchical label tree. The second-level label is constructed based on the potential behavior patterns of customers in the customer community in the behavior semantics layer and serves as the middle-level node of the hierarchical label tree. The third-level label is constructed based on the dynamic trajectory time series characteristics of customers in the customer community in the heterogeneous graph network and serves as the bottom-level node of the hierarchical label tree.
2. The method according to claim 1, characterized in that, The operation of constructing a heterogeneous graph network according to the multi-dimensional data includes: Based on the multi-dimensional data, determine the node attributes of six predefined types of nodes; wherein the six types of nodes include customer nodes, product nodes, area nodes, environmental nodes, customer behavior nodes, and product combination nodes; Analyze the multi-dimensional data and construct the edge relationships between the six types of nodes; Based on the node attributes and corresponding edge relationships of the customer nodes, the product nodes, and the area nodes, construct a basic physical layer for representing the transaction association relationship between customers and products and the physical attributes of products; On top of the basic physical layer, based on the node attributes and corresponding edge relationships of the customer behavior nodes and the product combination nodes, construct a behavior semantics layer for representing the potential behavior patterns of customers and the semantic association characteristics between products; On top of the behavior semantics layer, based on the node attributes and corresponding edge relationships of the environmental nodes, construct an environmental association layer for representing the dynamic influence of external environmental factors on the interaction behavior between products and customers, and obtain a preliminary heterogeneous graph network; and Establish a cross-layer interaction mechanism for the preliminary heterogeneous graph network and perform multi-granularity attribute fusion processing to generate a final heterogeneous graph network.
3. The method according to claim 2, wherein The operation of determining the node attributes of six predefined types of nodes based on the multi-dimensional data includes: Extract customer static features from the customer transaction data as the node attributes of the customer nodes; wherein, the customer static features include customer ID, membership level, and life cycle stage; Extract the static features of the product from the product attribute data as the node attributes of the product node; wherein, the static features of the product include product ID, physical attributes, and price range; Extract the regional features from the customer behavior trajectory data as the node attributes of the regional node; wherein, the regional features include sales area coordinates and delivery range; Extract the environmental features from the environmental data as the node attributes of the environmental node; wherein, the environmental features include time dimension, weather condition, and seasonal factor; Extract the customer behavior features from the customer behavior trajectory data as the node attributes of the customer behavior node; wherein, the customer behavior features include behavior type, behavior timestamp, and stay duration; and Extract the product association features from the product attribute data as the node attributes of the product combination node; wherein, the product association features include product co-occurrence frequency, complementarity coefficient, and combination recommendation weight.
4. The method according to claim 2, characterized in that, The operation of analyzing the multi-dimensional data and constructing the edge relationships between the six types of nodes includes: Analyze the transaction frequency and amount in the customer transaction data, and construct a transaction edge between the customer node and the product node; wherein, the transaction edge represents the weighted sum of the transaction frequency and the amount; Analyze the physical attributes and category information in the product attribute data, and construct an attribute association edge between the product nodes; wherein, the attribute association edge represents the similarity of product attributes; Analyze the behavior sequence information in the customer behavior trajectory data, and construct a behavior sequence edge between the customer behavior node and the customer node; wherein, the behavior sequence edge represents the frequency or conversion rate of behavior occurrence; Analyze the product combination sales data between the customer transaction data and the product attribute data, and construct a combination association edge between the product combination node and the product node; wherein, the combination association edge represents the contribution degree of the product to the combination sales; Analyze the environmental impact information among the environmental data, the customer transaction data, and the product attribute data, and construct an environmental impact edge between the environmental node and the customer node and the product node; wherein, the environmental impact edge represents the impact degree of environmental factors on customer behavior or product sales; Analyze the seasonal change information among the environmental data, the customer transaction data, and the product attribute data, and construct an environmental adjustment edge between the environmental node and the product combination node; wherein, the environmental adjustment edge represents the impact of the seasonal factor on the product combination sales.
5. The method according to claim 2, wherein The operation of establishing a cross-layer interaction mechanism for the preliminary heterogeneous graph network and performing multi-granularity attribute fusion processing to generate the final heterogeneous graph network includes: Establish a cross-layer interaction mechanism between the basic physical layer, the behavior semantic layer, and the environmental association layer in the preliminary heterogeneous graph network through hyperedges or meta-paths; Fuse the static attributes of the customer node and the dynamic attributes of the customer behavior node in the preliminary heterogeneous graph network to generate a customer comprehensive feature attribute, and add the customer comprehensive feature attribute to the node attributes of the customer node; Determine the semantic association characteristic attributes between products according to the node attributes of the product nodes and the node attributes of the product combination nodes, and add the semantic association characteristic attributes to the node attributes of the product nodes; and Dynamically adjust the edge weights in the preliminary heterogeneous graph network according to the node attributes of the environment nodes to obtain the final heterogeneous graph network.
6. The method according to claim 1, characterized in that The operation of dividing the customer groups into multiple customer communities through the community discovery algorithm based on the heterogeneous graph network includes:[[]] Perform community division based on transaction associations at the basic physical layer of the heterogeneous graph network to form an initial community centered on product physical attributes and transaction frequencies; Perform community optimization based on behavior patterns at the behavioral semantics layer of the heterogeneous graph network, and adjust the community boundaries of the initial community in combination with customer behavior sequences and product combination association characteristics; and Introduce dynamic environmental factors at the environmental association layer of the heterogeneous graph network to perform temporal sensitivity correction on the adjusted initial community to obtain the corresponding customer community.
7. The method according to claim 6, characterized in that, The operation of generating a hierarchical label tree for each customer community includes:[[]] Based on the basic physical layer, determine the basic value attributes of each customer community; based on the environmental association layer, determine the value fluctuation coefficient of each customer community; according to the basic value attributes and the value fluctuation coefficient, determine the first-level label of each customer community as the top node of the hierarchical label tree; Based on the behavioral semantics layer, determine the behavior sequence information of the customers in each customer community; based on the product nodes and product combination nodes in the heterogeneous graph network, determine the product purchase situations of the customers in each customer community; according to the behavior sequence information and the product purchase situations, determine the second-level label of each customer community as the middle-level node of the hierarchical label tree; and Based on the basic physical layer, the behavioral semantics layer and the environmental association layer, calculate the temporal feature indicators of the customers in each customer community, and determine the third-level label of each customer community according to the temporal feature indicators as the bottom node of the hierarchical label tree; wherein, the temporal feature indicators include the volatility of the purchase cycle and the change rate of the activity participation frequency.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein the method according to any one of claims 1 to 7 is executed by a processor when the program runs.
9. A data annotation device based on a graph structure and community discovery, characterized in that, Including:[[]] A data acquisition module for acquiring multi-dimensional data of a customer group to be labeled with labels; wherein, the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; A graph construction module for constructing a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavioral semantics layer, and an environmental association layer, the basic physical layer is used to represent the transaction association relationship between customers and products and the product physical attributes, the behavioral semantics layer is used to represent the potential behavior patterns of customers and the semantic association characteristics between products, and the environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between products and customers; and A labeling module, which is used to divide the customer group into multiple customer communities through a community discovery algorithm based on the heterogeneous graph network, and generate a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes first-level labels, second-level labels and third-level labels, the first-level labels are constructed based on the interaction patterns of the basic physical layer and the environmental association layer, and serve as the top-level nodes of the hierarchical label tree; the second-level labels are constructed based on the potential behavior patterns of customers in the customer community in the behavior semantic layer, and serve as the middle-level nodes of the hierarchical label tree; the third-level labels are constructed based on the dynamic trajectory time series characteristics of customers in the customer community in the heterogeneous graph network, and serve as the bottom-level nodes of the hierarchical label tree.
10. A data annotation device based on a graph structure and community discovery, characterized in that, Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: Obtain multi-dimensional data of the customer group to be labeled with labels; wherein, the multi-dimensional data includes customer transaction data, product attribute data, customer behavior trajectory data, and environmental data; Construct a heterogeneous graph network according to the multi-dimensional data; wherein, the heterogeneous graph network includes a basic physical layer, a behavior semantic layer, and an environmental association layer, the basic physical layer is used to represent the transaction association relationship between customers and products and the physical attributes of products, the behavior semantic layer is used to represent the potential behavior patterns of customers and the semantic association characteristics between products, and the environmental association layer is used to represent the dynamic influence of external environmental factors on the interaction behavior between products and customers; and Based on the heterogeneous graph network, divide the customer group into multiple customer communities through a community discovery algorithm, and generate a hierarchical label tree for each customer community; wherein, the hierarchical label tree includes first-level labels, second-level labels and third-level labels, the first-level labels are constructed based on the interaction patterns of the basic physical layer and the environmental association layer, and serve as the top-level nodes of the hierarchical label tree; the second-level labels are constructed based on the potential behavior patterns of customers in the customer community in the behavior semantic layer, and serve as the middle-level nodes of the hierarchical label tree; the third-level labels are constructed based on the dynamic trajectory time series characteristics of customers in the customer community in the heterogeneous graph network, and serve as the bottom-level nodes of the hierarchical label tree.
Citation Information
Patent Citations
E-commerce recommendation method and system
CN110458641A
Overlapping community identification method and device, equipment, storage medium and program product
CN114329099A
Abnormal merchant group identification method and device, equipment and medium
CN117575627A
Customer portrait key data mining method and system based on space-time big data
CN118797542A
User portrait intelligent analysis system and method based on data visualization
CN119691245A
Cited By
Dynamic quantile filtering method for commodity co-occurrence network
CN122089367A