Data processing method and device, equipment and storage medium
By acquiring data from home gateway devices and using the RFM algorithm to analyze family attributes and calculate similarity values, the problem of relying on manual labeling for family relationship graphs is solved. This enables automatic discovery and real-time updates of family relationships, improving the accuracy and efficiency of family association identification.
Patent Information
- Application Number
- CN202410850361.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-06-27
AI Technical Summary
Existing family relationship mapping technologies rely on manual labeling, resulting in low accuracy in identifying relationships between families, inability to update in real time, and unsuitability for mining massive amounts of data and real-time family relationships.
By acquiring family data from home gateway devices, analyzing family attribute information using the RFM algorithm, filtering out target home gateway devices, and calculating similarity values based on the connection data of the access devices, the degree of association between families is determined, thereby achieving automatic discovery and real-time updates of family relationships.
It improves the accuracy and efficiency of identifying connections between households, is suitable for real-time analysis of large-scale, massive data, and supports the promotion and coverage of household services and the association of smart home devices.
Smart Images

Figure CN118827260B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a data processing method, apparatus, device and storage medium. Background Technology
[0002] With the widespread adoption of home broadband networks, home networks have become an indispensable part of people's lives, and related home services have grown accordingly. The promotion and development of these home services often rely on family relationship diagrams. A relatively complete family relationship diagram helps analyze the connections between families, facilitating the promotion and coverage of home services and enabling the connection of smart home devices.
[0003] In related technologies, multiple family relationship graphs can be aggregated according to family attributes manually marked by users to identify the relationships between families. However, the aforementioned method relies too heavily on manual labeling, resulting in low accuracy of family aggregation results and making it impossible to accurately identify the relationships between families. Summary of the Invention
[0004] This application provides a data processing method, apparatus, device, and storage medium that can solve the problem in related technologies of the inability to accurately identify the connections between households.
[0005] In a first aspect, embodiments of this application provide a data processing method, which may include:
[0006] Obtain home data for N home gateway devices. Home data includes device data and connection data. Device data includes gateway device data of the home gateway devices and access device data of access devices that are connected to the home gateway devices. Connection data is the data of access devices connecting to the home gateway devices. N is an integer greater than 1.
[0007] Based on the household data of N home gateway devices, the household where each of the N home gateway devices is located is analyzed to obtain the household attribute information of each household where the home gateway device is located.
[0008] Based on the family attribute information, select the first family gateway device from N family gateway devices. The target family attribute information of the family where the first family gateway device is located matches the preset family attribute information. The first family gateway device has a connection relationship with the first access device.
[0009] Based on the first access device data of the first access device and the target connection data of M target home gateway devices connected to the first access device, a similarity value is determined among the M target home gateway devices. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the homes to which each home gateway device in the M target home gateway devices belongs.
[0010] Secondly, embodiments of this application provide a data processing apparatus, which may include:
[0011] The acquisition module is used to acquire home data from N home gateway devices. The home data includes device data and connection data. The device data includes gateway device data of the home gateway devices and access device data of access devices that are connected to the home gateway devices. The connection data is the data of the access devices connecting to the home gateway devices. N is an integer greater than 1.
[0012] The analysis module is used to analyze the household data of N home gateway devices to obtain the household attribute information of each home gateway device.
[0013] The filtering module is used to filter the first home gateway device from N home gateway devices according to home attribute information. The target home attribute information of the home where the first home gateway device is located matches the preset home attribute information. The first home gateway device has a connection relationship with the first access device.
[0014] The determination module is used to determine the similarity value between the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the homes to which each home gateway device in the M target home gateway devices belongs.
[0015] Thirdly, embodiments of this application provide a computer device, which includes: a processor and a memory storing computer program instructions;
[0016] When the processor executes computer program instructions, it implements the data processing method as described in the first aspect.
[0017] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the data processing method as described in the first aspect.
[0018] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the data processing method as shown in the first aspect.
[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the data processing method as described in the first aspect.
[0020] The data processing method, apparatus, device, and storage medium of this application embodiment acquires household data from N household gateway devices. The household data includes device data and connection data. The device data includes gateway device data of the household gateway devices and access device data of access devices connected to the household gateway devices. The connection data is the data of the access devices connecting to the household gateway devices. Next, based on the household data of the N household gateway devices, the household where each of the N household gateway devices is located is analyzed to obtain household attribute information for each household. According to the household attribute information, a first household gateway device is selected from the N household gateway devices. The target household attribute information of the household where the first household gateway device is located matches preset household attribute information, and the first household gateway device has a connection relationship with a first access device. Then, based on the first access device data of the first access device and the target connection data of M target household gateway devices connected to the first access device, a similarity value is determined among the M target household gateway devices. The M target household gateway devices include the first household gateway device and at least one second household gateway device. The similarity value is used to determine the degree of association between the households where each of the M target household gateway devices is located. In a home network, the home gateway device, as the entry point, is the core of the entire network, carrying all access devices and their information. Therefore, real-time dynamic analysis and data collection based on the home gateway device's connections eliminates the need for manual labeling, solving the problems of difficult, inaccurate, and unreal-time data collection. Next, by analyzing the target home gateway devices and their connected access devices through sequential filtering of home data and access device data, the relationships between multiple homes are determined. Thus, by analyzing the connection behavior of access devices across different home gateway devices, proactive discovery of home relationships is achieved, avoiding manual intervention and improving the accuracy and efficiency of identifying relationships between homes. This approach is suitable for real-time analysis of large-scale, massive datasets, providing real-time updates on home relationships and enhancing the accuracy of data mining. It also helps in the promotion and coverage of home services and the association of smart home devices. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0023] Figure 2 A schematic diagram illustrating the process of extracting and storing family data according to an embodiment of this application;
[0024] Figure 3 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0027] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0028] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0029] A home network is a home information platform that integrates a home control network and a multimedia information network. It is a system that enables the interconnection and management of information devices, communication devices, entertainment devices, home appliances, automation devices, lighting devices, security (monitoring) devices, water, electricity, gas and heat meters, home emergency alarms, and other devices within the home, as well as the sharing of data and multimedia information.
[0030] With the widespread adoption of home broadband networks, home networks have become an indispensable part of people's lives, and related home services have grown accordingly. The promotion and development of these home services often rely on family relationship diagrams. A relatively complete family relationship diagram helps analyze the connections between families, facilitating the promotion and coverage of home services and enabling the connection of smart home devices.
[0031] In related technologies, families can be aggregated through multiple family relationship graphs based on manually labeled tags to determine the relationships between families. However, due to the difficulty in collecting family data and its reliance on user-generated labels, this method suffers from challenges in data collection, low accuracy, and an inability to perform real-time calculations, making it unsuitable for widespread adoption in family-related businesses. Specifically, this method is costly, has poor generalization ability, is unsuitable for mining massive amounts of data and real-time family relationships, and has limited user coverage. Furthermore, due to the single source of family data, much of the data is inaccurate, and errors from installation and maintenance personnel further complicate the identification accuracy. Moreover, family relationship data is complex; existing family relationship mining systems rely on family label data for relationship mining, such as proximity of family addresses, which does not accurately reflect relationships between multiple families, thus resulting in low accuracy. Finally, due to the wide coverage of family relationships and the large number of family users, the labeled data cannot be updated in real time, lacking the ability for real-time calculations and mining.
[0032] To address the aforementioned technical challenges such as difficulties in data collection, lack of real-time data analysis, and inaccurate data mining, this application provides a data processing method, apparatus, computer equipment, and storage medium for mining family relationships based on device connection behavior. This application proposes a real-time dynamic mining method that can discover extended family coverage in real-time under high-concurrency big data scenarios. In a home network, the home gateway device, as the entry point device, is the core of the entire home network, carrying all access devices and their information within the home intranet. Thus, real-time dynamic analysis and collection based on the connections between the home gateway device and the access devices connected to it solves the problems of difficult data collection and lack of real-time data. Furthermore, by using a multi-level hierarchical model for judgment and analysis, the relationships between multiple families are discovered, improving the accuracy of data mining and consequently enhancing the accuracy of business promotion and the support for home devices.
[0033] Therefore, in order to better illustrate the content of the embodiments of this application, the following will be combined with... Figures 1 to 5 The following describes a data processing system, method, apparatus, and computer device provided in the embodiments of this application.
[0034] like Figure 1 As shown in the figure, this application provides a data processing method.
[0035] Figure 1 This is a flowchart of a data processing method provided in an embodiment of this application.
[0036] The data processing method provided in this application can be applied to computer devices, and the data processing method may specifically include the following steps:
[0037] Step 110: Obtain household data from N home gateway devices. Household data includes device data and connection data. Device data includes gateway device data for each home gateway device and access device data for access devices connected to the home gateway devices. Connection data refers to the data showing the connection between the access devices and the home gateway devices. N is an integer greater than 1. Step 120: Based on the household data from the N home gateway devices, analyze the household where each of the N home gateway devices is located to obtain household attribute information for each household. Step 130: Based on the household attribute information, select the first household from the N home gateway devices. The gateway device, the target family attribute information of the family where the first home gateway device is located matches the preset family attribute information, and the first home gateway device has a connection relationship with the first access device; then, in step 140, based on the first access device data of the first access device and the target connection data of M target home gateway devices connected to the first access device, the similarity value between the M target home gateway devices is determined. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the families where each home gateway device in the M target home gateway devices is located, and M is an integer greater than 1.
[0038] In this way, since the home gateway device, as the entry point device in the home network, is the core of the entire home network and carries all access devices and their information, the connection dynamics of the home gateway device can be analyzed and collected in real time without relying on manual labeling. This solves the problems of difficult home data collection, inaccurate data, and lack of real-time data. Then, by sequentially filtering target home gateway devices and their connected access devices through home data and access devices, the relationships between multiple homes are determined. In this way, by analyzing the connection behavior of access devices between different home gateway devices, the relationships between homes are proactively discovered, avoiding manual intervention and improving the accuracy and efficiency of identifying the relationships between homes. This is suitable for real-time analysis of large-scale, massive data, and real-time updates of home relationships, improving the accuracy of data mining and helping to promote and cover home services and complete the association of smart home devices.
[0039] The steps described above are explained in detail below.
[0040] First, regarding step 110, in some embodiments of this application, the home gateway device includes a home gateway device activated by the user. Based on this, step 110 may specifically include steps 1101 and 1102.
[0041] Step 1101: If the home gateway device is detected to be activated, obtain the gateway device data of the activated home gateway device. The gateway device data includes the gateway LAN address and gateway model.
[0042] Step 1102: Through the long connection established with the home gateway device, obtain access device data of the access devices connected to the home gateway device and connection data between the home gateway device and the access devices; wherein, the access device data includes the device's local area network address, device type and device model; the connection data includes the device's online time, device offline time, number of access devices connected to the home gateway device, and access frequency of the access devices accessing the home gateway device within a preset time period.
[0043] For example, step 110 can specifically perform real-time collection of gateway device data from the home gateway device, access device data from access devices connected to the home gateway device, and connection data from access devices connecting to the home gateway device, and construct a home data database for later analysis of relationships between homes. Based on this, when the activation of a home gateway device is detected, the gateway LAN address (Media Access Control Address, MAC), gateway model, broadband information, physical address, mobile phone number, and other relevant information of the home gateway device are deduplicated by province and stored in the gateway details information database; using an event-triggered mechanism, the long-term connection established between home gateway devices reports device online / offline events in real time, obtaining device data from access devices connected to the home gateway device, such as device MAC, device type and model, device online time, device offline time, etc. After deduplicating the device MAC, a device information details database and a device operation information database are created, i.e. Figure 2 As shown, a mapping relationship is established between device MAC information and gateway MAC, forming a relational linked list. This linked list is connected using both the device MAC and gateway MAC as primary keys, creating two sets of mappings. Specifically, the gateway LAN address of the gateway device is associated with the device LAN address of the access device, resulting in an association linked list. This association linked list includes a first association linked list and a second association linked list. The first association linked list is stored using the gateway LAN address as the primary key, and the second association linked list is stored using the access device as the primary key. The association linked list storage structure is stored in Redis or local memory. Here, the first association linked list can be used to retrieve home data based on the home gateway device as the primary key, and the second association linked list can be used in step 130 to retrieve the home gateway device, such as the first home gateway device, according to home attribute information.
[0044] Next, in step 120, a customized RFM algorithm is used to filter high-value gateways and identify high-value home users. Since there are hundreds of millions of home gateway devices in the current network, and the online / offline messages of access devices connected to these home gateway devices reach billions of times per day, and some home gateway devices connect to a small number of access devices with low online frequency, collecting data from all devices would inevitably lead to high collection pressure and resource constraints. To address this issue, this application's embodiment proposes a region-based gateway filtering method. This method uses a customized RFM model based on device connection status to rank users by value, prioritizing the collection of data from the top-ranked users within each region. The specific details are as follows.
[0045] In some embodiments of this application, step 120 may specifically include:
[0046] Using the RFM data analysis algorithm, based on the household data of N home gateway devices, the household where each of the N home gateway devices is located is analyzed to obtain the household attribute information of each household where the home gateway device is located.
[0047] Among them, the family attribute information includes the attribute information corresponding to each of the P data analysis elements corresponding to the RFM data analysis algorithm, where P is an integer greater than 1.
[0048] Specifically, the P data analysis elements include a first data analysis element (R), a second data analysis element (M), and a third data analysis element (F). The first data analysis element is used to measure the proximity of the access device's online time to the current time. The second data analysis element is used to measure the importance of the home gateway device among N home gateway devices. The third data analysis element is used to measure the activity level of the access devices connected to the home gateway device. Based on this, step 120 may specifically include steps 1201 and 1202.
[0049] Step 1201: Using the RFM data analysis algorithm, based on the home data from N home gateway devices, calculate the first evaluation quantity corresponding to the first data analysis element, the second evaluation quantity corresponding to the second data analysis element, and the third evaluation quantity corresponding to the third data analysis element.
[0050] Furthermore, the home data of the N home gateway devices includes the home data of each of the N home gateway devices. Based on this, step 1201 may specifically include steps 12011 to 12014.
[0051] Step 12011: Extract from the home data the most recent online time of the access devices connected to each home gateway device, the number of access devices connected to each home gateway device, and the access frequency of the access devices connected to each home gateway device.
[0052] For example, data such as the most recent device online time, the total number of devices, and the access frequency are collected from the household data. In some embodiments, before step 12012, because the access device data and / or connection data of the downstream devices (i.e., connected access devices) collected by the home gateway device contain a lot of abnormal data, such as empty device names or empty strings, it is necessary to clean the access device data and / or connection data to remove some data such as null values, abnormal values, and special characters before proceeding to step 12012.
[0053] Step 12012: Sort the access devices connected to each of the N home gateway devices in ascending order of time to obtain the device online time sequence; determine the device online time corresponding to the center value in the device online time sequence as the first evaluation metric.
[0054] For example, in order to adapt to scenarios where gateway value is calculated based on device connection status and to increase the accuracy of the model, this application embodiment modifies the calculation of the center value of RFM. Since the connection between the access device and the home gateway device is periodic, the center value of the device online time of multiple home gateway devices can be used as the center value of the period, i.e., the first evaluation value.
[0055] Step 12013: Select at least two target quantities from the number of access devices connected to each of the N home gateway devices that fall within a preset device quantity range, with the at least two target quantities arranged in ascending order; determine the center value of the at least two target quantities within the preset device quantity range as the second evaluation metric. In this embodiment, the number of access devices connected to each home gateway device refers to the total number of access devices connected to each home gateway device.
[0056] For example, the total number of access devices connected to each home gateway device is filtered using the following formula (1). A preset device number range [TS, TE] is set, and the number of access devices connected to each home gateway device is compared using [TS, TE]. If the RFM value satisfies Formula 1, the data is retained; otherwise, the data is discarded. The second evaluation metric is calculated using the center value of [TS, TE], as shown below:
[0057] TS <= T(X) <= TE (1)
[0058] Where T(X) represents the total number of access devices connected to each home gateway device.
[0059] Step 12014: Using a distance algorithm, cluster the access frequencies of the access devices based on the similarity of access frequencies between every two home gateway devices in the N home gateway devices to obtain T clusters; using a clustering algorithm, perform clustering iteration calculation on the center point of each of the T clusters to obtain the target center value of the access frequency of the access devices in the N home gateway devices, and determine the target center value as the third evaluation metric.
[0060] Specifically, using distance algorithms (such as Euclidean distance, Manhattan distance, and Chebyshev distance), the access frequency similarity between any two home gateway devices is calculated based on the access frequency of each access device connected to each of the N home gateway devices. Based on this access frequency similarity, the N home gateway devices are clustered into T clusters, where T is an integer greater than 1. The centroid of each of the T clusters is calculated using a clustering algorithm. The centroid of each cluster is iteratively calculated using the clustering algorithm to obtain the target median value of the access frequency between any two home gateway devices, and this target median value is determined as the third evaluation metric.
[0061] For example, the center value of the access frequency is aggregated using the following formula (2) to find the center value and form the original data of the RFM model, as shown below:
[0062]
[0063] Where, x i and x j These represent the center points of two arbitrary clusters. Through continuous iteration, the target median value of the access frequency of the access device between every two home gateway devices in N home gateway devices is found and determined as the third evaluation metric.
[0064] Step 1202: Analyze the home where each home gateway device is located using the first evaluation metric, the second evaluation metric, and the third evaluation metric to obtain home attribute information.
[0065] The family attribute information includes a first evaluation identifier corresponding to the first data analysis element, a second evaluation identifier corresponding to the second data analysis element, and a third evaluation identifier corresponding to the third data analysis element. The first evaluation identifier is used to indicate the proximity of the device online time of the access device connected to each family gateway device to the current time. The second evaluation identifier is used to indicate the importance of each family gateway device among N family gateway devices. The third evaluation identifier is used to indicate the activity level of the access device connected to each family gateway device.
[0066] Furthermore, step 1202 may specifically include steps 12021 to 12013.
[0067] Step 12021: Compare the most recent online time of the access device connected to each home gateway device with the first evaluation quantity to obtain the first evaluation identifier.
[0068] For example, the last online time of the access device connected to each home gateway device is compared with a first evaluation value. If it is greater than the first evaluation value, it is set to 1; otherwise, it is set to 0.
[0069] Step 12022: Compare the number of access devices connected to each home gateway device with the second evaluation quantity to obtain the second evaluation identifier.
[0070] For example, the number of access devices connected to each home gateway device is compared with a second evaluation value; if it is greater than the second evaluation value, it is set to 1, otherwise it is set to 0.
[0071] Step 12023: Compare the access frequency of the access devices connected to each home gateway device with the third evaluation quantity to obtain the third evaluation identifier.
[0072] For example, the access frequency of the access device to each home gateway device is compared with a third evaluation value. If the frequency is greater than the third evaluation value, the value is set to 1; otherwise, it is set to 0.
[0073] Therefore, after dividing the three data analysis elements in detail, we can obtain the family attribute information of 8 types of families. For example, 111 are families with recent reporting time, high reporting frequency and total number of reports, which are important value families; 101 are families with recent reporting time, high reporting frequency but low total number of reports, which are important development families; 011 are important maintenance families; 110 are important retention families; 001 are general value families; 100 are general development families; 010 are general maintenance families; and 000 are general retention families.
[0074] Furthermore, regarding step 130, in some embodiments of this application, a first home gateway device that includes at least two 1s in the home attribute information can be selected, such as important value home 111, important development home 101, important retention home 011, and important retention home 110, as the first home gateway device for monitoring access devices. This can alleviate the pressure of data collection and reduce the waste of resources caused by low-value user data storage and computing.
[0075] Then, regarding step 140, in some embodiments of this application, the first access device data includes the first device type of the first access device, and the target connection data includes the access frequency of the first access device accessing the target home gateway device within a preset time period. Based on this, step 140 may specifically include steps 1401 to 1403.
[0076] Step 1401: Obtain the first device type weight value corresponding to the first device type based on the first device type of the first access device.
[0077] Step 1402: Extract the first minimum access frequency of the first access device accessing the M target home gateway devices within a preset time period from the target connection data of the M target home gateway devices connected to the first access device.
[0078] Step 1403: Based on the first device weight value and the first minimum access frequency, calculate the similarity value between the first home gateway device and each of the M target home gateway devices.
[0079] For example, preliminary correlation screening is performed by device type to complete the initial association of family relationships. After the screening of the first home gateway device is completed through the above step 130, the next step is to analyze the types of access devices connected to it. Because mobile devices such as mobile phones and tablets move frequently between homes, while desktop computers and routers move less frequently, in order to isolate the influence of different device connection frequencies and conduct preliminary mining of family relationships, the following method is designed in this application embodiment for preliminary relationship mining. First, the frequency index of device types is calculated to obtain the weight I(w) of each type of device, which can be calculated by the following formula (3):
[0080]
[0081] Where N is the total number of access frequencies of all access devices connecting to the home gateway device under a certain device type, and TF(w) is the access frequency of the device type.
[0082] Based on this, after calculating the weight I(w), the access devices of home gateway device A are a set. The set of access devices for home gateway device B is as follows: The similarity between the two can be calculated using formula (4):
[0083] P(X)=TF1(w)*I(w1)+TF2(w)*I(w2)+…+TF n (w)*I(w n (4)
[0084] Wherein, TF1(w) is the minimum access frequency of the device among the two home gateway devices.
[0085] In other embodiments of this application, the first access device data includes the first device type of the first access device, and the first home gateway device has a connection relationship with the second access device; the M target home gateway devices include a first type of target home gateway device and a second type of target home gateway device, the first type of target home gateway device also has a connection relationship with the second access device and the third access device, and the second type of target home gateway device also has a connection relationship with the third access device, wherein the device types of the first access device, the second access device and the third access device are different. Based on this, step 140 may specifically include steps 1404 to 1406.
[0086] Step 1404: Based on the first device type of the first access device, the second device type of the second access device, and the third device access type of the third access device, obtain the first device type weight value corresponding to the first device type, the second device type weight value corresponding to the second device type, and the third device type weight value corresponding to the third device type, respectively.
[0087] Step 1405: Extract the first minimum access frequency of the first access device accessing the M target home gateway devices within a preset time period from the target connection data of the M target home gateway devices connected to the first access device; and extract the second minimum access frequency of the second access device accessing the first type of target home gateway devices within a preset time period from the target connection data of the first type of target home gateway devices connected to the second access device; and extract the third minimum access frequency of the third access device accessing the first type of target home gateway devices and the first type of target home gateway devices within a preset time period from the target connection data of the first type of target home gateway devices and the target connection data of the second type of target home gateway devices connected to the third access device.
[0088] Step 1406: A weighted sum is performed on the first minimum access frequency and the first device weight value, the second minimum access frequency and the second device weight value to obtain a similarity value between the first home gateway device and the first type of target home gateway device; and a weighted sum is performed on the first minimum access frequency and the first device weight value, the third minimum access frequency and the third device weight value to obtain a similarity value between the first type of target home gateway device and the second type of target home gateway device; and a similarity value between the first home gateway device and the second type of target home gateway device is obtained based on the first minimum access frequency and the first device weight value.
[0089] In some other embodiments of this application, the target connection data includes first target connection data within a first preset time period and second target connection data within a second preset time period. Based on this, step 140 may specifically include steps 1407 and 1408.
[0090] Step 1407: Based on the similarity values between the M target home gateway devices in the first preset time period and the similarity values between the M target home gateway devices in the second preset time period, obtain the first average similarity value between the M target home gateway devices.
[0091] Step 1408: If the first average similarity value is greater than or equal to the preset similarity value, determine that at least two of the M target home gateway devices are associated with each other.
[0092] For example, by using connection frequency, further confirmation of associated families is achieved. Once a wide range of relationships between two families is determined, to ensure the accuracy of family relationship discovery, further confirmation of the discovered family relationships is needed to eliminate the influence of accidental device connections. Therefore, this application proposes a method to avoid coupled devices, thereby ensuring the accuracy of discovery. This is achieved by monitoring the associated devices for the subsequent N cycles.
[0093] In some other embodiments of this application, the data processing method may further include steps 1501 to 1504.
[0094] Step 1501: If the similarity value among the M target home gateway devices is less than the preset similarity value, obtain the latest target connection data for each of the M target home gateway devices.
[0095] Step 1502: Based on the latest target connection data of each target home gateway device, determine the latest similarity value among the M target home gateway devices.
[0096] Step 1503: Calculate the second average similarity value among the M target home gateway devices based on the similarity values among the M target home gateway devices and the latest similarity values among the M target home gateway devices.
[0097] Step 1504: If the second average similarity value is greater than or equal to the preset similarity value, determine that at least two of the M target home gateway devices are associated with each other.
[0098] For example, when P(X) is greater than a threshold, it is considered that there is a wide range of connections between the two gateways, and steps 1501 to 1504 need to be calculated to further confirm the associated families based on the connection frequency. After determining that there is a wide range of connections between the two families, in order to ensure the accuracy of family relationship discovery, it is necessary to further confirm the discovered family relationships and eliminate the influence of accidental device connections. Therefore, this application proposes a method to avoid coupled devices, thereby ensuring the accuracy of discovery. By monitoring the associated devices for the next N periods, the accuracy of family relationships is determined. In some embodiments of this application, the similarity value between the M target family gateway devices includes the similarity value between at least two target family gateways among the M target family gateway devices; based on this, if the similarity value is greater than or equal to a preset similarity value, it is determined that the families where at least two target family gateway devices are located have an association relationship; if the similarity value is less than the preset similarity value, it is determined that the families where at least two target family gateway devices are located do not have an association relationship.
[0099] Therefore, this application proposes a multi-level family identification model by analyzing the behavior of access devices to solve the problems of difficult data acquisition, incorrect identification of family relationships, and the lack of real-time updates in the family relationship mining system. This application proactively discovers family relationships by analyzing the connection behavior of smart devices across different family networks. In a family network, the smart gateway, as the entry point device, is the core of the entire network. The home gateway typically connects all devices in the family and can collect information such as the online / offline frequency, network time, device type, and MAC address of its connected devices. Currently, the number of family networks is enormous, and the frequency of device reporting is high. To discover valuable families, this application first provides an RFM algorithm based on the status of smart devices to filter family gateways, eliminating families in public areas and those with low-value services. Information such as the type of connected devices, internet access time, and online / offline frequency obtained from high-value gateways is then quantitatively calculated. Due to the frequent movement of devices and the massive computational load, this application's embodiments employ a multi-layered device search method to simplify the process. First, the importance of device types is analyzed to determine the weights of different device types. When multiple smart devices appear in different households, the relevance of the devices is calculated using a model. If the relevance exceeds a threshold, a third-layer household relationship determination is performed. This third layer analyzes the device's online / offline time series to determine the stable connection status of devices within a given interval. When the collected relevance falls below a certain threshold, the two households are considered to have a household relationship.
[0100] This application proposes a family clustering algorithm based on device connection status, which addresses the challenges of family relationship mining. It provides a one-stop solution for family relationship mining, resolving the issues of family relationship data collection and analysis. By analyzing the connection status of smart devices, it can automatically discover and output family relationship sets, effectively covering scenarios such as product deployment, business coverage, and family device association. After a user's smart device connects to the gateway, the connection patterns of user devices are used to associate them, thereby discovering the relationships between devices. The implementation of this method can be divided into four steps: real-time acquisition and storage of family gateway device connection data; first-level screening of high-value families using an RFM algorithm based on smart device status; second-level filtering of family relationships based on changes in smart device connection; and third-level identification of family relationships through detailed analysis of smart devices. Connection data is automatically acquired, stored, and calculated in real-time by users automatically connecting to the gateway. Through the analysis of historical data and the application of an improved RFM algorithm, the difficulties of data collection and real-time updates of family relationships are solved. Furthermore, a multi-layered analysis method is proposed to address the issue of accuracy in identifying family relationships.
[0101] It should be noted that the data processing method of this application embodiment can be applied to the system connected to this smart device, satisfying the needs of scenarios involving large-scale data analysis and mining. By mining family relationships, precise targeting of family services can be achieved, potential valuable users can be discovered, and the growth of family services can be facilitated. Simultaneously, by mining family relationships, full coverage of intelligent scenarios for family devices can be achieved, making smart devices even more intelligent. Furthermore, the family relationship mining method proposed in this application embodiment can also be applied to the mining of other related relationships, enabling real-time relationship mining.
[0102] To better illustrate the data processing method provided in the embodiments of this application, based on Figure 1 and Figure 2 The content shown is specifically combined with Figure 3 Please provide a detailed explanation.
[0103] like Figure 3 As shown, firstly, data is collected in real time, and a database of device information and gateway information details is built.
[0104] Specifically, based on the user's gateway activation information, the MAC address, broadband information, physical address, mobile phone number, and other relevant information of the home gateway device are deduplicated by province and stored in the gateway details database. An event-triggered mechanism is used to report device online / offline events in real time through the gateway's long connection, obtaining information such as the type and model of the connected device, the device's online time, offline time, and MAC address. After deduplicating the device MAC information, a device information details database and a device operation information database are created. A correspondence between device MAC information and gateway MAC is established, forming a relational linked list. The linked list is connected using the device MAC and gateway MAC as the primary keys, forming two sets of corresponding relationships. The linked list storage structure is stored in Redis or local memory.
[0105] Next, a customized RFM algorithm is used to filter high-value gateways and identify high-value home users. Specifically, since there are hundreds of millions of home gateway users in the existing network, and the number of device online / offline messages reaches billions every day, and some gateways have fewer devices with low online frequency, collecting data from all devices would inevitably lead to high collection pressure and resource constraints. To address this issue, this application proposes a region-based gateway filtering method. A customized RFM model based on device connection status is used to rank users by value, and users ranked higher are collected first by province.
[0106] The specific steps of the RFM model include: collecting data of devices connected to the gateway, mainly including the latest device online time, the total number of devices, and the access frequency of the devices; since there are many abnormal data in the data of devices connected to the gateway, such as empty device names, empty strings, etc., it is necessary to clean the device data and remove some data such as empty values, abnormal values, special characters, etc.; in order to adapt to the scenario of calculating the gateway value based on the device connection status and increase the accuracy of the model, the central value calculation of RFM is modified in this application embodiment. Since the device online is periodic, the central value of the device online time is directly used as the central value of the period; the total number of devices is filtered using the above formula (1), and the interval of device online within the period is set as [TS, TE]. The time of collecting R value, F value, and M value is compared through the threshold interval. If the RFM value meets the formula (1), the data is retained; otherwise, the data is discarded. The central value is calculated using the central value of [TS, TE]. The central value of the access frequency (M) is aggregated using formula (2) to find the central value and form the original data of the RFM model; according to the specific RFM of the user. The M-value is compared with the center value. If it is greater than the center value, it is set to 1; otherwise, it is set to 0. This completes the subdivision of specific indicators. After further detailed division of the three indicators, eight categories of households can be obtained. For example: 111 are users with recent reporting time, high reporting frequency, and high total reporting number, who are considered important value households; 101 are users with recent reporting time, high reporting frequency, but low total reporting number, who are considered important development households; 011 are important retention households; 001 are important retention households; 110 are general value households; 100 are general development households; 010 are general retention households; and 000 are general retention households. In this way, based on the important value households, important development households, and important retention households in the RFM model, the access devices are monitored as key customers. This can alleviate the pressure of data collection and reduce the resource waste caused by the storage and computation of low-value user data.
[0107] Furthermore, a preliminary screening of relevance is conducted by device type to complete the initial association of family relationships. After completing the screening of family data through the aforementioned steps, the next step is to analyze the types of connected devices. Because mobile devices such as mobile phones and tablets move frequently between families, while desktop computers and routers move less frequently, a preliminary exploration of family relationships is conducted in order to isolate the influence of different device connection frequencies. Based on this, the frequency index of device types is first calculated to obtain the weight of each type of device; after calculating the weight, the similarity between the two can be calculated using the above formula (4).
[0108] Then, the connection frequency is used to further confirm the related families. After determining that two families have a wide range of relationships, in order to ensure the accuracy of family relationship discovery, it is necessary to further confirm the discovered family relationships and eliminate the influence of accidental device connections. Therefore, this application proposes a method to avoid coupled devices, thereby ensuring the accuracy of discovery. By monitoring the associated devices for the next N cycles, the corresponding deviation is calculated using the above formula (5), thereby determining the accuracy of the family relationship.
[0109] Therefore, this application proposes a complete family clustering method, realizing a full-process solution from data collection, storage, filtering, processing, and relationship discovery. It proposes a method for aggregating family relationships based on real-time connected smart devices, solving the problems of difficult data acquisition and lack of real-time performance. Simultaneously, by improving and optimizing the MRF scheme, it makes it suitable for value analysis of smart access devices, helping to discover high-quality home gateway devices, and processing and storing the collected data in real time, solving the problem of automatic discovery and updating of family relationships in multi-device, high-concurrency environments. Furthermore, this application proposes multi-level data analysis and filtering based on device connection behavior, introducing the concept of connection stability, and performing multi-level filtering and analysis on smart devices, improving the accuracy of family relationship aggregation and reducing the impact of labeled data and accidental connections.
[0110] Compared to existing technologies, this application is more suitable for mining family relationship systems with large data volumes, solving the problem of real-time updates of family relationships. Furthermore, it addresses relationship mining in high-concurrency, big data scenarios through a multi-level hierarchical design, and improves the accuracy of family relationship mining. The embodiments of this application first discover high-value gateways using a customized MRF method, then conduct an initial search through the association process to discover connections between family relationships, and finally, further confirm the connections between family relationships by judging the stability of device connections. This solves the problem of family relationship mining in high-concurrency, massive data environments while ensuring the accuracy of family relationship mining.
[0111] This application also provides a data processing apparatus, specifically combined with... Figure 4 Please provide a detailed explanation.
[0112] Figure 4 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application.
[0113] In some embodiments of this application, Figure 4 The data processing device shown can be installed in the computer equipment provided in the embodiments of this application.
[0114] like Figure 4 As shown, the data processing device 40 may specifically include:
[0115] The acquisition module 401 is used to acquire home data from N home gateway devices. The home data includes device data and connection data. The device data includes gateway device data of the home gateway devices and access device data of access devices that are connected to the home gateway devices. The connection data is the data of the access devices connecting to the home gateway devices. N is an integer greater than 1.
[0116] Analysis module 402 is used to analyze the family where each of the N home gateway devices is located based on the family data of the N home gateway devices, and obtain the family attribute information of the family where each home gateway device is located.
[0117] The filtering module 403 is used to filter the first home gateway device from N home gateway devices according to home attribute information. The target home attribute information of the home where the first home gateway device is located matches the preset home attribute information. The first home gateway device has a connection relationship with the first access device.
[0118] The determining module 404 is used to determine the similarity value between the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the homes to which each home gateway device in the M target home gateway devices belongs.
[0119] Therefore, in a home network, the home gateway device, as the entry point device, is the core of the entire home network, carrying all access devices and their information within the home intranet. Thus, in this embodiment, the data processing device can dynamically analyze and collect data in real time based on the connections of the home gateway device, eliminating the need for manual labeling and solving the problems of difficult home data collection, inaccurate data, and lack of real-time data. Next, by analyzing the target home gateway devices and their connected access devices through home data and access devices, the relationships between multiple homes are determined. In this way, by analyzing the connection behavior of access devices between different home gateway devices, home relationships are proactively discovered, avoiding manual intervention and improving the accuracy and efficiency of identifying relationships between homes. This is suitable for real-time analysis of large-scale, massive data, real-time updates of home relationships, and improved data mining accuracy, which helps promote and cover home services and complete the association of smart home devices.
[0120] The data processing device 40 in the embodiments of this application will be described in detail below.
[0121] In some embodiments of this application, the acquisition module 401 may be specifically used to acquire the gateway device data of the activated home gateway device when it is detected that the home gateway device is activated and the home gateway device includes the home gateway device activated by the user. The gateway device data includes the gateway LAN address and the gateway model.
[0122] By establishing a long connection with the home gateway device, access device data of access devices that are connected to the home gateway device and connection data between the home gateway device and the access devices are obtained.
[0123] The access device data includes the device's LAN address, device type, and device model; the connection data includes the device's online time, device offline time, the number of access devices connected to the home gateway device, and the access frequency of the access device to the home gateway device within a preset time period.
[0124] In some embodiments of this application, the analysis module 402 can be specifically used to analyze the family where each of the N home gateway devices is located based on the family data of the N home gateway devices using the RFM data analysis algorithm, and obtain the family attribute information of the family where each home gateway device is located; wherein, the family attribute information includes the attribute information corresponding to each of the P data analysis elements corresponding to the RFM data analysis algorithm, where P is an integer greater than 1.
[0125] In some embodiments of this application, the analysis module 402 can specifically be used to calculate, based on the home data of N home gateway devices, a first evaluation value corresponding to the first data analysis value, a second evaluation value corresponding to the second data analysis value, and a third evaluation value corresponding to the third data analysis value, using the RFM data analysis algorithm, based on the home data of N home gateway devices, when there are P data analysis elements including a first data analysis element, a second data analysis element, and a third data analysis element, wherein the first data analysis element is used to measure the proximity of the device online time of the access device to the current time, the second data analysis element is used to measure the importance of the home gateway device among N home gateway devices, and the third data analysis element is used to measure the activity level of the access device connected to the home gateway device;
[0126] By analyzing the first, second, and third evaluation metrics, the home where each home gateway device is located is analyzed to obtain home attribute information;
[0127] The family attribute information includes a first evaluation identifier corresponding to the first data analysis element, a second evaluation identifier corresponding to the second data analysis element, and a third evaluation identifier corresponding to the third data analysis element. The first evaluation identifier is used to indicate the proximity of the device online time of the access device connected to each family gateway device to the current time. The second evaluation identifier is used to indicate the importance of each family gateway device among N family gateway devices. The third evaluation identifier is used to indicate the activity level of the access device connected to each family gateway device.
[0128] In some embodiments of this application, the analysis module 402 may be specifically used to extract, from the home data of N home gateway devices, the most recent online time of the access device connected to each home gateway device, the number of access devices connected to each home gateway device, and the access frequency of the access devices connected to each home gateway device, when the home data of N home gateway devices includes the home data of each of the N home gateway devices.
[0129] Sort the access devices connected to each of the N home gateway devices in ascending order of time to obtain the device online time sequence; determine the device online time corresponding to the center value in the device online time sequence as the first evaluation metric.
[0130] Furthermore, from the number of access devices connected to each of the N home gateway devices, at least two target quantities within a preset device quantity range are selected, and the at least two target quantities are arranged in ascending order of quantity; the center value of the at least two target quantities within the preset device quantity range is determined as the second evaluation quantity;
[0131] Furthermore, using a distance algorithm, based on the similarity of access frequencies between any two home gateway devices connected to N home gateway devices, the access frequencies of the access devices are clustered to obtain T clusters. Using a clustering algorithm, the center point of each of the T clusters is iteratively calculated to obtain the target center value of the access frequency of the access devices connected to N home gateway devices, and the target center value is determined as the third evaluation metric.
[0132] In some embodiments of this application, the analysis module 402 may specifically be used to compare the most recent online time of the access device connected to each home gateway device with the first evaluation quantity to obtain the first evaluation identifier;
[0133] In addition, the number of access devices connected to each home gateway device is compared with the second evaluation metric to obtain the second evaluation identifier;
[0134] Furthermore, the access frequency of each access device connected to the home gateway device is compared with the third evaluation metric to obtain the third evaluation identifier.
[0135] In some embodiments of this application, the determining module 404 may be specifically used to obtain a first device type weight value corresponding to the first device type when the first access device data includes the first device type of the first access device and the target connection data includes the access frequency of the first access device accessing the target home gateway device within a preset time period.
[0136] Extract the first minimum access frequency of the first access device to the M target home gateway devices within a preset time period from the target connection data of the M target home gateway devices connected to the first access device.
[0137] Based on the first device weight value and the first minimum access frequency, calculate the similarity value between the first home gateway device and each of the M target home gateway devices.
[0138] In some embodiments of this application, the determining module 404 can be specifically used to: First access device data includes a first device type of the first access device; a first home gateway device has a connection relationship with a second access device; M target home gateway devices include a first type of target home gateway device and a second type of target home gateway device; the first type of target home gateway device also has a connection relationship with the second access device and a third access device; the second type of target home gateway device also has a connection relationship with the third access device; wherein, when the device types of the first access device, the second access device, and the third access device are different, based on the first device type of the first access device, the second device type of the second access device, and the third device access type of the third access device, respectively, obtain a first device type weight value corresponding to the first device type, a second device type weight value corresponding to the second device type, and a third device type weight value corresponding to the third device type;
[0139] From the target connection data of M target home gateway devices connected to the first access device, extract the first minimum access frequency of the first access device accessing the M target home gateway devices within a preset time period; and from the target connection data of the first type of target home gateway devices connected to the second access device, extract the second minimum access frequency of the second access device accessing the first type of target home gateway devices within a preset time period; and from the target connection data of the first type of target home gateway devices and the second type of target home gateway devices connected to the third access device, extract the third minimum access frequency of the third access device accessing both the first type of target home gateway devices and the first type of target home gateway devices within a preset time period.
[0140] The similarity value between the first home gateway device and the first type of target home gateway device is obtained by weighted summation of the first minimum access frequency and the first device weight value, the second minimum access frequency and the second device weight value; and the similarity value between the first type of target home gateway device and the second type of target home gateway device is obtained by weighted summation of the first minimum access frequency and the first device weight value, the third minimum access frequency and the third device weight value; and the similarity value between the first home gateway device and the second type of target home gateway device is obtained based on the first minimum access frequency and the first device weight value.
[0141] In some embodiments of this application, the determining module 404 may be specifically used to obtain a first average similarity value among M target home gateway devices based on the similarity values among M target home gateway devices in the first preset time period and the similarity values among M target home gateway devices in the second preset time period, when the target connection data includes first target connection data in the first preset time period and second target connection data in the second preset time period.
[0142] If the first average similarity value is greater than or equal to the preset similarity value, it is determined that at least two of the M target home gateway devices are associated with each other.
[0143] In some embodiments of this application, the data processing device 40 may further include a computing module; wherein,
[0144] The acquisition module 401 can also be used to acquire the latest target connection data of each of the M target home gateway devices when it is determined that the similarity value between the M target home gateway devices is less than the preset similarity value.
[0145] The determination module 404 can also be used to determine the latest similarity value between M target home gateway devices based on the latest target connection data of each target home gateway device;
[0146] The calculation module is used to calculate the second average similarity value among the M target home gateway devices based on the similarity value among the M target home gateway devices and the latest similarity value among the M target home gateway devices;
[0147] The determination module 404 can also be used to determine, when the second average similarity value is greater than or equal to the preset similarity value, that at least two of the M target home gateway devices are associated with each other.
[0148] Based on the same inventive concept, this application also provides a computer device. (Specifically combined with...) Figure 5 Please provide a detailed explanation.
[0149] Figure 5 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application.
[0150] like Figure 5 As shown, the computer device may include at least one of the following as described in the embodiments of this application: an electronic device, a server. The computer device may include a processor 501 and a memory 502 storing computer program instructions.
[0151] Specifically, the processor 501 may include a central processing unit (CPU), or an application-specific integrated circuit (ASTC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0152] Memory 502 may include mass storage for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk drive, magneto-optical disk drive, magnetic tape drive, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 502 may include removable or non-removable (or fixed) media. Where appropriate, memory 502 may be internal or external to a computer device. In a particular embodiment, memory 502 is non-volatile solid-state memory. In a particular embodiment, memory 502 includes solid-state storage (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0153] The processor 501 implements any of the data processing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 502.
[0154] In one example, the computer device may also include a communication interface 503 and a bus 510. Wherein, as... Figure 5 As shown, the processor 501, memory 502, and communication interface 503 are connected through bus 510 and complete communication with each other.
[0155] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0156] Bus 510 includes hardware, software, or both, that couples components of a flow control device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard System (ETSA) bus, a Front Side Bus (FSB), an HyperTransport (HT) interconnect, an Industry Standard System (TSA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel System (MCA) bus, a Peripheral Component Interconnect (PCT) bus, a PCT-Express (PCT-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 510 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0157] The data processing device can execute the data processing method described in the embodiments of this application, thereby achieving the combination Figures 1 to 5 The data processing methods and apparatus described.
[0158] Furthermore, in conjunction with the data processing methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the data processing methods in the above embodiments.
[0159] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0160] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASTCs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0161] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0162] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A data processing method, characterized in that, include: Obtain home data for N home gateway devices. The home data includes device data and connection data. The device data includes gateway device data of the home gateway devices and access device data of access devices that are connected to the home gateway devices. The connection data is the data of the access devices connecting to the home gateway devices. N is an integer greater than 1. Based on the household data of the N household gateway devices, the household where each of the N household gateway devices is located is analyzed to obtain the household attribute information of the household where each household gateway device is located. According to the family attribute information, a first family gateway device is selected from the N family gateway devices. The target family attribute information of the family where the first family gateway device is located matches the preset family attribute information. The first family gateway device has a connection relationship with the first access device. Based on the first access device data of the first access device and the target connection data of M target home gateway devices connected to the first access device, a similarity value is determined among the M target home gateway devices. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the homes to which each home gateway device in the M target home gateway devices is located.
2. The method according to claim 1, characterized in that, The home gateway device includes a home gateway device activated by the user; The acquisition of home data from N home gateway devices includes: When a home gateway device is detected to be activated, the gateway device data of the activated home gateway device is obtained, including the gateway LAN address and gateway model. By establishing a long connection with the home gateway device, access device data of access devices that are connected to the home gateway device and connection data between the home gateway device and the access devices are obtained. The access device data includes the device's local area network address, device type, and device model; the connection data includes the device's online time, device's offline time, the number of access devices connected to the home gateway device, and the access frequency of the access device to the home gateway device within a preset time period.
3. The method according to claim 1, characterized in that, The method involves analyzing the household data from the N household gateway devices to obtain household attribute information for each household, including: Using the RFM data analysis algorithm, based on the household data of the N household gateway devices, the household where each of the N household gateway devices is located is analyzed to obtain the household attribute information of the household where each household gateway device is located; The family attribute information includes the attribute information corresponding to each of the P data analysis elements corresponding to the RFM data analysis algorithm, where P is an integer greater than 1.
4. The method according to claim 3, characterized in that, The P data analysis elements include a first data analysis element, a second data analysis element, and a third data analysis element. The first data analysis element is used to measure the proximity of the access device's online time to the current time. The second data analysis element is used to measure the importance of the home gateway device among the N home gateway devices. The third data analysis element is used to measure the activity level of the access devices connected to the home gateway device. The step involves using the RFM data analysis algorithm to analyze the household data of the N home gateway devices, thereby obtaining the household attribute information of each home gateway device's household, including: Using the RFM data analysis algorithm, based on the home data of the N home gateway devices, calculate the first evaluation quantity corresponding to the first data analysis element, the second evaluation quantity corresponding to the second data analysis element, and the third evaluation quantity corresponding to the third data analysis element; The family where each home gateway device is located is analyzed using the first evaluation metric, the second evaluation metric, and the third evaluation metric to obtain the family attribute information. The family attribute information includes a first evaluation identifier corresponding to the first data analysis element, a second evaluation identifier corresponding to the second data analysis element, and a third evaluation identifier corresponding to the third data analysis element. The first evaluation identifier is used to indicate the proximity of the online time of the access device connected to each family gateway device to the current time. The second evaluation identifier is used to indicate the importance of each family gateway device among the N family gateway devices. The third evaluation identifier is used to indicate the activity level of the access device connected to each family gateway device.
5. The method according to claim 4, characterized in that, The home data of the N home gateway devices includes the home data of each of the N home gateway devices; The step of calculating a first evaluation value corresponding to the first data analysis element, a second evaluation value corresponding to the second data analysis element, and a third evaluation value corresponding to the third data analysis element based on the home data of the N home gateway devices using the RFM data analysis algorithm includes: From the household data, extract the most recent online time of the access devices connected to each household gateway device, the number of access devices connected to each household gateway device, and the access frequency of the access devices connected to each household gateway device; Sort the access devices connected to each of the N home gateway devices in ascending order of time to obtain a device online time sequence; determine the device online time corresponding to the center value in the device online time sequence as the first evaluation value. Furthermore, from the number of access devices connected to each of the N home gateway devices, at least two target quantities within a preset device quantity range are selected, and the at least two target quantities are arranged in ascending order of quantity; the center value of the at least two target quantities within the preset device quantity range is determined as the second evaluation quantity; Furthermore, using a distance algorithm, based on the similarity of access frequencies between any two home gateway devices among the N home gateway devices, the access frequencies of the access devices are clustered to obtain T clusters; using a clustering algorithm, the center point of each of the T clusters is iteratively calculated to obtain the target center value of the access frequency of the access devices accessing the N home gateway devices, and the target center value is determined as the third evaluation metric.
6. The method according to claim 5, characterized in that, The analysis of the home where each home gateway device is located, using the first evaluation metric, the second evaluation metric, and the third evaluation metric, yields the home attribute information, including: The first evaluation identifier is obtained by comparing the most recent online time of the access device connected to each home gateway device with the first evaluation value. In addition, the number of access devices connected to each home gateway device is compared with the second evaluation value to obtain the second evaluation identifier; Furthermore, the access frequency of the access devices connected to each home gateway device is compared with the third evaluation metric to obtain the third evaluation identifier.
7. The method according to claim 1, characterized in that, The first access device data includes the first device type of the first access device, and the target connection data includes the access frequency of the first access device to the target home gateway device within a preset time period; The step of determining the similarity value among the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device includes: Based on the first device type of the first access device, obtain the first device type weight value corresponding to the first device type; From the target connection data of M target home gateway devices connected to the first access device, extract the first minimum access frequency of the first access device accessing the M target home gateway devices within a preset time period; Based on the first device weight value and the first minimum access frequency, calculate the similarity value between the first home gateway device and each of the M target home gateway devices.
8. The method according to claim 1, characterized in that, The first access device data includes the first device type of the first access device, and the first home gateway device has a connection relationship with the second access device; the M target home gateway devices include a first type of target home gateway device and a second type of target home gateway device, the first type of target home gateway device also has a connection relationship with the second access device and the third access device, the second type of target home gateway device also has a connection relationship with the third access device, wherein the device types of the first access device, the second access device and the third access device are different; The step of determining the similarity value among the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device includes: Based on the first device type of the first access device, the second device type of the second access device, and the third device access type of the third access device, obtain the first device type weight value corresponding to the first device type, the second device type weight value corresponding to the second device type, and the third device type weight value corresponding to the third device type, respectively. From the target connection data of M target home gateway devices connected to the first access device, extract the first minimum access frequency of the first access device accessing the M target home gateway devices within a preset time period; and from the target connection data of the first type of target home gateway devices connected to the second access device, extract the second minimum access frequency of the second access device accessing the first type of target home gateway devices within a preset time period; and from the target connection data of the first type of target home gateway devices and the second type of target home gateway devices connected to the third access device, extract the third minimum access frequency of the third access device accessing the first type of target home gateway devices and the first type of target home gateway devices within a preset time period. The first home gateway device and the first device weight value, and the second minimum access frequency and the second device weight value are weighted and summed to obtain a similarity value between the first home gateway device and the first type of target home gateway device; the first minimum access frequency and the first device weight value, and the third minimum access frequency and the third device weight value are weighted and summed to obtain a similarity value between the first type of target home gateway device and the second type of target home gateway device; and the first home gateway device and the second type of target home gateway device are obtained based on the first minimum access frequency and the first device weight value.
9. The method according to claim 1, characterized in that, The target connection data includes first target connection data within a first preset time period and second target connection data within a second preset time period; The step of determining the similarity value among the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device includes: Based on the similarity values among the M target home gateway devices within the first preset time period and the similarity values among the M target home gateway devices within the second preset time period, a first average similarity value among the M target home gateway devices is obtained. If the first average similarity value is greater than or equal to the preset similarity value, it is determined that at least two of the M target home gateway devices are associated with each other.
10. The method according to claim 1, characterized in that, The method further includes: If the similarity value among the M target home gateway devices is determined to be less than a preset similarity value, the latest target connection data of each of the M target home gateway devices is obtained; Based on the latest target connection data of each target home gateway device, determine the latest similarity value among the M target home gateway devices; Calculate the second average similarity value among the M target home gateway devices based on the similarity values among the M target home gateway devices and the latest similarity values among the M target home gateway devices; If the second average similarity value is greater than or equal to the preset similarity value, it is determined that at least two of the M target home gateway devices are associated with each other.
11. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire home data from N home gateway devices. The home data includes device data and connection data. The device data includes gateway device data of the home gateway devices and access device data of access devices that are connected to the home gateway devices. The connection data is the data of the access devices connecting to the home gateway devices. N is an integer greater than 1. The analysis module is used to analyze the family of each of the N home gateway devices based on the family data of the N home gateway devices, and obtain the family attribute information of the family of each home gateway device. The filtering module is used to filter a first home gateway device from the N home gateway devices according to the home attribute information. The target home attribute information of the home where the first home gateway device is located matches the preset home attribute information. The first home gateway device has a connection relationship with the first access device. The determining module is configured to determine a similarity value among the M target home gateway devices based on the first access device data of the first access device and the target connection data of the M target home gateway devices connected to the first access device. The M target home gateway devices include the first home gateway device and at least one second home gateway device. The similarity value is used to determine the degree of association between the homes to which each home gateway device in the M target home gateway devices belongs.
12. A computer device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the steps of the data processing method as described in any one of claims 1-10.
13. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the data processing method as described in any one of claims 1-10.
14. A computer program product, characterized in that, The program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the data processing method as described in any one of claims 1-10.
Citation Information
Patent Citations
Data analysis method and system, electronic equipment and storage medium
CN112506063A
Intelligent terminal label labeling method and system based on family relationship and storage medium
CN116910601A