Family identifier generation method and device, equipment, storage medium and program product
Through high-dimensional sparse representation and approximation algorithm processing of adaptive feature sets, dynamic family identifiers are generated, which solves the problem of static and in real-time home unique identifiers in the prior art, and realizes dynamic traceability and rapid retrieval of family identifiers.
Patent Information
- Application Number
- CN202311580513.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the unique household identifiers are usually static and cannot be updated in real time, resulting in the inability to effectively track down under user information changes or house rentals, and the retrieval efficiency is inefficient under massive household data.
By obtaining the adaptive feature set associated with the home identification and the target features of the user's family to be characterized, the adaptive feature set is used for high-dimensional sparse representation, and feature processing is performed through the approximation algorithm to generate the target family identifier used to characterize the user's family to be characterized.
It realizes dynamic traceability and rapid retrieval of home logos, and can effectively match and track home logos in changes in user information or massive data environments.
Smart Images

Figure CN120030194A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of intelligent information technology, and in particular, relates to a method, device, equipment, storage medium and program product for generating a family identifier. Background Art
[0002] At present, the unique identification of a family is usually a static unique identification, for example, it is strongly bound to the address and the householder's information. Once the information changes, the unique identification will be invalidated. Only one user identification and positioning work can automatically obtain the user's home physical location information, user basic information and user permission information. The features used are relatively fixed and do not change, including home physical location information (longitude, latitude and altitude information), user basic information (house number and householder's contact phone number), user permission information (user level, permission type and permission validity period), etc. Since the existing solutions usually only combine static information such as geographic information, and the matching method is fixed and single, it can often only characterize static customers. Once the house is rented or the address changes, the home service provider or public department cannot update in real time and continue to track. In addition, in the case of a large number of families, the direct unique family identification matching operation is inefficient, and there is a problem of slow direct retrieval under massive family data. Summary of the invention
[0003] The embodiments of the present application provide a method, apparatus, device, storage medium and program product for generating a family identification, which help to achieve traceability and rapid retrieval of the family identification.
[0004] In a first aspect, an embodiment of the present application provides a method for generating a family identifier, the method comprising:
[0005] Acquire an adaptive feature set associated with a family identifier and a target feature of a user family to be characterized; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes device type features corresponding to the user family;
[0006] The target features are represented in a high-dimensional sparse manner using an adaptive feature set to obtain the initial representation information of the user family to be represented;
[0007] An approximate algorithm is used to perform approximate feature processing on the initial characterization information to generate a target family identifier for characterizing the family of the user to be characterized.
[0008] In some possible implementations, before using the adaptive feature set to perform high-dimensional sparse representation on the target feature to obtain initial representation information of the user family to be represented, the family identifier generation method further includes:
[0009] The adaptive feature set is processed by the bag-of-words algorithm to obtain a dictionary of family identification features;
[0010] The adaptive feature set is used to perform high-dimensional sparse representation of the target features to obtain the initial representation information of the user family to be represented, including:
[0011] According to the family identification feature dictionary, the features in the target adaptive feature set are labeled to obtain initial representation information.
[0012] In some possible implementations, the initial characterization information is processed with an approximate algorithm to generate a target family identifier for characterizing the family of the user to be characterized, including:
[0013] Perform feature dimension reduction processing on the initial representation information according to the minimum hash algorithm to obtain the target family identification;
[0014] Alternatively, the initial representation information is hashed and bucketed according to a local sensitive hashing algorithm to generate a bucketed target family identifier;
[0015] Alternatively, the initial representation information is subjected to feature dimensionality reduction processing according to the minimum hash algorithm to obtain intermediate representation information after dimensionality reduction processing; the intermediate representation information is subjected to hash bucketing processing according to the local sensitive hashing algorithm to obtain the target family identification after bucketing.
[0016] In some possible implementations, after obtaining the bucketed target household identifier, the household identifier generation method further includes:
[0017] Receiving a search input for a target family identifier, the search input including an address of a target hash bucket, where the target hash bucket is a hash bucket where the target family identifier is located;
[0018] Based on the address of the target hash bucket, the target family identifier is retrieved in the target hash bucket.
[0019] In some possible implementations, after generating a target family identifier for representing a family of a user to be represented, the family identifier generation method further includes:
[0020] Calculate the similarity between the target family ID and N family IDs in the target hash bucket, where N is a positive integer; the N family IDs do not include the target family ID; the target hash bucket is the hash bucket where the target family ID is located;
[0021] When the similarity between the target family identifier and the first family identifier is greater than a preset threshold, the target family identifier is determined as the identifier of the target user family, and the target user family is the user family corresponding to the first family identifier.
[0022] In some possible implementations, before obtaining the adaptive feature set associated with the family identifier and the target feature to be used to characterize the user's family, the family identifier generation method includes:
[0023] Obtain a historical user family feature set, the historical user family feature set includes M types of initial candidate features, where M is a positive integer; the M types of initial candidate features include initial device type features corresponding to the user family;
[0024] The feature selection algorithm is used to adaptively select features in the historical user family feature set to obtain an adaptive feature set.
[0025] In some possible implementations, a feature selection algorithm is used to adaptively select features in a historical user family feature set to obtain an adaptive feature set, including:
[0026] Perform feature aggregation and dimensionality reduction on the initial device type features in the historical user family feature set to obtain candidate device type features corresponding to the user family;
[0027] The feature selection algorithm is used to adaptively select the candidate device type features and the features other than the initial device type in the M types of initial candidate features to obtain an adaptive feature set.
[0028] Based on the same inventive concept, in a second aspect, an embodiment of the present application provides a family identification generation device, the family identification generation device comprising:
[0029] A first acquisition module is used to acquire an adaptive feature set associated with a family identifier and a target feature of a user family to be characterized; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes a device type feature corresponding to the user family;
[0030] The first characterization module is used to perform high-dimensional sparse characterization on the target features using an adaptive feature set to obtain initial characterization information of the user family to be characterized;
[0031] The first processing module is used to perform approximate feature processing on the initial characterization information using an approximate algorithm to generate a target family identifier for characterizing the family of the user to be characterized.
[0032] In a third aspect, an embodiment of the present application provides a family identification generation device, the family identification generation device comprising:
[0033] a processor and a memory storing computer program instructions;
[0034] When the processor executes the computer program instructions, it implements the family identification generation method provided in any one of the above-mentioned embodiments of the present application.
[0035] In a fourth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, a family identification generation method as provided in any one of the above-mentioned embodiments of the present application is implemented.
[0036] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes a family identification generation method provided in any one of the above-mentioned embodiments of the present application.
[0037] It can be seen from the above description that the family identification generation method, device, equipment, storage medium and program product provided in the embodiment of the present application obtains an adaptive feature set associated with the family identification and a target feature of the user family to be characterized, and the adaptive feature set includes the device type feature corresponding to the user family. Then, the adaptive feature set is used to perform high-dimensional sparse representation on the target feature to obtain the initial representation information of the user family to be characterized. In this way, after obtaining the above-mentioned high-dimensional sparse initial representation information, an approximate algorithm is used to perform approximate feature processing on the initial representation information, and finally a target family identification for representing the above-mentioned user family to be characterized is generated. Compared with the prior art, the family identification generation method, device, equipment, storage medium and program product of the embodiment of the present application, on the one hand, uses an adaptive feature set including device type features to perform high-dimensional sparse representation on the user family to be characterized, so as to introduce device type information in the feature design of the family identification, which helps to achieve dynamic traceability of the family identification. On the other hand, after obtaining the initial representation information of the above-mentioned high-dimensional sparse representation, an approximate algorithm is used to perform approximate simplification processing on the initial representation information, so that the family identification after the approximate simplification processing has rapid retrieval in subsequent applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0039] Figure 1 It is a flowchart of a method for generating a family identifier provided in an embodiment of the present application;
[0040] Figure 2 It is a schematic diagram of a process for generating an adaptive feature set provided by an embodiment of the present application;
[0041] Figure 3 This is a schematic diagram of a scenario flow of a method for generating a family identifier provided in an embodiment of the present application;
[0042] Figure 4 It is a structural diagram of a family identification generating device provided in an embodiment of the present application;
[0043] Figure 5 It is a structural diagram of a family identification generation device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0044] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0045] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0046] As mentioned in the background technology section, the related technology often generates a static household unique identifier based on fixed information such as geographic location. However, users may change their home address, change their mobile phone number / broadband / device, add new family members, and landlords may rent out their houses, and tenants may use the landlord's broadband and landlord's appliances. Static household unique identifiers are no longer practical. In addition, in the case of a large number of households, directly performing unique identifier matching operations is too slow and needs to be optimized.
[0047] Based on this, in order to fully serve the cloud-based home customer operations, how to more reasonably determine the user's family identification so that operators can continue to track users and provide customers with longer-term, more continuous, and higher-quality services on a family basis is one of the issues that need to be urgently addressed in the industry.
[0048] In view of the above, in order to solve the problems of the prior art, the embodiments of the present application provide a method, device, equipment, computer storage medium and program product for generating a family identifier. It should be noted that the embodiments provided in the present application are not intended to limit the scope of the present application.
[0049] The following first introduces the family identification generation method provided in the embodiment of the present application.
[0050] Figure 1 The flowchart of the method for generating a family identifier provided by an embodiment of the present application is shown. The method for generating a family identifier is applied to an electronic device, which may include a server or a user terminal. Figure 1 As shown, the family identification generation method includes the following steps:
[0051] S110, obtaining an adaptive feature set associated with a family identifier and a target feature of a user family to be characterized; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes device type features corresponding to the user family;
[0052] S120, performing high-dimensional sparse representation of the target feature using the adaptive feature set to obtain initial representation information of the user family to be represented;
[0053] S130, performing approximate feature processing on the initial characterization information using an approximate algorithm to generate a target family identifier for characterizing the family of the user to be characterized.
[0054] From the above description, it can be seen that the embodiment of the present application provides a method for generating a family identifier, by obtaining an adaptive feature set associated with the family identifier and a target feature of the user family to be characterized, wherein the adaptive feature set includes the device type feature corresponding to the user family. Then, the adaptive feature set is used to perform high-dimensional sparse representation of the target feature to obtain the initial representation information of the user family to be characterized. In this way, after obtaining the above-mentioned high-dimensional sparse initial representation information, an approximate algorithm is used to perform approximate feature processing on the initial representation information, and finally a target family identifier for representing the above-mentioned user family to be characterized is generated.
[0055] Compared with the prior art, a method for generating a family identification in an embodiment of the present application, on the one hand, uses an adaptive feature set including device type features to perform high-dimensional sparse representation of the user family to be represented, so as to introduce device type information into the feature design of the family identification, which helps to achieve dynamic traceability of the family identification. On the other hand, after obtaining the initial representation information of the above-mentioned high-dimensional sparse representation, an approximate algorithm is used to perform approximate simplification processing on the initial representation information, so that the family identification after the approximate simplification processing has rapid retrieval in subsequent applications.
[0056] Before introducing the specific implementation methods of the above steps in detail, the significance of the core evaluation indicators for family identification in this application is first explained.
[0057] The characteristics of the family identification (family dynamic unique identification) in this application are its stability, uniqueness and rapid retrieval in massive data sets. Among them, stability and uniqueness are two core evaluation indicators of dynamic unique identification (see the device dynamic unique identification part in the known technology for details). Specifically for the family dynamic unique identification, its practical significance is:
[0058] (1) Stability: When some information changes slowly, such as when a user changes their home address, changes their mobile phone number / broadband / device, or adds new family members, the dynamic unique identifier can still identify the new family as the same family as the existing family.
[0059] (2) Uniqueness: Even if some features of the two households overlap (the tenant uses the landlord's broadband, the landlord's home appliances, etc.), the two households can still be distinguished. At the operator level, the household dynamic unique identifier can be applied to cloud-based household customer operations, allowing operators to continuously track users and provide customers with longer-term, more continuous, and higher-quality services on a household basis. Therefore, this application refers to the stability and uniqueness of the household dynamic unique identifier as traceability.
[0060] (3) Fast retrieval: High-dimensional features are used to ensure the uniqueness of the dynamic unique identifier of a family. The higher the feature dimension, the lower the possibility that different families will generate the same identifier. However, representations with such high-dimensional features usually have a very slow matching speed when directly retrieved and are not suitable for use in large data sets. In the case of a large number of families, the efficiency of family identifier matching based on high-dimensional sparse features is low, and there is a problem that direct retrieval is too slow under massive family data.
[0061] The specific implementation of the above steps 110 to 130 is described in detail below.
[0062] In S110, in specific implementation, an adaptive feature set associated with the family identifier and a target feature to characterize the user family are obtained. The adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set may include device type features corresponding to the user family.
[0063] In this embodiment, considering that the number and type of devices in each household are different, and each household generally has few devices, the subsequent household identification characterization can be performed through the adaptive feature set including the device type feature, which can effectively achieve the traceability of the user's household.
[0064] According to some embodiments of the present application, more specifically, the above-mentioned device type characteristics may include at least one of device manufacturer characteristic information, device model characteristic information and device system characteristic information.
[0065] It should be noted that, in some other feasible methods, the above-mentioned device type characteristics may also include other relevant device characteristic information, such as device factory time characteristic information, device coding characteristic information, etc., and this application does not impose specific restrictions on this.
[0066] According to some embodiments of the present application, optionally, in order to ensure the accuracy and validity of the characteristic information under the device type characteristics, the characteristic information under the above device type characteristics are all characteristic information under the same communication protocol. The communication protocol can be, for example, a WiFi (wireless fidelity, wireless connection) protocol or a Bluetooth protocol, which is not strictly limited in this embodiment.
[0067] According to some embodiments of the present application, optionally, in order to more reasonably obtain the above-mentioned adaptive feature set, please refer to the following Figure 2 , Figure 2 It is a schematic diagram of a process for generating an adaptive feature set provided in an embodiment of the present application.
[0068] Specific as Figure 2 As shown, before obtaining the adaptive feature set associated with the family identifier and the target feature to be characterized for the user's family, the family identifier generation method may include:
[0069] Obtain a historical user family feature set, where the historical user family feature set may include M types of initial candidate features, where M is a positive integer; the M types of initial candidate features may include initial device type features corresponding to the user family;
[0070] The feature selection algorithm is used to adaptively select features in the historical user family feature set to obtain an adaptive feature set.
[0071] by Figure 2 For example, the M-type initial candidate features may include user-type features, family-type features, initial device type features, and other class features. In some other possible implementations, the above-mentioned M-type initial candidate features may also be other class feature combinations, which is not specifically limited in this application.
[0072] The user class characteristics may include at least one of the user's personal characteristic information and the user's service characteristic information. For example, the user's personal characteristic information may include the user's age, name, gender and other basic personal information, and the user's service characteristic information may include whether the user has subscribed to broadband, the use of the family network and other service information.
[0073] The above-mentioned family characteristics may include at least one of family relationship characteristic information and family service characteristic information. For example, the family relationship characteristic information may include known family relationship information (parent-child relationship, etc.), inference relationship information such as frequent connection and co-occurrence under the same broadband. The above-mentioned family service characteristic information may include service information such as broadband payer change, etc., which is not limited in this embodiment.
[0074] The above-mentioned initial device type characteristics may include at least one of device manufacturer characteristic information, device model characteristic information, device system characteristic information, device production time characteristic information, device coding characteristic information and other related device characteristic information, etc. This application does not impose any specific restrictions on this.
[0075] Furthermore, the feature information under the above initial device type feature can all be feature information under the same communication protocol. In this way, by introducing the device type under the same home communication protocol in the process of building the adaptive feature set, both the uniqueness of traceability (many feature types, not easy to repeat) and the stability of traceability (devices change slowly, similarities remain unchanged) are guaranteed.
[0076] In a specific implementation, the above historical user family feature set can be obtained by collecting relevant information of historical user families, cleaning, analyzing, and designing a family feature set. The historical user family feature set can specifically include M types of initial candidate features.
[0077] Thus, after obtaining the above historical user family feature set, the feature selection algorithm can be used to adaptively select the M types of initial candidate features in the historical user family feature set to obtain an adaptive feature set. The above feature selection algorithm can be, for example, a random forest selection algorithm, etc., and this application does not make strict restrictions here.
[0078] It should be noted that, in other embodiments, the features under the above-mentioned M types of initial candidate features can also be customized and selected through manual screening, so as to more flexibly realize the design of an adaptive feature set associated with the family identification.
[0079] In this embodiment, adaptive selection of features is performed based on the acquired historical user family feature set, and some of the designed features are adaptively selected based on actual conditions such as historical information and correlation. The design and selection of different features can make the subsequently generated family identification have different emphasis on stability and uniqueness, and the feature design and adaptive selection that can balance stability and uniqueness will ensure that different user family identifications have good traceability.
[0080] According to some embodiments of the present application, optionally, in order to reduce the difficulty of adaptive feature selection and to better ensure that different user family identifiers have good traceability, the feature selection algorithm is used to adaptively select features in the historical user family feature set to obtain an adaptive feature set, which may include:
[0081] Perform feature aggregation and dimensionality reduction on the initial device type features in the historical user family feature set to obtain candidate device type features corresponding to the user family;
[0082] The feature selection algorithm is used to adaptively select the candidate device type features and the features other than the initial device type in the M types of initial candidate features to obtain an adaptive feature set.
[0083] In specific implementation, when performing feature aggregation and dimensionality reduction on the above-mentioned initial device type features, for example, historically devices of the same brand with the same changes can be aggregated into the same device, etc., to achieve dimensionality reduction processing of complex features, thereby simplifying the screening difficulty during adaptive selection, and facilitating a more concise and efficient characterization of user households in the future.
[0084] In S120, when specifically implemented, the above-mentioned adaptive feature set can be used to perform high-dimensional sparse representation on the target features of the user to be represented, and obtain the initial representation information of the user's family to be represented. For example, the features in the adaptive feature set can be sorted to obtain an ordered code representing the above-mentioned target features, etc. This embodiment does not limit this, and other feasible feature representation methods can also be used in this application.
[0085] According to some embodiments of the present application, more specifically, in order to more reasonably characterize the target features of the user family to be characterized, before the target features are subjected to high-dimensional sparse characterization using the adaptive feature set to obtain initial characterization information of the user family to be characterized, the family identifier generation method may further include:
[0086] The adaptive feature set is processed by the bag-of-words algorithm to obtain a family identification feature dictionary;
[0087] The target features are sparsely represented using an adaptive feature set in high dimension to obtain the initial representation information of the user family to be represented, which may include:
[0088] According to the family identification feature dictionary, the features in the target adaptive feature set are labeled to obtain initial representation information.
[0089] The bag-of-words algorithm is an algorithm that unifies feature dimensions so that vectors of different dimensions can be compared with each other and similarity calculations can be achieved. In this embodiment, after obtaining the above-mentioned adaptive feature set, the above-mentioned family identification feature dictionary is constructed based on the bag-of-words algorithm, and finally the target features of the user family to be represented are sparsely represented in a high dimension based on the generated family identification feature dictionary, thereby obtaining the initial representation information of the user family to be represented that is traceable, and the initial representation information can be specifically embodied as a high-dimensional sparse matrix.
[0090] In this way, when characterizing the characteristics of different user families, the bag-of-words algorithm can unify the representation dimensions of different user families through feature extraction and matrix expression, thereby calculating the distance between two families. In other words, the bag-of-words algorithm unifies the feature dimensions of family identification, solving the problem that the dimensions of family representation are different due to the different number and type of devices in each family, which makes it impossible to compare and track them.
[0091] In S130, during the specific implementation, it is considered that if the initial representation information obtained by the above high-dimensional sparse representation is directly used as the family identifier, then in the subsequent actual business scenarios, there will be a problem of slow retrieval under high-dimensional sparse features and massive data sets.
[0092] Based on this, after obtaining the initial characterization information of the user family to be characterized, an approximate feature processing is performed on the initial characterization information using an approximate algorithm to generate a target family identifier for characterizing the user family to be characterized.
[0093] According to some embodiments of the present application, in order to effectively ensure that the family identifier can be quickly retrieved in subsequent application scenarios, it is possible to consider reducing the dimension of the high-dimensional and sparse initial representation information. Based on this, the above-mentioned use of an approximate algorithm to perform approximate feature processing on the initial representation information to generate a target family identifier for representing the user family to be represented may include:
[0094] The initial representation information is processed for feature dimension reduction according to the minimum hash algorithm to generate the target family identification.
[0095] According to some embodiments of the present application, optionally, based on similar reasons, in order to effectively ensure that the family identifier can be quickly retrieved in subsequent application scenarios, the above-mentioned use of an approximate algorithm to perform approximate feature processing on the initial characterization information to generate a target family identifier for characterizing the user family to be characterized may include:
[0096] The initial representation information is hashed and bucketed according to the local sensitive hashing algorithm to generate the bucketed target family identifier.
[0097] According to some embodiments of the present application, more specifically, the above-mentioned using an approximate algorithm to perform approximate feature processing on the initial characterization information to generate a target family identifier for characterizing the user family to be characterized may include:
[0098] Performing feature dimensionality reduction processing on the initial representation information according to the minimum hash algorithm to obtain intermediate representation information after dimensionality reduction processing;
[0099] The intermediate representation information is hashed and bucketed according to the local sensitive hashing algorithm to obtain the target family identification after bucketing.
[0100] This embodiment designs a method that combines MinHash, Locality-Sensitive Hashing (LSH) with high-dimensional and sparse initial representation information to reduce the dimension and bucket the initial representation information, thereby greatly improving the rapid retrieval of the target family identifier in subsequent business scenarios.
[0101] According to some embodiments of the present application, in combination with actual application scenarios, after generating the bucketed target household identifier, the household identifier generation method may further include:
[0102] Receiving a search input for a target family identifier, the search input may include an address of a target hash bucket, where the target hash bucket is a hash bucket where the target family identifier is located;
[0103] Based on the address of the target hash bucket, the target family identifier is retrieved in the target hash bucket.
[0104] In specific implementation, after the above-mentioned target family identifier is hashed and bucketed, in actual retrieval, the target family identifier can be compared only with other family identifiers in the target hash bucket through "bucket" matching, which fully narrows the retrieval scope of the above-mentioned target family identifier, thereby greatly improving the retrieval efficiency.
[0105] According to some embodiments of the present application, optionally, after generating the target family identifier for representing the family of the user to be represented, the family identifier generation method may further include:
[0106] Calculate the similarity between the target family ID and N family IDs in the target hash bucket, where N is a positive integer; the N family IDs cannot include the target family ID; the target hash bucket is the hash bucket where the target family ID is located;
[0107] When the similarity between the target family identifier and the first family identifier is greater than a preset threshold, the target family identifier is determined as the identifier of the target user family, and the target user family is the user family corresponding to the first family identifier.
[0108] It should be noted that, when performing similarity calculation, the similarity calculation methods that can be used include Euclidean distance, Jaccard distance, sine / cosine distance or Manhattan distance, etc., and this embodiment does not impose strict restrictions on this.
[0109] In specific implementation, the similarity between the target family identifier and the N family identifiers in the target hash bucket is calculated respectively, and the target hash bucket is the hash bucket where the target family identifier is located. When the similarity between the target family identifier and the first family identifier is greater than a preset threshold, the target family identifier is determined as the identifier of the target user family, and the target user family is the user family corresponding to the first family identifier. In other words, when the similarity between the target family identifier and the first family identifier is high, it can be considered that the two family identifiers correspond to the same user family.
[0110] In a specific example, considering that the changes in family devices are often purchased in batches, even if the same family purchases new devices, the similarity of their devices is still very high. Whether it is family user information or device information, when it is not a sudden change, the family similarity before and after the change is very high, so two families with extremely high similarity can be merged into the same family, thus realizing the characteristic that the family dynamic unique identification is not affected by the slow changes in user information.
[0111] In order to facilitate understanding of the family identification generation method provided in the above embodiment, the above method is described below using a specific scenario embodiment. Figure 3 It is a scenario flow diagram of a family identification generation method provided in one embodiment of the present application.
[0112] In this scenario embodiment, the historical user family feature set can be obtained in advance, and the family identification feature dictionary can be constructed using the bag-of-words algorithm. When it is necessary to characterize the user family to be characterized, the family identification feature dictionary is used to perform feature standardization on the target features to be characterized for the user family, thereby obtaining the original representation of the user family based on the above family identification feature dictionary (high-dimensional sparse matrix), which corresponds to the initial representation information in the above embodiment and has traceability.
[0113] In some embodiments, after obtaining the above historical user family feature set, the features in the historical user family feature set can also be adaptively aggregated and selected to obtain an adaptive feature set. In this way, after obtaining the above adaptive feature set, the bag-of-words algorithm is used to determine the above family identification feature dictionary to perform feature labeling on the user family to be characterized, and obtain the initial characterization information of the user family to be characterized.
[0114] After obtaining the high-dimensional sparse initial representation information, in order to ensure the fast retrieval of the finally generated family identification, the high-dimensional sparse representation representing the family can be further compiled based on the minimum hash and local sensitive hashing algorithms to generate the family identification of the user's family to be represented. In this way, since the compiled family identification dimension is much lower than the high-dimensional sparse initial representation information, and when searching, only group comparison and search within the corresponding hash bucket are required, it has fast retrieval.
[0115] Based on the family identification generation method provided in the above embodiment, out of the same inventive concept, the present application also provides a family identification generation device corresponding to the above family identification generation method. Figure 4 A detailed introduction to the family identification generating device is given.
[0116] Figure 4 A schematic diagram of the structure of a family identification generating device provided in one embodiment of the present application is shown.
[0117] Figure 4 The illustrated family identification generating device 400 comprises:
[0118] A first acquisition module 410 is used to acquire an adaptive feature set associated with a family identifier and a target feature to be characterized for a user family; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes a device type feature corresponding to the user family;
[0119] A first characterization module 420 is used to perform high-dimensional sparse characterization on the target feature using an adaptive feature set to obtain initial characterization information of the user family to be characterized;
[0120] The first processing module 430 is used to perform approximate feature processing on the initial characterization information using an approximate algorithm to generate a target family identifier for characterizing the family of the user to be characterized.
[0121] The embodiment of the present application provides a family identification generation device, which sets corresponding functional modules, obtains an adaptive feature set associated with the family identification and a target feature of the user family to be characterized, wherein the adaptive feature set includes the device type feature corresponding to the user family. Then, the adaptive feature set is used to perform high-dimensional sparse characterization on the target feature to obtain the initial characterization information of the user family to be characterized. In this way, after obtaining the above-mentioned high-dimensional sparse initial characterization information, an approximate algorithm is used to perform approximate feature processing on the initial characterization information, and finally a target family identification for characterizing the above-mentioned user family to be characterized is generated.
[0122] Compared with the prior art, a family identification generation device in an embodiment of the present application, on the one hand, uses an adaptive feature set including device type features to perform high-dimensional sparse representation of the user family to be represented, so as to introduce device type information into the feature design of the family identification, which helps to achieve dynamic traceability of the family identification. On the other hand, after obtaining the initial representation information of the above-mentioned high-dimensional sparse representation, an approximate algorithm is used to perform approximate simplification processing on the initial representation information, so that the family identification after the approximate simplification processing has rapid retrieval in subsequent applications.
[0123] According to some embodiments of the present application, optionally, before performing high-dimensional sparse characterization of the target feature using the adaptive feature set to obtain initial characterization information of the user family to be characterized, the family identifier generation device may further include:
[0124] The second processing module can be used to process the adaptive feature set through a bag-of-words algorithm to obtain a family identification feature dictionary;
[0125] The first characterization module 420 uses the adaptive feature set to perform high-dimensional sparse characterization on the target feature to obtain initial characterization information of the user family to be characterized, which may include:
[0126] According to the family identification feature dictionary, the features in the target adaptive feature set are labeled to obtain initial representation information.
[0127] According to some embodiments of the present application, optionally, the first processing module 430 uses an approximate algorithm to perform approximate feature processing on the initial characterization information to generate a target family identifier for characterizing the user family to be characterized, which may include:
[0128] Perform feature dimension reduction processing on the initial representation information according to the minimum hash algorithm to generate the target family identification;
[0129] Alternatively, the initial representation information is hashed and bucketed according to a local sensitive hashing algorithm to obtain a bucketed target family identifier;
[0130] Alternatively, the initial representation information is subjected to feature dimensionality reduction processing according to the minimum hash algorithm to obtain intermediate representation information after dimensionality reduction processing; the intermediate representation information is subjected to hash bucketing processing according to the local sensitive hashing algorithm to obtain the target family identification after bucketing.
[0131] According to some embodiments of the present application, optionally, after obtaining the bucketed target household identifier, the household identifier generation device may further include:
[0132] A first receiving module may be used to receive a search input for a target family identifier, where the search input may include an address of a target hash bucket, where the target hash bucket is a hash bucket where the target family identifier is located;
[0133] The first retrieval module can be used to retrieve the target family identifier in the target hash bucket based on the address of the target hash bucket.
[0134] According to some embodiments of the present application, optionally, after generating the target family identifier for representing the family of the user to be represented, the family identifier generating device may further include:
[0135] The similarity calculation module can be used to calculate the similarity between the target family identifier and N family identifiers in the target hash bucket, where N is a positive integer; the N family identifiers cannot include the target family identifier; the target hash bucket is the hash bucket where the target family identifier is located;
[0136] The first determination module can be used to determine the target family identifier as the identifier of the target user family when the similarity between the target family identifier and the first family identifier is greater than a preset threshold, and the target user family is the user family corresponding to the first family identifier.
[0137] According to some embodiments of the present application, optionally, before acquiring the adaptive feature set associated with the family identifier and the target feature to be characterized for the user's family, the family identifier generating device may include:
[0138] The second acquisition module may be used to acquire a historical user family feature set, where the historical user family feature set may include M types of initial candidate features, where M is a positive integer; the M types of initial candidate features may include initial device type features corresponding to the user family;
[0139] The adaptive selection module can be used to adaptively select features in the historical user family feature set using a feature selection algorithm to obtain an adaptive feature set.
[0140] According to some embodiments of the present application, optionally, the above-mentioned adaptive selection module uses a feature selection algorithm to adaptively select features in the historical user family feature set to obtain an adaptive feature set, which may include:
[0141] The first dimensionality reduction submodule can be used to perform feature aggregation and dimensionality reduction on the initial device type features in the historical user family feature set to obtain candidate device type features corresponding to the user family;
[0142] The first selection submodule can be used to adaptively select the candidate device type features and the features other than the initial device type in the M types of initial candidate features using a feature selection algorithm to obtain an adaptive feature set.
[0143] Based on the family identification generation method provided in the above embodiment, out of the same inventive concept, the present application also provides a family identification generation device corresponding to the above family identification generation method. Figure 5 A detailed introduction to the home identity generation device.
[0144] See below Figure 5 , Figure 5 It is a structural diagram of a family identification generation device provided in one embodiment of the present application.
[0145] The family identification generating device may include a processor 501 and a memory 502 storing computer program instructions.
[0146] Specifically, the processor 501 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0147] The memory 502 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 502 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 502 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 502 is a non-volatile solid-state memory.
[0148] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0149] The processor 501 implements any one of the family identification generation methods in the above embodiments by reading and executing computer program instructions stored in the memory 502.
[0150] In one example, the data family identifier generating device may further include a communication interface 503 and a bus 510. Figure 5 As shown, the processor 501, the memory 502, and the communication interface 503 are connected via a bus 510 and communicate with each other.
[0151] The communication interface 503 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0152] Bus 510 includes hardware, software or both, and the parts of the family identification generating device are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransmission (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 510 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0153] The family identification generation device executes the family identification generation method in the embodiment of the present application, thereby realizing the family identification generation method described in the embodiment of the present application.
[0154] In addition, in combination with the family identification generation method in the above embodiment, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any family identification generation method in the above embodiment is implemented.
[0155] Based on the family identification generation method in the above-mentioned embodiment, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the family identification generation method provided in any one of the above-mentioned embodiments of the present application.
[0156] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0157] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0158] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0159] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0160] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for generating a family identifier, It is characterized in that The method comprises: Acquire an adaptive feature set associated with a family identifier and a target feature of a user family to be characterized; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes device type features corresponding to the user family; Using the adaptive feature set to perform high-dimensional sparse representation on the target feature, to obtain initial representation information of the user family to be represented; An approximate feature processing is performed on the initial characterization information using an approximate algorithm to generate a target family identifier for characterizing the user family to be characterized.
2. The method according to claim 1, It is characterized in that Before performing high-dimensional sparse characterization on the target feature using the adaptive feature set to obtain initial characterization information of the user family to be characterized, the method further includes: Processing the adaptive feature set by using a bag-of-words algorithm to obtain a family identification feature dictionary; The step of performing high-dimensional sparse characterization on the target feature by using the adaptive feature set to obtain initial characterization information of the user family to be characterized includes: The features in the target adaptive feature set are labeled according to the family identification feature dictionary to obtain the initial characterization information.
3. The method according to claim 1, It is characterized in that The using an approximate algorithm to perform approximate feature processing on the initial characterization information to generate a target family identifier for characterizing the user family to be characterized includes: Performing feature dimension reduction processing on the initial representation information according to a minimum hash algorithm to generate the target family identifier; Alternatively, the initial representation information is hashed and bucketed according to a local sensitive hashing algorithm to obtain the target household identifier after bucketing; Alternatively, the initial representation information is subjected to feature dimensionality reduction processing according to a minimum hash algorithm to obtain intermediate representation information after dimensionality reduction processing; the intermediate representation information is subjected to hash bucketing processing according to a locally sensitive hashing algorithm to obtain the target family identifier after bucketing.
4. The method according to claim 3, It is characterized in that After obtaining the bucketed target household identifier, the method further includes: Receiving a search input for the target family identifier, wherein the search input includes an address of a target hash bucket, where the target hash bucket is a hash bucket where the target family identifier is located; Based on the address of the target hash bucket, the target household identifier is retrieved from the target hash bucket.
5. The method according to claim 3, It is characterized in that After generating the target household identifier for characterizing the household of the user to be characterized, the method further includes: Calculate the similarity between the target family identifier and N family identifiers in a target hash bucket, where N is a positive integer; the N family identifiers do not include the target family identifier; the target hash bucket is the hash bucket where the target family identifier is located; When the similarity between the target family identifier and the first family identifier is greater than a preset threshold, the target family identifier is determined as the identifier of the target user family, and the target user family is the user family corresponding to the first family identifier.
6. The method according to claim 1, It is characterized in that Before acquiring the adaptive feature set associated with the family identifier and the target feature to be characterized for the user family, the method includes: Acquire the historical user family feature set, wherein the historical user family feature set includes M types of initial candidate features, where M is a positive integer; the M types of initial candidate features include initial device type features corresponding to the user family; The features in the historical user family feature set are adaptively selected using a feature selection algorithm to obtain the adaptive feature set.
7. The method according to claim 6, It is characterized in that The features in the historical user family feature set are adaptively selected using a feature selection algorithm to obtain the adaptive feature set, including: Performing feature aggregation and dimensionality reduction on the initial device type features in the historical user family feature set to obtain candidate device type features corresponding to the user family; The feature selection algorithm is used to adaptively select the features of the candidate device type and the features of the M types of initial candidate features except the initial device type to obtain the adaptive feature set.
8. A family identification generating device, It is characterized in that The device comprises: A first acquisition module is used to acquire an adaptive feature set associated with a family identifier and a target feature to characterize a user family; the adaptive feature set is determined based on a historical user family feature set, and the adaptive feature set includes a device type feature corresponding to the user family; A first characterization module, configured to perform high-dimensional sparse characterization on the target feature using the adaptive feature set to obtain initial characterization information of the user family to be characterized; The first processing module is used to perform approximate feature processing on the initial characterization information by using an approximate algorithm to generate a target family identifier for characterizing the user family to be characterized.
9. A family identification generating device, It is characterized in that The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the family identification generation method as described in any one of claims 1-7.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the family identification generation method as described in any one of claims 1-7.
11. A computer program product, It is characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the family identification generation method as described in any one of claims 1-7.