A user tag construction method and device, electronic equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU XINGYUN INTELLIGENT TECH CO LTD
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-04
AI Technical Summary
该类建模划分逻辑固化,难以充分挖掘数据内在关联规律,生成的标签精细化程度较差,标签准确度受限
[0013] This disclosed method, apparatus, electronic device, and storage medium for constructing user tags integrate static basic user data and dynamic business data to build a user feature dataset. Based on multi-source data, it can more realistically reconstruct a comprehensive user profile. A target analysis model extracts shallow explicit features and deep implicit features respectively, generating candidate user feature categories with confidence levels. Then, an independently set target verification model verifies and optimizes the candidate results, filtering out compliant target user feature categories. Finally, it combines these with a standard user tag library to match and output target user tags. This approach ensures comprehensive data sources and rich feature mining dimensions, while relying on the division of labor between the front and rear models for verification and the elimination of unreasonable classification results at each level. It avoids tag bias caused by errors in single-model calculations, effectively improving the accuracy and standardization of the final generated user tags.
Smart Images

Figure CN122509952A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for constructing user tags. Background Technology
[0002] In customer relationship management (CRM) and precision marketing, customer profiling technology is a crucial technological support for implementing personalized services and improving marketing conversion efficiency. The industry generally relies on mining user attributes from massive amounts of customer data to generate profile tags that closely match users' actual behavioral characteristics. Currently, most mainstream customer tagging solutions employ rule engines or traditional statistical models, relying on manually preset business rules or mathematical statistical results to complete customer segmentation and tag assignment. This type of modeling and segmentation logic is rigid, making it difficult to fully explore the inherent correlations within the data, resulting in tags with poor granularity and limited accuracy. Summary of the Invention
[0003] This disclosure provides a user tag construction method, apparatus, electronic device, and storage medium to at least solve the above-mentioned technical problems existing in the prior art.
[0004] According to a first aspect of this disclosure, a method for constructing user tags is provided, the method comprising: Obtain the static basic data and dynamic business data corresponding to the user to obtain the user feature dataset; The user feature dataset is subjected to feature extraction using a target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, candidate user feature categories and their corresponding confidence levels are obtained. The target verification model verifies the candidate user feature categories based on preset verification rules and the confidence levels corresponding to each candidate user feature category, and obtains the verification results. If the verification results meet the first verification condition, the candidate user feature categories are retained or modified to obtain the target user feature categories. The target user feature categories are matched with a standard user tag library to obtain target user tags.
[0005] In one possible implementation, obtaining the candidate user feature category and corresponding confidence level based on the shallow explicit features and the deep implicit features includes: The target analysis model is used to generate feature weights for each feature. The shallow explicit features and the deep implicit features are weighted and fused based on the feature weights to obtain the initial user feature categories and their corresponding confidence levels. Based on a preset confidence threshold, the initial user feature categories and their corresponding confidence levels that reach the preset confidence threshold are determined as candidate user feature categories and their corresponding confidence levels.
[0006] In one possible implementation, the method further includes: If the verification result meets the second verification condition, the new candidate user feature category and its corresponding confidence level are determined by the target analysis model.
[0007] In one possible implementation, determining new candidate user feature categories and their corresponding confidence levels through the target analysis model includes: Obtain weight constraint information, which is generated based on anomaly information in the verification results; Based on the weight constraint information, the target analysis model determines new candidate user feature categories and their corresponding confidence levels.
[0008] In one possible implementation, the method further includes: Generate corresponding tag categories based on the target business area; Determine the standard tag terms corresponding to each tag category to form the standard user tag library.
[0009] In one possible implementation, obtaining the user's static basic data and dynamic business data to obtain a user feature dataset includes: Obtain the user's static basic data and dynamic business data; The static basic data and the dynamic business data are cleaned and standardized to obtain standardized data. The standardized data is integrated to obtain a user feature dataset.
[0010] According to a second aspect of this disclosure, a user tagging apparatus is provided, the apparatus comprising: The data acquisition module is used to acquire the user's static basic data and dynamic business data to obtain the user feature dataset; The analysis module is used to extract features from the user feature dataset using a target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, candidate user feature categories and their corresponding confidence scores are obtained. The verification module is used to verify the candidate user feature categories using the target verification model, based on preset verification rules and the confidence level corresponding to each candidate user feature category, and obtain the verification result; if the verification result meets the first verification condition, the candidate user feature category is retained or modified to obtain the target user feature category. The tag generation module is used to match the target user feature categories with a standard user tag library to obtain target user tags.
[0011] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0012] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described in this disclosure.
[0013] This disclosed method, apparatus, electronic device, and storage medium for constructing user tags integrate static basic user data and dynamic business data to build a user feature dataset. Based on multi-source data, it can more realistically reconstruct a comprehensive user profile. A target analysis model extracts shallow explicit features and deep implicit features respectively, generating candidate user feature categories with confidence levels. Then, an independently set target verification model verifies and optimizes the candidate results, filtering out compliant target user feature categories. Finally, it combines these with a standard user tag library to match and output target user tags. This approach ensures comprehensive data sources and rich feature mining dimensions, while relying on the division of labor between the front and rear models for verification and the elimination of unreasonable classification results at each level. It avoids tag bias caused by errors in single-model calculations, effectively improving the accuracy and standardization of the final generated user tags.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which: In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0016] Figure 1 A schematic diagram illustrating the implementation flow of a user tag construction method according to an embodiment of this disclosure is shown; Figure 2 A schematic diagram of the composition structure of a user tag construction apparatus according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0017] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0018] A first aspect of this disclosure provides a user tag construction method, such as... Figure 1 As shown, the method includes the following steps: Step 101: Obtain the static basic data and dynamic business data corresponding to the user to obtain the user feature dataset.
[0019] Static basic data refers to a user's long-term, fixed profile information, including inherent personal attributes such as age, gender, place of residence, and occupation. Dynamic business data refers to data that continuously changes with user business behavior, and can be mainly divided into behavioral data, interaction data, and transaction data. Behavioral data represents user browsing, clicking, page dwell time, and other product operation trajectory information; interaction data represents user interaction records in scenarios such as community posts, customer service inquiries, and content comments; transaction data represents order and financial records such as order payments, returns and exchanges, and consumption amounts.
[0020] Static foundational data is accessed via a data interface to an authorized and open customer management database, retrieving archived user profile information. Dynamic business data is collected by integrating with various business data sources, such as consumer business systems and user interaction systems, using methods like program tracking and system log parsing. All types of raw data are stored in a standardized format such as JSON (JavaScript Object Notation), and then undergo data cleaning, format standardization, and multi-field fusion processing to ultimately obtain a standardized user feature dataset.
[0021] Step 102: Extract features from the user feature dataset using the target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, obtain candidate user feature categories and their corresponding confidence levels.
[0022] The target analysis model is a machine learning model suitable for data modeling and analysis. First, it directly reads raw attributes and intuitive behavioral indicators from the user feature dataset, generating shallow explicit features. These features directly correspond to the user's basic information and intuitive behavioral performance already clearly recorded in the raw data, such as user age, region, cumulative order count, and average daily page view duration—indicators directly obtainable from the raw records. Then, the model performs correlation operations on multiple sets of data from different sources to uncover data association patterns not directly reflected in the raw fields, obtaining deep implicit features. These features reflect the user's potential preferences implied by the interaction of multiple data points, such as user preferred product categories and price ranges derived from a combination of browsed categories, price ranges, and interactive comments—information that cannot be directly obtained from a single raw field.
[0023] After feature extraction, the target analysis model autonomously learns from historical samples to calculate the feature weights corresponding to each feature item. These weights are then used to calculate and summarize the shallow explicit features and deep implicit features to obtain comprehensive features. The comprehensive features are then classified according to a preset classification standard to generate initial user feature categories and select candidate user feature categories. Simultaneously, the confidence score for each feature category is output. The candidate user feature categories are the user attribute and preference classification results formed through induction and refinement. The confidence score is used to quantify the degree of credibility between the candidate user feature categories and the actual user profile; for example, the confidence score ranges from 0 to 1.
[0024] Step 103: Using the target verification model, the candidate user feature categories are verified based on preset verification rules and the confidence levels corresponding to each candidate user feature category to obtain verification results; if the verification results meet the first verification condition, the candidate user feature categories are retained or modified to obtain the target user feature categories.
[0025] The target validation model is a machine learning model with a different architecture and computational logic than the target analysis model. First, the target validation model retrieves candidate user feature categories and their corresponding confidence levels for secondary screening. On one hand, it verifies the reasonableness of the confidence level matching for each feature category based on historical labeled samples, eliminating invalid feature categories with high matching deviations. On the other hand, it filters out feature categories with content logic conflicts or that do not conform to common business sense by combining the established business rules corresponding to the tag landing (i.e., preset validation rules). After these two layers of screening and validation, candidate user feature categories that meet the first validation condition are directly determined as target user feature categories or, after correction, determined as target user feature categories.
[0026] The preset verification rules are business constraints and data verification specifications adapted to customer profile tag construction scenarios. They can be set based on actual marketing business logic, historical tag annotation samples, and user profile matching standards. For example, they may include logically mutually exclusive rules for feature categories, confidence level matching compliance rules, and user behavior and attribute association consistency rules. Logically mutually exclusive rules define mutually exclusive feature categories that cannot coexist; the same user can only match at most one of the mutually exclusive categories. Confidence level matching compliance rules stipulate that feature categories of different importance levels must match the corresponding minimum confidence level threshold, with the confidence level benchmark for core profile tags being higher than that of secondary tags. User behavior and attribute association consistency rules constrain the inherent basic attributes of users and the derived behavioral preferences to avoid contradictory matching relationships that defy common sense. For example, logically mutually exclusive rules limit "low-spending students" and "high-end, high-spending users" to not simultaneously belong to the same user; confidence level matching compliance rules require that the confidence level of core consumption value tags not be lower than 0.7, and the confidence level of secondary preference tags not be lower than 0.6; behavior and attribute association consistency rules constrain underage users from generating profile tags indicating a preference for high-value luxury goods.
[0027] In this embodiment, the first verification condition includes two compliance scenarios: First, the candidate user feature category fully complies with the preset verification rules, the confidence level is reasonably matched, and the overall verification passes; second, the candidate user feature category fails partial verification, but there is no fundamental business logic conflict, and there is room for correction.
[0028] If the verification result falls under the first compliance scenario (i.e., complete verification pass), the target verification model directly retains the candidate user feature category. If the verification result falls under the second compliance scenario (i.e., partial failure but with room for correction), the target verification model combines preset verification rules with historical high-quality sample data to adaptively correct and optimize key information such as feature content. The compliant feature category obtained after retention or correction is the final target user feature category.
[0029] Step 104: Match the target user feature category with the standard user tag library to obtain the target user tag.
[0030] The standard user tag library is a pre-organized and standardized tag set that stores various standardized user tags and their corresponding mappings to user characteristic categories. It can be built based on historical business data, product classification systems, and operational tag specifications, and includes multiple standard entries such as user attribute tags, consumption preference tags, and behavioral characteristic tags. During matching, the target user characteristic category is used as the matching basis. Standardized tags with matching mappings in the standard user tag library are retrieved, and the successfully matched standard tags are identified as the target user tags for that user, thus completing the user profile tag determination.
[0031] This embodiment integrates static user data and dynamic business data to construct a user feature dataset, leveraging multi-source data to more accurately reconstruct comprehensive user profiles. A target analysis model extracts both shallow explicit features and deep implicit features, generating candidate user feature categories with confidence levels. An independently configured target verification model then reviews and optimizes the candidate results, filtering out compliant target user feature categories. Finally, these categories are matched against a standard user tag library to output target user tags. This layered and progressive tag construction process ensures comprehensive data sources and rich feature mining dimensions. Furthermore, the division of labor between the preceding and following model stages for verification and the elimination of unreasonable classification results at each level avoids tag bias caused by errors in a single model, effectively improving the accuracy and standardization of the final generated user tags.
[0032] In another embodiment of this disclosure, obtaining static basic data and dynamic business data corresponding to a user to obtain a user feature dataset includes: obtaining static basic data and dynamic business data corresponding to a user; cleaning and standardizing the static basic data and the dynamic business data to obtain standardized data; and integrating the standardized data to obtain a user feature dataset.
[0033] First, static basic data and dynamic business data are collected through database interface retrieval, service tracking point collection, and system log parsing. Then, data cleaning is performed to remove missing fields and abnormal data from the original data, such as removing age information with values exceeding reasonable thresholds, and incorrectly formatted or logically invalid consumption documents. For critical missing fields, interpolation and other methods are used to fill in the data, ensuring the data source is complete and usable.
[0034] The original data for different categories comes from scattered sources and varies greatly in units and value ranges. Indicators such as age and cumulative consumption amount have different units of measurement and numerical magnitudes, making them unsuitable for direct use in subsequent model calculations. Therefore, after cleaning, normalization and other standardization methods are used to perform a unified conversion, mapping multi-dimensional and different-magnitude indicators to a unified numerical range to generate standardized data.
[0035] Finally, the standardized full data is combined with field concatenation, dimension aggregation, and fusion to complete the construction of the user feature dataset.
[0036] The solution in this embodiment involves sequentially cleaning and completing the collected static basic data and dynamic business data, standardizing their dimensions, and integrating the data. On the one hand, it removes abnormal and dirty data, fills in missing fields, and avoids interference from invalid data in subsequent calculations, ensuring the integrity and reliability of the basic data source. On the other hand, it unifies the data volume of each dimension through standardization methods such as normalization, eliminating calculation errors caused by inconsistencies in the dimensions of different indicators, which is beneficial for the subsequent target analysis model to accurately complete feature weighting calculations and feature category determination.
[0037] In another embodiment of this disclosure, obtaining candidate user feature categories and corresponding confidence levels based on the shallow explicit features and the deep implicit features includes: generating feature weights corresponding to each feature through the target analysis model; performing weighted fusion of the shallow explicit features and the deep implicit features based on the feature weights to obtain initial user feature categories and corresponding confidence levels; and determining the initial user feature categories and corresponding confidence levels that reach the preset confidence threshold as candidate user feature categories and corresponding confidence levels based on a preset confidence threshold.
[0038] First, a target analysis model is used to learn and fit based on historical sample data to solve for the feature weights corresponding to each feature item. These feature weights represent the contribution ratio of a single feature in the user profile determination process. Then, based on the weight coefficients corresponding to each feature, shallow explicit features and deep implicit features are weighted, converted, and summed to generate a comprehensive feature vector that can comprehensively describe the user profile. The comprehensive feature vector is then compared and matched item by item with predefined classification reference standards. Initial user feature categories are determined based on the matching results, and the confidence scores for each category are output simultaneously. The pre-set confidence thresholds can be configured in advance based on practical business experience and historical sample statistics. The confidence scores of all initial user feature categories are compared with the pre-set confidence thresholds. Initial user feature categories with confidence scores greater than or equal to the threshold are retained as candidate user feature categories, and their corresponding confidence scores are stored. After filtering, the candidate user feature categories and their matched confidence scores are stored in an intermediate result storage unit for easy data retrieval in subsequent verification steps.
[0039] The scheme in this embodiment relies on the model to autonomously learn and fit the weight coefficients of various features, accurately quantifying the contribution ratio of different features in the profile determination. Then, it performs a weighted fusion of shallow explicit features and deep implicit features, and uses a pre-set reliability threshold to filter the initial classification results. On the one hand, this optimizes the calculation logic of comprehensive features and improves the scientific nature of feature aggregation results; on the other hand, it eliminates classification content with insufficient reliability from the source, effectively ensuring the overall accuracy of candidate user feature categories.
[0040] In another embodiment of this disclosure, the method further includes: if the verification result satisfies the second verification condition, determining a new candidate user feature category and its corresponding confidence level through the target analysis model.
[0041] The second verification condition refers to situations where the candidate user feature category verification fails, and there are fundamental problems such as serious data bias, business logic conflicts, or severely low confidence levels, leaving no room for correction or optimization, and making it impossible to obtain a valid feature category through correction. In this scenario, the target analysis model is triggered to re-execute the feature extraction, weight fusion, and category determination process. Based on the original user feature dataset, it re-mines deep and shallow user features, recalculates the weights of each feature, and completes feature weighted fusion, thereby generating new candidate user feature categories and corresponding confidence levels. This provides effective data support for subsequent secondary verification and feature selection, avoiding the impact of invalid features on the accuracy of user profile construction.
[0042] Taking the construction of e-commerce user profile tags as an example, if the candidate user feature category generated by the model is "high consumption preference, preference for light luxury products", and the corresponding confidence level is 0.85, which meets the preset confidence threshold and preset verification rules, and satisfies the verification pass condition in the first verification condition, this feature category is directly retained as the target user feature category. If the confidence level of the candidate tag meets the standard, but there is a slight deviation in the tag judgment, such as being misclassified as a high-value customer instead of a mid-value customer based on the total consumption amount, then the correction condition within the first verification condition is met. The target verification model, based on historical labeled samples and business rules, corrects "high-value customer" to "mid-value customer", and the corrected feature category is directly used as the target user feature category. If the confidence level of the candidate feature category does not meet the standard, or there are logical conflicts between tags, seriously violate business common sense, and there is no room for correction, and the second verification condition is met, then the target analysis model is triggered to re-extract user features, perform weighted fusion calculations, generate a new candidate user feature category and corresponding confidence level, and enter a new round of verification iteration.
[0043] This embodiment's solution uses a target verification model combined with preset verification rules and confidence levels to perform secondary verification on candidate user feature categories. It differentiates between first and second verification conditions, handling verification results in a tiered manner. Candidate categories meeting the first verification condition can be directly retained or partially modified to form target user feature categories. For unqualified results meeting the second verification condition but uncorrectable, the target analysis model is driven to regenerate candidate feature categories. This approach allows for flexible correction and utilization of labels with minor deviations, reducing data waste, and also enables recalculation of classification results with fundamental defects. Layer-by-layer control over the quality of user profile classification effectively improves the accuracy and business adaptability of the final target user feature categories.
[0044] In another embodiment of this disclosure, determining new candidate user feature categories and their corresponding confidence levels through the target analysis model includes: obtaining weight constraint information, which is generated based on anomaly information in the verification results; and the target analysis model determining new candidate user feature categories and their corresponding confidence levels based on the weight constraint information.
[0045] After the target verification model completes the verification and determines that the current candidate feature category meets the second verification condition, it generates exclusive weight constraint information based on the logical contradictions, classification misalignments, and other anomalies detected in this verification, and sends this constraint information to the target analysis model. The weight constraint information refers to the adjustment parameters such as the upper and lower limits of the weights, attenuation coefficients, and masking coefficients set for the corresponding feature items causing classification anomalies, used to limit the role of problematic features in modeling calculations. When the target analysis model performs feature weight calculation and feature fusion operations again, it incorporates the aforementioned weight constraint information into the calculation process. Based on the constraint information, it lowers the weight values of anomalous features, reasonably adjusts the contribution ratio of various normal features, and then completes the weighted fusion of deep and shallow features based on the corrected weights, reclassifies the features, and finally outputs the optimized new candidate user feature categories and corresponding confidence scores.
[0046] The solution in this embodiment utilizes verification anomaly information to generate weight constraint information to guide the iterative modeling of the target analysis model, thereby achieving targeted correction of error root causes, avoiding indiscriminate repetitive calculations, and effectively improving the accuracy and iteration efficiency of secondary generation of candidate feature categories.
[0047] In another embodiment of this disclosure, the method further includes: generating corresponding tag categories based on the target business domain; determining standard tag entries corresponding to each tag category to form the standard user tag library.
[0048] Target business areas can be flexibly defined based on actual operational scenarios, such as differentiated business segments like e-commerce retail and online education and training. A dedicated tagging system is built based on the operational needs of the corresponding business area and industry classification standards, dividing the data into multi-level tag categories based on user basic attributes, consumption habits, and product preferences. Standardized tag entries with uniform format are then configured for each category, with content calibrated against industry-standard definitions and historical manually labeled samples. All tag categories and their subordinate standard tag entries are integrated to construct a complete standard user tag library, providing a unified and standardized benchmark for subsequent feature matching and tag output.
[0049] For example, if e-commerce retail is selected as the target business area, three main tag categories can be defined: the first category is basic user attributes, which includes standard terms such as adults, minors, students, and working professionals; the second category is consumer behavior habits, which includes standard terms such as high-value customers, medium-value customers, low-frequency small-amount purchases, and high-frequency stockpiling purchases; and the third category is product preferences, which includes standard terms such as accessible luxury apparel, affordable daily necessities, digital appliances, and beauty and skincare products. This entire set of terms is standardized and unambiguous, forming a standard user tag library specifically for e-commerce scenarios. When the verified target user characteristics are classified as working professionals, high-value customers, or those who prefer accessible luxury apparel, the corresponding standard terms in the library can be accurately matched to generate compliant target user tags.
[0050] The verified target user feature categories will be compared and matched one by one with the standard user tag library to map and output standardized target user tags. This will eliminate the problems of inconsistent and messy custom tags, ensure that the user tag system across the platform is unified and comparable, and greatly improve the reusability and interoperability of user profile data.
[0051] In another embodiment of this disclosure, target user tags are generated and persistently stored in a tag library. Tag data can be pushed to downstream business systems such as customer profiling systems, customer relationship management systems, and marketing operations systems via standardized external interfaces. Each system can retrieve tag data as needed, supporting business functions such as building complete customer profiles, customer churn risk warning, and targeted precision marketing. Simultaneously, the tag output stage can perform targeted filtering and priority sorting of existing tags according to the business needs of different target systems, outputting a subset of tags suitable for the corresponding business scenario, adapting to the differentiated usage needs of multiple systems.
[0052] A second aspect of this disclosure provides a user tag building apparatus, such as Figure 2 As shown, the device includes: The data acquisition module is used to acquire the user's static basic data and dynamic business data to obtain the user feature dataset; The analysis module is used to extract features from the user feature dataset using a target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, candidate user feature categories and their corresponding confidence scores are obtained. The verification module is used to verify the candidate user feature categories using the target verification model, based on preset verification rules and the confidence level corresponding to each candidate user feature category, and obtain the verification result; if the verification result meets the first verification condition, the candidate user feature category is retained or modified to obtain the target user feature category. The tag generation module is used to match the target user feature categories with a standard user tag library to obtain target user tags.
[0053] In another embodiment of this disclosure, the analysis module is further configured to generate feature weights corresponding to each feature through the target analysis model; perform weighted fusion of the shallow explicit features and the deep implicit features based on the feature weights to obtain initial user feature categories and corresponding confidence levels; and determine the initial user feature categories and corresponding confidence levels that reach the preset confidence threshold as candidate user feature categories and corresponding confidence levels based on a preset confidence threshold.
[0054] In another embodiment of this disclosure, the analysis module is further configured to determine new candidate user feature categories and their corresponding confidence levels through the target analysis model if the verification result satisfies the second verification condition.
[0055] In another embodiment of this disclosure, the analysis module is further configured to obtain weight constraint information, which is generated based on anomaly information in the verification results; and to determine new candidate user feature categories and their corresponding confidence levels based on the weight constraint information through the target analysis model.
[0056] In another embodiment of this disclosure, the device further includes a tag library generation module, used to generate corresponding tag categories based on the target business domain; determine the standard tag entries corresponding to each tag category, and form the standard user tag library.
[0057] In another embodiment of this disclosure, the data acquisition module is further configured to acquire static basic data and dynamic business data corresponding to the user; clean and standardize the static basic data and the dynamic business data to obtain standardized data; and integrate the standardized data to obtain a user feature dataset.
[0058] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0059] Figure 3 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0060] like Figure 3As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0061] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0062] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the user tag construction method. For example, in some embodiments, the user tag construction method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the user tag construction method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the user tag construction method by any other suitable means (e.g., by means of firmware).
[0063] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0064] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0065] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0066] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0067] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0068] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0069] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0070] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0071] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for constructing user tags, characterized in that, The method includes: Obtain the static basic data and dynamic business data corresponding to the user to obtain the user feature dataset; The user feature dataset is subjected to feature extraction using a target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, candidate user feature categories and their corresponding confidence levels are obtained. The target verification model verifies the candidate user feature categories based on preset verification rules and the confidence levels corresponding to each candidate user feature category, and obtains the verification results. If the verification results meet the first verification condition, the candidate user feature categories are retained or modified to obtain the target user feature categories. The target user feature categories are matched with a standard user tag library to obtain target user tags.
2. The method according to claim 1, characterized in that, The process of obtaining candidate user feature categories and corresponding confidence levels based on the shallow explicit features and the deep implicit features includes: The target analysis model is used to generate feature weights for each feature. The shallow explicit features and the deep implicit features are weighted and fused based on the feature weights to obtain the initial user feature categories and their corresponding confidence levels. Based on a preset confidence threshold, the initial user feature categories and their corresponding confidence levels that reach the preset confidence threshold are determined as candidate user feature categories and their corresponding confidence levels.
3. The method according to claim 1, characterized in that, The method further includes: If the verification result meets the second verification condition, the new candidate user feature category and its corresponding confidence level are determined by the target analysis model.
4. The method according to claim 3, characterized in that, The step of determining new candidate user feature categories and their corresponding confidence levels through the target analysis model includes: Obtain weight constraint information, which is generated based on anomaly information in the verification results; Based on the weight constraint information, the target analysis model determines new candidate user feature categories and their corresponding confidence levels.
5. The method according to claim 1, characterized in that, The method further includes: Generate corresponding tag categories based on the target business area; Determine the standard tag terms corresponding to each tag category to form the standard user tag library.
6. The method according to claim 1, characterized in that, The process of obtaining static basic data and dynamic business data corresponding to users to obtain user feature datasets includes: Obtain the user's static basic data and dynamic business data; The static basic data and the dynamic business data are cleaned and standardized to obtain standardized data. The standardized data is integrated to obtain a user feature dataset.
7. A user tag construction apparatus, characterized in that, The device includes: The data acquisition module is used to acquire the user's static basic data and dynamic business data to obtain the user feature dataset; The analysis module is used to extract features from the user feature dataset using a target analysis model to obtain shallow explicit features and deep implicit features; based on the shallow explicit features and the deep implicit features, candidate user feature categories and their corresponding confidence scores are obtained. The verification module is used to verify the candidate user feature categories using the target verification model, based on preset verification rules and the confidence level corresponding to each candidate user feature category, and obtain the verification result; if the verification result meets the first verification condition, the candidate user feature category is retained or modified to obtain the target user feature category. The tag generation module is used to match the target user feature categories with a standard user tag library to obtain target user tags.
8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.