A method and apparatus for image processing based on continuous and discrete features
By using portrait processing methods and devices based on continuous and discrete features, the limitations of professional and general portrait methods in application scenarios have been overcome. This has enabled the expansion of general applicability and feature quantification while ensuring professionalism, thereby improving the utilization rate and analysis depth of portrait data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING NXIN TECH GRP CO LTD
- Filing Date
- 2023-01-28
- Publication Date
- 2026-05-12
AI Technical Summary
Existing professional profiling methods are difficult to meet the needs of rapid iteration, while general profiling methods are difficult to meet specific complex needs, thus limiting the application scenarios of profiling methods.
The method and apparatus for image processing based on continuous and discrete features are adopted. Through steps such as entity management, dimension management, fact management, bitmap pattern calculation, effective range inference, completeness detection and missing value supplementation, combined with multi-channel delivery method, the image data is output to achieve a balance between professionalism and versatility.
While maintaining professionalism, it expands the general application scenarios of portraits and enables a more comprehensive and quantitative judgment of portrait features, thereby improving the utilization rate and analysis depth of portrait data.
Smart Images

Figure CN116383180B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to big data analytics, and in particular to a method and apparatus for image processing based on continuous and discrete features. Background Technology
[0002] User or enterprise profiling involves identifying various characteristics of users or enterprises, assigning tags to them, and then segmenting users or enterprises into different groups for targeted product / operational actions. Current profiling methods can be broadly categorized into two types: specialized and general-purpose. Specialized profiling methods typically process data from specific sources, focusing on specific user needs and generating targeted tags based on established rules. General-purpose profiling methods, on the other hand, rarely involve specialized fields. A better approach is to design independent modules for data processing, tagging, profile generation, and group segmentation to suit different profiling scenarios.
[0003] However, the actual needs for profiling are complex and ever-changing. Professional profiling is generally scenario-specific, and when requirements change (such as different data formats, different quality, or changes in the profiling target), existing professional profiling methods can hardly meet the needs of rapid iteration, to the point that versatility becomes an obstacle; while existing general-purpose profiling is generally a functional design, limited by specific implementation scenarios, and can hardly meet specific complex needs. Summary of the Invention
[0004] To address the challenge that existing professional profiling methods struggle to meet the demands of rapid iteration, hindering their versatility, and that general-purpose profiling methods, typically functionally focused and limited by specific implementation scenarios, fail to meet complex requirements, this invention proposes a profiling method and apparatus based on continuous and discrete features. This method and apparatus balance the needs of both professional and general-purpose profiling, integrating key elements to solve the problems of breadth in professional profiling and depth in general-purpose profiling. To achieve this objective, the invention employs the following technical solution.
[0005] A portrait processing method based on continuous and discrete features, the method comprising the following steps:
[0006] A. Perform entity management, dimension management, and fact management for data assets;
[0007] B. Perform bitmap-based profiling calculations on data assets and extract profiling data from the data assets;
[0008] C. Detect and evaluate the portrait data, including inference and processing of the effective range, as well as completeness detection and missing value supplementation;
[0009] D. Output profile data based on multi-channel delivery methods.
[0010] Furthermore, in the portrait processing method based on continuous and discrete features of the present invention, the bitmap-based portrait calculation for data assets includes factual data calculation, bitmap calculation and parsing, wherein...
[0011] Fact data computation cleans and unifies data modeling for data assets, including encapsulating data computation logic, scheduling corresponding scripts, and extracting model-based profile data.
[0012] Bitmap computation and parsing compresses the modeled portrait data and stores it in a relational database, and decompresses it when querying the relational database.
[0013] In addition, in the portrait processing method based on continuous and discrete features of the present invention, the effective range inference and processing includes filtering portrait data that is effective for the current business through different feature combinations, and judging portrait data with large outliers as abnormal values in the data portraits that are effective for the current business, and excluding or clearing them.
[0014] Furthermore, in the portrait processing method based on continuous and discrete features of the present invention, the completeness detection and missing value supplementation include:
[0015] Completeness detection includes detecting the content of each field of the portrait data, summarizing the data to identify the proportion of illegal and null values in each field, and obtaining the completeness rate of each field. If the completeness rate of a field in the portrait data is greater than a first predetermined threshold, it is used as the portrait data for that field. If the completeness rate of a field in the portrait data is between the first and second predetermined thresholds, it is used as the auxiliary portrait data for that field.
[0016] Missing value imputation includes filling in illegal and null values in the various fields of the profile data, using one or more methods such as field relevance method, business rule method, randomness method, and random matching method.
[0017] Furthermore, in the portrait processing method based on continuous and discrete features of the present invention, the portrait data output based on the multi-channel delivery method includes single entity output and group delineation. Group delineation includes selecting a portrait data group, displaying the statistical proportion of each field in the portrait data group, converting the portrait data stored in the relational database into an index in the distributed full-text search, and performing visualization output. Group delineation also includes setting custom labels for different entities in the portrait data group and performing custom group delineation output.
[0018] The present invention also includes an image processing device based on continuous and discrete features, the device comprising a management and control unit, a calculation unit, a detection and evaluation unit, and an output unit, wherein,
[0019] The management and control unit is used for entity management, dimension management, and fact management of data assets;
[0020] The computing unit is used to perform bitmap-based profiling calculations on data assets and extract profiling data from the data assets.
[0021] The detection and evaluation unit is used to detect and evaluate the profile data, including effective range inference and processing, as well as completeness detection and missing value filling;
[0022] The output unit is used to output profile data based on a multi-channel delivery method.
[0023] Furthermore, in the portrait processing apparatus based on continuous and discrete features of the present invention, the computing unit includes a fact data computing unit, a bitmap computing and parsing unit, and a query cache unit, wherein,
[0024] The fact data calculation unit is used to clean data assets and perform unified data modeling, including encapsulating data calculation logic, scheduling corresponding scripts, and extracting model-based profile data.
[0025] The bitmap calculation and parsing unit is used to compress the modeled portrait data and store it in a relational database, as well as to decompress it when querying the relational database.
[0026] The query cache unit is used to provide profile data to the detection and evaluation unit.
[0027] In addition, in the portrait processing device based on continuous and discrete features of the present invention, the detection and evaluation unit includes an effective range judgment unit and a discrete value processing unit. The effective range judgment unit is used to filter portrait data that is effective for the current business through different feature combinations. The discrete value processing unit is used to determine that portrait data with large outliers in the data portraits that are effective for the current business are outliers and exclude or clear them.
[0028] Furthermore, in the image processing apparatus based on continuous and discrete features of the present invention, the detection and evaluation unit includes a completeness detection unit and a missing value completion unit, wherein,
[0029] The completeness detection unit is used to detect the content of each field of the portrait data. By summarizing, it identifies the proportion of illegal and null values in each field of the portrait data and obtains the completeness rate of each field of the portrait data. If the completeness rate of a field in the portrait data is greater than a first predetermined threshold, it is used as the portrait data for that field. If the completeness rate of a field in the portrait data is between the first and second predetermined thresholds, it is used as the auxiliary portrait data for that field.
[0030] The missing value filling unit is used to fill in illegal and null values in the various fields of the profile data, and the filling is performed according to one or more of the following methods: field correlation method, business rule method, random method, and random matching method.
[0031] Furthermore, in the portrait processing device based on continuous and discrete features of the present invention, the detection and evaluation unit includes a completeness detection unit and a missing value supplementation unit. The output unit includes a single entity output unit and a group delineation unit. The group delineation unit is used to select a portrait data group, display the statistical proportion of each field in the portrait data group, convert the portrait data stored in the relational database into an index in the distributed full-text search, and perform visualization output. The group delineation also includes setting custom labels for different entities in the portrait data group and performing custom group delineation output.
[0032] The technical solution of the present invention can achieve the following technical effects.
[0033] 1. It solves the problems of the universality of professional profiles and the analytical depth of general profiles, expanding the general application scenarios of profiles while ensuring professionalism.
[0034] 2. It can make more comprehensive and quantitative judgments on the characteristics of the portrait. Attached Figure Description
[0035] Figure 1 This is a schematic block diagram of an image processing device based on continuous and discrete features according to a specific embodiment of the present invention.
[0036] Figure 2 This is a flowchart illustrating a portrait processing method based on continuous and discrete features according to a specific embodiment of the present invention.
[0037] Figure 3 This is a schematic diagram of entity relationships in portrait data based on a portrait processing method using continuous and discrete features according to a specific embodiment of the present invention.
[0038] Figure 4 This is a schematic diagram illustrating the image group selection method based on continuous and discrete features according to a specific embodiment of the present invention.
[0039] Figure 5 This is a schematic diagram of portrait group statistics based on the portrait processing method using continuous and discrete features, which is a specific embodiment of the present invention.
[0040] Figure 6 This is a schematic diagram of portrait group statistics based on the portrait processing method using continuous and discrete features, which is a specific embodiment of the present invention.
[0041] Figure 7 This is a schematic diagram of a single-entity image based on the image processing method of continuous and discrete features, which is a specific embodiment of the present invention. Detailed Implementation
[0042] The present invention will now be described in detail with reference to the accompanying drawings.
[0043] The following detailed exemplary embodiments are disclosed. However, the specific structural and functional details disclosed herein are merely for the purpose of describing exemplary embodiments.
[0044] However, it should be understood that the present invention is not limited to the specific exemplary embodiments disclosed, but covers all modifications, equivalents, and substitutions falling within the scope of this disclosure. Throughout the description of the drawings, the same reference numerals denote the same elements.
[0045] Referring to the accompanying drawings, the structures, proportions, sizes, etc., depicted in the drawings are merely for illustrative purposes to aid those skilled in the art in understanding and reading the content disclosed herein. They are not intended to limit the conditions under which the invention can be implemented and therefore have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to the size, without affecting the effects and objectives achieved by the invention, should still fall within the scope of the technical content disclosed herein. Furthermore, the positional limitations used in this specification are merely for clarity of description and are not intended to limit the scope of the invention. Changes or adjustments to their relative relationships, without substantially altering the technical content, should also be considered within the scope of the invention's implementation.
[0046] It should also be understood that the term “and / or” as used herein includes any and all combinations of one or more of the related listed items. Furthermore, it should be understood that when a component or unit is referred to as “connected” or “coupled” to another component or unit, it may be directly connected or coupled to the other component or unit, or there may be intermediate components or units. In addition, other words used to describe the relationship between components or units should be understood in the same manner (e.g., “between” versus “directly between,” “adjacent” versus “directly adjacent,” etc.).
[0047] Figure 1 This is a schematic block diagram of a portrait processing device based on continuous and discrete features according to a specific embodiment of the present invention. As shown in the figure, the specific embodiment of the present invention includes a portrait processing device based on continuous and discrete features. The device includes a management and control unit, a calculation unit, a detection and evaluation unit, and an output unit. The management and control unit is used for entity management, dimension management, and fact management of data assets; the calculation unit is used for performing portrait calculations on data assets based on bitmap patterns and extracting portrait data from the data assets; the detection and evaluation unit is used for detecting and evaluating the portrait data, including effective range inference and processing, as well as completeness detection and missing value supplementation; the output unit is used for outputting the portrait data based on a multi-channel delivery method.
[0048] Entity management specifies an object for analysis: such as a person, enterprise, organization, product, customer, etc. Each entity must have a unique ID for data identification. Dimension management, on the other hand, refers to a data approach used for analysis.
[0049] Corresponding to the image processing device based on continuous and discrete features in the specific embodiments of the present invention, the specific embodiments of the present invention include an image processing method based on continuous and discrete features, such as... Figure 2 As shown, the method includes the following steps:
[0050] A. Perform entity management, dimension management, and fact management for data assets;
[0051] B. Perform bitmap-based profiling calculations on data assets and extract profiling data from the data assets;
[0052] C. Detect and evaluate the portrait data, including inference and processing of the effective range, as well as completeness detection and missing value supplementation;
[0053] D. Output profile data based on multi-channel delivery methods.
[0054] Furthermore, in the portrait processing method based on continuous and discrete features in the specific embodiments of the present invention, the portrait calculation based on bitmap pattern for data assets includes fact data calculation, bitmap calculation and parsing, wherein,
[0055] Fact data computation cleans and unifies data modeling for data assets, including encapsulating data computation logic, scheduling corresponding scripts, and extracting model-based profile data.
[0056] Bitmap computation and parsing compresses the modeled portrait data and stores it in a relational database, and decompresses it when querying the relational database.
[0057] The factual data calculation unit defines data profiles from data assets. The entity relationship diagram involved is as follows: Figure 3 As shown. The main contents of a relational database include the following.
[0058] (1) Data semantic generation, data is stored in object tables, physical tables, etc.
[0059] (2) Define the portrait dimensions and store the data in the dimension table.
[0060] (3) Entity definition, data is stored in entity table.
[0061] (4) Defining the facts of the profile, the data is stored in the fact table.
[0062] (5) Combined data storage, the data is stored in the portrait result related table.
[0063] (6) Data retrieval.
[0064] Therefore, by storing object tables, physical tables, entity tables, fact tables, and tables related to profiling results in relational databases, relational databases can provide data retrieval functions.
[0065] The feature calculation based on the profile, as well as the data storage and data parsing, include the following steps:
[0066] (1) Calculate the specific statistical value of a single entity through basic data (referred to as a certain dimension).
[0067] Dimension attributes: dimension type, whether it can be combined with time period, related tables, value columns, name columns, filter conditions, etc.;
[0068] (2) Dimension types are divided into three categories according to different data characteristics: monetary, integer, and BIT.
[0069] (3) Fact time window, such as 7 days, 30 days, 3 months, one year, cumulative, etc.
[0070] The tables and columns related to the dimension settings correspond to the basic data in the data assets, while the data items in the value columns serve as the final stored data for the profile. Based on different table sizes and counting requirements, the tables are categorized into four types: enumeration, set, table data, and custom.
[0071] (1) Enumerated data can be used to select specific values from the table as the calculation unit for data statistics. Only one specific value is valid in the calculation unit in the table. For example, the user's current location or the source of enterprise registration.
[0072] (2) For aggregate data, multiple values are retrieved from the table as the calculation unit for data statistics. The calculation unit has an indefinite number of valid values in the table. For example, systems that a user has logged into, industries that they follow, etc.
[0073] (3) The characteristic of table data types is that referenced tables have a massive number of records as units of calculation. This is used to analyze concrete characteristics of user profiles, such as the top 5 most frequently purchased items by a user, or the 5 most active employees of a company.
[0074] (4) Custom class data stores users as independent calculation units in a dictionary table.
[0075] The dictionary table has the following data structure:
[0076] Dictionary categories, dictionary encoding, dictionary names, dictionary values
[0077] Dictionary categories are used to distinguish different dictionary types, corresponding to the configuration conditions in the dimensions. Dictionary values are used for data storage, and dictionary names are retrieved by dictionary category and dictionary value.
[0078] A fact is a unit of profile calculation. To facilitate data aggregation, it is stored in a relational database as a fact table. The fact table has the following attributes:
[0079] Fact ID, Fact Name, Dimension ID, Measure ID, Profile Storage Table;
[0080] Whether it is available, the user, whether it is periodized, and whether it is analyzed.
[0081] The fact table and the calculated results are displayed in different profile result tables according to the above configuration.
[0082] Data is stored in a relational database (RDBMS). The profile result table is categorized into the following types based on data type: bit data, integer data, floating-point data, periodized bit data (Bit2), periodized integer data (Int2), and periodized floating-point data (Decimal2).
[0083] (1) Bit type data structure: object ID, data ID, fact ID, item ID of the fact.
[0084] (2) Int type data structure: object ID, data ID, fact ID, item ID of fact, item value of fact (integer value).
[0085] (3) Decimal data structure: object ID, data ID, fact ID, item ID of fact, item value of fact (floating-point number).
[0086] (4) Bit2 data structure: object ID, data ID, periodized type ID, fact ID, item ID of fact
[0087] (5) Int2 type data structure: object ID data ID periodized type ID fact ID fact item ID fact item value (integer value).
[0088] (6) Decimal data structure: object ID, data ID, periodized type ID, fact ID, item ID of fact, item value of fact (floating-point value).
[0089] (7) Dates are approximately "continuous data". To simplify data analysis and make the analysis more intuitive, "discrete" processing is required. Based on experience from actual application scenarios, the periodization type can be classified as follows: to date, last 7 days, last 3 months, last 1 year, etc.
[0090] Fact data is primarily used for cleaning raw data and standardizing data modeling. For example, the data computation logic can be encapsulated using open-source tools like Kettle, scheduling corresponding scripts such as JS, SQL, Java code, and regular expressions for processing. This should ideally align with the analysis objectives. The goal is to aggregate data to minimize its size, thereby enabling more efficient feature calculations and improved computational efficiency. Fact data computation allows for the modeling and extraction of data from different sources, types, and formats. This provides reliable material for further data processing and applications.
[0091] Furthermore, analyzing factual data typically involves large amounts of data. Therefore, controlling the data scale and improving data retrieval efficiency become particularly important. Bitmap computation effectively compresses data, and bitmap interpretation is used for decompression during data retrieval. Data storage is implemented in relational databases using a retrieval routing table. The retrieval routing table data structure is: object ID, data ID, fragment, and routing value; fragment + routing value is used to determine whether fact storage is sufficient. In this specific embodiment, the int type is used to store bit values. For example, when the maximum stored value is 2^31-1, the system fragments according to 30 bits to process the fact ID. Fragment value: Fact ID / 30. The routing value refers to filling each bit from right to left with the remainder after fragmentation. Since the remainder range is (0~29), it needs to be shifted one bit to the left to become (1~30). The routing value calculation formula is: ∑2^(Fact ID%30+1), and the routing value parsing method is to inverse the above method. Therefore, the structure and design of bitmap computation and bitmap interpretation save IO costs and improve the efficiency of portrait data retrieval.
[0092] In addition, in the portrait processing method based on continuous and discrete features in the specific embodiments of the present invention, the effective range inference and processing includes filtering portrait data that is effective for the current business through different feature combinations, and judging portrait data with large outliers as abnormal values in the data portraits that are effective for the current business, and excluding or clearing them.
[0093] In a specific embodiment of this invention, after data from the same source is incorporated into the user profile, the standard definition of the profile data differs for different use scenarios. To address the issue of profile data reuse, an effective range inference function is introduced. By combining different features, data valid for the current business can be filtered out. For each profile use scenario, the data to be used needs to be selected again. The effective range inference function solves the data reuse problem.
[0094] Furthermore, outliers refer to values that deviate significantly from other data points and are generally considered as anomalies. To improve data quality, they are typically excluded or cleared. More specific embodiments of this invention may employ methods such as linear regression, logistic regression, support vector machines, K-means, KNN, tree-based methods, or other custom methods. Outlier processing effectively filters outliers, eliminating abnormal data that could negatively impact project analysis.
[0095] Furthermore, in the portrait processing method based on continuous and discrete features in the specific embodiments of the present invention, the completeness detection and missing value supplementation include:
[0096] Completeness detection involves inspecting the content of each field in the profile data, summarizing the data to identify the proportion of illegal and null values in each field, and obtaining the completeness rate of each field. If the completeness rate of a field in the profile data is greater than a first predetermined threshold, for example, the first predetermined threshold is 90%, it is used as the profile data for that field. If the completeness rate of a field in the profile data is between the first predetermined threshold and a second predetermined threshold, for example, the second predetermined threshold is 80%, it is used as auxiliary profile data for that field.
[0097] Missing value imputation includes filling in illegal and null values in the various fields of the profile data, using one or more methods such as field relevance method, business rule method, randomness method, and random matching method.
[0098] Specifically, several methods are used for missing value imputation, including:
[0099] 1. Field Relevance Imputation: Using Bayesian methods, multiple dimensions are input. Different combinations of dimensions are analyzed. The dimension combination that best reflects the target value is obtained, and this combination is used to impute missing values.
[0100] 2. Business rule population: Filling missing values using certain business attributes. For example, the value of a field can be inferred from the case of certain fields.
[0101] 3. Random Imputation: A small number of missing data points can be imputed using random methods. The goal is to minimize the impact on the original data ratios. For example, if the male-to-female ratio before imputation is 1:1.04, the ratio should be maintained as much as possible after imputation.
[0102] 4. Randomized Filling: For example, if the original male-to-female ratio is 1:0.95, and we need to fill in missing values to make the final male-to-female ratio 1:1, then the filling of missing values should be tilted towards females.
[0103] Missing value imputation relaxes the data access quality requirements, greatly improves the utilization rate of profile data, and enables more data to participate in profile data analysis and calculation.
[0104] Furthermore, in the portrait processing method based on continuous and discrete features in the specific embodiments of the present invention, the portrait data output based on the multi-channel delivery method includes single entity output and group delineation. Group delineation includes selecting a portrait data group, displaying the statistical proportion of each field in the portrait data group, converting the portrait data stored in the relational database into an index in the distributed full-text search, and performing visualization output. Group delineation also includes setting custom labels for different entities in the portrait data group and performing custom group delineation output.
[0105] In a specific embodiment of this invention, data parsing is performed in two ways based on the metadata configuration in the fact table: A) data parsing based on relational database routing values; B) data parsing based on ElasticSearch query caching. Two types of data are parsed: a) profile group statistics; b) single entity output.
[0106] Data parsing based on relational database routing values is primarily used for single-entity output and is a suitable option for applications with high real-time data requirements. ElasticSearch caching can be used for user profiling and group analysis, as well as single-entity output. This option is suitable for applications with high analytical requirements that need to slice down to specific entities for viewing.
[0107] (1) Data parsing based on relational database routing values directly parses profile data for retrieval through data structures: for example, retrieving the routing table by inputting object ID and data ID, parsing out the fact ID that the data matches bit by bit; finding the corresponding profile result table through the fact table; parsing the corresponding fact items through the dimension table; determining whether the parsed data can be used for display and summarization through the type field of the dimension table, and obtaining single entity output, such as... Figure 7 As shown.
[0108] (2) Query caching based on Elasticsearch involves pre-aggregating data into Elasticsearch. The data structure in Elasticsearch is as follows: Object ID, Data ID, and Profile Data. The Profile Data is aggregated by field in the format: Fact Name + Item Name + Periodization Type. For example... Figure 6 The diagram shown is a statistical illustration of the portrait group based on the portrait processing method using continuous and discrete features according to a specific embodiment of the present invention.
[0109] Based on the above specific implementation methods, profile data can be used in two main ways: individual analysis and group analysis. Individual analysis mainly provides a comprehensive view of the individuals within the profile, including pre-defined tags and user-defined tags. For example... Figure 7 This is a schematic diagram of a single image, illustrating a specific embodiment of the image processing method based on continuous and discrete features according to the present invention.
[0110] Besides grouping by relevant characteristics, population analysis can also combine different sets (defined as profiling) for merging, a process known as "framing". For example... Figure 4 The conditions for delineation in the portrait processing method based on continuous and discrete features in a specific embodiment of the present invention are as follows.
[0111] For example, in a specific application instance of this invention, the delineation logic is as follows:
[0112] (1) Set the group name, such as credit users who have recently logged into a certain system and whose transaction volume exceeds 1 million within the past three months. (2) Select various data features. (3) Analyze the data and count the data scale of each different feature. (4) Generate the profile group statistical results. (5) Export the data or generate it via an interface. Figure 5 and Figure 6 This is a schematic diagram of portrait group statistics based on the portrait processing method using continuous and discrete features, which is a specific embodiment of the present invention. Figure 6 The image group was further visualized in the text.
[0113] For the output of user profile groups, secondary tags can be added for further use and tracking. For example, tags such as target customers, potential customers, and converted customers can be set. A specific implementation flow of this invention is as follows: 1. Enter the user profile entity data list based on the user's selection of the user profile group; 2. Receive the selection items prompted by clicking on a specific entity, selecting multiple entities via checkboxes, or selecting all; 3. Receive the tag input results; 4. Query entities with set tags by selecting tags.
[0114] In addition, such as Figure 1As shown in the figure, in the portrait processing device based on continuous and discrete features according to a specific embodiment of the present invention, the computing unit includes a fact data computing unit, a bitmap computing and parsing unit, and a query cache unit, wherein,
[0115] The fact data calculation unit is used to clean data assets and perform unified data modeling, including encapsulating data calculation logic, scheduling corresponding scripts, and extracting model-based profile data.
[0116] The bitmap calculation and parsing unit is used to compress the modeled portrait data and store it in a relational database, as well as to decompress it when querying the relational database.
[0117] The query cache unit is used to provide profile data to the detection and evaluation unit.
[0118] In addition, such as Figure 1 As shown in the figure, in the portrait processing device based on continuous and discrete features of the present invention, the detection and evaluation unit includes an effective range judgment unit and a discrete value processing unit. The effective range judgment unit is used to filter portrait data that are effective for the current business through different feature combinations. The discrete value processing unit is used to determine that portrait data with large outliers in the data portraits that are effective for the current business are outliers and exclude or clear them.
[0119] In addition, such as Figure 1 As shown, in the portrait processing device based on continuous and discrete features according to a specific embodiment of the present invention, the detection and evaluation unit includes a completeness detection unit and a missing value completion unit, wherein,
[0120] The completeness detection unit is used to detect the content of each field of the portrait data. By summarizing, it identifies the proportion of illegal and null values in each field of the portrait data and obtains the completeness rate of each field of the portrait data. If the completeness rate of a field in the portrait data is greater than a first predetermined threshold, it is used as the portrait data for that field. If the completeness rate of a field in the portrait data is between the first and second predetermined thresholds, it is used as the auxiliary portrait data for that field.
[0121] The missing value filling unit is used to fill in illegal and null values in the various fields of the profile data, and the filling is performed according to one or more of the following methods: field correlation method, business rule method, random method, and random matching method.
[0122] In addition, such as Figure 1As shown in the embodiment of the portrait processing device based on continuous and discrete features of the present invention, the detection and evaluation unit includes a completeness detection unit and a missing value supplementation unit. The output unit includes a single entity output unit and a group delineation unit. The group delineation unit is used to select a portrait data group, display the statistical proportion of each field in the portrait data group, convert the portrait data stored in the relational database into an index in the distributed full-text search, and perform visualization output. The group delineation also includes setting custom labels for different entities in the portrait data group and performing custom group delineation output.
[0123] The technical solution of the present invention can achieve the following technical effects.
[0124] 1. It solves the problems of the universality of professional profiles and the analytical depth of general profiles, expanding the general application scenarios of profiles while ensuring professionalism.
[0125] 2. It can make more comprehensive and quantitative judgments on the characteristics of the portrait.
[0126] The foregoing description illustrates and describes several preferred embodiments of the present invention. The embodiments described are merely typical examples of the technical process of the present invention under current technical conditions. Without departing from the technical principles, steps, functions, applications, and implementation framework of the present invention, there is considerable room for optimization and improvement. These improvements and optimizations are also considered within the scope of protection of this patent. Therefore, as mentioned above, it should be understood that the present invention is not limited to the forms disclosed in this specification and should not be considered as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be modified within the scope of the inventive concept described in this specification through the above teachings or related technologies or knowledge. Modifications and changes made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A portrait processing method based on continuous and discrete features, characterized in that, The method includes the following steps: A. Perform entity management, dimension management, and fact management for data assets; B. Perform bitmap-based profiling calculations on data assets and extract profiling data from the data assets; C. Detect and evaluate the portrait data, including inference and processing of the effective range, as well as completeness detection and missing value supplementation; D. Outputting profile data based on multi-channel delivery methods; Among them, the bitmap-based profiling calculation for data assets includes fact data calculation, bitmap calculation and parsing. The fact data calculation cleans the data assets and unifies data modeling, including encapsulating the data calculation logic, scheduling the corresponding scripts, and extracting model-based profiling data. Bitmap computation and parsing compresses the modeled portrait data and stores it in a relational database, and decompresses it when querying the relational database. Effective range inference and processing includes filtering profile data that is effective for the current business through different feature combinations, and judging outliers in the profile data that is effective for the current business based on the size of outliers, and then excluding or clearing them. The completeness check and missing value imputation include: Completeness detection includes detecting the content of each field of the portrait data, summarizing the data to identify the proportion of illegal and null values in each field, and obtaining the completeness rate of each field. If the completeness rate of a field in the portrait data is greater than a first predetermined threshold, it is used as the portrait data for that field. If the completeness rate of a field in the portrait data is between the first and second predetermined thresholds, it is used as the auxiliary portrait data for that field. Missing value imputation includes filling in illegal and null values in the various fields of the profile data, using one or more methods such as field relevance method, business rule method, randomness method, and random matching method. The multi-channel delivery method for outputting profile data includes single entity output and group delineation. Group delineation includes selecting a profile data group, displaying the statistical proportion of each field in the profile data group, converting the profile data stored in the relational database into an index in the distributed full-text search, and outputting the visualization. Group delineation also includes setting custom labels for different entities in the profile data group and outputting custom group delineation.
2. A portrait processing device based on continuous and discrete features, characterized in that, The device includes a management and control unit, a computing unit, a detection and evaluation unit, and an output unit. The management and control unit is used for entity management, dimension management, and fact management of data assets; The computing unit is used to perform bitmap-based profiling calculations on data assets and extract profiling data from the data assets. The detection and evaluation unit is used to detect and evaluate the profile data, including effective range inference and processing, as well as completeness detection and missing value filling; The output unit is used to output profile data based on a multi-channel delivery method; The calculation unit includes a fact data calculation unit, a bitmap calculation and parsing unit, and a query cache unit. The fact data calculation unit is used to clean data assets and perform unified data modeling, including encapsulating data calculation logic, scheduling corresponding scripts, and extracting model-based profile data. The bitmap calculation and parsing unit is used to compress the modeled portrait data and store it in a relational database, as well as to decompress it when querying the relational database. The query cache unit is used to provide profile data to the detection and evaluation unit; The detection and evaluation unit includes an effective range judgment unit and a discrete value processing unit. The effective range judgment unit is used to filter the profile data that is valid for the current business through different feature combinations. The discrete value processing unit is used to determine outliers in the profile data that is valid for the current business based on the size of outliers, and then exclude or clear them. The detection and evaluation unit includes a completeness detection unit and a missing value completion unit. The completeness detection unit is used to detect the content of each field of the portrait data. By summarizing, it identifies the proportion of illegal and null values in each field of the portrait data and obtains the completeness rate of each field of the portrait data. If the completeness rate of a field in the portrait data is greater than a first predetermined threshold, it is used as the portrait data for that field. If the completeness rate of a field in the portrait data is between the first and second predetermined thresholds, it is used as the auxiliary portrait data for that field. The missing value imputation unit is used to fill in illegal and null values in the various fields of the profile data, and the filling is performed according to one or more of the following methods: field correlation method, business rule method, random method, and random matching method. The detection and evaluation unit includes a completeness detection unit and a missing value completion unit. The output unit includes a single entity output unit and a group delineation unit. The group delineation unit is used to select a profile data group, display the statistical proportion of each field in the profile data group, convert the profile data stored in the relational database into an index in the distributed full-text search, and perform visualization output. The group delineation also includes setting custom labels for different entities in the profile data group to perform custom group delineation output.