User portraying method and system based on big data

By analyzing the user page browsing trajectory, quasi-interest array sets are generated, combined with high-frequency array statistics and timestamp mechanisms, the user portrait is dynamically updated, solving the problems of insufficient individualization and timeliness of user portraits, and achieving the accuracy of accurate capture and recommendation of user interests.

CN120471672AActive Publication Date: 2025-08-12GANZHOU DEVELOPMENT CREDIT REPORTING CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510610576.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-12
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing user portrait technology is insufficiently individualized, and it is difficult to capture dynamic changes in user interests in real time, resulting in insufficient timeliness and accuracy.

Method used

By analyzing the user page browsing trajectory, generating a set of quasi-interest arrays, combining high-frequency array statistics and timestamp mechanisms, updating user basic portraits, filtering browsing records of similar user groups, and dynamically optimizing user portraits.

Benefits of technology

It improves the individualization and timeliness of user portraits, ensures the sensitivity of changes in user interests, realizes automatic update and cleaning of user portraits, and improves the accuracy and efficiency of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471672A_ABST
    Figure CN120471672A_ABST
Patent Text Reader

Abstract

The invention discloses a user portraying method and system based on big data, and relates to the technical field of data analysis, and the method comprises the steps: obtaining a quasi-interest array set through analyzing a page browsing track of a target user; retrieving a first browsing record set of a first sample user group meeting the basic portrait of the target user, and traversing the quasi interest array set based on the product type and the plurality of advertisement element types to perform high-frequency array statistics to obtain an interest array set; retrieving a second browsing record set of a second sample user group meeting the interest array set to obtain a secondary quasi-interest array set, and performing high-frequency array statistics on non-intersection quasi-interest arrays of the secondary quasi-interest array set and the quasi-interest array set to obtain a secondary interest array set; the basic portrait of the target user is updated according to the interest array set and the secondary interest array set, and the technical problems that the pertinence and individualization degree of the user portrait are insufficient, and the timeliness and accuracy errors of the user portrait are caused due to the fact that the dynamic change of the user interest is difficult to capture in real time are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a user profiling method and system based on big data. Background Art

[0002] In today's digital age, with the rapid development of internet technology and the widespread application of big data, user profiling, as a core technology in areas such as precision marketing and personalized recommendations, has attracted widespread attention and research. Existing user profiling technologies often rely on user clustering. This involves clustering large amounts of user behavior data, grouping users with similar behavioral characteristics, and then providing recommendation services for the entire user group.

[0003] However, while user clustering offers the advantage of strong generalization and the ability to provide recommendations based on shared characteristics across a group, it lacks individualization. Due to significant differences in interests and needs among users, recommendations based on shared characteristics often struggle to accurately match each user's unique needs, resulting in unsatisfactory recommendations. Furthermore, as user interests evolve dynamically, existing user profiling technologies often struggle to capture these changes in real time, impacting the timeliness and accuracy of user profiling. Summary of the Invention

[0004] This application provides a user profiling method and system based on big data, which is used to solve the technical problems in the existing technology that user profiling is not targeted and individualized enough, and it is difficult to capture the dynamic changes of user interests in real time, resulting in errors in the timeliness and accuracy of user profiling.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] In a first aspect, the present invention provides a user profiling method based on big data, comprising: parsing the page browsing trajectory of the target user to obtain a quasi-interest array set, wherein any quasi-interest array includes a product type and several advertising element types; retrieving a first browsing record set of a first sample user group that meets the basic profile of the target user, and based on the product type and the several advertising element types, traversing the quasi-interest array set to perform high-frequency array statistics to obtain an interest array set; retrieving a second browsing record set of a second sample user group that meets the interest array set to obtain a secondary quasi-interest array set, and performing high-frequency array statistics on the secondary quasi-interest array set and the non-intersecting quasi-interest array of the quasi-interest array set to obtain a secondary interest array set; and updating the basic profile of the target user according to the interest array set and the secondary interest array set.

[0007] In a second aspect, the present invention provides a user portrait system based on big data, the system comprising: an information parsing module for parsing the page browsing trajectory of the target user to obtain a quasi-interest array set, wherein any quasi-interest array includes a product type and several advertising element types; a first data statistics module for retrieving a first browsing record set of a first sample user group that meets the basic portrait of the target user, and based on the product type and the several advertising element types, traversing the quasi-interest array set to perform high-frequency array statistics to obtain an interest array set; a second data statistics module for retrieving a second browsing record set of a second sample user group that meets the interest array set to obtain a secondary quasi-interest array set, and performing high-frequency array statistics on the secondary quasi-interest array set and the non-intersection quasi-interest array of the quasi-interest array set to obtain a secondary interest array set; a portrait update module for updating the basic portrait of the target user based on the interest array set and the secondary interest array set.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: The big data-based user profiling method provided in the embodiment of the present application obtains a quasi-interest array set by analyzing the page browsing trajectory of the target user, and continuously optimizes and updates the basic portrait of the target user by combining the hierarchical interest array statistics and update mechanism. By comparing with the user group with similar basic portraits, it can more accurately capture the individual interests of the user, while maintaining sensitivity to changes in user interests, ensuring the timeliness and accuracy of the user portrait. It not only overcomes the shortcomings of the existing user clustering method of insufficient individualization, but also realizes the automatic update and cleaning of the user portrait by introducing the timestamp mechanism, further improving the simplicity and efficiency of the user portrait. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart of a user profiling method based on big data is provided for this application.

[0010] Figure 2 A structural diagram of a user portrait system based on big data is provided for this application.

[0011] Description of the accompanying drawings: information analysis module 11, first data statistics module 12, second data statistics module 13, portrait updating module 14. DETAILED DESCRIPTION

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0013] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0014] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.

[0015] Example 1:

[0016] like Figure 1 As shown, the embodiment of the present invention provides a user profiling method based on big data, including: S10: Analyze the page browsing trajectory of the target user to obtain a set of quasi-interest arrays, wherein any quasi-interest array includes a product type and several advertising element types.

[0017] Specifically, a user profile is a virtual image or collection of tags that reflects a user's interests, preferences, needs, and other characteristics, constructed through data analysis and mining techniques based on their internet behavior data (such as page browsing history, click behavior, and search history) and related attribute information (such as age, gender, and region). This is used to support personalized recommendations, precision marketing, and other application scenarios. Existing technologies often use user clustering for group recommendations, which has strong generalization capabilities. This results in insufficient targeting and individualization of user profiles, and the difficulty in capturing dynamic changes in user interests in real time, resulting in errors in the timeliness and accuracy of user profiles. This solution, by incorporating a hierarchical update mechanism, improves the targeting of user profiles, enabling more precise capture of individual user interests while maintaining sensitivity to changes in user interests, ensuring the timeliness and accuracy of user profiles. This approach not only overcomes the shortcomings of existing user clustering methods in terms of insufficient individualization, but also introduces a timestamp mechanism that enables automatic updating and cleanup of user profiles, further improving the streamlining and efficiency of user profiles.

[0018] For example, when building a user profile, we first analyze the target user's page browsing history, aiming to extract potentially interesting features from the user's behavioral data. Specifically, we conduct an in-depth analysis of the target user's page browsing history, extracting key click windows within the history to capture the user's interactive behavior within a specific time period. These click windows not only record the pages the user visited but also imply the user's interest in the page content.

[0019] Furthermore, UI design information within the click window is extracted. Product types describe the category attributes of the products displayed on the page, such as "electronics" or "clothing and accessories." Ad element types encompass a range of ad display-related element types, such as "image ads," "video ads," or "text link ads." Based on this extracted information, the system constructs quasi-interest arrays, each containing a specific product type and several associated ad element types, such as "electronics - image ads, video ads." This process systematically collects and organizes the interest characteristics exhibited by target users during different page browsing sessions, forming a quasi-interest array set. This lays the foundation for subsequent interest array statistics and user profile updates. For example, if a user browses electronics pages in multiple consecutive click windows, and image and video ads frequently appear on these pages, a quasi-interest array containing "electronics - image ads, video ads" will be constructed and included in the quasi-interest array set, reflecting the user's strong interest in electronics and their advertising formats.

[0020] S20: Retrieve a first browsing record set of a first sample user group that meets the basic profile of the target user, and based on the product type and the multiple advertising element types, traverse the quasi-interest array set to perform high-frequency array statistics to obtain an interest array set.

[0021] Furthermore, retrieving the first set of browsing records for a first sample user group that meets the target user's basic profile and performing high-frequency array statistics based on product type and advertising element type is a key step in accurately characterizing user interests. Specifically, based on the target user's basic profile, a group with similar characteristics to the target user is selected, i.e., the first sample user group. The basic profile typically includes general profile tags (such as age, gender, and region) and individual profile tags (such as an initial interest array set).

[0022] Subsequently, a set of first browsing records for the sample user is collected. This set records the sample user's browsing behavior on the internet, including pages visited, ads clicked, and so on. Based on the product types and advertising element types analyzed from the target user's page browsing history, the system traverses the set of quasi-interest arrays and counts the triggering frequency of each quasi-interest array in the first browsing record set. The triggering frequency refers to the number or proportion of occurrences of a particular quasi-interest array (e.g., "electronic products - picture ads") in the sample user's browsing history. By counting high-frequency arrays, the system can identify frequently occurring interest arrays within the sample user group. These arrays reflect common characteristics that the target user may be interested in. Finally, quasi-interest arrays with a triggering frequency exceeding a preset threshold are aggregated to form an interest array set. This set serves as a concentrated reflection of the target user's interest characteristics and provides an important basis for subsequent user profile updates. For example, if the quasi-interest array "electronic products - picture ads" appears frequently in the first browsing record set, it is included in the interest array set, indicating that the target user and similar sample users have a high interest in electronic products and their picture ads.

[0023] S30: Retrieve a second browsing record set of a second sample user group that satisfies the interest array set to obtain a secondary quasi-interest array set, perform high-frequency array statistics on the secondary quasi-interest array set and the non-intersecting quasi-interest arrays of the quasi-interest array set to obtain a secondary interest array set.

[0024] Next, further searching for a second sample user group that meets the identified interest group set and collecting their second browsing history set is a step to deepen user interest mining. Specifically, the previously obtained interest group set is first used to reflect the product types and advertising element type combinations that reflect the common interests of the target user and similar groups, such as "electronics products - picture ads." Then, based on these interest characteristics, another group of users with similar interest patterns is selected, namely the second sample user group. For this user group, their browsing behavior data is collected to form a second browsing history set, which records the detailed browsing trajectory of these users on the internet.

[0025] The second set of browsing records is then parsed to extract the secondary quasi-interest arrays contained therein. These arrays contain interest features similar to, but not identical to, the initial interest array set, such as "electronics products - video ads" or "household goods - text link ads." To accurately identify potential new interests of the target user, the secondary quasi-interest array set is compared with the previously obtained quasi-interest array set to extract the non-intersecting portions between the two. These interest arrays are specifically those that appear only in the second sample user group and were not significantly reflected in the initial analysis. Finally, a high-frequency array statistics is performed on these non-intersecting quasi-interest arrays. By calculating their triggering frequency in the second set of browsing records, arrays with a triggering frequency exceeding a preset threshold are selected to form the secondary interest array set.

[0026] This process not only enriches the interest dimensions of the user profile but also increases the profile's sensitivity to changes in user interests. For example, if the non-intersecting quasi-interest group "household goods - video ads" appears frequently in the second set of browsing records, it will be included in the secondary interest group set, indicating that the target user may also have a strong interest in video ads for household goods, even though this interest was not significantly reflected in the initial analysis.

[0027] S40: Update the target user basic portrait according to the interest group set and the secondary interest group set.

[0028] Specifically, the target user's basic profile is updated based on the acquired interest array set and secondary interest array set. This process aims to integrate newly discovered user interest features into the existing profile to improve the accuracy of the profile's reflection of the user's true interests. Specifically, the interest array set (such as "electronic products - picture ads") and the secondary interest array set (such as "household goods - video ads") are integrated. These array sets represent the interest preferences shown by the target user and similar groups at different stages. Subsequently, these interest features are added as new profile tags to the target user's basic profile, which originally contained general profile tags (such as age and gender) and individual profile tags (such as the initial interest array).

[0029] In the specific update process, pay attention to the application of timestamps, mark the first timestamp of the newly added interest array, and record the time point when it was first identified. When the target user clicks on the content related to a certain interest array again, the timestamp of the array is updated with the latest click timestamp to reflect the continuity and variability of the user's interest. If the time interval between the timestamp of a certain interest array and the current time exceeds the preset threshold, it is determined that the interest may have faded or is no longer active, and it will be deleted from the basic portrait of the target user to ensure the timeliness and accuracy of the portrait information. For example, if it is found in subsequent monitoring that the target user's interest in "electronic products-picture ads" persists, while the interest in "household items-video ads" gradually weakens, the former will be retained in the basic portrait and its timestamp will be updated, while the latter will be removed, so as to dynamically optimize the user portrait and provide more accurate data support for applications such as personalized recommendations.

[0030] In a specific embodiment, the page browsing trajectory of the target user is parsed to obtain a quasi-interest array set, including: extracting the first click window of the page browsing trajectory; extracting the first UI design information of the first click window, wherein the first UI design information includes a first product design type and a first advertising element design type set, and constructing a first quasi-interest array; and adding the first quasi-interest array to the quasi-interest array set.

[0031] Optionally, the system first implements time series slicing on the target user's page browsing trajectory, and extracts the first click window in the trajectory by setting a reasonable sliding window length (such as 30 seconds or 1 page). This window serves as the basic unit of user behavior analysis, covering the time period and interaction range from page loading to the first click operation of the user.

[0032] Next, UI design information is extracted for the first click window, identifying and recording the first product design type (e.g., "smart wearable device") and the first ad element design type set (e.g., "dynamic pop-up ad" or "bottom floating ad"). The product design type identifies the product category of the core content displayed on the page, while the ad element design type set describes the visual presentation and interactive methods of the page ad.

[0033] Based on the extracted UI design information, a first quasi-interest array is constructed. This array uses product design type as the baseline dimension and the advertising element design type set as the feature set, forming a binary structure of "product design type-advertising element design type", such as "smart wearable device-dynamic pop-up ads, bottom floating ads". Finally, the first quasi-interest array is incorporated into the quasi-interest array set, which serves as a temporary container for the user's potential interest characteristics and provides a data basis for subsequent high-frequency array statistics and profile updates. For example, if a user frequently clicks on the click windows containing "smart wearable device", "dynamic pop-up ads" and "bottom floating ads" while browsing a technology product page, the system will construct a corresponding quasi-interest array and add it to the set to reflect the user's interest in technology products and their specific advertising forms.

[0034] In a specific embodiment, the target user basic portrait includes a general portrait tag and an individual portrait tag, the individual portrait tag represents an initial interest array set, the first browsing record set of the first sample user group that meets the target user basic portrait is retrieved, and based on the product type and the several advertising element types, the quasi-interest array set is traversed to perform high-frequency array statistics to obtain an interest array set, including: retrieving the first browsing record set of the first sample user group that meets the general portrait tag and the individual portrait tag; extracting the first product design type and the first advertising element design type set of the first quasi-interest array of the quasi-interest array set; counting the triggering frequency ratio of the first product design type and the first advertising element design type set in the first browsing record set, and setting it as the first quasi-interest array confidence; when the first quasi-interest array confidence is greater than or equal to the confidence threshold, the first quasi-interest array is added to the interest array set.

[0035] Preferably, in the process of dynamic optimization of user portraits, implementing sample user group screening and high-frequency interest array mining for the target user's basic portrait (including general portrait tags and individual portrait tags) is an important step in accurately identifying the user's potential interests.

[0036] First, based on the target users' general portrait tags (such as age, occupation, region, etc., such as "25-34 years old", "Internet practitioners", "first-tier cities") and individual portrait tags (that is, the initial interest group array set, such as "sports equipment-video advertising", "beauty products-picture and text recommendations"), multi-dimensional screening conditions are constructed to retrieve the first sample user group that meets all the tags, and collect the first browsing record set of this user group, which records the complete browsing trajectory of the sample users on the Internet.

[0037] Subsequently, feature extraction is performed on the first quasi-interest array in the target user's quasi-interest array set (e.g., "sports equipment - video ads"), parsing out the first product design type ("sports equipment") and the first advertising element design type set ("video ads") it contains. Based on the extracted features, the first browsing record set is traversed, the number of times the first quasi-interest array is triggered in the browsing records of the sample user is counted, and its proportion of the total number of views in the set is calculated, which is defined as the confidence level of the first quasi-interest array. This indicator is used to quantify the representativeness of the array's interests in the sample user group. For example, if a quasi-interest array is triggered 100 times in the first browsing record set, and the total number of views in the set is 1000, its confidence level is 10%.

[0038] Finally, the confidence level of the first quasi-interest array is compared with a preset confidence threshold (e.g., 15%). If the confidence level is greater than or equal to the threshold, the array is determined to be significantly representative of the sample user group's interests and is added to the interest array set, dynamically updating the target user's interest feature set. For example, if the confidence level of the quasi-interest array "sports equipment - video ads" reaches 20%, exceeding the preset threshold, the system will add it to the interest array set, indicating that the target user and similar sample users have a high interest in video ads for sports equipment, providing data support for subsequent personalized recommendations.

[0039] In a specific embodiment, the triggering frequency ratio of the first product design type and the first advertising element design type set in the first browsing record set is counted and set as the first quasi-interest array confidence, including: counting the first triggering frequency ratio of the first product design type in the first browsing record set; counting the second triggering frequency ratio set of the first advertising element design type set in the first browsing record set respectively; calculating the mean of the first triggering frequency ratio and the second triggering frequency ratio set, and setting it as the first quasi-interest array confidence.

[0040] Exemplarily, in the process of constructing the confidence of the interest array based on the user browsing behavior data, the system first carries out multi-dimensional trigger frequency statistics for the first quasi-interest array in the target user's quasi-interest array set (such as "sports equipment-video ads, graphic pop-ups"). Specifically, the system first counts the first trigger frequency ratio of the first product design type ("sports equipment") in the first browsing record set. This ratio is obtained by calculating the ratio of the number of page browsing records containing "sports equipment" to the total number of the first browsing record set. The first trigger frequency ratio is used to quantify the breadth of interest coverage of the product type in the sample user group. For example, if there are 300 browsing records involving "sports equipment" in the first browsing record set and the total number of the set is 2,000, then the first trigger frequency ratio of "sports equipment" is 15%.

[0041] Next, we perform a statistical analysis of each of the first ad element design types ("video ads, pop-up ads"), calculating the percentage of each ad element type's second trigger frequency in the first set of browsing records to form a second trigger frequency percentage set. This second trigger frequency percentage set reflects the interest activation strength of different ad formats in the sample user group. For example, the trigger percentage for "video ads" is 10%, while the trigger percentage for "pop-up ads" is 8%.

[0042] Finally, the confidence level of the first quasi-interest array is calculated by averaging the percentages of the first and second trigger frequency percentages. This metric combines the strength of interest associations between product type and ad element type. The confidence level is used to assess the representativeness of the quasi-interest array within the sample user base. For example, if the triggering percentage for "sports equipment" is 15%, and the triggering percentages for "video ads" and "text and image pop-ups" are 10% and 8%, respectively, the average is calculated as (15% + 10% + 8%) / 3 ≈ 11%. This confidence level serves as the key basis for determining whether the first quasi-interest array should be included in the interest array set.

[0043] Through the above process, a quantitative evaluation of user interest characteristics is achieved, providing data support for subsequent portrait updates and recommendation strategy optimization.

[0044] In a specific embodiment, the universal portrait tag includes a universal type tag and a universal quantitative tag, and retrieving a first browsing record set of a first sample user group that meets the universal portrait tag and the individual portrait tag includes: configuring a universal quantitative attribute deviation threshold, combining the universal type tag and the universal quantitative tag, and constructing a universal portrait similarity evaluation function, wherein the universal portrait similarity evaluation function calculates the ratio of the sum of the number of first attributes that meet the universal quantitative attribute deviation threshold in the universal quantitative tag deviation between the universal quantitative tag and the sample universal quantitative tag, and the number of attributes of the intersection of the universal type tag and the sample universal type tag, to the total number of attributes of the universal portrait tag; constructing an individual portrait tag similarity evaluation function, wherein the individual portrait tag similarity evaluation function is used to calculate the ratio of the intersection label and the union label of the sample individual portrait tag and the individual portrait tag; calling the universal portrait similarity evaluation function and the individual portrait tag similarity evaluation function, and retrieving the first sample user group that meets both the universal similarity threshold and the individual similarity threshold.

[0045] Specifically, in the process of screening sample user groups driven by user portraits, we first conduct structured analysis on the general portrait labels of target users, breaking them down into general type labels (such as "age level - youth", "occupation category - Internet practitioners") and general quantitative labels (such as "age value - 28 years old", "consumption capacity score - 75 points"), and build a two-layer screening mechanism based on this.

[0046] In the general portrait similarity assessment phase, a general quantitative attribute deviation threshold is pre-configured (e.g., age deviation must not exceed ±3 years, and spending power score deviation must not exceed ±10 points). A general portrait similarity assessment function is then designed, which achieves sample matching through a two-step calculation: First, for each general quantitative tag, the number of attributes between the sample user and the target user whose deviation values fall within the threshold range is counted (e.g., if both age and spending power score meet the deviation threshold, the count is 2); second, the number of intersection attributes of the general type tags is calculated (e.g., if both the target user and the sample user are "youth" and "internet practitioners," the number of intersection attributes is 2). Finally, the sum of the number of matching attributes of the quantitative tags and the number of intersection attributes of the type tags is divided by the total number of general portrait tag attributes (e.g., if the total number of attributes is 5 and the number of matching attributes is 4), resulting in a general portrait similarity score, which quantifies the degree of similarity between the target user and the sample user at the general feature level.

[0047] For example, if the general portrait of the target user contains five labels: "age 28 years old (quantitative)", "youth (type)", "Internet practitioner (type)", "first-tier city (type)", and "consumption ability score 75 points (quantitative)", and the sample user portrait is "age 29 years old", "youth", "Internet practitioner", "first-tier city", and "consumption ability score 80 points", after calculation, the age deviation of 1 year and the consumption ability score deviation of 5 points in the quantitative labels both meet the threshold, and the three type labels are completely matched, then the similarity score is (2+3) / 5=100%.

[0048] In the individual portrait label similarity evaluation stage, a similarity evaluation function based on set operations is constructed for the target user's individual portrait labels (such as the initial interest array set "sports equipment-video advertising" and "beauty products-picture and text recommendations") and the sample user's individual portrait labels. This function quantifies the degree of overlap between the two at the individual interest feature level by calculating the ratio of the number of intersection labels and the number of union labels of the sample individual portrait labels and the target user's individual portrait labels (for example, the target user has 2 interest arrays and the sample user has 3 interest arrays, and 1 of the interest arrays overlaps, then the similarity is 1 / 3≈33.3%).

[0049] Based on the results of the two-tiered evaluation, the system calls both the general profile similarity evaluation function and the individual profile tag similarity evaluation function, sets a general similarity threshold (e.g., similarity score ≥ 80%) and an individual similarity threshold (e.g., similarity score ≥ 50%), retrieves a first sample user group that satisfies both thresholds, and extracts the first set of browsing records for that user group, providing a data foundation for subsequent high-frequency interest group mining. For example, if a sample user's similarity score with the target user at the general profile level is 90% and at the individual profile level is 60%, both exceeding the preset thresholds, the user is included in the first sample user group, and their browsing records will be used to analyze the target user's potential interest preferences.

[0050] In a specific embodiment, retrieving a second sample user group that meets the interest group array set includes: updating the individual portrait label similarity evaluation function according to the interest group array set to obtain an updated individual portrait label similarity evaluation function; calling the updated individual portrait label similarity evaluation function to retrieve the second sample user group that meets the individual similarity threshold.

[0051] Furthermore, in the dynamic screening mechanism of sample user groups driven by user portraits, iterative optimization of the individual portrait label similarity evaluation function is implemented for the constructed interest group array set (such as "sports equipment-video advertising" and "beauty products-live streaming") to improve the matching accuracy of the sample user group and the target user's interest characteristics.

[0052] Specifically, the interest array set is integrated as a core feature into the individual profile tag system, and a structured update is performed on the original evaluation function. Through a feature mapping mechanism, each element in the interest array set (e.g., "sports equipment - video ads") is converted into a quantifiable interest feature vector. Similarity calculation rules are then defined for this vector and the sample user's individual profile tag. For example, a cosine similarity algorithm is used to quantify the angular relationship between interest arrays in vector space, or the overlap of interest tag sets is calculated based on the Jaccard coefficient. This results in an updated individual profile tag similarity evaluation function. While retaining the original similarity calculation logic for individual profile tags (e.g., the initial interest array set and browsing preference type), this function adds an interest array set matching module. This function uses a weighted summation method to comprehensively evaluate the degree of match between the sample user and the target user across both basic and dynamic interests. The updated individual profile tag similarity evaluation function refers to this new evaluation model that incorporates dynamic interest features. For example, if the target user's individual portrait label includes the initial interest array "sports equipment-picture and text recommendation" and the newly mined interest array "sports equipment-video advertising", and the sample user's individual portrait label includes "sports equipment-video advertising" and "beauty products-picture and text recommendation", the updated evaluation function will calculate the similarity of the two sets of labels separately (such as through the intersection and union of the label sets, the similarity is 1 / 2=50%), and combine the preset weight coefficients (such as the initial interest weight is 0.4, the dynamic interest weight is 0.6) to comprehensively obtain the final similarity score.

[0053] The system then calls the updated evaluation function to perform a traversal search of the entire user sample library, selecting users whose individual similarity scores exceed a preset threshold (e.g., similarity ≥ 60%) as the second sample user group. This threshold is determined through cross-validation to ensure that the sample user group maintains high-dimensional consistency with the target user in terms of interest characteristics. For example, if a sample user and the target user score 65% under the updated individual portrait label similarity evaluation function, exceeding the threshold, the user is included in the second sample user group, and their subsequent browsing history will be used to further explore the target user's potential interest trends, thus forming a closed-loop optimization mechanism of "interest feature extraction - evaluation function update - sample user screening."

[0054] In a specific embodiment, updating the target user basic portrait according to the interest array set and the secondary interest array set also includes: after the target user basic portrait is updated, identifying a first timestamp for the newly added first interest array, wherein when the first interest array is clicked again, the click timestamp is used to update the first timestamp; when the time interval between the first timestamp and the current timestamp is greater than or equal to a time interval threshold, deleting the first interest array from the target user basic portrait.

[0055] Specifically, for the basic portrait of the target user, iterative updates are implemented based on interest group sets (such as "sports equipment-video advertising", "digital products-open screen pop-ups") and secondary interest group sets (such as "outdoor products-information flow pictures and text", "household products-short video recommendations"), and a time-sensitive array validity management mechanism is introduced to ensure the timeliness and accuracy of the portrait data.

[0056] After the system completes the fusion and update of the interest array set and the secondary interest array set, it timestamps the first interest array (such as "Outdoor Products - Information Flow Images and Text") added to the target user's basic profile, and uses the first timestamp to record the moment when the array was first included in the profile. For example, when "Outdoor Products - Information Flow Images and Text" is added to the profile at 14:00:00 on March 1, 2024, it is assigned this timestamp. If the first interest array is triggered again in subsequent user behavior (such as the user clicks on a similar ad or browses a related page), the system captures the real-time interaction moment through the click timestamp mechanism and overwrites and updates its first timestamp. For example, if the user clicks on the "Outdoor Products - Information Flow Images and Text" related content again at 10:30:00 on March 5, 2024, the timestamp of the array is synchronously updated to this moment.

[0057] At the same time, a preset time interval threshold (e.g., 30 days) is used as a criterion for determining the validity of interest arrays. By calculating the time interval between the current timestamp and the first timestamp of the first interest array, the system assesses whether it still reflects the user's true interests. If the time interval exceeds the threshold (e.g., if the current date is April 2, 2024, and the first timestamp of "Outdoor Products - Information Flow Graphics and Text" is March 1, 2024, the interval has reached 32 days), the interest array is deemed invalid and removed from the target user's basic profile through a deletion mechanism to prevent outdated data from interfering with the profile's accuracy. For example, if a user expressed interest in "Home Supplies - Short Video Recommendations" 30 days ago but has not interacted with it since, the system will automatically purge the array on the 31st day. Conversely, if a user continuously clicks on "Sports Equipment - Video Ads" within the past 30 days, its timestamp will be continuously updated to ensure that the profile remains focused on the user's active interests.

[0058] Through the above-mentioned time-sensitive dynamic update mechanism, the user portrait has evolved from a static feature set to a timely knowledge base, providing more valuable data support for scenarios such as personalized recommendations and advertising.

[0059] The user profiling method based on big data provided by the embodiment of the present invention has at least the following technical effects:

[0060] 1. By analyzing user page browsing trajectories, a set of quasi-interest arrays is generated, including product types and advertising element types. The browsing history of the sample user group is filtered based on the basic profile of the target user. The high-frequency interest arrays are filtered in combination with the confidence threshold. Finally, the primary and secondary interest array sets are integrated to complete the portrait update, forming a multi-dimensional interest feature set. This significantly improves the portrait's ability to cover users' complex interest preferences and provides richer feature input for personalized recommendations.

[0061] 2. A two-layer screening mechanism is constructed using general portrait tags and individual portrait tags. The general portrait similarity evaluation function calculates the basic similarity by quantifying the label deviation threshold and the number of intersections between type tags. The individual portrait tag similarity evaluation function quantifies the interest overlap through label set operations. The combination of the two achieves accurate matching of sample user groups, ensuring that portrait updates are based on user groups with highly similar behaviors, thereby improving data credibility.

[0062] 3. Dynamic validity management of interest arrays is achieved through a timestamp mechanism. New interest arrays are marked with a timestamp for the first time they are included. Interactions trigger timestamp updates, and arrays that exceed the time interval threshold are automatically deleted to avoid recommendation bias due to outdated interest features in the portrait. This mechanism ensures that the portrait always reflects the user's active interests, significantly improving recommendation accuracy in interest migration scenarios while reducing computing resource consumption.

[0063] Example 2:

[0064] like Figure 2 As shown, based on the same inventive concept as the user profiling method based on big data provided in Example 1, an embodiment of the present invention also provides a user profiling system based on big data, the system comprising: The information parsing module 11 is used to parse the page browsing trajectory of the target user to obtain a set of quasi-interest arrays, wherein any quasi-interest array includes a product type and several advertising element types.

[0065] The first data statistics module 12 is used to retrieve a first browsing record set of a first sample user group that meets the basic profile of the target user, and based on the product type and the multiple advertising element types, traverse the quasi-interest array set to perform high-frequency array statistics to obtain an interest array set.

[0066] The second data statistics module 13 is used to retrieve a second browsing record set of a second sample user group that meets the interest array set to obtain a secondary quasi-interest array set, and perform high-frequency array statistics on the secondary quasi-interest array set and the non-intersecting quasi-interest array of the quasi-interest array set to obtain a secondary interest array set.

[0067] The portrait updating module 14 is configured to update the target user basic portrait according to the interest group set and the secondary interest group set.

[0068] Furthermore, the information parsing module 11 is further configured to perform the following steps: Extract the first click window of the page browsing trajectory; extract the first UI design information of the first click window, wherein the first UI design information includes a first product design type and a first advertising element design type set, and construct a first quasi-interest array; add the first quasi-interest array to the quasi-interest array set.

[0069] Furthermore, the first data statistics module 12 is further configured to perform the following steps: Retrieve a first browsing record set of a first sample user group that meets the general portrait tag and the individual portrait tag; extract the first product design type and the first advertising element design type set of the first quasi-interest array of the quasi-interest array set; count the proportion of the triggering frequency of the first product design type and the first advertising element design type set in the first browsing record set, and set it as the first quasi-interest array confidence; when the first quasi-interest array confidence is greater than or equal to the confidence threshold, add the first quasi-interest array to the interest array set.

[0070] Furthermore, the first data statistics module 12 is further configured to perform the following steps: Count the first trigger frequency ratio of the first product design type in the first browsing record set; count the second trigger frequency ratio of the first advertising element design type set in the first browsing record set respectively; calculate the mean of the first trigger frequency ratio and the second trigger frequency ratio, and set it as the first quasi-interest array confidence level.

[0071] Furthermore, the first data statistics module 12 is further configured to perform the following steps: A universal quantitative attribute deviation threshold is configured, and a universal portrait similarity evaluation function is constructed in combination with the universal type label and the universal quantitative label, wherein the universal portrait similarity evaluation function calculates the ratio of the sum of the number of first attributes that meet the universal quantitative attribute deviation threshold and the number of attributes of the intersection of the universal type label and the sample universal type label, to the total number of universal portrait label attributes; an individual portrait label similarity evaluation function is constructed, wherein the individual portrait label similarity evaluation function is used to calculate the ratio of the intersection label and the union label of the sample individual portrait label and the individual portrait label; the universal portrait similarity evaluation function and the individual portrait label similarity evaluation function are called to retrieve the first sample user group that meets both the universal similarity threshold and the individual similarity threshold.

[0072] Furthermore, the second data statistics module 13 is further configured to perform the following steps: According to the interest group array set, the individual portrait label similarity evaluation function is updated to obtain an updated individual portrait label similarity evaluation function; the updated individual portrait label similarity evaluation function is called to retrieve the second sample user group that meets the individual similarity threshold.

[0073] Furthermore, the portrait updating module 14 is further configured to perform the following steps: When the basic portrait of the target user is updated, the first timestamp is identified for the newly added first interest array, wherein when the first interest array is clicked again, the click timestamp is used to update the first timestamp; when the time interval between the first timestamp and the current timestamp is greater than or equal to the time interval threshold, the first interest array is deleted from the basic portrait of the target user.

[0074] Through the above detailed description of the user profiling method based on big data in this specification, those skilled in the art can clearly understand the user profiling system based on big data in this embodiment. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0075] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. The user profiling method based on big data is characterized by: include: Analyze the target user's page browsing history to obtain a set of quasi-interest arrays, where any quasi-interest array includes a product type and several advertising element types; Retrieving a first browsing record set of a first sample user group that meets the target user's basic profile, and traversing the quasi-interest array set to perform high-frequency array statistics based on the product type and the multiple advertising element types to obtain an interest array set; Retrieving a second browsing record set of a second sample user group that satisfies the interest array set to obtain a secondary quasi-interest array set, performing high-frequency array statistics on the secondary quasi-interest array set and a non-intersecting quasi-interest array of the quasi-interest array set to obtain a secondary interest array set; The target user basic portrait is updated according to the interest group set and the secondary interest group set.

2. The method according to claim 1, wherein Analyze the target user's page browsing history to obtain a set of quasi-interest arrays, including: Extracting the first click window of the page browsing trajectory; Extracting first UI design information of the first click window, wherein the first UI design information includes a first product design type and a first advertisement element design type set, and constructing a first quasi-interest array; Add the first quasi-interest array to the quasi-interest array set.

3. The method according to claim 1, wherein The target user basic profile includes a general profile tag and an individual profile tag, wherein the individual profile tag represents an initial interest array set, a first browsing record set of a first sample user group that meets the target user basic profile is retrieved, and based on the product type and the multiple advertising element types, the quasi-interest array set is traversed to perform high-frequency array statistics to obtain an interest array set, including: Retrieving a first browsing record set of a first sample user group that meets the general portrait tag and the individual portrait tag; Extracting a first product design type and a first advertisement element design type set of a first quasi-interest array of the quasi-interest array set; Counting the triggering frequency ratios of the first product design type and the first advertising element design type set in the first browsing record set, and setting the ratio as the first quasi-interest group confidence level; When the confidence of the first quasi-interest array is greater than or equal to a confidence threshold, the first quasi-interest array is added to the interest array set.

4. The method according to claim 3, wherein Counting the triggering frequency ratios of the first product design type and the first advertising element design type set in the first browsing record set, and setting the ratio as the first quasi-interest group confidence level, includes: Counting the first trigger frequency ratio of the first product design type in the first browsing record set; respectively counting a second trigger frequency ratio set of the first advertisement element design type set in the first browsing record set; Calculate the mean of the first trigger frequency ratio and the second trigger frequency ratio, and set it as the first quasi-interest array confidence level.

5. The method according to claim 3, wherein The general portrait tag includes a general type tag and a general quantitative tag, and retrieving a first browsing record set of a first sample user group that meets the general portrait tag and the individual portrait tag includes: Configuring a universal quantitative attribute deviation threshold, and constructing a universal portrait similarity evaluation function in combination with the universal type label and the universal quantitative label, wherein the universal portrait similarity evaluation function calculates the ratio of the sum of the number of first attributes that meet the universal quantitative attribute deviation threshold in the universal quantitative label deviation between the universal quantitative label and the sample universal quantitative label, the number of attributes in the intersection of the universal type label and the sample universal type label, and the total number of attributes in the universal portrait label; Constructing an individual portrait label similarity evaluation function, wherein the individual portrait label similarity evaluation function is used to calculate the ratio of the intersection label to the union label of the sample individual portrait label and the individual portrait label; The general portrait similarity evaluation function and the individual portrait tag similarity evaluation function are retrieved to retrieve the first sample user group that satisfies both the general similarity threshold and the individual similarity threshold.

6. The method according to claim 5, wherein Retrieving a second sample user group that satisfies the interest group set includes: According to the interest group set, updating the individual portrait label similarity evaluation function to obtain an updated individual portrait label similarity evaluation function; The updated individual portrait label similarity evaluation function is called to retrieve the second sample user group that meets the individual similarity threshold.

7. The method according to claim 1, wherein Updating the target user basic profile according to the interest group set and the secondary interest group set further includes: After the target user basic profile is updated, a first timestamp is marked for the newly added first interest array, wherein when the first interest array is clicked again, the first timestamp is updated using the click timestamp; When the time interval between the first timestamp and the current timestamp is greater than or equal to a time interval threshold, the first interest array is deleted from the target user basic portrait.

8. The user portrait system based on big data is characterized by: For implementing the user profiling method based on big data according to any one of claims 1 to 7, the system comprises: An information parsing module is used to parse the target user's page browsing trajectory to obtain a set of quasi-interest arrays, wherein any quasi-interest array includes a product type and several advertising element types; A first data statistics module is configured to retrieve a first browsing record set of a first sample user group that meets the basic profile of the target user, and based on the product type and the multiple advertising element types, traverse the quasi-interest array set to perform high-frequency array statistics to obtain an interest array set; a second data statistics module, configured to retrieve a second browsing record set of a second sample user group that satisfies the interest array set, obtain a secondary quasi-interest array set, and perform high-frequency array statistics on the secondary quasi-interest array set and a non-intersecting quasi-interest array of the quasi-interest array set, to obtain a secondary interest array set; A portrait updating module is used to update the target user basic portrait according to the interest group set and the secondary interest group set.

Citation Information

Patent Citations

  • Construction method and device based on user portrait and storage medium

    CN109815386A

  • User similarity determination method and information recommendation method

    CN110223186A

  • Accurate advertisement putting method and system based on Internet big data

    CN118229362A

  • Intelligent broadcast ecological system and method based on artificial intelligence multi-terminal application

    CN119089052A

  • User portrait enhancement method based on dynamic tag

    CN119539910A