A processing method, device and medium for generating word cloud
By dividing the value range of the data type label into equal-length sub-intervals and adjusting the relative index, the problem of unclear differences in user characteristics is solved, and more specific target user group characteristics are displayed on the user interface.
Patent Information
- Application Number
- CN202411969414.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the prior art, improper setting of the value range of user data type labels results in unclear feature differences among specific target user groups, making it difficult to obtain sufficient feature information.
The value range of the data type label is divided into several equal-length sub-intervals, and the relative index of each attribute value is obtained through multiple adjustments to determine the optimal value range, which is used to judge the user's label attribute value and generate a word cloud to show the feature differences.
By adjusting equal-length subintervals and relative indexes, the characteristic differences among user groups are highlighted, and the number of characteristics and information display of specific target user groups are increased.
Smart Images

Figure CN119862222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a processing method, device and medium for generating a word cloud. Background Art
[0002] Existing databases store attribute values of tags of different dimensions for a large number of users. The attribute values of user tags can reflect the user's basic information and preferences, etc. By comparing the differences in the attribute values of the tags of a specific target user group and all users stored in the database, the characteristics of the specific target user group can be obtained, and then matching information can be pushed to the target user group based on the characteristics of the specific target user group.
[0003] User labels are generally divided into two types: data type labels and non-data type labels. Data type labels are labels whose corresponding attribute values correspond to value ranges. For example, a data type label is a consumption level label, and its corresponding attribute values are low consumption level, medium consumption level, and high consumption level. The corresponding value range for low consumption level is less than 300, indicating that a single consumption amount of less than 300 is considered low consumption level; the corresponding value range for medium consumption level is 300-700, indicating that a single consumption amount of 300-700 is considered medium consumption level; and the corresponding value range for high consumption level is greater than 700, indicating that a single consumption amount of more than 700 is considered high consumption level. Non-data type labels are labels whose corresponding attribute values do not correspond to value ranges. For example, a non-data type label is a gender label, and its corresponding attribute values are male and female.
[0004] The attribute values of a user's data type tag are related to the corresponding value range of each attribute value. If the value range of each attribute value of the data type tag is not set properly, the difference in the attribute values of the data type tag for a specific target user group compared to all users will not be obvious, making it difficult to obtain the characteristics of this specific target user group. How to increase the number of characteristics of specific target user groups to display more information about their characteristics is an urgent problem to be solved. Summary of the Invention
[0005] The present invention aims to provide a processing method, device and medium for generating a word cloud, so as to increase the number of features of a specific target user group obtained and display more information about the features of the specific target user group.
[0006] According to a first aspect of the present invention, a method for generating a word cloud is provided, the method comprising the following steps:
[0007] A value range of each data type label corresponding to the second user group is obtained; any data type label belongs to a preset label set, and the preset label set includes a plurality of data type labels and a plurality of non-data type labels.
[0008] The value range of each data type label is divided into several sub-intervals of equal length; wherein, the number of sub-intervals corresponding to any data type label is the number of attribute values corresponding to the data type label.
[0009] The equal-length subintervals corresponding to each data type label are adjusted a first preset number of times, and the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment is obtained; any user in the first user group belongs to the second user group, and the number of users included in the first user group is less than the number of users included in the second user group; the relative index of any attribute value of any data type label corresponding to the first user group is the ratio of the proportion of users corresponding to the attribute value of the data type label in the first user group to the proportion of users corresponding to the attribute value of the data type label in the second user group.
[0010] The preferred value range of each attribute value corresponding to each data type label is obtained according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment, and the attribute value of each data type label corresponding to each user in the first user group and the second user group is determined according to the preferred value range of each attribute value corresponding to each data type label.
[0011] The target word set displayed on the user interface is determined based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group; the target word set includes several target words, any target word is an attribute value of a tag, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, and the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group.
[0012] According to a second aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for generating a word cloud when executing the computer program.
[0013] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the above-mentioned processing method for generating a word cloud.
[0014] Compared with the prior art, the present invention has at least the following beneficial effects:
[0015] The present invention obtains the value range of each data type label corresponding to the second user group (i.e., the group including all users in the database). For the value range of any data type label, it is first divided into several equal-length sub-intervals, and then the equal-length sub-intervals are adjusted a first preset number of times, and the relative index of each attribute value of the data type label of the first user group (i.e., the specific target user group) corresponding to each adjustment is obtained. The relative index can reflect the difference in attribute values of the data type label between the first user group and the second user group; by adjusting the equal-length sub-intervals a first preset number of times, the present invention obtains the preferred value range corresponding to each attribute of the data type label. Using the preferred value range as the basis for judging the attribute value of the data type label for each user can highlight the difference in attribute values of the data type label between the first user group and the second user group, which is conducive to increasing the number of features of the first user group obtained, and further conducive to displaying more information about the features of the first user group on the user interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 A flowchart of a method for generating a word cloud according to the first embodiment of the present invention;
[0018] Figure 2 A flowchart of the steps of obtaining a preferred value range for each attribute value corresponding to each data type tag provided in the first embodiment of the present invention;
[0019] Figure 3 A flowchart of the steps of determining a target word set displayed on a user interface provided in the first embodiment of the present invention;
[0020] Figure 4 A flowchart of the steps of determining an initial word set provided in the first embodiment of the present invention;
[0021] Figure 5 A flowchart of the steps of screening the initial word set provided in the first embodiment of the present invention;
[0022] Figure 6 A flowchart of steps performed when the number of target words included in the intermediate word set is greater than a preset target display number is provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0024] Example 1:
[0025] According to this embodiment, Figure 1 As shown, a processing method for generating a word cloud is provided, and the method includes the following steps:
[0026] S100, obtaining a value range of each data type label corresponding to the second user group; any data type label belongs to a preset label set, and the preset label set includes a plurality of data type labels and a plurality of non-data type labels.
[0027] In this embodiment, the preset tag set is a pre-constructed tag set, and any tag in the preset tag set is a data type tag or a non-data type tag. The data type tag refers to a tag whose corresponding different attribute values correspond to a value range, such as a consumption level tag; the non-data type tag refers to a tag whose corresponding different attribute values do not correspond to a value range, such as a gender tag.
[0028] In this embodiment, the second user group is a group composed of users stored in a pre-constructed database, which is a larger group. The value range of each data type label of each user in the second user group is known; for example, the values corresponding to the consumption levels of all users in the second user group are known, so the minimum value (for example, 1) and the maximum value (for example, 1000) corresponding to the consumption levels of all users in the second user group are also known. Therefore, the value range of the consumption level labels of all users in the second user group (for example, [1,1000]) is also determined.
[0029] S200 , dividing the value range of each data type label into a number of sub-intervals of equal length; wherein the number of sub-intervals corresponding to any data type label is the number of attribute values corresponding to the data type label.
[0030] In this embodiment, the number of attribute values corresponding to each data type label is an empirical value; for example, the number of attribute values corresponding to the consumption level label is 3, so the value range corresponding to the consumption level label is divided into 3 equal-length sub-intervals. If the value range corresponding to the consumption level label is [1,1000], then the 3 equal-length sub-intervals obtained by division are: [1,334), [334,667) and [667,1000].
[0031] S300, adjust the equal-length subintervals corresponding to each data type label a first preset number of times, and obtain the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment; any user in the first user group belongs to the second user group, and the number of users included in the first user group is less than the number of users included in the second user group; the relative index of any attribute value of any data type label corresponding to the first user group is the ratio of the proportion of users corresponding to the attribute value of the data type label in the first user group to the proportion of users corresponding to the attribute value of the data type label in the second user group.
[0032] In this embodiment, the first preset number of times is an empirical value, and the magnitude of each adjustment to the equal-length subinterval corresponding to each data type label is also an empirical value. The magnitudes of the adjustments to the equal-length subintervals corresponding to different data type labels can be the same or different. For any adjustment, a relative index of each attribute value of each data type label corresponding to the first user group after the adjustment is obtained; the process of obtaining the relative index of any attribute value of any data type label corresponding to the first user group after the adjustment includes: obtaining the ratio of the number of users corresponding to the attribute value of the data type label in the first user group after the adjustment to the number of users included in the first user group (i.e., the percentage of users corresponding to the attribute value of the data type label in the first user group), obtaining the ratio of the number of users corresponding to the attribute value of the data type label in the second user group after the adjustment to the number of users included in the second user group (i.e., the percentage of users corresponding to the attribute value of the data type label in the second user group), and determining the ratio of the two as the relative index of the attribute value of the data type label corresponding to the first user group after the adjustment.
[0033] As an optional specific implementation method, the value range corresponding to the consumption level label is divided into three equal-length subintervals: [1,334), [334,667), and [667,1000]. These three equal-length subintervals are adjusted four times. The three subintervals obtained after the first adjustment are: [1,335), [335,665), and [665,1000], so that the division points of the adjusted subintervals are the nearest multiples of 5; the three subintervals obtained after the second adjustment are: [1,330), [ The three subintervals obtained after the third adjustment are [1,350), [350,650) and [650,1000], so that the split points of the adjusted subintervals are the nearest multiples of 50; the three subintervals obtained after the fourth adjustment are [1,300), [300,700) and [700,1000], so that the split points of the adjusted subintervals are the nearest multiples of 100.
[0034] S400, obtaining the preferred value range of each attribute value corresponding to each data type label according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment, and determining the attribute value of each data type label corresponding to each user in the first user group and the second user group according to the preferred value range of each attribute value corresponding to each data type label.
[0035] As a preferred embodiment, Figure 2 As shown, obtaining the preferred value range of each attribute value corresponding to each data type label according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment includes:
[0036] S410, obtaining a relative index set of each data type label corresponding to the first user group after each adjustment based on the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment; the relative index set of any data type label includes the relative index of each attribute value of the data type label.
[0037] S420, obtaining the number of target relative indexes included in the relative index set of each data type label corresponding to the first user group after each adjustment; the target relative index is a relative index that does not belong to a preset relative index range, the minimum value of the preset relative index range is less than 1 and the maximum value of the preset relative index range is greater than 1.
[0038] In this embodiment, the minimum value and the maximum value of the preset relative index range are both empirical values; preferably, the minimum value of the preset relative index range is 0.8, and the maximum value of the preset relative index range is 1.2.
[0039] S430 , determining the adjustment corresponding to the relative index set including the largest number of target relative indices corresponding to each data type label as the optimal adjustment corresponding to each data type label.
[0040] As a specific implementation, if there are more than two relative index sets that include the largest target relative index, the adjustment corresponding to any relative index set that includes the largest target relative index can be determined as the optimal adjustment corresponding to each data type label.
[0041] S440 , determining a preferred value range for each attribute value corresponding to each data type label according to the value range of each subinterval corresponding to the optimal adjustment corresponding to each data type label.
[0042] As a specific implementation method, the relative index set corresponding to the third adjustment of the three equal-length sub-intervals corresponding to the consumption level label includes the largest number of target relative indexes, then the sub-intervals [1,350), [350,650) and [650,1000] corresponding to the third adjustment are respectively determined as the preferred value ranges corresponding to the first attribute value (i.e., low consumption level), the second attribute value (medium consumption level) and the third attribute value (high consumption level) of the consumption level label.
[0043] Based on S410-S440, the preferred value range of each attribute value corresponding to each data type label can be obtained. Using the preferred value range as the basis for judging the attribute value of the data type label of each user can highlight the difference in the attribute value of the data type label between the first user group and the second user group, which is conducive to increasing the number of features (i.e., target words) of the first user group obtained subsequently.
[0044] S500, determining a target word set displayed on the user interface based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group; the target word set includes a plurality of target words, any target word is an attribute value of a tag, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, and the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group.
[0045] In this embodiment, the target word set is a word cloud; when each target word included in the target word set is displayed on the user interface, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, that is, the larger the relative index of a target word corresponding to the first user group is, the darker the display color of the target word is; the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group, that is, the larger the number of users corresponding to a target word corresponding to the first user group is, the larger the display area of the target word on the user interface is.
[0046] In this embodiment, based on S400, the attribute value of each data type label corresponding to each user in the first user group and the second user group can be obtained. On this basis, the relative index of each attribute value of each data type label corresponding to the first user group in S500 can be further obtained.
[0047] In this embodiment, the attribute value of each non-data type tag corresponding to each user in the first user group and the second user group is also known, so the relative index of each attribute value of each non-data type tag corresponding to the first user group can be obtained. The process of obtaining the relative index of each attribute value of each non-data type tag corresponding to the first user group is similar to the process of obtaining the relative index of each attribute value of each data type tag corresponding to the first user group, and will not be repeated here.
[0048] As a preferred embodiment, Figure 3 As shown, determining the target word set displayed on the user interface according to the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group includes:
[0049] S510, determining an initial word set based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group and a preset relative index range; the initial word set includes a plurality of initial words, any initial word is an attribute value of a tag, the relative index of any initial word corresponding to the first user group does not fall within the preset relative index range, the minimum value of the preset relative index range is less than 1, and the maximum value of the preset relative index range is greater than 1.
[0050] As a specific embodiment, Figure 4 As shown, determining the initial word set according to the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group and the preset relative index range includes:
[0051] S511, determine whether the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group belongs to the preset relative index range; if not, add the attribute value of the tag corresponding to the relative index to the first preset set, and the first preset set is initialized to an empty set; otherwise, do not add the attribute value of the tag corresponding to the relative index to the first preset set.
[0052] S512: After the determination is completed, the first preset set is determined as the initial word set.
[0053] Based on S511 - S512 , the attribute values of all tags that meet the condition that the relative index does not fall within the preset relative index range can be obtained.
[0054] S520, filtering the initial word set according to the number of users corresponding to each initial word in the first user group and a preset user number threshold to obtain an intermediate word set; the intermediate word set includes a plurality of intermediate words, and the number of users corresponding to any intermediate word in the first user group is greater than or equal to the preset user number threshold.
[0055] As a specific embodiment, Figure 5 As shown, screening the initial word set according to the number of users corresponding to each initial word in the first user group and a preset user number threshold includes:
[0056] S521, determine whether the number of users corresponding to each initial word in the first user group is greater than or equal to a preset user number threshold. If so, add the initial word to a second preset set, which is initialized to an empty set; otherwise, do not add the initial word to the second preset set.
[0057] In this embodiment, the preset user quantity threshold is an empirical value.
[0058] S522: After the determination is completed, the second preset set is determined as the intermediate word set.
[0059] Based on S521 - S522 , all initial words that meet the condition that the corresponding number of users is greater than or equal to a preset user number threshold can be obtained, and these initial words constitute an intermediate word set.
[0060] S530: If the number of intermediate words included in the intermediate word set is less than or equal to the preset target display number, the intermediate word set is determined as the target word set.
[0061] In this embodiment, the preset target display quantity is an empirical value, which corresponds to the maximum number of words that can be displayed on the user interface.
[0062] As a preferred embodiment, Figure 6As shown, if the number of target words included in the intermediate word set is greater than the preset target display number, the process proceeds to S540.
[0063] S540, updating the preset relative index range to obtain an updated relative index range; the minimum value of the updated relative index range is smaller than the minimum value of the preset relative index range, and the maximum value of the updated relative index range is larger than the maximum value of the preset relative index range.
[0064] As a specific implementation, the preset relative index range is [0.8, 1.2], and the updated relative index range is [0.75, 1.25].
[0065] S550, obtain the number of intermediate words included in the updated intermediate word set based on the updated relative index range. If the number of intermediate words included in the updated intermediate word set is less than or equal to the preset target display number, the updated intermediate word set is determined as the target word set; otherwise, continue to update the updated relative index range until the number of intermediate words included in the updated intermediate word set is less than or equal to the preset target display number.
[0066] Based on S540-S550, when the number of target words included in the intermediate word set is greater than the preset target display number, further screening of the target words displayed on the user interface can be achieved, and the characteristics of the first user group after further screening are more prominent.
[0067] This embodiment obtains the value range of each data type label corresponding to the second user group (i.e., the group including all users in the database). For the value range of any data type label, it is first divided into several equal-length sub-intervals, and then the equal-length sub-intervals are adjusted a first preset number of times, and the relative index of each attribute value of the data type label of the first user group (i.e., the specific target user group) corresponding to each adjustment is obtained. The relative index can reflect the difference in attribute values of the data type label between the first user group and the second user group; by adjusting the equal-length sub-intervals a first preset number of times, this embodiment obtains the preferred value range corresponding to each attribute of the data type label. Using this preferred value range as the basis for judging the attribute value of the data type label for each user can highlight the difference in attribute values of the data type label between the first user group and the second user group, which is conducive to increasing the number of features of the first user group obtained, and further conducive to displaying more information about the features of the first user group on the user interface.
[0068] As a preferred embodiment, after S500, the method further includes:
[0069] S10, obtain the user quantity list Q of the target word set; Q=(q1,q2,…,qj , m ,…,q n ), q i is the number of users corresponding to the i-th target word in the first user group, where the value range of i is from 1 to n, and n is the number of target words included in the target word set.
[0070] S20, obtain the initial display area list E of the target word set according to Q; E = (e1, e2,…, e i ,…, e n ), e i is the initial display area of the i-th target word on the user interface; e i = c0 × q i / (∑ n i=1 q i ), c0 is the area of the region on the user interface for displaying the word cloud, c0 > n × c’, e i ≥ c’, c’ is the preset minimum display area of a single target word.
[0071] In this embodiment, c0 is a preset value.
[0072] S30, obtain the number of users q0 corresponding to the specified word in the first user group; the specified word belongs to the target word set.
[0073] In this embodiment, the specified word is the word pre-input by the user.
[0074] S40, obtain the target display area e’ of the i-th target word i ; when q i ≥ q0, e’ i = k1 × e i ; when q i < q0, e’ i = (c0 - ∑ b j=1 f j ) × (q i / ∑ M m=1 h m ); f j is the target display area of the j-th target word in the target word set that satisfies the condition that the number of corresponding users in the first user group is greater than or equal to q0, where the value range of j is from 1 to b, and b is the number of target words in the target word set that satisfy the condition that the number of corresponding users in the first user group is greater than or equal to q0; h mis the number of users corresponding to the mth target word in the target word set in the first user group that satisfies the condition that the number of users corresponding to the first user group is less than q0, the value range of m is 1 to M, M is the number of target words in the target word set that satisfies the condition that the number of users corresponding to the first user group is less than q0; k1 is the preset amplification coefficient, k1>1, c0-∑ b j=1 f j ≥M×c',e' i ≥c'.
[0075] Based on S10-S40, the designated word can be highlighted on the user interface, and the relative size relationship of each target word in the target word set is maintained unchanged, which is conducive to users quickly locating the designated word and the target word with a larger number of corresponding users in the first user group.
[0076] Example 2:
[0077] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0078] A value range of each data type label corresponding to the second user group is obtained; any data type label belongs to a preset label set, and the preset label set includes a plurality of data type labels and a plurality of non-data type labels.
[0079] The value range of each data type label is divided into several sub-intervals of equal length; wherein, the number of sub-intervals corresponding to any data type label is the number of attribute values corresponding to the data type label.
[0080] The equal-length subintervals corresponding to each data type label are adjusted a first preset number of times, and the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment is obtained; any user in the first user group belongs to the second user group, and the number of users included in the first user group is less than the number of users included in the second user group; the relative index of any attribute value of any data type label corresponding to the first user group is the ratio of the proportion of users corresponding to the attribute value of the data type label in the first user group to the proportion of users corresponding to the attribute value of the data type label in the second user group.
[0081] The preferred value range of each attribute value corresponding to each data type label is obtained according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment, and the attribute value of each data type label corresponding to each user in the first user group and the second user group is determined according to the preferred value range of each attribute value corresponding to each data type label.
[0082] The target word set displayed on the user interface is determined based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group; the target word set includes several target words, any target word is an attribute value of a tag, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, and the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group.
[0083] Example 3:
[0084] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following steps are implemented:
[0085] A value range of each data type label corresponding to the second user group is obtained; any data type label belongs to a preset label set, and the preset label set includes a plurality of data type labels and a plurality of non-data type labels.
[0086] The value range of each data type label is divided into several sub-intervals of equal length; wherein, the number of sub-intervals corresponding to any data type label is the number of attribute values corresponding to the data type label.
[0087] The equal-length subintervals corresponding to each data type label are adjusted a first preset number of times, and the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment is obtained; any user in the first user group belongs to the second user group, and the number of users included in the first user group is less than the number of users included in the second user group; the relative index of any attribute value of any data type label corresponding to the first user group is the ratio of the proportion of users corresponding to the attribute value of the data type label in the first user group to the proportion of users corresponding to the attribute value of the data type label in the second user group.
[0088] The preferred value range of each attribute value corresponding to each data type label is obtained according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment, and the attribute value of each data type label corresponding to each user in the first user group and the second user group is determined according to the preferred value range of each attribute value corresponding to each data type label.
[0089] The target word set displayed on the user interface is determined based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group; the target word set includes several target words, any target word is an attribute value of a tag, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, and the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group.
[0090] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0091] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A processing method for generating a word cloud, characterized in that: The processing method comprises the following steps: Obtaining a value range of each data type tag corresponding to the second user group; any data type tag belongs to a preset tag set, the preset tag set including a plurality of data type tags and a plurality of non-data type tags; Divide the value range of each data type label into several equal-length subintervals; the number of subintervals corresponding to any data type label is the number of attribute values corresponding to the data type label; Adjusting the equal-length subintervals corresponding to each data type label a first preset number of times, and obtaining a relative index of each attribute value of each data type label corresponding to the first user group after each adjustment; any user in the first user group belongs to the second user group, and the number of users included in the first user group is less than the number of users included in the second user group; the relative index of any attribute value of any data type label corresponding to the first user group is a ratio of the proportion of users corresponding to the attribute value of the data type label in the first user group to the proportion of users corresponding to the attribute value of the data type label in the second user group; Obtaining a preferred value range for each attribute value corresponding to each data type label based on the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment, and determining an attribute value for each data type label corresponding to each user in the first user group and the second user group based on the preferred value range for each attribute value corresponding to each data type label; The target word set displayed on the user interface is determined based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group; the target word set includes several target words, any target word is an attribute value of a tag, the display color depth of any target word is positively correlated with the relative index of the target word corresponding to the first user group, and the display size of any target word is positively correlated with the number of users corresponding to the target word in the first user group.
2. The method for generating a word cloud according to claim 1, wherein: Obtaining the preferred value range of each attribute value corresponding to each data type label according to the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment includes: Obtaining a relative index set of each data type label corresponding to the first user group after each adjustment based on the relative index of each attribute value of each data type label corresponding to the first user group after each adjustment; the relative index set of any data type label includes the relative index of each attribute value of the data type label; Obtaining the number of target relative indexes included in the relative index set for each data type label corresponding to the first user group after each adjustment; the target relative index is a relative index that does not fall within a preset relative index range, the minimum value of the preset relative index range is less than 1 and the maximum value of the preset relative index range is greater than 1; Determine the adjustment corresponding to the relative index set including the largest target relative index corresponding to each data type label as the optimal adjustment corresponding to each data type label; The preferred value range of each attribute value corresponding to each data type label is determined according to the value range of each subinterval corresponding to the optimal adjustment corresponding to each data type label.
3. The method for generating a word cloud according to claim 1, wherein: Determining the target word set displayed on the user interface according to the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group includes: An initial word set is determined based on the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group and a preset relative index range; the initial word set includes a plurality of initial words, each initial word is an attribute value of a tag, the relative index of any initial word corresponding to the first user group does not fall within the preset relative index range, the minimum value of the preset relative index range is less than 1, and the maximum value of the preset relative index range is greater than 1; The initial word set is screened based on the number of users corresponding to each initial word in the first user group and a preset user number threshold to obtain an intermediate word set; the intermediate word set includes a plurality of intermediate words, and the number of users corresponding to any intermediate word in the first user group is greater than or equal to the preset user number threshold; If the number of intermediate words included in the intermediate word set is less than or equal to the preset target display number, the intermediate word set is determined as the target word set.
4. The method for generating a word cloud according to claim 3, wherein: The processing method further includes the following steps: if the number of target words included in the intermediate word set is greater than a preset target display number, updating the preset relative index range to obtain an updated relative index range; wherein the minimum value of the updated relative index range is less than the preset minimum value of the relative index range, and the maximum value of the updated relative index range is greater than the preset maximum value of the relative index range; The number of intermediate words included in the updated intermediate word set is obtained according to the updated relative index range. If the number of intermediate words included in the updated intermediate word set is less than or equal to the preset target display number, the updated intermediate word set is determined as the target word set; otherwise, the updated relative index range is continued to be updated until the number of intermediate words included in the updated intermediate word set is less than or equal to the preset target display number.
5. The method for generating a word cloud according to claim 3, wherein: Determining the initial word set according to the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group and the preset relative index range includes: Determine whether the relative index of each attribute value of each tag in the preset tag set corresponding to the first user group falls within a preset relative index range; if not, add the attribute value of the tag corresponding to the relative index to a first preset set, where the first preset set is initialized to an empty set; otherwise, do not add the attribute value of the tag corresponding to the relative index to the first preset set; After the determination is completed, the first preset set is determined as the initial word set.
6. The method for generating a word cloud according to claim 3, wherein: Screening the initial word set according to the number of users corresponding to each initial word in the first user group and a preset user number threshold includes: Determine whether the number of users corresponding to each initial word in the first user group is greater than or equal to a preset user number threshold; if so, add the initial word to a second preset set, which is initialized to an empty set; otherwise, do not add the initial word to the second preset set; After the determination is completed, the second preset set is determined as the intermediate word set.
7. The method for generating a word cloud according to claim 2, wherein: The preset minimum value of the relative index range is 0.8, and the preset maximum value of the relative index range is 1.
2.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating a word cloud according to any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a word cloud according to any one of claims 1 to 7 is implemented.