Information Recommendation Method, Information Recommendation Device, Electronic Device, and Storage Medium
By combining random access data and the initial recommendation data of real virtual users, the similarity value is calculated to determine the target recommendation data, the problem of information cocoon and digital divide in the prior art is solved, multi-dimensional information recommendation is achieved, and the diversity and depth of user knowledge acquisition is enhanced.
Patent Information
- Application Number
- CN202210931509.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-04
AI Technical Summary
The existing algorithm recommendations have problems with information cocoons and digital divides, which leads to users being easily trapped in the singularization of knowledge acquisition and may even fall into the "three kilometers around the surroundings of habit".
By obtaining random access data sets generated by random access to different types of information information, as well as initial recommendation data generated for real users and virtual users, the similarity value is calculated, the target recommendation data is determined, and information recommendation is then recommended to real users.
This method can not only retain a certain degree of recommendation advantages, but also solve the problem of information cocoon, avoid falling into information homogeneity and one-sidedness, and enhance users' active thinking ability and social cognition.
Smart Images

Figure CN115238186B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a method for information recommendation, an information recommendation device, an electronic device, a storage medium, and a program product. Background Art
[0002] With the rapid development of the Internet era, we are now in the era of information explosion. People are increasingly inseparable from the network, watching current affairs news, industry hotspots, information videos, pictures and other various types of content. The amount of information they come into contact with every day is in the tens of thousands. We have entered an era of information overload from the era of information scarcity. A large number of information flow platforms use algorithm preferences to cater to users' desires to obtain the information they are interested in from numerous pieces of information.
[0003] In the process of implementing the present disclosure, it is found that existing algorithm recommendations still have the problems of information cocoons and digital divides, resulting in users being prone to single - sourced knowledge acquisition and even potentially falling into the "peripheral three kilometers of habit". Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a method for information recommendation, an information recommendation device, an electronic device, a storage medium, and a program product.
[0005] According to a first aspect of the present disclosure, there is provided a method for information recommendation, including:
[0006] Obtaining a random access data set generated by randomly accessing different types of information, where the random access data set includes at least one random access data;
[0007] Obtaining first initial recommendation data generated for real users and second initial recommendation data generated for N virtual users, where the N virtual users are fictionalized based on the attribute characteristics of real users, and N is an integer greater than or equal to 1;
[0008] Determining target recommendation data according to the random access data set and the initial recommendation data set, where the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data; and
[0009] Performing information recommendation for real users according to the target recommendation data.
[0010] According to an embodiment of the present disclosure, determining target recommendation data according to the random access data set and the initial recommendation data set includes:
[0011] Calculating the similarity between the random access data and the first initial recommendation data and the second initial recommendation data to obtain similarity values;
[0012] Determining an initial data group when it is determined that the similarity value exceeds a first threshold;
[0013] Divide the initial data group into M target data groups according to a second threshold, where M is an integer greater than or equal to 1; and
[0014] Screen from the M target data groups respectively to obtain target recommended data.
[0015] According to an embodiment of the present disclosure, determining target recommended data according to a random access data set and an initial recommended data set further includes:
[0016] Determine hot information data according to the random access data set and the initial recommended data set;
[0017] Use the initial recommended data set to determine user access tags of the hot information data;
[0018] Determine a first threshold according to the hot information data, the user access tags and a similarity value; and
[0019] Determine a second threshold according to the first threshold.
[0020] According to an embodiment of the present disclosure, calculating the similarity between the random access data and the first initial recommended data and the second initial recommended data to obtain a similarity value includes:
[0021] Obtain algorithm keywords from the random access data, the first initial recommended data and the second initial recommended data; and
[0022] Calculate the similarity according to the algorithm keywords to obtain the similarity value.
[0023] According to an embodiment of the present disclosure, the attribute characteristics of a real user include the user behavior characteristics of the real user;
[0024] N virtual users are fictionalized according to the attribute characteristics of the real user, including:
[0025] Obtain the user behavior characteristics of the real user; and
[0026] Fictionalize N virtual users according to the user behavior characteristics.
[0027] According to an embodiment of the present disclosure, the attribute characteristics of a real user further include the user attribute characteristics of the real user;
[0028] N virtual users are fictionalized according to the attribute characteristics of the real user, and further include:
[0029] Fictionalize virtual users by changing the user attribute characteristics of the real user.
[0030] According to an embodiment of the present disclosure, before obtaining the initial recommended data of N virtual users and real users, it further includes:
[0031] Preprocess the real access data of real users to obtain the preprocessed real access data;
[0032] Generate the first initial recommendation data according to the real access data;
[0033] Generate the second initial recommendation data according to the attribute characteristics of N virtual users.
[0034] The second aspect of the present disclosure provides an information recommendation device, including:
[0035] A first acquisition module, configured to acquire a random access data set generated by randomly accessing different types of information, where the random access data set includes at least one random access data;
[0036] A second acquisition module, configured to acquire the first initial recommendation data generated for real users and the second initial recommendation data generated for N virtual users, where the N virtual users are fictionalized according to the attribute characteristics of real users, and N is an integer greater than or equal to 1;
[0037] A determination module, configured to determine target recommendation data according to the random access data set and the initial recommendation data set, where the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data; and
[0038] A recommendation module, configured to perform information recommendation for real users according to the target recommendation data.
[0039] The third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above information recommendation method.
[0040] The fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above information recommendation method.
[0041] The fifth aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above information recommendation method is implemented.
[0042] According to an embodiment of the present disclosure, by obtaining random access data of different types of information and obtaining initial recommendation data for real and virtual users, recommendation data is obtained from multiple dimensions, and by using the random access data set and the initial recommendation data set to determine the target recommendation data, it is not only possible to retain a certain degree of recommendation advantage, but also solve the problem of information cocoons, avoid falling into information homogenization and one-sidedness, and even avoid the limitations of users' active thinking ability and social cognition due to being addicted to the surrounding "three kilometers". BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0044] Figure 1 Schematically shows an application scenario diagram of an information recommendation method, an information recommendation device, an electronic device, a storage medium, and a program product according to an embodiment of the present disclosure;
[0045] Figure 2 Schematically shows a flowchart of an information recommendation method according to an embodiment of the present disclosure;
[0046] Figure 3 Schematically shows a method flowchart for determining target recommendation data according to a random access data set and an initial recommendation data set according to an embodiment of the present disclosure;
[0047] Figure 4 Schematically shows an information recommendation system architecture diagram according to an embodiment of the present disclosure;
[0048] Figure 5 Schematically shows a structural block diagram of an information recommendation device according to an embodiment of the present disclosure; and
[0049] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing an information recommendation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0051] The terms used herein are for describing specific embodiments only and are not intended to limit the present disclosure. Terms such as "including" and "comprising" used herein indicate the presence of the described features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.
[0052] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0053] In the case of using expressions such as "at least one of A, B, and C, etc.", generally it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0054] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs.
[0055] In the technical solutions of the embodiments of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0056] In the process of implementing the present disclosure, it is found that people of all ages and genders like to use mobile phones for entertainment in their spare time. A large number of information flow platforms cater to the public through algorithm preferences and recommend similar content to the public. Over time, on the one hand, it will face being fixed in a certain information circle, and even continuously strengthening views on certain issues, and finally forming values; on the other hand, with the increase of information, it is increasingly difficult to concentrate on complex issues, and the judgment is prone to decline. Both aspects are likely to fall into the information cocoon unconsciously. If the public does not know that they have fallen into the information cocoon, it may form a state of continuous victimization without awareness, which is likely to fall into the echo chamber effect. Among them, the echo chamber effect has four key impacts: isolation, extremeization of views, homogenization of views, and repeated dissemination of the same information.
[0057] Embodiments of the present disclosure provide an information recommendation method, including: obtaining a random access data set generated by randomly accessing different types of information, where the random access data set includes at least one random access data; obtaining first initial recommendation data generated for a real user and second initial recommendation data generated for N virtual users, where the N virtual users are fictionalized based on the attribute characteristics of the real user, and N is an integer greater than or equal to 1; determining target recommendation data according to the random access data set and the initial recommendation data set, where the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data; and performing information recommendation for the real user according to the target recommendation data.
[0058] Figure 1 Schematically shows an application scenario diagram of an information recommendation method, an information recommendation device, an electronic device, a storage medium, and a program product according to an embodiment of the present disclosure.
[0059] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0060] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0061] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0062] The server 105 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal devices 101, 102, 103 (only for example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0063] It should be noted that the information recommendation method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the information recommendation device provided by the embodiments of the present disclosure can generally be set in the server 105. The information recommendation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the information recommendation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0064] The information recommendation method provided by the embodiments of the present disclosure can also be executed by the terminal devices 101, 102, 103. Correspondingly, the information recommendation device provided by the embodiments of the present disclosure can generally also be set in the terminal devices 101, 102, 103. The information recommendation method provided by the embodiments of the present disclosure can also be executed by other terminals different from the terminal devices 101, 102, 103. Correspondingly, the information recommendation device provided by the embodiments of the present disclosure can also be set in other terminals different from the terminal devices 101, 102, 103.
[0065] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0066] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figures 2 to 4 The following will be based on the
[0067] Figure 2 scenario described below, and will describe in detail the information recommendation method of the disclosed embodiments through
[0068] As Figure 2 shown, the information recommendation method 200 of this embodiment includes operations S201 to S204.
[0069] In operation S201, a random access data set generated by randomly accessing different types of information is obtained, where the random access data set includes at least one random access data.
[0070] According to the embodiments of the present disclosure, different types of information can include: current affairs hot news information, financial information, entertainment information, sports information, technology information, game information, novel information, and location information, etc. Random access can automatically simulate real user behavior through a crawler, and can create requests containing questions of user-session control to simulate the random behavior of real users.
[0071] Among them, random behaviors may include, but are not limited to: random clicks, random page browsing (swiping the interface up, down, left, or right), random keyword searches, random likes, random collections, etc.
[0072] It should be noted that random behaviors such as randomly clicking on different types of information and randomly searching for keywords can diversify the access data, thus disrupting the collection of real user behavior data by the recommendation system.
[0073] In operation S202, obtain the first initial recommendation data generated for real users and the second initial recommendation data generated for N virtual users, where the N virtual users are fictionalized based on the attribute characteristics of real users, and N is an integer greater than or equal to 1.
[0074] According to an embodiment of the present disclosure, the first initial recommendation data may be generated based on the attribute characteristics of real users. The second initial recommendation data may be generated based on the attribute characteristics of virtual users. Among them, the fact that the N virtual users are fictionalized based on the attribute characteristics of real users may include first obtaining the attribute characteristics similar to or related to the attribute characteristics of real users based on the attribute characteristics of real users, and obtaining virtual users based on these attribute characteristics.
[0075] In operation S203, determine the target recommendation data according to the random access data set and the initial recommendation data set, where the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data.
[0076] According to an embodiment of the present disclosure, the target recommendation data may be generated according to the random access data set and the initial recommendation data set. It is also possible to calculate the similarity of the data in the random access data set and the initial recommendation data set, and retain the data with a high repetition rate to generate the target recommendation data.
[0077] In operation S204, perform information recommendation for real users according to the target recommendation data.
[0078] According to an embodiment of the present disclosure, information recommendation for real users may be performed according to the target recommendation data obtained in operation S203.
[0079] According to an embodiment of the present disclosure, through the acquisition of random access data of different types of information and the acquisition of initial recommendation data of real and virtual users, recommendation data is obtained from multiple dimensions, and the target recommendation data is determined through the random access data set and the initial recommendation data set. This can not only retain a certain degree of recommendation advantages, but also solve the problem of information cocoons, avoid information homogenization and one-sidedness, and even avoid the limitations of users' active thinking ability and social cognition due to being addicted to the surrounding "three kilometers".
[0080] Figure 3 Schematically shown is a flowchart of a method for determining target recommendation data based on a random access data set and an initial recommendation data set according to an embodiment of the present disclosure.
[0081] As Figure 3 shown, the method 300 for determining target recommendation data based on a random access data set and an initial recommendation data set according to this embodiment includes operation S301 to operation S304.
[0082] In operation S301, calculate the similarity between the random access data and the first initial recommendation data and the second initial recommendation data to obtain a similarity value.
[0083] According to an embodiment of the present disclosure, algorithms such as TF-IDF, BM25, and TextRank can be used to obtain algorithm keywords for the random access data and the first initial recommendation data and the second initial recommendation data, and then algorithms such as cosine distance, Jaccard coefficient, Dice coefficient, Jaro-Winkler, edit distance, and Hamming distance based on keywords can be used to calculate the similarity to obtain a similarity value.
[0084] For example, the similarity between every two of the random access data and the first initial recommendation data and the second initial recommendation data can be calculated and stored as key-value pairs, such as, keyword: similarity value. For example, A-B: 0.8, A-C: 0.32, A-D: 0.48, A-E: 0.73, A-F: 0.16, B-C: 0.26, B-D: 0.45, B-E: 0.67, B:F: 0.08, C-D: 0.6, C-E: 0.42, D-E: 0.34, D-F: 0.12, E-F: 0.28, etc.
[0085] In operation S302, when it is determined that the similarity value exceeds a first threshold, determine an initial data group.
[0086] According to an embodiment of the present disclosure, the first threshold can be determined based on user access tags obtained from actual information recommendations, random access data and initial recommendation data, similarity values, etc. The first threshold can also be set dynamically. The initial data group can include the random access data and the first initial recommendation data and the second initial recommendation data whose similarity values exceed the first threshold.
[0087] For example, the first threshold can be set to 0.3, then the initial data group can be {(A-B: 0.8), (A-C: 0.32), (A-D: 0.48), (A-E: 0.73), (B-D: 0.45), (B-E: 0.67), (C-D: 0.6), (C-E: 0.42), (D-E: 0.34)}.
[0088] In operation S303, according to the second threshold, the initial data group is divided into M target data groups, where M is an integer greater than or equal to 1.
[0089] According to the embodiments of the present disclosure, the second threshold can be determined according to the set first threshold. The second threshold can also be dynamically set according to the dynamically set first threshold.
[0090] For example, if the second threshold is 0.5, then one target data group can be {(A - B: 0.8), (A - E: 0.73), (B - E: 0.67)}; one target data group can be {(C - D: 0.6)}.
[0091] In operation S304, screening is performed respectively from the M target data groups to obtain target recommended data.
[0092] According to the embodiments of the present disclosure, the data corresponding to the keyword with the largest keyword score can be screened from each target data group as the target recommended data. Wherein, the keyword score can be the sum of similarity values related to the keyword.
[0093] For example, the score of keyword A can be the sum of similarity values in (A - B: 0.8) and (A - E: 0.73). The score of keyword B can be the sum of similarity values in (A - B: 0.8) and (B - E: 0.67). The score of keyword E can be the sum of similarity values in (A - E: 0.73) and (B - E: 0.67). The target recommended data can be the data corresponding to keyword A.
[0094] According to the embodiments of the present disclosure, by calculating similarity values, and then according to the set double-layer threshold, target recommended data is determined from the random access data set and the initial recommendation data set. It can not only retain a certain degree of recommendation advantage, but also solve the problem of information cocoons and avoid falling into the echo chamber effect.
[0095] According to the embodiments of the present disclosure, determining target recommended data according to the random access data set and the initial recommendation data set may further include: determining hot information data according to the random access data set and the initial recommendation data set; using the initial recommendation data set to determine user access tags of the hot information data; determining a first threshold according to the hot information data, the user access tags and the similarity values; and determining a second threshold according to the first threshold.
[0096] According to an embodiment of the present disclosure, by utilizing the characteristics of information, the one with more information of the same type in the random access data set and the initial recommendation data set can be determined as the hot information data. Based on the initial recommendation data set, the user access label of the hot information data can be determined according to the preference degree of a certain type of information user in the initial recommendation data. According to the hot information data, the user access label, and the similarity value, a first threshold can be determined. For example, the first threshold F1(threshold) can be obtained according to the calculation method of the following formula (1):
[0097]
[0098] where similarity represents the similarity value; hot represents the number of hot information data; prefer represents the user access label; time represents the time after the information is generated; sum represents the total number of data in the random access data set and the initial recommendation data set within the time of time; e, f, k, d, b all represent adjustable parameters; a i represents the i-th group of random access data sets and the initial recommendation data set.
[0099] According to the first threshold, determining the second threshold may include iterative calculation based on the first threshold and being positively correlated with the proportion of the hot information data. For example, the second threshold F2(threshold) can be obtained according to the calculation method of the following formula (2):
[0100] F2(threshold) = F1(threshold) * (g * hot / sum + time) (2)
[0101] where g represents an adjustable parameter.
[0102] According to an embodiment of the present disclosure, the double-layer threshold is a dynamic threshold, which is considered dynamically from three dimensions: one is the hot information data, and the first threshold is relatively reduced, so that more data of this type of information are selected; the second is through the user's preferences. If the content that the user prefers is corresponding to the user access label, similarly, the first threshold is relatively reduced, appropriately increasing the content that the user is interested in, while retaining a certain recommendation function; the third is timeliness. As the timestamp of the information data increases, the value of the information may gradually decrease, and the second threshold is gradually increased to reduce the selection of this type of information data. While retaining a certain degree of recommendation advantage, it can also avoid information cocoons, and even avoid the limitations of the user's active thinking ability and social cognition due to being addicted to the surrounding "three kilometers".
[0103] According to an embodiment of the present disclosure, calculating the similarity between the randomly accessed data and the first initial recommendation data and the second initial recommendation data to obtain a similarity value may include: obtaining algorithm keywords from the randomly accessed data, the first initial recommendation data, and the second initial recommendation data; and calculating the similarity according to the algorithm keywords to obtain the similarity value.
[0104] According to an embodiment of the present disclosure, the algorithm keywords can be obtained through algorithms such as TF-IDF, BM25, and TextRank. Algorithms such as cosine distance based on keywords, Jaccard coefficient, Dice coefficient, Jaro-Winkler, edit distance, and Hamming distance can be used to calculate the similarity to obtain the similarity value. It should be noted that using such methods to calculate the similarity value can reduce the amount of calculation.
[0105] According to an embodiment of the present disclosure, the attribute characteristics of real users may include the user behavior characteristics of real users. Among them, the N virtual users are fictionalized according to the attribute characteristics of real users, and may include: obtaining the user behavior characteristics of real users; and fictionalizing N virtual users according to the user behavior characteristics.
[0106] According to an embodiment of the present disclosure, the user behavior characteristics may include behavior characteristics such as clicking, browsing pages (swiping the interface up, down, left, or right), searching keywords, liking, and collecting. The user behavior characteristics of different real users can be obtained, and by changing some behavior characteristics, the corresponding users may be changed, and the changed users are used as the obtained virtual users.
[0107] According to an embodiment of the present disclosure, by fictionalizing virtual users and increasing user behavior characteristics, a more real and random multi-angle content display can be provided for users, which is beneficial to breaking the information cocoon.
[0108] According to an embodiment of the present disclosure, the attribute characteristics of real users may further include the user attribute characteristics of real users. Among them, the N virtual users are fictionalized according to the attribute characteristics of real users, and may further include: fictionalizing virtual users by changing the user attribute characteristics of real users.
[0109] According to an embodiment of the present disclosure, the user attribute characteristics may include, for example, user address, age, name, gender, etc.
[0110] According to an embodiment of the present disclosure, since, for example, news apps often push military news to men; push game and entertainment news to children, etc., fictionalizing virtual users can be used to break the fixed label settings for users, and can avoid the occurrence of push based on the inherent feature labels of users, which is beneficial to breaking the information cocoon.
[0111] According to an embodiment of the present disclosure, before obtaining the initial recommendation data of N virtual users and real users, it may further include: preprocessing the real access data of the real users to obtain the preprocessed real access data; generating the first initial recommendation data according to the real access data; and generating the second initial recommendation data according to the attribute characteristics of the N virtual users.
[0112] According to an embodiment of the present disclosure, preprocessing the real access data of the real users may include deleting or changing a part of the real access data of the real users. For example, the real access data may be data generated for the favorites, shopping cart, browser, historical browsing traces, etc. The preprocessed real access data can be obtained by cleaning or changing these data.
[0113] Based on the information recommendation method provided above, the present disclosure also provides an information recommendation system. In the past, mass media recommended the same information to everyone without distinction through radio or television, while now the APP will actively screen the information and create a unique "personal digest" for the user through the interface. When a user uses a certain APP, the APP will habitually capture each picture or the content below that the user clicks, and then classify it in the background by tagging. Its essence is sorting, that is, the platform can collect the sorting of multiple tags of an account. Most of the recommendation systems of large platforms adopt automatic recommendation such as supervised learning algorithms. Simply put, it is: who are you, where are you, and what do you like to watch? However, there are still problems of information cocoons and digital divides in this kind of recommendation, resulting in users being prone to single knowledge acquisition and even possibly falling into the "perimeter three kilometers of habits".
[0114] Figure 4 Schematically shows an information recommendation system architecture diagram according to an embodiment of the present disclosure.
[0115] As Figure 4 shown, this information recommendation system architecture diagram can be divided into three layers according to system functions, namely the support layer, the interactive service layer, and the display layer. The support layer may include a common resource sharing pool, namely a user resource pool, a policy pool, a user behavior logic unit pool, and a computing cache pool. The interactive service layer may integrate an interface management subsystem, a user data management subsystem, a user management subsystem, a user behavior subsystem, a price comparison subsystem, a price discrimination reporting subsystem, etc. The display layer can realize the page display of the client.
[0116] Among them, the user resource pool: can be used to store, manage, create, maintain, modify, and destroy user information. Fictitious virtual users in the embodiments of the present disclosure can be implemented in this layer.
[0117] Policy Pool: It can store adaptation policies for different APPs. The information recommendation policies adopted by different APPs can be different. For example, in a search engine app, when searching for how to install a light bulb, shopping apps such as Taobao and JD.com will push shopping links for light bulbs. To prohibit such unnecessary data collection between different apps, measures such as closing the collection of user input and prohibiting the inter - app communication of user data can be taken. This can solve the information cocoon problem caused by "where are you" in the domain algorithm. The policy pool can store the policies of the information recommendation method provided in the embodiments of the present disclosure.
[0118] User Behavior Logic Unit Pool: Since the behavior of users has a certain logic. Assembling user behaviors into user behavior logic units facilitates being called.
[0119] Computing Cache Pool: It can calculate the data similarity and cache it. The similarity is stored in the cache pool, and different users can directly call it without repeated calculation, reducing the computational amount. The comparison between the similarity and the threshold provides a data basis for the judgment of random and pseudo - random selection. At the same time, caching the display of the content of different users provides basic data for randomly mixing the display results of different users. The embodiments of the present disclosure provide calculating the similarity between the randomly accessed data and the first initial recommendation data and the second initial recommendation data to obtain a similarity value. The computing cache pool of this layer can be used to calculate the data similarity and cache it. According to different types of APPs, different similarity calculation methods can be adopted to achieve more accurate personalized recommendations.
[0120] Integrated Interface Management Sub - system: It has the function of a user interface interface. Different strategies are selected according to the APP type (such as information, shopping, transportation, travel, etc.) for integration and adaptation of different APPs. The Policy Manager, StrategyManager, manages the policy repository in a fine - grained manner. The integrated interface management sub - system, as the brain of the system, uniformly schedules and coordinates the functions of other sub - modules. At the same time, the plug - and - play interface, plugin, design is used to shield the support of different underlying APP architectures as much as possible. Set a policy to uninstall the APP regularly and reinstall it. If the usage time exceeds a certain threshold, it will not be reinstalled after uninstallation, which also reduces the burden on the mobile phone.
[0121] User Data Management Subsystem: Manages user data. Non-essential data is not provided to the APP. Data cannot be shared between APPs, and the behavior of collecting and obtaining user input to other APPs is prohibited. By controlling the data requests of the APP at different levels, data is only collected during use, and the requested data is also filtered. This is achieved through the Information Filter Manager FilterManager. Only some essential data can be displayed through the filter interface filter of FilterManager. To enable users to obtain different types of information content, the similarity is calculated through the compare method of the FilterManager's comparison interface.
[0122] User Management Subsystem: It mainly includes a user creation sub-module and a user management sub-module. The user creation sub-module can randomly create different virtual users, including non-registered users and registered users. The user management sub-module is responsible for registered users to fill in basic information according to the principle of minimal filling, and user characteristics are filled in randomly. Frequently modify personal information, such as address, age, name, to break the label setting for users. The created users are placed in the user pool for reuse. It is responsible for deleting abandoned users. The fictional virtual users provided in the embodiments of the present disclosure can be managed through this user management subsystem.
[0123] User Behavior Subsystem: It can be divided into a page element discovery unit, a user operation unit, a user behavior cleaning unit, and a time control unit.
[0124] Among them, the page element discovery unit: Non-registered users can operate on user behavior without logging in. Registered users first log in and then operate. Use crawler tools such as python to crawl the URLs of the app, loop through web page elements. There are three levels of polling, insert random numbers, and multiple users poll page elements simultaneously. Use methods like find_element_by_id or find_element_by_tag_name in the browser method to locate tags. Through specific splitting conditions: 1) Loop through the URL list 2) Loop through the text list 3) Loop through the fixed element list.
[0125] User operation unit: Use a crawler to simulate random clicks by users. For example, objects of Python or Selenium's webdriver are used to generate Web requests of requests and responses according to specific platform adaptations, and create user-sessions to simulate different users. User-sessions contain important fields such as user-id, user-password, user-agent, etc. Three necessary elements, namely user-agent, cookie, and token, are constructed according to different request requests. The token and cookie can be obtained through browser plugins, and the requests for the obtained token and cookie are reassembled. To prevent anti-crawling, on the one hand, the time interval of the crawler needs to be controlled. Generally, the 20-second mechanism is adopted in the industry and can be appropriately increased. At the same time, multiple tokens and cookies are created to form a token pool and a cookie pool respectively, and different simulated users can directly call them to achieve reuse. Different users can use libraries like fake_useragent to generate various user-agents by randomly calling random functions to simulate different browsers. The user operations on random pages can be randomly generated by random. Randomly click on different types of item information to disrupt the data collection and randomly search for various words. The implementation of crawler automation has certain logical requirements. Because there is a certain sequential logic in some user behaviors. For example, after clicking on a certain content, it is necessary to scroll to the bottom of the page at a certain time to perform like or favorite operations, and then return to the main page after completion. Therefore, individual user behaviors need to be assembled into user behavior units, and the user behavior units are used as the smallest logical units for calling. The design of the smallest logical unit can draw on the model of playing cards. The shuffle function shuffle(strategy) can be used to randomly call logical units. The strategy of the shuffle(strategy) function can be true randomness and pseudo-randomness. True randomness uses random numbers for items. Pseudo-randomness can use the item selector itemSelector to randomly select different types of items and then make a secondary random selection within the type to evenly distribute various items, which is used to deliberately disrupt user behaviors and avoid the situation of information cocoons. The random access dataset obtained by randomly accessing different types of information in the embodiments of the present disclosure can be obtained by generating a random access dataset through the execution of this user operation unit.
[0126] It should be noted that some accesses to websites with https can bypass the certificate verification process. For example, not performing certificate verification can be divided into two cases. One is purchasing a certificate from an authoritative CA. The other is that although the website performs certificate verification, it can still log in normally without using the https protocol (verify = false). When using a crawler proxy, to prevent the IP from being blocked by a certain website due to excessive crawling frequency of that website, authentication settings are used.
[0127] User behavior cleaning unit: Set certain policies to regularly clean the user's browsing history, like records, favorites, and shopping cart. Different apps can execute different policies to meet the personalized needs of users. By clearing the browser cache, simulating user behavior, cleaning the user's browsing history, reversing like records, cleaning the favorites, and cleaning the shopping cart, etc. The preprocessing of the real access data of real users provided by the embodiments of the present disclosure to obtain the preprocessed real access data can be implemented, for example, by executing this user behavior cleaning unit.
[0128] Time control unit: Set time intervals and timeout settings. Among them, on the one hand, the time interval is such that there is no need for real-time updates, which can reduce the consumption of data traffic. Using a certain amount of cached calculations can improve the calculation and display efficiency. On the other hand, it can avoid the IP being blocked due to overly frequent crawling. In addition, a proxy can be used as a method to prevent the IP from being blocked. For requests that exceed the duration of 20s, the timeout setting disconnects the connection to avoid reaching the upper limit of the number of connections.
[0129] Price comparison subsystem: Refer to existing price comparison systems or directly integrate existing price comparison systems. Horizontally, price comparison across different platforms can be achieved by searching for the same item or through image recognition (only as an auxiliary). Vertically, the prices of different users on the same platform can be compared. It is also possible to cross the above two methods and randomly select.
[0130] Price discrimination reporting subsystem: Obtain the situation of price discrimination through the price comparison system. Only for the situation where the prices are different for different users on the same platform, it interfaces with the regulatory department's API to report the situation. The reported data can be in the form of json or html, etc. The data includes the ID of this application, app name, web page link URL, store name, item name, price comparison, timestamp, and remarks, etc. It should be noted that the data format can be, for example:
[0131] AppID: 798123
[0132] AppName: MaiDuoDuo
[0133] Url: http: / / qgifhva.com.cn
[0134] ShopName: myshop
[0135] Item Name: oneItem
[0136] Price: 58,39,81
[0137] Time Stamp: 2022031207123610
[0138] Note: Upload the price picture.
[0139] According to an embodiment of the present disclosure, the information recommendation system is universal. It can integrate multiple APPs of multiple types, has a wide range of applications, uses a pluggable Plugin mode, shields the underlying details of different APPs, can adjust the included APPs at any time, and is transparent to the upper layer. Among them, the types of APPs can include information and news, video, music, shopping, catering, tourism, transportation, etc.
[0140] It should be noted that the strategies for breaking the information cocoon of different types of apps have slightly different emphases. For example, if it is for shopping information, in the calculation scheme of the shopping APP, pictures are the key calculation objects. The automatic detection of text regions and text automatic recognition can be achieved through the existing OCR technology using paddleocr, and the item pictures on the page can be directly crawled and the names can be extracted for weighted ratio hybrid calculation of similarity. Existing technologies such as Euclidean distance, cosine distance, Hamming distance, and structural similarity index SSIM (extracting brightness, contrast, and structure) can be used to calculate the picture similarity. Since the requirement for picture similarity is not too high, the Hamming distance is recommended because of its small calculation amount and fast speed.
[0141] Based on the above information recommendation method, the present disclosure also provides an information recommendation device. The following will be combined with Figure 5 to describe this device in detail.
[0142] Figure 5 Schematically shows a structural block diagram of an information recommendation device according to an embodiment of the present disclosure.
[0143] As Figure 5 shown, the information recommendation device 500 of this embodiment includes a first acquisition module 510, a second acquisition module 520, a determination module 530, and a recommendation module 540.
[0144] The first acquisition module 510 is used to acquire a random access data set generated by randomly accessing different types of information. Among them, the random access data set includes at least one random access data. In one embodiment, the first acquisition module 510 can be used to perform the operation S201 described above, which will not be elaborated here.
[0145] The second acquisition module 520 is configured to acquire first initial recommendation data generated for a real user and second initial recommendation data generated for N virtual users, where the N virtual users are fictionalized based on the attribute characteristics of the real user, and N is an integer greater than or equal to 1. In one embodiment, the second acquisition module 520 may be configured to perform the operation S202 described above, which will not be elaborated here.
[0146] The determination module 530 is configured to determine target recommendation data according to the random access data set and the initial recommendation data set, where the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data. In one embodiment, the determination module 530 may be configured to perform the operation S203 described above, which will not be elaborated here.
[0147] The recommendation module 540 is configured to perform information recommendation for the real user according to the target recommendation data. In one embodiment, the recommendation module 540 may be configured to perform the operation S204 described above, which will not be elaborated here.
[0148] According to an embodiment of the present disclosure, the information recommendation device 500 may further include a preprocessing module, a first generation module, and a second generation module.
[0149] The preprocessing module is configured to preprocess the real access data of the real user to obtain the preprocessed real access data.
[0150] The first generation module is configured to generate first initial recommendation data according to the real access data.
[0151] The second generation module is configured to generate second initial recommendation data according to the attribute characteristics of the N virtual users.
[0152] According to an embodiment of the present disclosure, any of the first acquisition module 510, the second acquisition module 520, the determination module 530, and the recommendation module 540 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first acquisition module 510, the second acquisition module 520, the determination module 530, and the recommendation module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware by integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first acquisition module 510, the second acquisition module 520, the determination module 530, and the recommendation module 540 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.
[0153] Figure 6 A block diagram of an electronic device suitable for implementing an information recommendation method according to an embodiment of the present disclosure is schematically shown.
[0154] As Figure 6 shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0155] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0156] According to an embodiment of the present disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 610 as needed so that a computer program read therefrom is installed into the storage portion 608 as needed.
[0157] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0158] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0159] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present disclosure.
[0160] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0161] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code contained in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0162] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0163] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, programming languages such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0165] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0166] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. An information recommendation method, comprising: Obtaining a random access data set generated by randomly accessing different types of information, wherein the random access data set includes at least one random access data; Obtaining first initial recommendation data generated for a real user and second initial recommendation data generated for N virtual users, wherein the N virtual users are fictionalized based on the attribute characteristics of the real user, and N is an integer greater than or equal to 1; Determining target recommendation data according to the random access data set and the initial recommendation data set, wherein the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data; and Performing information recommendation for the real user according to the target recommendation data.
2. The method according to claim 1, wherein, The determining of the target recommendation data according to the random access data set and the initial recommendation data set includes: Calculating the similarity between the random access data and the first initial recommendation data and the second initial recommendation data to obtain a similarity value; Determining an initial data group when it is determined that the similarity value exceeds a first threshold; Dividing the initial data group into M target data groups according to a second threshold, wherein M is an integer greater than or equal to 1; and Screening from each of the M target data groups to obtain the target recommendation data.
3. The method according to claim 2, further comprising: Determining hot information data according to the random access data set and the initial recommendation data set; Using the initial recommendation data set to determine user access tags for the hot information data; Determining the first threshold according to the hot information data, the user access tags, and the similarity value; and Determining the second threshold according to the first threshold.
4. The method according to claim 2, wherein, The calculating of the similarity between the random access data and the first initial recommendation data and the second initial recommendation data to obtain a similarity value includes: Obtaining algorithm keywords from the random access data, the first initial recommendation data, and the second initial recommendation data; and Calculating the similarity according to the algorithm keywords to obtain the similarity value.
5. The method according to claim 1, wherein, The attribute characteristics of the real user include the user behavior characteristics of the real user; The N virtual users are fictionalized based on the attribute characteristics of the real user, including: Obtaining the user behavior characteristics of the real user; and Fictionalizing N virtual users according to the user behavior characteristics.
6. The method according to claim 1, wherein, The attribute characteristics of the real user further include the user attribute characteristics of the real user; The N virtual users are fictionalized based on the attribute characteristics of the real user, further including: Fictionalizing the virtual users by changing the user attribute characteristics of the real user.
7. The method according to claim 1, before obtaining the initial recommendation data of the N virtual users and the real user, further comprising: Preprocess the real access data of the real user to obtain the preprocessed real access data; Generate the first initial recommendation data according to the real access data; And Generate the second initial recommendation data according to the attribute characteristics of N virtual users.
8. An information recommendation device, Comprising: A first acquisition module, configured to acquire a random access data set generated by randomly accessing different types of information, wherein the random access data set includes at least one random access data; A second acquisition module, configured to acquire the first initial recommendation data generated for a real user and the second initial recommendation data generated for N virtual users, wherein the N virtual users are fictitiously obtained according to the attribute characteristics of the real user, and N is an integer greater than or equal to 1; A determination module, configured to determine target recommendation data according to the random access data set and the initial recommendation data set, wherein the initial recommendation data set includes the first initial recommendation data and the second initial recommendation data; and A recommendation module, configured to perform information recommendation for the real user according to the target recommendation data.
9. An electronic device, Comprising: One or more processors; A storage device, configured to store one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Information recommendation method and device, recommendation server and storage medium
CN112100221A
Information data processing method and device, equipment and storage medium
CN112231571A