Mobile Internet Big Data User Behavior Analysis System Based on Cloud Computing

Through the user behavior analysis system based on cloud computing, user reading records are used for secondary screening and sorting, which solves the problem of users needing secondary screening in the existing technology and improves the retrieval fluency and experience.

CN116521988BActive Publication Date: 2025-09-16HUANJU SHIDAI MEDIA BEIJING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310391349.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-09-16
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

Existing technologies do not take users' behavioral habits into consideration during retrieval, which results in users having to perform a large amount of secondary screening, affecting the smoothness of retrieval.

Method used

Through the cloud computing-based mobile Internet big data user behavior analysis system, the user's reading records are used to calculate the sensitivity coefficient and composite fit coefficient, and the preliminary information is screened and sorted for the second time to recommend information that meets user needs.

Benefits of technology

It improves the fluency of retrieval, reduces the time and difficulty of self-screening for users, and enhances the retrieval experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0004175878110000011
    Figure FDA0004175878110000011
Patent Text Reader

Abstract

The present invention discloses a mobile Internet big data user behavior analysis system based on cloud computing, which belongs to the field of Internet technology. The system performs secondary screening on each primary selected material, and the secondary screening is performed based on the corresponding user's reading records in the past period of time. Therefore, it is possible to provide users with a corresponding appropriate material recommendation sequence as much as possible, so that users can quickly obtain ideal search results, reduce the user's self-screening time and screening difficulty, and improve the search experience. Specifically, the present invention redistributes the weights of the praise rate, click-through rate, and duration of each primary selected material through the corresponding sensitivity coefficient, reduces the influence of the parameters such as the praise rate, click-through rate, and material duration of each corresponding material on the recommendation results, and is conducive to recommending the most appropriate video material to users, improving the user's search experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet technology, and specifically relates to a mobile Internet big data user behavior analysis system based on cloud computing. Background Art

[0002] The rapid development of Internet technology has brought convenience to people's lives. In the process of Internet application and development, a large amount of information is stored and retained. When users need to obtain the information they want, it will be very difficult. The emergence of the search function can obtain some of the information that users may want based on keywords, and then users can filter according to the search results.

[0003] However, with the increasing amount of information on the Internet, when entering keywords for searching, there may be a problem of too many search results. Although many systems now sort the search results according to certain preset rules, so that users can give priority to obtaining high-quality objects under the corresponding preset rules, this screening method does not take into account the user's own habits when screening data. As a result, in actual operation, the user still needs to perform more secondary screening according to his own needs, which is not conducive to the smoothness of the user's search work. In order to solve the above problem, a method is provided that can analyze the user's behavior when watching short video materials, and recommend more suitable target objects to the user according to the user's behavioral habits when the user performs the search work. The present invention provides the following technical solutions. Summary of the Invention

[0004] The purpose of the present invention is to provide a mobile Internet big data user behavior analysis system based on cloud computing to solve the problem that the existing technology does not take into account the user's own habits in data screening when performing retrieval, resulting in that in actual operation, the user still needs to perform more secondary screening work according to his own needs, which is not conducive to the smoothness of the user's retrieval work.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] The cloud computing-based mobile Internet big data user behavior analysis system includes:

[0007] A retrieval unit, which obtains preliminary information from the data storage unit by searching keywords;

[0008] A data storage unit for storing information and each user's reading records;

[0009] User login unit, through which users log in to the system;

[0010] The control center is used to sort the preliminary materials according to the user's reading history and preliminary materials, and give priority to recommending the preliminary materials that meet the user's needs;

[0011] The working method of the control center includes the following steps:

[0012] The steps include:

[0013] S1. Mark a user as a target user and obtain the target user's browsing history within the past preset time T1;

[0014] The reading record includes the praise rate, click rate, length of the data and the field to which the target user belongs when reading the data;

[0015] Obtain the data that the target user has completed reading within the same field within the past T1 period, and mark these completed reading data as historical reference data;

[0016] Obtain reading records of historical comparison materials;

[0017] Calculate the target user's sensitivity coefficient G1 to the praise rate, G2 to the click rate, and G3 to the length of the material within the corresponding field.

[0018] The calculation method of the sensitivity coefficient G1 is:

[0019] Obtain the favorable comment rate hi of each historical reference material, where 1≤i≤n, and n is the number of historical reference materials;

[0020] According to the formula Calculate the dispersion value F of the set of parameters from h1 to hn;

[0021] Where hp=(h1+h2+…+hn) / n;

[0022] The target user's sensitivity coefficient G1 to the favorable review rate is calculated using the formula G1 = α3 / (α1*F+α2*hp), where α1, α2, and α3 are all preset values, and α1+α2=1;

[0023] The sensitivity coefficient G2 is calculated based on the click rate di of each historical control data;

[0024] The sensitivity coefficient G3 is calculated based on the data length ti of each historical control data;

[0025] The calculation methods for G2 and G3 are the same as G1;

[0026] S2. The target user inputs a search keyword through the search unit. The search unit obtains corresponding information from the data storage unit based on the search keyword and marks the corresponding information as preliminary selected information.

[0027] S3, obtaining the keyword fit R1 of each preliminary selection material;

[0028] Obtain the field added value β corresponding to each preliminary selection data;

[0029] Obtain the topic strength R2 of each primary selection data within the past preset time T1;

[0030] Get the praise rate hk and click rate dk corresponding to each preliminary selection information at the current moment;

[0031] Get the duration tk corresponding to each primary selection data;

[0032] According to the formula:

[0033] U=γ1*R1+γ2*R2+γ3*hk G1 / α4 +γ4*dk G2 / α4 +γ5*|tk-tp| G3 / α4 +β calculation

[0034] Obtain the composite fitting coefficient U of each primary selected material;

[0035] Wherein α4 is a preset value, and when G1 / α4 < σ, G1 / α4 takes the value σ, where σ is a parameter greater than 0 and less than 1; when G1 / α4 > 1, G1 / α4 takes the value 1;

[0036] Wherein, γ1, γ2, γ3, γ4 and |γ5| are all preset parameters, and γ1, γ2, γ3 and γ4 are all positive values. When tk-tp is greater than or equal to 0, γ5 takes a negative value, and when tk-tp is less than 0, γ5 takes a positive value.

[0037] tp=(t1+t2+…+tn) / n;

[0038] S4. Recommend the preliminary selected materials to the target users in descending order of the composite fitting coefficient U.

[0039] As a further solution of the present invention, the completed viewing material refers to the material for which the ratio of the duration of the actual playback portion of the corresponding material by the corresponding user to the total duration of the corresponding material is greater than a preset ratio value θ.

[0040] As a further solution of the present invention, the value of θ is 70%.

[0041] As a further solution of the present invention, the value of σ is 0.25.

[0042] As a further solution of the present invention, the calculation method of the keyword fit R1 of each preliminary selected data is:

[0043] For a preliminary selection document, obtain the number q1 of search keywords that appear in the preliminary selection document;

[0044] Get the total number of search keywords q;

[0045] The keyword fit R1 of the corresponding preliminary selection information is calculated according to the formula R1=q1 / q.

[0046] As a further solution of the present invention, the calculation method of the field added value β is:

[0047] Obtain all the data that the target user has completed reading in the past T1 time, and mark these completed reading data as classified sub-data;

[0048] Obtain the fields corresponding to each category sub-data, and group each category sub-data according to the fields they belong to;

[0049] Get the number of classification sub-data corresponding to each field e1;

[0050] Calculate and obtain the domain added value β corresponding to each domain for the target user, β=β1*e1 / e, where β1 is the preset value and e is the number of classification sub-data.

[0051] As a further solution of the present invention, the topic strength R2 of each preliminary selected data is calculated as follows:

[0052] Obtain the number of times r a primary selection document has been cited or forwarded within a preset time period T1, and mark the corresponding primary selection document after citation or forwarding as secondary information;

[0053] The first-level impact value of the corresponding primary selection data is calculated according to the formula r*μ1;

[0054] Obtain the number of times r2j each secondary document has been cited or forwarded within the past preset time T1, and mark the corresponding secondary document after citation or forwarding as a third-level document, where 1≤j≤m, and m is the number of secondary documents;

[0055] The secondary impact value of the corresponding primary data is calculated according to the formula μ2*(r21+r22+,…,+r2m);

[0056] Obtain the number of times r3j that each third-level document has been cited or forwarded within the past preset time T1, and then calculate the third-level influence value μ3*(r31+r32+,…,+r3m) of the corresponding primary selected document;

[0057] Calculate the subsequent impact values ​​of the corresponding primary selected information in sequence until the information is no longer cited or forwarded;

[0058] Then the sum of the influence values ​​of each level of the corresponding preliminary selection data is taken as the topic strength R2 of the corresponding preliminary selection data in the past preset time T1;

[0059] Among them, μ1>μ2>μ3>….

[0060] Beneficial effects of the present invention:

[0061] 1. The present invention performs a secondary screening on each primary selected document, and the secondary screening is performed based on the corresponding user's reading records in the past period of time. Therefore, it can provide users with a corresponding appropriate document recommendation order as much as possible, enabling users to quickly obtain ideal search results, reducing users' self-screening time and screening difficulty, and improving the search experience.

[0062] 2. The present invention determines the degree of fit between each preliminary selected material and the target user by calculating the composite fit coefficient U, wherein the weights of the praise rate, click-through rate and duration of each preliminary selected material are redistributed and calculated by the ratio of the sensitivity coefficient to the preset value α4, thereby reducing the influence of parameters such as the praise rate, click-through rate and material duration of each corresponding material on the recommendation results, which is conducive to recommending the most suitable video material to the user and improving the user's search experience. DETAILED DESCRIPTION

[0063] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0064] The cloud computing-based mobile Internet big data user behavior analysis system includes:

[0065] A retrieval unit, configured to input a retrieval keyword through the retrieval unit, and the retrieval unit obtains preliminary selected data from the data storage unit through the retrieval keyword;

[0066] A data storage unit for storing information and each user's reading records;

[0067] A user login unit, through which a user enters his or her account number and logs into the system after identity verification;

[0068] The control center is used to sort the preliminary selected materials according to the user's reading history and preliminary selected materials, and recommend the preliminary selected materials that meet the user's needs to the corresponding user first;

[0069] The method for the control center to sort the preliminary selected materials and recommend the preliminary selected materials that meet the user's needs to the corresponding user includes the following steps:

[0070] The steps include:

[0071] S1. Mark a user as a target user and obtain the target user's browsing history within the past preset time T1;

[0072] The reading record includes the praise rate, click rate, length of the data and the field to which the target user belongs when reading the data;

[0073] The fields mentioned include food, dance, animation, animals, comics, sports, cars, movies, music, etc.;

[0074] The material described is video material;

[0075] Obtain the data that the target user has completed reading within the same field within the past T1 period, and mark these completed reading data as historical reference data;

[0076] Obtain reading records of historical comparison materials;

[0077] The completed viewing material refers to the material for which the ratio of the actual playback duration of the corresponding user to the total playback duration of the corresponding material is greater than a preset ratio value θ;

[0078] In one embodiment of the present invention, the value of θ is 70%;

[0079] Calculate the target user's sensitivity coefficient G1 to the praise rate, G2 to the click rate, and G3 to the length of the material within the corresponding field.

[0080] The calculation method of the sensitivity coefficient G1 is:

[0081] Obtain the favorable comment rate hi of each historical reference material, where 1≤i≤n, and n is the number of historical reference materials;

[0082] According to the formula Calculate the scattered values ​​of the parameters from h1 to hn;

[0083] Where hp=(h1+h2+…+hn) / n;

[0084] The target user's sensitivity coefficient G1 to the favorable review rate is calculated using the formula G1 = α3 / (α1*F+α2*hp), where α1, α2, and α3 are all preset values, and α1+α2=1;

[0085] The calculation method of the sensitivity coefficient G2 is:

[0086] Obtain the click rate di of each historical reference material, where 1≤i≤n, and n is the number of historical reference materials;

[0087] According to the formula Calculate the scattered values ​​of the set of parameters d1 to dn;

[0088] Where dp = (d1 + d2 + ... + dn) / n;

[0089] The target user's sensitivity coefficient G2 to click rate is calculated using the formula G2 = α3 / (α1*F1+α2*dp);

[0090] The calculation method of the sensitivity coefficient G3 is:

[0091] The data duration ti of each historical control data is obtained, where 1≤i≤n, and n is the number of historical control data;

[0092] According to the formula Calculate the scattered values ​​of the parameters from t1 to tn;

[0093] Where tp=(t1+t2+…+tn) / n;

[0094] The target user's sensitivity coefficient G3 to the data duration is calculated using the formula G3 = α3 / (α1*F2+α2*tp);

[0095] The larger G1, G2, and G3 are, the more sensitive the target user is to the positive review rate, duration, and click-through rate of the material. This step can intuitively express the target user's sensitivity to the positive review rate, duration, and popularity of the video material through the sensitivity coefficient.

[0096] S2. The target user inputs a search keyword through the search unit. The search unit obtains corresponding information from the data storage unit based on the search keyword and marks the corresponding information as preliminary selected information.

[0097] S3, obtaining the keyword fit R1 of each preliminary selection material;

[0098] Obtain the field added value β corresponding to each preliminary selection data;

[0099] Obtain the topic strength R2 of each primary selection data within the past preset time T1;

[0100] Get the praise rate hk and click rate dk corresponding to each preliminary selection information at the current moment;

[0101] Get the duration tk corresponding to each primary selection data;

[0102] According to the formula:

[0103] U=γ1*R1+γ2*R2+γ3*hk G1 / α4 +γ4*dk G2 / α4 +γ5*|tk-tp| G3 / α4 +β calculation

[0104] Obtain the composite fitting coefficient U of each primary selected material;

[0105] Wherein α4 is a preset value, and when G1 / α4 < σ, G1 / α4 takes the value σ, where σ is a parameter greater than 0 and less than 1; when G1 / α4 > 1, G1 / α4 takes the value 1;

[0106] In one embodiment of the present invention, the value of σ is 0.25;

[0107] Wherein, γ1, γ2, γ3, γ4 and |γ5| are all preset parameters, and γ1, γ2, γ3 and γ4 are all positive values. When tk-tp is greater than or equal to 0, γ5 takes a negative value, and when tk-tp is less than 0, γ5 takes a positive value.

[0108] In one embodiment of the present invention, the keyword compatibility R1 of each preliminary selected data is calculated as follows:

[0109] For a preliminary selection document, obtain the number q1 of search keywords that appear in the preliminary selection document;

[0110] Get the total number of search keywords q;

[0111] The keyword compatibility R1 of the corresponding preliminary selection data is calculated according to the formula R1=q1 / q;

[0112] The calculation method of the field added value β is:

[0113] Obtain all the data that the target user has completed reading in the past T1 time, and mark these completed reading data as classified sub-data;

[0114] Obtain the fields corresponding to each category sub-data, and group each category sub-data according to the fields they belong to;

[0115] Get the number of classification sub-data corresponding to each field e1;

[0116] Calculate and obtain the domain added value β corresponding to each domain for the target user, β=β1*e1 / e, where β1 is the preset value and e is the number of classification sub-data.

[0117] The calculation method of the topic strength R2 of each preliminary selection data is:

[0118] Obtain the number of times r a primary selection document has been cited or forwarded within a preset time period T1, and mark the corresponding primary selection document after citation or forwarding as secondary information;

[0119] The first-level impact value of the corresponding primary selection data is calculated according to the formula r*μ1;

[0120] Obtain the number of times r2j each secondary document has been cited or forwarded within the past preset time T1, and mark the corresponding secondary document after citation or forwarding as a third-level document, where 1≤j≤m, and m is the number of secondary documents;

[0121] The secondary impact value of the corresponding primary data is calculated according to the formula μ2*(r21+r22+,…,+r2m);

[0122] Obtain the number of times r3j that each third-level document has been cited or forwarded within the past preset time T1, and then calculate the third-level influence value μ3*(r31+r32+,…,+r3m) of the corresponding primary selected document;

[0123] Calculate the subsequent impact values ​​of the corresponding primary selected information in sequence until the information is no longer cited or forwarded;

[0124] Then the sum of the influence values ​​of each level of the corresponding preliminary selection data is taken as the topic strength R2 of the corresponding preliminary selection data in the past preset time T1;

[0125] where μ1>μ2>μ3>…;

[0126] S4. Recommend the preliminary selected materials to the target users in descending order of the composite fitting coefficient U.

[0127] The present invention performs a secondary screening of each initially selected document based on the corresponding user's browsing history over the past period of time. This allows the user to be provided with a corresponding and appropriate order of document recommendations as much as possible, enabling the user to quickly obtain ideal search results, reducing the user's self-screening time and difficulty, and improving the search experience.

[0128] The present invention determines the degree of fit between each preliminary selected material and the target user by calculating the composite fit coefficient U, wherein the weights of the praise rate, click-through rate and duration of each preliminary selected material are redistributed and calculated by the ratio of the sensitivity coefficient to the preset value α4, thereby reducing the influence of parameters such as the praise rate, click-through rate and material duration of each corresponding material on the recommendation results, which is conducive to recommending the most suitable video material to the user and improving the user's search experience.

[0129] Throughout the specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0130] The above contents are merely examples and explanations of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in similar ways. As long as they do not deviate from the invention or exceed the scope defined by the claims, they should all fall within the scope of protection of the present invention.

Claims

1. A cloud computing-based mobile Internet big data user behavior analysis system, characterized by: include: A retrieval unit, which obtains preliminary information from the data storage unit by searching keywords; A data storage unit for storing information and each user's reading records; User login unit, through which users log in to the system; The control center is used to sort the preliminary materials according to the user's reading history and preliminary materials, and give priority to recommending the preliminary materials that meet the user's needs; The working method of the control center includes the following steps: The steps include: S1. Mark a user as a target user and obtain the target user's browsing history within the past preset time T1; The reading record includes the praise rate, click rate, length of the data and the field to which the target user belongs when reading the data; Obtain the data that the target user has completed reading within the same field within the past T1 period, and mark these completed reading data as historical reference data; Obtain reading records of historical comparison materials; Calculate the target user's sensitivity coefficient G1 to the praise rate, G2 to the click rate, and G3 to the length of the material within the corresponding field. The calculation method of the sensitivity coefficient G1 is: Obtain the favorable comment rate hi of each historical reference material, where 1≤i≤n, and n is the number of historical reference materials; According to the formula Calculate the dispersion value F of the set of parameters from h1 to hn; Where hp=(h1+h2+…+hn) / n; The target user's sensitivity coefficient G1 to the favorable review rate is calculated using the formula G1 = α3 / (α1*F+α2*hp), where α1, α2, and α3 are all preset values, and α1+α2=1; The sensitivity coefficient G2 is calculated based on the click rate di of each historical control data; The sensitivity coefficient G3 is calculated based on the data length ti of each historical control data; The calculation methods for G2 and G3 are the same as G1; S2. The target user inputs a search keyword through the search unit. The search unit obtains corresponding information from the data storage unit based on the search keyword and marks the corresponding information as preliminary selected information. S3, obtaining the keyword fit R1 of each preliminary selection material; Obtain the field added value β corresponding to each preliminary selection data; Obtain the topic strength R2 of each primary selection data within the past preset time T1; Get the praise rate hk and click rate dk corresponding to each preliminary selection information at the current moment; Get the duration tk corresponding to each primary selection data; According to the formula: U=γ1*R1+γ2*R2+γ3*hk G1 / α4 +γ4*dk G2 / α4 +γ5*|tk-tp| G3 / α4 +β calculation to obtain the composite fitting coefficient U of each primary selected material; Wherein α4 is a preset value, and when G1 / α4 < σ, G1 / α4 takes the value σ, where σ is a parameter greater than 0 and less than 1; when G1 / α4 > 1, G1 / α4 takes the value 1; Wherein, γ1, γ2, γ3, γ4 and |γ5| are all preset parameters, and γ1, γ2, γ3 and γ4 are all positive values. When tk-tp is greater than or equal to 0, γ5 takes a negative value, and when tk-tp is less than 0, γ5 takes a positive value. tp=(t1+t2+…+tn) / n; S4. Recommend the preliminary selected materials to the target users in descending order of the composite fitting coefficient U.

2. The cloud computing-based mobile Internet big data user behavior analysis system according to claim 1, characterized in that: The completed viewing material refers to a material in which the ratio of the duration of the actual playback portion of the corresponding material by the corresponding user to the total duration of the corresponding material is greater than a preset ratio value θ.

3. The cloud computing-based mobile Internet big data user behavior analysis system according to claim 2, characterized in that: The value of θ is 70%.

4. The cloud computing-based mobile Internet big data user behavior analysis system according to claim 1, characterized in that: The value of σ is 0.

25.

5. The cloud computing-based mobile Internet big data user behavior analysis system according to claim 1, characterized in that: The calculation method of the keyword fit R1 of each preliminary selection data is as follows: For a preliminary selection document, obtain the number q1 of search keywords that appear in the preliminary selection document; Get the total number of search keywords q; The keyword fit R1 of the corresponding preliminary selection information is calculated according to the formula R1=q1 / q.

6. The cloud computing-based mobile Internet big data user behavior analysis system according to claim 1, characterized in that: The calculation method of the field added value β is: Obtain all the data that the target user has completed reading in the past T1 time, and mark these completed reading data as classified sub-data; Obtain the fields corresponding to each category sub-data, and group each category sub-data according to the fields they belong to; Get the number of classification sub-data corresponding to each field e1; Calculate and obtain the domain added value β corresponding to each domain for the target user, β=β1*e1 / e, where β1 is the preset value and e is the number of classification sub-data.

7. The mobile Internet big data user behavior analysis system based on cloud computing according to claim 1 is characterized in that: The calculation method of the topic strength R2 of each preliminary selection data is: Obtain the number of times r a primary selection document has been cited or forwarded within a preset time period T1, and mark the corresponding primary selection document after citation or forwarding as secondary information; The first-level impact value of the corresponding primary selection data is calculated according to the formula r*μ1; Obtain the number of times r2j each secondary document has been cited or forwarded within the past preset time T1, and mark the corresponding secondary document after citation or forwarding as a third-level document, where 1≤j≤m, and m is the number of secondary documents; The secondary impact value of the corresponding primary data is calculated according to the formula μ2*(r21+r22+,…,+r2m); Obtain the number of times r3j that each third-level document has been cited or forwarded within the past preset time T1, and then calculate the third-level influence value μ3*(r31+r32+,…,+r3m) of the corresponding primary selected document; Calculate the subsequent impact values ​​of the corresponding primary selected information in sequence until the information is no longer cited or forwarded; Then the sum of the influence values ​​of each level of the corresponding preliminary selection data is taken as the topic strength R2 of the corresponding preliminary selection data in the past preset time T1; Among them, μ1>μ2>μ3>….

Citation Information

Patent Citations

  • Mobile internet browsing content intelligent recommendation method, system and device based on feature recognition and storage medium

    CN113158048A

  • Information recommendation method and apparatus, and electronic device and storage medium

    WO2022142519A1