A method and system for analyzing the popularity of big data information
By obtaining the propagation index of the propagation file and comparing the reference file library in real time, the prediction popularity of uploaded files is determined, and the problem of unbalanced resource allocation in information software is solved, and the rational promotion of original high-traffic content is achieved.
Patent Information
- Application Number
- CN202111511676.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-06
AI Technical Summary
When existing information software allocates promotion resources, it is easy to miss original high-traffic content, resulting in uneven resource allocation.
By obtaining the propagation index of the propagation file in real time, comparing it with the preset index threshold, copying it to the reference file library, determining the predicted popularity of uploaded files based on the reference file library, and correcting the promotion resources based on the actual popularity, including push range and frequency.
It has achieved more accurate allocation of promotion resources, ensured that original high-traffic content is reasonably promoted, and improved resource utilization efficiency.
Smart Images

Figure CN114417128B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and specifically to a method and system for analyzing the popularity of big data information. Background Art
[0002] In today's society, online media has gradually become the main way for most people to obtain information. Therefore, many information-based software has emerged. Most of these information-based software will have a file upload function and will also push some popularity data in real time to obtain traffic. However, the file push process requires the consumption of promotion resources. For example, there is only one C position. When receiving a new file, how to judge its promotion resources requires the help of popularity analysis. In the existing technical solutions, most of them determine the promotion resources according to different accounts. The accounts with more fans have more promotion resources, and the accounts with fewer fans have fewer promotion resources. It can be imagined that the result of doing so is that it is easy to miss some original high-traffic content. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for analyzing the popularity of big data information to solve the problems raised in the above background art.
[0004] To achieve the above purpose, the present invention provides the following technical solutions:
[0005] A method for analyzing the popularity of big data information, the method includes:
[0006] Obtain the propagation index of the propagated file in real time, compare the propagation index with at least one preset index threshold, and when the propagation index reaches the index threshold, copy the propagated file to the reference file library; wherein, each reference file library corresponds to an index threshold;
[0007] Receive the uploaded file sent by the user, read the corresponding reference file library according to the descending order of the index threshold, and determine the predicted popularity of the uploaded file according to the reference file library;
[0008] Allocate promotion resources according to the predicted popularity, and monitor the actual popularity of the file in real time, and correct the promotion resources according to the actual popularity and the predicted popularity;
[0009] Wherein, the promotion resources include the push range and the push frequency.
[0010] As a further solution of the present invention: the step of obtaining the propagation index of the propagated file in real time includes:
[0011] Obtain the operation record of the propagated file, and convert the operation record into an equivalent view count according to a preset conversion formula; the operation record at least includes like operations, favorite operations, download operations and share operations containing quantities;
[0012] Obtain the actual number of views of the propagated file, and calculate the propagation index according to the actual number of views and the equivalent number of views;
[0013] Among them, the propagation index is a decreasing function of time.
[0014] As a further solution of the present invention: the steps of receiving the uploaded file sent by the user, reading the corresponding reference file library in descending order of the index threshold, and determining the predicted popularity of the uploaded file according to the reference file library include:
[0015] Receive the uploaded file sent by the user, and read the corresponding reference file library in descending order of the index threshold;
[0016] Traverse the reference file library, compare the uploaded file with the reference files in the reference file library, and obtain the similarity;
[0017] Generate a similarity array according to the similarity, perform feature analysis on the similarity array, and query the corresponding index threshold according to the feature analysis result;
[0018] Determine the predicted popularity according to the index threshold.
[0019] As a further solution of the present invention: the steps of traversing the reference file library, comparing the uploaded file with the reference files in the reference file library, and obtaining the similarity include:
[0020] Read the reference files in the reference file library in sequence, and convert the reference files into reference images and reference texts;
[0021] Convert the uploaded file into a to-be-inspected image and to-be-inspected text, compare the to-be-inspected text with the reference text, and obtain the text similarity;
[0022] When the text similarity reaches the preset similarity threshold, use the text similarity as the similarity of the reference file;
[0023] When the text similarity does not reach the preset similarity threshold, traverse the reference images in units of the to-be-inspected image to obtain the image similarity, and calculate the similarity of the reference file according to the image similarity and the text similarity.
[0024] As a further solution of the present invention: the steps of converting the reference file or the uploaded file into a reference image and a reference text include:
[0025] Obtain the file suffix name, and determine the file type according to the file suffix name;
[0026] When the file is an audio file, performing content recognition on the audio file to obtain a reference text;
[0027] When the file is a video file, converting the video file into an audio file and an image group;
[0028] Duplicate images in the image group are eliminated, and the remaining images are connected to obtain a reference image that is in a mapping relationship with the video file.
[0029] As a further solution of the present invention: the step of removing duplicate images in the image group includes:
[0030] Traversing the image group, converting the images in the image group into an image array according to a preset conversion formula, and calculating the mean of the image array;
[0031] Calculating the deviation rate of the mean of adjacent image arrays, and when the deviation rate of the mean is less than a preset deviation threshold, calculating the variance of the adjacent image arrays;
[0032] The offset rate of the variance of adjacent image arrays is calculated, and when the offset rate of the variance is less than a preset offset threshold, an image corresponding to one of the image arrays is deleted from the image group.
[0033] As a further solution of the present invention, the steps of generating a similarity array according to the similarity, performing feature analysis on the similarity array, and querying a corresponding index threshold according to the feature analysis result include:
[0034] Generate a feature map according to the elements and array subscripts in the similarity array;
[0035] Generate a separation line according to a preset similarity threshold, intercept the feature graph according to the separation line to obtain a feature segment, and query a corresponding index threshold according to the feature segment;
[0036] When the feature segment intercepted according to the dividing line is empty, converting the feature graph into a fluctuation function, inputting the fluctuation function into a trained fluctuation analysis model to obtain a matching degree;
[0037] When the matching degree reaches the preset matching threshold, the corresponding index threshold is queried;
[0038] The index threshold is the index threshold corresponding to the reference file library corresponding to the similarity array.
[0039] The technical solution of the present invention also provides a big data information heat analysis system, the system comprising:
[0040] A reference file determination module, configured to obtain the propagation index of a propagated file in real time, compare the propagation index with at least one preset index threshold, and when the propagation index reaches the index threshold, copy the propagated file to a reference file library; wherein, each reference file library corresponds to an index threshold;
[0041] A popularity prediction module, configured to receive an uploaded file sent by a user, read the corresponding reference file library according to the descending order of the index threshold, and determine the predicted popularity of the uploaded file according to the reference file library;
[0042] A promotion module, configured to allocate promotion resources according to the predicted popularity, and monitor the actual popularity of the file in real time, and correct the promotion resources according to the actual popularity and the predicted popularity;
[0043] Wherein, the promotion resources include the push scope and the push frequency.
[0044] As a further solution of the present invention: the reference file determination module includes:
[0045] An equivalent calculation unit, configured to obtain the operation record of the propagated file, and convert the operation record into an equivalent view count according to a preset conversion formula; the operation record at least includes like operations, favorite operations, download operations, and share operations containing quantities;
[0046] An index calculation unit, configured to obtain the actual view count of the propagated file, and calculate the propagation index according to the actual view count and the equivalent view count;
[0047] Wherein, the propagation index is a decreasing function of time.
[0048] As a further solution of the present invention: the prediction module includes:
[0049] A reference file reading unit, configured to receive an uploaded file sent by a user, and read the corresponding reference file library according to the descending order of the index threshold;
[0050] A similarity calculation unit, configured to traverse the reference file library, compare the uploaded file with the reference files in the reference file library, and obtain the similarity;
[0051] A feature analysis unit, configured to generate a similarity array according to the similarity, perform feature analysis on the similarity array, and query the corresponding index threshold according to the feature analysis result;
[0052] A processing execution unit, configured to determine the predicted popularity according to the index threshold.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows: when the present invention receives a file, it compares the file with the existing hotspot data, and then determines the promotion resources of the file, changing from "recognizing" the account to "recognizing" the content, which can better allocate the promotion resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0055] Figure 1 The flowchart of the big data information heat analysis method is shown.
[0056] Figure 2 The first sub-flowchart of the big data information heat analysis method is shown.
[0057] Figure 3 The second sub-flowchart of the big data information heat analysis method is shown.
[0058] Figure 4 The third sub-flowchart of the big data information heat analysis method is shown.
[0059] Figure 5 The fourth sub-flowchart of the big data information heat analysis method is shown.
[0060] Figure 6 The block diagram of the composition structure of the big data information heat analysis system is shown.
[0061] Figure 7 The block diagram of the composition structure of the reference file determination module in the big data information heat analysis system is shown.
[0062] Figure 8 The block diagram of the composition structure of the prediction module in the big data information heat analysis system is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] Embodiment 1
[0065] Figure 1 The flowchart of the big data information heat analysis method is shown. In the embodiment of the present invention, a big data information heat analysis method includes steps S100 to S300:
[0066] Step S100: Obtain the propagation index of the propagated file in real time, compare the propagation index with at least one preset index threshold, and when the propagation index reaches the index threshold, copy the propagated file to the reference file library; wherein, each reference file library corresponds to an index threshold;
[0067] The purpose of step S100 is to generate a reference file library, and files with different propagation indexes are classified into one category. It should be noted that if traversing starts from a high index threshold, the propagated file can be directly copied into the reference file library. If traversing starts from a low index threshold, when the propagated file meets a higher index threshold, the propagated file needs to be deleted from the reference file library corresponding to the lower index threshold.
[0068] Step S200: Receive the uploaded file sent by the user, read the corresponding reference file library according to the descending order of the index threshold, and determine the predicted popularity of the uploaded file according to the reference file library;
[0069] When receiving the file uploaded by the user, it is judged which category the file belongs to in turn to determine a predicted popularity. It is worth mentioning that the comparison process starts from the reference file with a high index threshold.
[0070] Step S300: Allocate promotion resources according to the predicted popularity, and monitor the actual popularity of the file in real time, and correct the promotion resources according to the actual popularity and the predicted popularity;
[0071] A high predicted popularity indicates that it is very likely to become a new hot spot. Therefore, on the premise of limited promotion resources, we should deliberately allocate promotion resources to files with a high predicted popularity. Among them, there are many promotion resources, and some sorting orders belong to one type of promotion resources. In addition, the promotion resources at least include the push range and the push frequency.
[0072] Figure 2 Shows the first sub - process block diagram of the big data information popularity analysis method. The step of obtaining the propagation index of the propagated file in real time includes step S101 to step S102:
[0073] Step S101: Obtain the operation record of the propagated file, and convert the operation record into an equivalent view count according to a preset conversion formula; the operation record at least includes like operations, favorite operations, download operations, and share operations containing quantities;
[0074] Step S102: Obtain the actual view count of the propagated file, and calculate the propagation index according to the actual view count and the equivalent view count;
[0075] Among them, the propagation index is a decreasing function of time.
[0076] Steps S101 to S102 provide a specific calculation process for the propagation index. The key point is that it unifies the "ranges" of like operations, favorite operations, download operations, and share operations, greatly simplifying the calculation process of the propagation index. It should be noted that the propagation index decreases over time. Of course, the rate of decrease may vary, which is related to the actual situation, such as the category of the propagated file, etc. In addition, it is not excluded that the propagation index of some propagated files reaches the maximum value on the second or third day. This situation belongs to a special case and needs to be specifically explained.
[0077] Figure 3 Shows the second sub-process block diagram of the big data information heat analysis method. The steps of receiving the uploaded file sent by the user, reading the corresponding reference file library in descending order of the index threshold, and determining the predicted heat of the uploaded file according to the reference file library include steps S201 to S204:
[0078] Step S201: Receive the uploaded file sent by the user and read the corresponding reference file library in descending order of the index threshold;
[0079] Step S202: Traverse the reference file library, compare the uploaded file with the reference files in the reference file library, and obtain the similarity;
[0080] Step S203: Generate a similarity array according to the similarity, perform feature analysis on the similarity array, and query the corresponding index threshold according to the feature analysis result;
[0081] Step S204: Determine the predicted heat according to the index threshold.
[0082] Steps S201 to S204 provide a calculation method for the predicted heat. The principle is to traverse and compare with the reference file libraries corresponding to different index thresholds in turn to obtain a similarity array. The similarity array corresponds to the entire reference database. Analyzing the similarity array can determine the part of the reference database that is similar to the uploaded file, and finally determine which reference database the uploaded file belongs to, or which index threshold it corresponds to; the index threshold represents the predicted heat.
[0083] Figure 4 Shows the third sub-process block diagram of the big data information heat analysis method. The steps of traversing the reference file library, comparing the uploaded file with the reference files in the reference file library, and obtaining the similarity include steps S2021 to S2024:
[0084] Step S2021: Read the reference files in the reference file library in turn, and convert the reference files into reference images and reference texts;
[0085] Step S2022: converting the uploaded file into an image to be inspected and a text to be inspected, and comparing the text to be inspected with the reference text to obtain text similarity;
[0086] Step S2023: when the text similarity reaches a preset similarity threshold, the text similarity is used as the similarity of the reference file;
[0087] Step S2024: when the text similarity does not reach a preset similarity threshold, traverse the reference image in units of the image to be checked to obtain image similarity, and calculate the similarity of the reference file according to the image similarity and the text similarity.
[0088] Steps S2021 to S2024 further limit the similarity calculation process. The elements of the similarity array are similarities, and the similarities correspond to each reference file in the reference file library. The specific comparison process is: first, the uploaded file and the reference file are converted into images to be inspected and texts to be inspected. If there are none, they are replaced by the "empty" symbol, and they do not participate in subsequent comparisons or are directly defined as dissimilar in the subsequent comparison process.
[0089] Then, the text to be checked is compared. If the similarity of the text to be checked is high enough, there is no need to compare the image to be checked. If the similarity of the text to be checked is not high enough, it is necessary to compare the image to be checked. Finally, a similarity is determined based on the image similarity and the text similarity. If both are very high but do not reach their respective thresholds, it can also be considered to be the same file as the corresponding reference file library.
[0090] Furthermore, the step of converting the reference file or the uploaded file into a reference image and a reference text includes:
[0091] Obtain a file extension, and determine the file type according to the file extension;
[0092] When the file is an audio file, performing content recognition on the audio file to obtain a reference text;
[0093] When the file is a video file, converting the video file into an audio file and an image group;
[0094] Duplicate images in the image group are eliminated, and the remaining images are connected to obtain a reference image that is in a mapping relationship with the video file.
[0095] Specifically, the step of removing duplicate images from the image group includes:
[0096] Traversing the image group, converting the images in the image group into an image array according to a preset conversion formula, and calculating the mean of the image array;
[0097] Calculate the offset rate of the mean value of adjacent image arrays. When the offset rate of the mean value is less than a preset offset threshold, calculate the variance of the adjacent image arrays.
[0098] Calculate the offset rate of the variance of adjacent image arrays. When the offset rate of the variance is less than a preset offset threshold, delete the image corresponding to one of the image arrays in the image group.
[0099] The above content provides a specific conversion method for reference images and reference files. Videos can be converted into images plus audio, and audio can be converted into text. Therefore, it can be said that files are divided into two categories, one is text and the other is images.
[0100] It should be specially noted that for video files, the process of converting video files into reference text mostly belongs to the prior art, while the process of converting image frames into reference images is somewhat special. The reference image that the technical solution of the present invention wants to convert is a "large" image, which directly corresponds to the video and can greatly improve the image comparison efficiency. During the process of connecting image frames into a "large" image, there will be many repeated parts, which is also a feature of video files. Therefore, it is necessary to detect and screen the repeated parts of video files.
[0101] The screening process is based on the mean value and variance. Only when the mean values are similar will the variance be calculated. If both are similar, it can be considered that two adjacent images are similar.
[0102] Figure 5 The fourth sub-process block diagram of the big data information heat analysis method is shown. The steps of generating a similarity array according to the similarity, performing feature analysis on the similarity array, and querying a corresponding index threshold according to the feature analysis result include steps S2031 to S2034:
[0103] Step S2031: Generate a feature map according to the elements and array subscripts in the similarity array;
[0104] Step S2032: Generate a dividing line according to a preset similarity threshold, intercept the feature map according to the dividing line to obtain a feature segment, and query a corresponding index threshold according to the feature segment;
[0105] Step S2033: When the feature segment intercepted according to the dividing line is empty, convert the feature map into a wave function, input the wave function into a trained wave analysis model to obtain a matching degree;
[0106] Step S2034: When the matching degree reaches a preset matching threshold, query a corresponding index threshold;
[0107] Among them, the exponential threshold is the exponential threshold corresponding to the reference file library corresponding to the similarity array.
[0108] Steps S2031 to S2034 define the analysis process of the similarity array. If there are some data with extremely high similarities in the similarity array, it can be directly considered that the uploaded file is in the same category as the corresponding reference image library. If not, further analysis is required.
[0109] The fluctuation function is a function fitted based on the similarity array. The fluctuation analysis model can be understood as some functions or operations for extracting the characteristics of the fluctuation function. For example, if multiple elements in the similarity array reach a certain threshold lower than the similarity threshold, it can also be considered a match. At this time, the fluctuation analysis model will involve the process of obtaining derivatives to identify the fluctuation situation and then determine the extreme value. Of course, this extreme value is an extreme value in terms of mathematical probability and is different from the maximum or minimum value.
[0110] Embodiment 2
[0111] Figure 6 The block diagram of the composition structure of the big data information heat analysis system is shown. In an embodiment of the present invention, a big data information heat analysis system, the system 10 includes:
[0112] A reference file determination module 11, configured to obtain the propagation index of the propagation file in real time, compare the propagation index with at least one preset exponential threshold, and when the propagation index reaches the exponential threshold, copy the propagation file to the reference file library; wherein, each reference file library corresponds to an exponential threshold;
[0113] A heat prediction module 12, configured to receive the uploaded file sent by the user, read the corresponding reference file library according to the descending order of the exponential threshold, and determine the predicted heat of the uploaded file according to the reference file library;
[0114] A promotion module 13, configured to allocate promotion resources according to the predicted heat, and monitor the actual heat of the file in real time, and correct the promotion resources according to the actual heat and the predicted heat;
[0115] Among them, the promotion resources include the push range and the push frequency.
[0116] Figure 7 The block diagram of the composition structure of the reference file determination module in the big data information heat analysis system is shown. The reference file determination module 11 includes:
[0117] The equivalent calculation unit 111 is configured to obtain the operation records of the propagated file and convert the operation records into equivalent view counts according to a preset conversion formula; the operation records at least include like operations, favorite operations, download operations, and share operations containing quantities;
[0118] The exponential calculation unit 112 is configured to obtain the actual view count of the propagated file and calculate a propagation index according to the actual view count and the equivalent view count;
[0119] Wherein, the propagation index is a decreasing function of time.
[0120] Figure 8 The block diagram showing the composition structure of the prediction module in the big data information heat analysis system, the heat prediction module 12 includes:
[0121] The reference file reading unit 121 is configured to receive the uploaded file sent by the user and read the corresponding reference file library in descending order of the exponential threshold;
[0122] The similarity calculation unit 122 is configured to traverse the reference file library, compare the uploaded file with the reference files in the reference file library, and obtain a similarity;
[0123] The feature analysis unit 123 is configured to generate a similarity array according to the similarity, perform feature analysis on the similarity array, and query the corresponding exponential threshold according to the feature analysis result;
[0124] The processing execution unit 124 is configured to determine the predicted heat according to the exponential threshold.
[0125] All functions that the big data information heat analysis method can achieve are completed by a computer device, the computer device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the program code is loaded and executed by the one or more processors to implement the functions of the big data information heat analysis method.
[0126] The processor fetches instructions one by one from the memory, analyzes the instructions, and then completes corresponding operations according to the requirements of the instructions, generating a series of control commands to make each part of the computer move automatically, continuously and coordinately, becoming an organic whole, realizing the input of the program, the input of data, as well as arithmetic operations and output results. All arithmetic operations or logical operations generated in this process are completed by the arithmetic unit; the memory includes a read-only memory (ROM), and the read-only memory is used to store computer programs, and a protection device is provided outside the memory.
[0127] Exemplarily, a computer program can be divided into one or more modules. One or more modules are stored in a memory and executed by a processor to implement the present invention. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in a terminal device.
[0128] Those skilled in the art can understand that the description of the above service device is only an example and does not constitute a limitation on the terminal device. It may include more or fewer components than the above description, or combine some components, or different components. For example, it may include input / output devices, network access devices, buses, etc.
[0129] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above processor is the control center of the above terminal device, and connects various parts of the entire user terminal through various interfaces and lines.
[0130] The above memory can be used to store computer programs and / or modules. The above processor realizes various functions of the above terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as an information collection template display function, a product information release function, etc.); the data storage area can store data created according to the use of the berth status display system (such as product information collection templates corresponding to different product types, product information to be released by different product providers, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMCs), secure digital (SD) cards, flash cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0131] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the modules / units in the above-described embodiment system of the present invention, it can also be completed by instructing relevant hardware through a computer program. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the functions of the above various system embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0132] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0133] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for analyzing the popularity of big data information, characterized in that, The method includes: Obtaining the propagation index of the propagated file in real time, comparing the propagation index with at least one preset index threshold, and when the propagation index reaches the index threshold, copying the propagated file to the reference file library; wherein, each reference file library corresponds to an index threshold; Receiving the uploaded file sent by the user, reading the corresponding reference file library in descending order of the index threshold, and determining the predicted popularity of the uploaded file according to the reference file library; Allocating promotion resources according to the predicted popularity, and monitoring the actual popularity of the file in real time, and correcting the promotion resources according to the actual popularity and the predicted popularity; Wherein, the promotion resources include the push range and the push frequency; The step of receiving the uploaded file sent by the user, reading the corresponding reference file library in descending order of the index threshold, and determining the predicted popularity of the uploaded file according to the reference file library includes: Receiving the uploaded file sent by the user, and reading the corresponding reference file library in descending order of the index threshold; Traversing the reference file library, comparing the uploaded file with the reference files in the reference file library to obtain the similarity; Generating a similarity array according to the similarity, performing feature analysis on the similarity array, and querying the corresponding index threshold according to the feature analysis result; Determining the predicted popularity according to the index threshold; The step of traversing the reference file library, comparing the uploaded file with the reference files in the reference file library to obtain the similarity includes: Sequentially reading the reference files in the reference file library, and converting the reference files into reference images and reference texts; Converting the uploaded file into a to-be-inspected image and to-be-inspected text, comparing the to-be-inspected text with the reference text to obtain the text similarity; When the text similarity reaches the preset similarity threshold, using the text similarity as the similarity of the reference file; When the text similarity does not reach the preset similarity threshold, traversing the reference images in units of the to-be-inspected image to obtain the image similarity, and calculating the similarity of the reference file according to the image similarity and the text similarity; The step of generating a similarity array according to the similarity, performing feature analysis on the similarity array, and querying the corresponding index threshold according to the feature analysis result includes: Generating a feature map according to the elements and array subscripts in the similarity array; Generating a dividing line according to the preset similarity threshold, intercepting the feature map according to the dividing line to obtain a feature segment, and querying the corresponding index threshold according to the feature segment; When the feature segment intercepted according to the dividing line is empty, converting the feature map into a fluctuation function, and inputting the fluctuation function into a trained fluctuation analysis model to obtain the matching degree; When the matching degree reaches the preset matching threshold, querying the corresponding index threshold; Wherein, the index threshold is the index threshold corresponding to the reference file library corresponding to the similarity array.
2. The big data information heat analysis method according to claim 1, wherein The step of obtaining the propagation index of the propagated file in real time includes: Obtain the operation record of the propagated file, and convert the operation record into an equivalent view count according to a preset conversion formula; the operation record at least includes like operations, favorite operations, download operations, and share operations containing quantities. Obtain the actual view count of the propagated file, and calculate the propagation index according to the actual view count and the equivalent view count. Among them, the propagation index is a decreasing function of time.
3. The big data information heat analysis method according to claim 2, characterized in that The steps of converting the reference file or the uploaded file into a reference image and a reference text include: Obtain the file suffix name, and determine the file type according to the file suffix name. When the file is an audio file, perform content recognition on the audio file to obtain a reference text. When the file is a video file, convert the video file into an audio file and an image group. Eliminate duplicate images in the image group, and connect the remaining images to obtain a reference image in a mapping relationship with the video file.
4. The big data information popularity analysis method according to claim 3, wherein The steps of eliminating duplicate images in the image group include: Traverse the image group, convert the images in the image group into an image array according to a preset conversion formula, and calculate the mean value of the image array. Calculate the offset rate of the mean values of adjacent image arrays. When the offset rate of the mean value is less than a preset offset threshold, calculate the variance of adjacent image arrays. Calculate the offset rate of the variances of adjacent image arrays. When the offset rate of the variance is less than a preset offset threshold, delete the image corresponding to one of the image arrays in the image group.
5. A big data information heat analysis system for implementing the method according to any one of claims 1-4, characterized in that, The system includes: A reference file determination module, which is used to obtain the propagation index of the propagated file in real time, compare the propagation index with at least one preset index threshold, and when the propagation index reaches the index threshold, copy the propagated file to the reference file library; among them, each reference file library corresponds to an index threshold. A popularity prediction module, which is used to receive the uploaded file sent by the user, read the corresponding reference file library in descending order of the index threshold, and determine the predicted popularity of the uploaded file according to the reference file library. A promotion module, which is used to allocate promotion resources according to the predicted popularity, and monitor the actual popularity of the file in real time, and correct the promotion resources according to the actual popularity and the predicted popularity. Among them, the promotion resources include the push range and the push frequency.
6. The big data information popularity analysis system according to claim 5, characterized in that, The reference file determination module includes: An equivalent calculation unit, which is used to obtain the operation record of the propagated file, and convert the operation record into an equivalent view count according to a preset conversion formula; the operation record at least includes like operations, favorite operations, download operations, and share operations containing quantities. An index calculation unit, which is used to obtain the actual view count of the propagated file, and calculate the propagation index according to the actual view count and the equivalent view count. Among them, the propagation index is a decreasing function of time.
7. The big data information heat analysis system according to claim 6, wherein The prediction module includes: A reference file reading unit, which is used to receive the uploaded file sent by the user, and read the corresponding reference file library in descending order of the index threshold. A similarity calculation unit, which is used to traverse the reference file library, compare the uploaded file with the reference files in the reference file library, and obtain the similarity. A feature analysis unit, configured to generate a similarity array according to the similarity, perform feature analysis on the similarity array, and query a corresponding exponential threshold according to the feature analysis result; A processing execution unit, configured to determine a predicted popularity according to the exponential threshold.
Citation Information
Patent Citations
Multimedia content data processing method and device
CN111833083A
Online public opinion monitoring and early warning method
CN112434226A