Traffic anomaly evaluation method and device, electronic equipment and storage medium
By obtaining the relevance matching degree of the content description text and comment text of multimedia data and combining it with abnormal word analysis, the problems of misjudgment and delayed identification of traffic cheating behavior in existing technologies are solved, and more accurate and timely traffic anomaly detection is achieved.
Patent Information
- Application Number
- CN202410316020.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-19
AI Technical Summary
When identifying video data traffic cheating, existing technologies are easily affected by time period values, leading to misjudgment, and are difficult to effectively identify traffic anomalies, which affects the promotion of normal video data.
By obtaining the content description text and N groups of comment texts associated with the target multimedia data, the correlation between the comment texts and the matching degree with the content description text are calculated. Combined with the abnormal word hit degree, the abnormal evaluation value is determined to determine the traffic abnormality.
It improves the accuracy of traffic anomaly assessment, reduces misjudgment of normal multimedia data, enhances the timeliness of detection, and avoids missed identification due to dispersed growth of heat.
Smart Images

Figure CN120671664A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a flow anomaly assessment method, device, electronic device, and storage medium. Background Art
[0002] Under existing technologies, after a video publisher publishes video data on a data publishing platform, as the number of views, likes, and comments on the video data increases, the video publisher can obtain more creative income; moreover, the data publishing platform will provide more promotional traffic for popular video data so that the video data can be recommended to more viewers. The popularity of video data increases with the increase in views, likes, and comments.
[0003] At present, in order to deal with the problem of video publishers using illegal means to increase the popularity of video data in order to obtain high promotion traffic, it is usually necessary to count the changes in the popularity of video data over multiple time periods. When it is determined that the popularity of video data has continuous sudden increases or decreases, it is determined that the video data has traffic cheating behavior.
[0004] However, existing traffic anomaly assessment methods may misjudge normal video data due to the influence of time period values, affecting the promotion of normal video data. In addition, the identification method based on popularity changes is extremely easy to circumvent, making it difficult to effectively identify traffic cheating behavior. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, electronic device, and storage medium for identifying traffic fraud behavior, so as to improve the accuracy of traffic anomaly assessment.
[0006] First, a traffic anomaly assessment method is proposed, including:
[0007] In response to an evaluation indication triggered for target multimedia data, obtaining content description text and N groups of comment texts associated with the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from a set of comment texts associated with the target multimedia data;
[0008] For each group of comment texts, the following operations are performed: obtaining each comment text contained in the group of comment texts; performing content relevance detection based on each comment text to obtain comment relevance corresponding to the group of comment texts; performing content relevance detection based on each comment text and the content description text to obtain content matching between the group of comment texts and the target multimedia data;
[0009] Calculating the degree of abnormality based on the N comment relevances and the N content matching degrees to obtain an abnormality evaluation value corresponding to the target multimedia data;
[0010] When it is determined that the abnormality evaluation value exceeds a set threshold, it is determined that a traffic abnormality exists in the target multimedia data.
[0011] In a second aspect, a flow anomaly assessment device is proposed, comprising:
[0012] a response unit, configured to, in response to an evaluation indication triggered for target multimedia data, obtain content description text and N groups of comment texts associated with the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from a set of comment texts associated with the target multimedia data;
[0013] An execution unit is configured to perform the following operations for each group of comment texts: obtaining each comment text contained in the group of comment texts; performing content relevance detection based on each comment text to obtain comment relevance corresponding to the group of comment texts; performing content relevance detection based on each comment text and the content description text to obtain a content matching degree between the group of comment texts and the target multimedia data;
[0014] an obtaining unit, configured to calculate an abnormality degree based on the N comment relevances and the N content matching degrees, and obtain an abnormality evaluation value corresponding to the target multimedia data;
[0015] The determination unit is configured to determine that traffic cheating behavior exists in the target multimedia data when the abnormal evaluation value exceeds a set threshold.
[0016] Optionally, before responding to the evaluation instruction triggered for the target multimedia data, the apparatus further includes a screening unit, the screening unit being configured to:
[0017] Obtain each candidate multimedia data, and for each candidate multimedia data, perform the following operations:
[0018] Performing a quantitative comparison of the operation heat according to the historical operation data of the candidate multimedia data in two adjacent time windows to obtain a heat change value of the candidate multimedia data;
[0019] performing a quantitative comparison of operation heats based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data to obtain a heat comparison value between the candidate multimedia data and each historical multimedia data; each historical multimedia data and the candidate multimedia data are associated with the same publishing object;
[0020] When it is determined that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data is used as the target multimedia data, and an evaluation indication is triggered.
[0021] Optionally, when obtaining the corresponding popularity change value based on the change of the historical operation data in the two time windows, the screening unit is configured to:
[0022] For each time window, the following operations are performed: the historical occurrence counts of various target operations in the time window are counted; according to the preset weight values for the various target operations, the historical occurrence counts are weighted and fused to obtain the operation heat value corresponding to the time window;
[0023] Calculating a heat value increment and a heat value ratio between two operation heat values; and determining the heat value increment and the heat value ratio as a heat change value of the candidate multimedia data.
[0024] Optionally, when obtaining the heat comparison value between the candidate multimedia data and each historical multimedia data based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, the screening unit is configured to:
[0025] According to the historical operation data of the candidate multimedia data, the cumulative occurrence times of various target operations are counted to obtain the operation heat value obtained by weighted fusion of the cumulative occurrence times;
[0026] Based on the historical operation data of each historical multimedia data, the cumulative number of occurrences of each type of target operation is counted to obtain an operation heat value obtained by weighted fusion of the cumulative number of occurrences; and based on the operation heat values, the average operation heat value of each historical multimedia data is calculated;
[0027] Calculate the heat value increment and heat value ratio between the operation heat value of the candidate multimedia data and the operation heat mean; and use the heat value increment and the heat value ratio as the heat comparison value between the candidate multimedia data and each historical multimedia data.
[0028] Optionally, when obtaining the heat comparison value between the candidate multimedia data and each historical multimedia data based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, the screening unit is configured to:
[0029] Determining corresponding time intervals for the candidate multimedia data and each historical multimedia data;
[0030] For the candidate multimedia data and each historical multimedia data, the following operations are performed respectively:
[0031] Count the number of occurrences of each target operation in the corresponding concurrent time interval to obtain statistical results of each operation; obtain the corresponding concurrent operation heat based on the statistical results of each operation; obtain the historical average heat based on the concurrent operation heat of each historical multimedia data; determine the heat value increment and heat value ratio between the concurrent operation heat of the candidate multimedia data and the historical average heat as the corresponding heat comparison value.
[0032] Optionally, after obtaining the content description text and N groups of comment texts associated with the target multimedia data, and before calculating the abnormality degree according to the N comment relevances and N content matching degrees, the execution unit is further configured to:
[0033] Obtain various abnormal words pre-constructed for traffic fraud behavior;
[0034] Determine the number of hits for each abnormal word in the content description text and the N groups of comment texts;
[0035] According to the number of hits of each abnormal word and the abnormal indicators determined in advance for each abnormal word, the abnormal word hit degree corresponding to the target multimedia data is determined, wherein the abnormal indicator of an abnormal word is determined by calculating the inclusion of the abnormal word in the multimedia data with and without traffic cheating behavior.
[0036] Optionally, the abnormal words and corresponding abnormal indicators are determined by the execution unit in the following manner:
[0037] Obtaining a comment text set associated with each sample multimedia data, wherein each sample multimedia data includes: positive sample multimedia data associated with popularity cheating behavior, and negative sample multimedia data not associated with popularity cheating behavior;
[0038] From each sample comment related to traffic fraud, word segmentation is performed to obtain candidate words;
[0039] For each candidate word, the following operations are performed: based on the sample comment sets associated with each positive sample multimedia data and the sample comment sets associated with each negative sample multimedia data, the inclusion difference of the candidate word is calculated to obtain the abnormality index of the candidate word; when it is determined that the abnormality index exceeds the set value, the candidate word is determined as an abnormal word.
[0040] Optionally, when calculating the abnormality degree according to the N comment relevances and the N content matching degrees to obtain the abnormality evaluation value corresponding to the target multimedia data, the obtaining unit is configured to:
[0041] According to the N comment relevances, the N content matching degrees, and the abnormal word hit degree of the target multimedia data, a data fusion calculation is performed to obtain an abnormal evaluation value of the target multimedia data.
[0042] Optionally, after determining that traffic cheating occurs in the target multimedia data, the determining unit is further configured to:
[0043] Obtaining pre-set cheating handling strategies at each level, wherein each cheating handling strategy has an associated evaluation value range;
[0044] Determine the target evaluation value interval corresponding to the abnormal evaluation value according to the inclusion of the abnormal evaluation value in each evaluation value interval;
[0045] The target multimedia data is processed using a cheating processing strategy corresponding to the target evaluation value interval.
[0046] Optionally, when performing content relevance detection based on the comment texts to obtain comment relevance corresponding to the group of comment texts, the execution unit is configured to:
[0047] Inputting each text feature obtained for each comment text into a trained first evaluation model to obtain a first prediction value corresponding to the popularity cheating behavior output by the first evaluation model after performing content relevance detection on each text feature, wherein the first evaluation model is obtained by performing multiple rounds of iterative training based on a first training sample set, wherein a first training sample includes: each sample comment associated with a sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs;
[0048] The first prediction value is used as the comment relevance corresponding to the group of comment texts.
[0049] Optionally, when performing content relevance detection based on the comment texts and the content description texts to obtain a content matching degree between the set of comment texts and the target multimedia data, the execution unit is configured to:
[0050] Inputting the text features obtained for each of the comment texts and the content description text into a trained second evaluation model to obtain a second prediction value corresponding to the popularity cheating behavior output by the second evaluation model after performing content relevance detection on each of the text features, wherein the second evaluation model is obtained by performing multiple rounds of iterative training based on a second training sample set, wherein a second training sample includes: each sample comment and sample description text associated with the sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs;
[0051] The second prediction value is used as the content matching degree corresponding to the group of comment texts.
[0052] In a third aspect, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.
[0053] In a fourth aspect, a computer-readable storage medium is proposed, on which a computer program is stored, and the computer program implements the above method when executed by a processor.
[0054] In a fifth aspect, a computer program product is proposed, comprising a computer program, which implements the above method when executed by a processor.
[0055] The beneficial effects of this application are as follows:
[0056] In an embodiment of the present application, a traffic anomaly assessment method, apparatus, electronic device, and storage medium are proposed. In response to an assessment indication triggered for target multimedia data, content description text and N groups of comment texts associated with the target multimedia data are obtained. When determining to perform a traffic anomaly assessment on the target multimedia data, the text content based on which the identification is based can be obtained from the perspective of the data content.
[0057] Afterwards, for each group of comment texts, the comment relevance between the comment texts is obtained based on the content relevance between the included comment texts, and the content matching between the comment texts and the content description text of the target multimedia data is obtained based on the content relevance between the comment texts and the content description text of the target multimedia data. Based on this, when evaluating traffic anomalies for the target multimedia data, combined with the specific manifestation of traffic anomalies in the comment texts, by identifying the correlation between the comment contents and between the comment contents and the data content of the target multimedia data, the anomalies of the comment texts associated with the target multimedia data can be expressed with the help of comment relevance and content matching.
[0058] Furthermore, after calculating N comment relevances and N content matching degrees for the target multimedia data, an anomaly assessment value is determined based on the N comment relevances and N content matching degrees, and when it is determined that the anomaly assessment value exceeds the set threshold, it is determined that there is a traffic anomaly; this allows the traffic anomaly to be determined based on the anomaly of the comment text, rather than identifying the changes in the popularity of the multimedia data in each time period. Therefore, the relevant identification process is not affected by the value of the time period on the one hand, and on the other hand, it cannot miss identification due to the dispersed growth of popularity; this can not only greatly reduce the misjudgment of normal multimedia data, but also improve the accuracy of traffic anomaly assessment. In addition, since there is no need to collect long-term historical data, the timeliness of traffic anomaly detection can be improved, thereby ensuring the detection effect of traffic anomalies. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a schematic diagram of the process of flow anomaly assessment under the relevant technology in the embodiment of this application;
[0060] Figure 2 A schematic diagram of possible application scenarios in the embodiments of the present application;
[0061] Figure 3A This is a schematic diagram of the flow anomaly assessment process in an embodiment of the present application;
[0062] Figure 3B This is a schematic diagram of the process of determining the target duration range in an embodiment of the present application;
[0063] Figure 4A Schematic diagram of the process of determining target multimedia data in an embodiment of the present application;
[0064] Figure 4B This is a schematic diagram of the time covered by two adjacent time windows in an embodiment of the present application;
[0065] Figure 4C This is a schematic diagram of the process of determining the same period time interval in the embodiment of the present application;
[0066] Figure 5 A schematic diagram of the process of constructing description content for video data in an embodiment of the present application;
[0067] Figure 6 A schematic diagram of the process of determining a cheating handling strategy in an embodiment of the present application;
[0068] Figure 7A This is a schematic diagram of the timing of evaluating flow anomalies in an embodiment of the present application;
[0069] Figure 7B This is a schematic diagram of the process of evaluating traffic anomalies for a video work in an embodiment of the present application;
[0070] Figure 7C This is a schematic diagram of the process of constructing a text vector in an embodiment of the present application;
[0071] Figure 7D A schematic diagram of the process of determining the features of each sentence in an embodiment of the present application;
[0072] Figure 7E Schematic diagram of the process of determining the TGI index of review segmentation in an embodiment of the present application;
[0073] Figure 7F This is a schematic diagram of the flow anomaly assessment process in an embodiment of the present application;
[0074] Figure 8This is a schematic diagram of the logical structure of the flow anomaly assessment device in an embodiment of the present application;
[0075] Figure 9 A schematic diagram of the hardware structure of an electronic device to which the embodiments of the present application are applied;
[0076] Figure 10 The figure is a schematic diagram of the hardware structure of another electronic device to which the embodiments of the present application are applied. DETAILED DESCRIPTION
[0077] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.
[0078] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the invention described herein can be practiced in sequences other than those illustrated or described herein.
[0079] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0080] The following explains some of the terms used in the embodiments of the present application to facilitate understanding by those skilled in the art.
[0081] Application (APP): mainly refers to software installed on terminal devices to improve the deficiencies of the original system and meet personalized needs.
[0082] Historical operation data: used to describe the various operations performed by related objects on multimedia data in the app. The types of operations involved include but are not limited to any one or a combination of the following: like, comment, browse, forward, and favorite.
[0083] Brushing behavior mainly refers to the behavior of increasing the popularity of a specified work (or specified multimedia data) on the Internet through abnormal means or techniques. The popularity of a multimedia data is usually determined by any one or a combination of the number of views and interactions with the multimedia data. The number of interactions is usually determined by any one or a combination of the number of likes, comments, collections, and forwarding. It should be understood that since brushing groups usually do not pay attention to the content of the multimedia data itself, but only mechanically use machines to perform operations such as "browsing, liking, and commenting" in batches, the comments generated by brushing are generally meaningless comments, and the content of the comments usually has nothing to do with the content of the multimedia data.
[0084] Multimedia data: refers to data constructed by at least one of text content, video content, and audio content, or a combination thereof.
[0085] Traffic anomaly assessment: It can also be understood as the identification of traffic cheating behavior, which is used to assess whether there is a situation where promotional traffic is illegally obtained by means of brushing behavior; in the embodiment of the present application, since the purpose of brushing behavior is to cheat popularity, and popularity cheating will further cause traffic cheating and illegally obtain promotional traffic, traffic anomaly assessment can be achieved by identifying the occurrence of brushing behavior.
[0086] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0087] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0088] The following is a brief introduction to the design concept of the embodiment of this application:
[0089] Under existing technology, video publishers can create and publish videos in platforms such as video apps. Other viewers will notice these videos and then perform actions such as browsing, liking, commenting, forwarding, and adding to favorites. Furthermore, the more actions performed on a video, the higher its popularity and popularity, meaning it becomes more popular. Furthermore, upon sensing the video's popularity, the video server will provide more promotional traffic to the video, allowing it to be recommended to more viewers, resulting in more views, likes, comments, forwarding, and adding to favorites, creating a positive cycle of increasing popularity and popularity. Furthermore, as the video's popularity and popularity increase, the creative revenue earned by the video publisher also gradually increases.
[0090] In this case, the following malicious process exists: some video publishers will purchase a brushing service in order to quickly gain popularity for their published works. After purchasing the brushing service, a team will browse, like, and comment on the designated video (assuming it is Video A) of the video publisher in batches, thereby making the video gain high popularity in a short period of time; then, the video server will determine Video A as a high-quality video, thereby allocating more promotional traffic to Video A, so that Video A can be recommended to more viewers; then, Video A of the video publisher will gain a lot of exposure, which will quickly attract attention, so that the video publisher can ultimately benefit from it.
[0091] At present, in order to deal with the problem of video publishers using illegal means to increase the popularity of video data in order to obtain high promotion traffic, traffic anomaly assessment is generally performed based on the popularity changes of video works (or video data). By detecting changes in values such as the number of views and comments of video works, video publishing objects and video works with abnormal changes are extracted to determine whether they have traffic anomalies.
[0092] For example, see Figure 1 As shown, it is a schematic diagram of the process of flow anomaly assessment under the relevant technology in the embodiment of this application. Figure 1 As shown in the following content, the process of identifying traffic fraud involves the following steps:
[0093] Step 101: The video server captures behaviors such as viewing, commenting, and liking of a video work.
[0094] Step 102: The video server obtains popularity data of the video works by statistically analyzing the time periods.
[0095] The time periods used may be: 5 minutes, 30 minutes, 1 hour, 3 hours, 12 hours, etc.
[0096] Step 103: After the video server determines that the time length corresponding to each piece of heat data reaches the set time length, it sorts the pieces of heat data in chronological order.
[0097] Step 104: The video server determines that the popularity of the video work has a continuous sudden increase or decrease based on the sorted popularity data, and then determines that the video work has a traffic anomaly.
[0098] However, since it is necessary to count the heat data for a period of time when evaluating traffic anomalies based on the long-term heat data of video works, it is often a long time before the heat data of a video work is judged to be abnormal. At this time, the brushing operation has been completed for the video work. Therefore, the existing evaluation method can only identify it after the brushing operation is performed on the video work, so the evaluation of traffic anomalies has a strong lag; moreover, when evaluating traffic anomalies based on the heat data of video works, since suddenly popular video works also have a sudden increase in heat, suddenly popular video works may be mistakenly damaged, affecting the dissemination of normal video works; furthermore, judging only by the changes in heat data is easy to be bypassed by brushing groups. When the brushing groups control the slow increase in the heat of the video works when brushing, it is impossible to effectively evaluate the traffic anomaly.
[0099] In view of this, embodiments of the present application provide a method, apparatus, electronic device, and storage medium for evaluating traffic anomaly. These methods, in response to an evaluation instruction triggered for target multimedia data, obtain content description text and N groups of comment texts associated with the target multimedia data. This allows the user to obtain text content for identification based on the data content when determining to perform traffic anomaly evaluation on the target multimedia data.
[0100] Afterwards, for each group of comment texts, the comment relevance between the comment texts is obtained based on the content relevance between the included comment texts, and the content matching between the comment texts and the content description text of the target multimedia data is obtained based on the content relevance between the comment texts and the content description text of the target multimedia data. Based on this, when evaluating traffic anomalies for the target multimedia data, combined with the specific manifestation of traffic anomalies in the comment texts, by identifying the correlation between the comment contents and between the comment contents and the data content of the target multimedia data, the anomalies of the comment texts associated with the target multimedia data can be expressed with the help of comment relevance and content matching.
[0101] Furthermore, after calculating N comment relevances and N content matching degrees for the target multimedia data, an anomaly assessment value is determined based on the N comment relevances and N content matching degrees, and when it is determined that the anomaly assessment value exceeds the set threshold, it is determined that there is a traffic anomaly; this allows the traffic anomaly to be determined based on the anomaly of the comment text, rather than identifying the changes in the popularity of the multimedia data in each time period. Therefore, the relevant identification process is not affected by the value of the time period on the one hand, and on the other hand, it cannot miss identification due to the dispersed growth of popularity; this can not only greatly reduce the misjudgment of normal multimedia data, but also improve the accuracy of traffic anomaly assessment. In addition, since there is no need to collect long-term historical data, the timeliness of traffic anomaly detection can be improved, thereby ensuring the detection effect of traffic anomalies.
[0102] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.
[0103] See Figure 2 , which is a schematic diagram of a possible application scenario in an embodiment of the present application. The application scenario schematic diagram includes a terminal device 210 and a processing device 220.
[0104] In a feasible embodiment of the present application, the terminal device 210 provides the publishing object with the functions of editing and generating multimedia data, as well as publishing multimedia data. In the specific processing process, the terminal device 210 publishes the multimedia data to the processing device 220 in response to the publishing object's publishing instruction for the multimedia data; furthermore, the processing device 220 can recommend the received multimedia data to other browsing objects according to a preset recommendation strategy, wherein the preset recommendation strategy is set according to actual processing needs, and the present application does not impose specific restrictions on this. For example, when the location information is obtained in compliance with regulations, it is recommended to browsing objects in the same region, etc.; in addition, the processing device 220 can allocate promotional traffic to the multimedia data according to the popularity of the multimedia data published by the publishing object, and provide the publishing object with creative income, wherein the higher the popularity of the multimedia data, the greater the promotional traffic allocated to it, and the greater the probability that the multimedia data will be browsed subsequently.
[0105] In addition, after the processing device 220 determines the target multimedia data that needs to be evaluated for traffic anomalies, it uses the comment texts associated with the target multimedia data and the content description text of the target multimedia data to identify the comment relevance between the comment texts and the content matching between the comment texts and the content description text, and determines the anomaly evaluation value of the target multimedia data from the perspective of the associated comment texts, and then determines whether the target multimedia data has traffic anomalies based on the anomaly evaluation value; further, when the processing device 220 determines that the target multimedia data has traffic anomalies, it can perform restriction processing on the target multimedia data and the corresponding publishing object.
[0106] It should be noted that a target application that provides multimedia data publishing and browsing functions can be installed on the terminal device 210, so that the publishing object can edit, generate and publish multimedia data, and display the creative benefits obtained to the publishing object; the processing device 220 can specifically be the server end of the target application.
[0107] The terminal device 210 includes but is not limited to mobile phones, tablet computers, notebooks, e-book readers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.
[0108] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.
[0109] The processing device 220 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0110] In the embodiment of the present application, the terminal device 210 and the processing device 220 can communicate through a wired network or a wireless network.
[0111] The following is a schematic illustration of traffic anomaly assessment scenarios based on possible application scenarios:
[0112] Application scenario 1: Detect video data with fake views.
[0113] In the scenario corresponding to application scenario 1, the processing device obtains the content description text associated with the target video data for traffic anomaly assessment, as well as N groups of comment texts. Then, for each group of comment texts, the processing device calculates the comment relevance, which indicates the correlation between the comment texts, and the content matching, which indicates the content correlation between the text content of the comment texts and the data content of the target video data.
[0114] Furthermore, based on the obtained N comment relevances and N content matching degrees, an anomaly assessment value is calculated, and then based on the relationship between the anomaly assessment value and the set threshold, it is determined whether there is traffic anomaly in the target video data.
[0115] Application scenario 2: Detecting audio data with fake traffic.
[0116] In the scenario corresponding to application scenario 2, the processing device obtains the content description text associated with the target audio data for traffic anomaly assessment, as well as N groups of comment texts. Then, for each group of comment texts, the processing device calculates the comment relevance, which indicates the correlation between the comment texts, and the content match, which indicates the content correlation between the text content of the comment texts and the data content of the target audio data.
[0117] Furthermore, based on the obtained N comment relevances and N content matching degrees, an anomaly evaluation value is calculated, and then based on the relationship between the anomaly evaluation value and the set threshold, it is determined whether there is traffic anomaly in the target audio data.
[0118] Application scenario three: Detect text data that contains fake traffic.
[0119] In the scenario corresponding to application scenario three, the processing device obtains the content description text associated with the target text data and N groups of comment texts for the target text data to be evaluated for traffic anomalies. Then, for each group of comment texts, the processing device calculates the comment relevance, which indicates the correlation between the comment texts, and the content matching, which indicates the content correlation between the text content of the comment texts and the data content of the target text data.
[0120] Furthermore, based on the obtained N comment relevances and N content matching degrees, an anomaly evaluation value is calculated, and then based on the relationship between the anomaly evaluation value and the set threshold, it is determined whether there is traffic anomaly in the target text data.
[0121] In addition, it should be understood that in the specific implementation of this application, it involves flow anomaly assessment. When the embodiments recorded in this application are applied to specific products or technologies, the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0122] The following describes the flow anomaly assessment process from the perspective of processing equipment, with reference to the accompanying figures:
[0123] See Figure 3A As shown in the figure, it is a schematic diagram of the flow anomaly evaluation process in the embodiment of the present application. Figure 3A , explains the process of handling device evaluation flow anomalies:
[0124] Step 301: The processing device obtains content description text and N groups of comment texts associated with the target multimedia data in response to an evaluation indication triggered for the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from the comment text set associated with the target multimedia data.
[0125] In an embodiment of the present application, before the processing device performs specific processing in response to the evaluation indication triggered for the target multimedia data, it is necessary to first determine the target multimedia data for which the traffic anomaly evaluation is to be performed; in addition, according to actual processing needs, the traffic anomaly evaluation process performed by the processing device can be performed periodically, that is, each target multimedia data is periodically determined, and the evaluation process is performed for each target multimedia data.
[0126] When determining target multimedia data for traffic anomaly assessment, the processing device may use any of the following feasible determination methods:
[0127] Method 1: Determine the target duration range based on the peak period of traffic fraud, and identify multimedia data with a published duration within the target duration range as target multimedia data that needs to be evaluated for traffic anomalies.
[0128] In the process of determining the target multimedia data indicated in method one, it is taken into account that there is usually a correlation between traffic cheating behavior and the release time of multimedia data. For example, according to actual statistics, the release object is most likely to purchase a brushing service within 1-2 days of the release of multimedia data; therefore, the processing device can determine the target duration range in combination with the time period when traffic cheating behavior is highly prevalent, and determine the multimedia data with a release duration within the target duration range as the target multimedia data that needs to be evaluated for traffic anomalies, and trigger an evaluation indication for the target multimedia data, wherein the specific boundary value of the target duration range is set according to actual processing needs, and this application does not impose specific restrictions on this.
[0129] It should be noted that when determining the target duration range, the following operations can be performed for each abnormal multimedia data item that has been determined to have a traffic anomaly: determining the time at which the abnormal multimedia data item was determined to have a traffic anomaly, and calculating the time between the abnormal multimedia data item's release time and the determination time. Subsequently, after obtaining the respective time lengths for each abnormal multimedia data item, the union of the respective time lengths is determined as the target duration range.
[0130] In addition, the ways in which abnormal multimedia data is judged as traffic anomaly include but are not limited to: manual verification of reported multimedia data; manual verification of multimedia data identified as having traffic anomalies using existing technologies; and determination using the traffic anomaly assessment method proposed in this application.
[0131] For example, see Figure 3B As shown, it is a schematic diagram of the process of determining the target duration range in an embodiment of the present application. Assume that there are m abnormal multimedia data that are determined to have traffic abnormalities, recorded as multimedia data 1-m; then the interval duration between the release time of each abnormal multimedia data and the corresponding traffic abnormality determination time is calculated respectively, and T1.2-T1.1, T2.2-T2.1, T3.2-T3.1, T4.2-T4.1, ..., Tm.2-Tm.1 are obtained; then, the maximum value △t1 and the minimum value △t2 are determined in each interval duration, and △t2 to △t1 after the multimedia data is released are determined as the target duration range.
[0132] For another example, in a feasible implementation, after determining the maximum value of the interval duration, the target duration range can be directly determined based on the obtained maximum value; assuming that the maximum value of the interval duration △t1 is 3 days, 5 hours, 20 minutes and 15 seconds, the target duration range can be determined as: starting from the associated release time to 3 days, 5 hours, 20 minutes and 15 seconds after the release time.
[0133] In this way, we can combine the high-incidence stage of traffic anomalies and, from the perspective of published time, filter out the target multimedia data that need to be evaluated for traffic anomalies from among the multimedia data in the published state, thereby greatly reducing the amount of multimedia data that needs to be processed.
[0134] Method 2: Determine candidate multimedia data with abnormal popularity changes as target multimedia data.
[0135] When the determination method indicated by the second method is used for processing, the processing device first determines each candidate multimedia data, and then determines the candidate multimedia data with abnormal heat change among the candidate multimedia data as the target multimedia data.
[0136] In some feasible implementation methods, the candidate multimedia data may be the target multimedia data determined by the processing method of method one. At this time, after the target multimedia data is screened by the processing method of method one, the processing method of method two is used for further screening to obtain the target multimedia data that ultimately needs to be used for traffic cheating behavior identification; in other feasible implementation methods, the candidate multimedia data may be all multimedia data that is currently in the publishing state.
[0137] When determining whether each candidate multimedia data can be used as the target multimedia data, refer to Figure 4A As shown in FIG, it is a schematic diagram of the process of determining target multimedia data in an embodiment of the present application. Figure 4A Taking the determination of a single candidate multimedia data as an example, the process of determining whether it can be used as the target multimedia data is explained:
[0138] Step 401: The processing device performs a quantitative comparison of the operation heat of the candidate multimedia data based on the historical operation data of the candidate multimedia data in two adjacent time windows to obtain a heat change value of the candidate multimedia data.
[0139] In an embodiment of the present application, the processing device selects the time length corresponding to the time window according to actual processing needs, and then obtains the historical operation data of the candidate multimedia data in two adjacent time windows for the candidate multimedia data, wherein the historical operation data records the occurrence time and details of various target operations.
[0140] It should be noted that in the embodiments of this application, the time lengths corresponding to the two time windows can be the same or different, and this application does not impose any specific restrictions on this. Furthermore, based on a preset time length, the two time windows closest to the current time can be determined with the current time as the right endpoint; alternatively, the time positions of the two adjacent time windows can be specified in the historical time according to actual processing needs.
[0141] For example, see Figure 4B As shown, it is a schematic diagram of the time covered by two adjacent time windows in the embodiment of the present application. Figure 4B As shown in the content, when determining the time range covered by the two time windows, the time when the candidate multimedia data is judged to be the target multimedia data can be taken as the current time, and then the current time is taken as the right edge point of the time window 2 to obtain Figure 4B The time range covered by two adjacent time windows is shown.
[0142] For another example, assuming that a round of traffic anomaly assessment is conducted every two hours with a two-hour cycle, and the set time window corresponds to a duration of 0.5 hours; then, in the new round of traffic anomaly assessment, the duration of two time windows can be arbitrarily specified within the past 2 hours, that is, a time position with a duration of 1 hour can be arbitrarily specified.
[0143] Furthermore, the processing device performs the following operations for each time window: counting the historical occurrence times of each type of target operation in the time window; performing weighted fusion processing on each historical occurrence time according to each weight value preset for each type of target operation to obtain the operation heat value corresponding to the time window; calculating the heat value increment and heat value ratio between two operation heat values; and determining the heat value increment and heat value ratio as the heat change value of the candidate multimedia data, wherein the types of target operations counted in this application include but are not limited to any one or combination of the following: likes, comments, browsing, forwarding, and collections, etc.
[0144] When calculating the operation heat value corresponding to each time window, the historical occurrence times of each type of target operation can be weighted and superimposed. The size of each weight value set for different operation types is set according to actual processing needs, and this application does not impose specific restrictions on this.
[0145] For example, suppose that when calculating the operation heat value, the target operations considered include: likes, views, and comments; the weight value set for likes is α1, the weight value set for views is α2, and the weight value set for comments is α3. Then, when calculating the operation heat value for a time window, the number of likes, views, and comments that occurred in that time window can be counted to obtain the corresponding number of likes, views, and comments for that time window; the calculated operation heat value is: likes * α1 + views * α2 + comments * α3.
[0146] Furthermore, after the processing device obtains the operation heat values for the two time windows respectively, it can calculate the heat value increment and heat value ratio between the two operation heat values; then, the obtained heat value increment and heat value ratio are determined as the heat change value obtained corresponding to the current candidate multimedia data.
[0147] For example, assuming that the two heat change values determined for two adjacent time windows are H1 and H2 respectively, then the calculated heat change values are H2-H1 and H2 / H1.
[0148] In this way, by starting from two adjacent time windows, counting the changes in the operation heat values in the two time windows, and obtaining the heat change value, it is possible to effectively describe the heat change between the two time windows, quantify the degree of heat change, and provide a selection basis for the selection of target multimedia data.
[0149] Step 402: The processing device performs a quantitative comparison of the operation heat based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, and obtains a heat comparison value between the candidate multimedia data and each historical multimedia data, wherein each historical multimedia data and the candidate multimedia data are associated with the same publishing object.
[0150] In an embodiment of the present application, when the processing device executes step 402, after determining the candidate multimedia data and each historical multimedia data associated with the same publishing object, it can obtain a heat comparison value by calculating the cumulative heat difference between the candidate multimedia data and each historical multimedia data; or, it can obtain a heat comparison value by calculating the heat difference between the candidate multimedia data and each historical multimedia data in the same historical period.
[0151] In some feasible implementations, when obtaining a heat comparison value by calculating the cumulative heat difference between the candidate multimedia data and each historical multimedia data, the processing device can count the cumulative number of occurrences of each type of target operation based on the historical operation data of the candidate multimedia data, and obtain an operation heat value obtained by weighted fusion of each cumulative number of occurrences; then count the cumulative number of occurrences of each type of target operation based on the historical operation data of each historical multimedia data, and obtain an operation heat value obtained by weighted fusion of each cumulative number of occurrences; and, based on the operation heat value of each historical multimedia data, calculate the average operation heat of each historical multimedia data; then, calculate the heat value increment and heat value ratio between the operation heat value of the candidate multimedia data and the operation heat average; and, use the heat value increment and heat value ratio as the heat comparison value between the candidate multimedia data and each historical multimedia data.
[0152] Specifically, the processing device can obtain all historical operation data of the candidate multimedia data, and determine the cumulative number of occurrences of various target operations based on all historical operation data, and weightedly determine the operation heat value of the candidate multimedia data based on each cumulative number of occurrences.
[0153] Similarly, the processing device performs the following operations for each historical multimedia data item: Based on all historical operation data for a historical multimedia data item, the device determines the cumulative number of occurrences of various target operations, and then weights the operation heat value of the historical multimedia data item based on each cumulative number of occurrences. Subsequently, the average operation heat value of each historical multimedia data item is calculated to obtain the average operation heat value of each historical multimedia data item. Furthermore, the heat value increment and heat value ratio between the operation heat value of the candidate multimedia data item and the average operation heat value of each historical multimedia data item can be determined as the heat comparison value between the candidate multimedia data item and each historical multimedia data item.
[0154] For example, assuming that the calculated operation heat value is H3 and the operation heat average is H4; then, the heat value increment is H3-H4, and the heat value ratio is H3 / H4.
[0155] In this way, by calculating the heat comparison value, the heat difference between the candidate multimedia data and the historical multimedia data published by the same publishing object can be quantified, and the heat anomaly of the candidate multimedia data compared with the historical multimedia data can be effectively indicated.
[0156] In other feasible implementations, when obtaining a heat comparison value by calculating the heat difference between the candidate multimedia data and each historical multimedia data in the same historical period, the processing device can determine the same period time interval for the candidate multimedia data and each historical multimedia data respectively; then perform the following operations for the candidate multimedia data and each historical multimedia data respectively: count the number of occurrences of each target operation in the corresponding same period time interval to obtain the statistical results of each operation; obtain the corresponding same period operation heat based on the statistical results of each operation; obtain the historical average heat based on the same period operation heat of each historical multimedia data; determine the heat value increment and heat value ratio between the same period operation heat of the candidate multimedia data and the historical average heat as the corresponding heat comparison value.
[0157] Specifically, when determining the concurrent time intervals for candidate multimedia data and each historical multimedia data, the processing device may determine the concurrent time intervals for the candidate multimedia data based on the time span between the current time and the release time of the candidate multimedia data, combined with a preset duration; thereafter, for each historical multimedia data associated with the same release object as the candidate multimedia data, determine each concurrent time interval for each historical multimedia data based on the release time of each historical multimedia data, combined with the time span and the preset duration.
[0158] It should be noted that in the embodiment of the present application, the value of the preset duration is set according to actual processing needs, and the preset duration is used to indicate the interval duration of the concurrent time interval; in addition, the concurrent time interval is used to indicate the time range corresponding to "concurrent" when determining the concurrent status of each candidate multimedia data and the candidate multimedia data, wherein each concurrent time interval has the same interval length.
[0159] For example, see Figure 4C As shown, it is a schematic diagram of the process of determining the same period time interval in the embodiment of the present application. Figure 4C As can be seen from the illustrated process, assuming that the time span between the release time of the candidate multimedia data and the current time is DT, and the preset duration is L; then, for the historical multimedia data 1 associated with the same release object as the candidate multimedia data, assuming that the release time of the historical multimedia data 1 is A1, the corresponding time interval of the historical multimedia data is [A1+DT-L, A1+DT].
[0160] In this way, by determining each synchronization time interval for the candidate multimedia data and each historical multimedia data, it is possible to determine a suitable time range for the heat comparison process, making the final heat comparison result more referenceable; moreover, after determining the heat comparison value based on the heat value difference between the operation heat of the same period and the historical average heat, it is possible to effectively measure the heat difference within a certain period of time after the multimedia data is released for the candidate multimedia data and the historical multimedia data.
[0161] Step 403: When the processing device determines that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data is used as the target multimedia data and an evaluation indication is triggered.
[0162] Specifically, after obtaining the heat change value and heat comparison value of the corresponding candidate multimedia data, when the heat change value specifically includes: the heat value increment and heat value ratio determined for the operation heat change of the candidate multimedia data in two adjacent time windows, and the heat comparison value includes: the heat value increment and heat value ratio determined for the candidate multimedia data and each historical multimedia data published by the same publishing object, the preset heat anomaly screening condition can be: the heat value increment in the heat change value reaches the first set value, the heat value ratio in the heat change value reaches the second set value, the heat value increment in the heat comparison value reaches the third set value, and the heat value ratio in the heat comparison value obtains the fourth set value, wherein the values of the first set value, the second set value, the third set value, and the fourth set value are set according to actual processing needs, and this application does not impose specific restrictions on this.
[0163] Afterwards, when the processing device determines that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data can be used as the target multimedia data and trigger an evaluation indication.
[0164] In this way, with the help of the processing process of steps 401-403, it is possible to screen each candidate multimedia data by calculating the heat anomaly of the candidate multimedia data, which can reduce the number of target multimedia data that need to be evaluated for traffic anomaly to a certain extent, and help improve the utilization rate of evaluation resources.
[0165] Furthermore, after the processing device determines the target multimedia data and triggers an evaluation indication for the target multimedia data, it obtains content description text and N groups of comment texts associated with the target multimedia data in response to the evaluation indication triggered for the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from the comment text set associated with the target multimedia data.
[0166] Specifically, for the content description text associated with the target multimedia data, different types of multimedia data may correspond to different ways of constructing the content description text, wherein the content description text is used to describe the content in the multimedia data.
[0167] When the multimedia data is video data, a content description text of video data generally includes at least one item or combination of the following: the title of the video data, the label added by the publisher to the video data, and the description content obtained for the video data using a large-scale language model, wherein the large-scale language model can be a generative pre-trained transform model (Chat Generative Pre-trained Transformer, ChatGPT); when the multimedia data is audio data, a content description text of audio data generally includes at least one item or combination of the following: the title of the audio data, the label added by the publisher to the audio data, and the text content obtained after speech recognition of the audio data, wherein this application does not impose specific restrictions on the speech recognition algorithm adopted, for example, the speech recognition algorithm adopted can be an automatic speech recognition algorithm (Automatic Speech Recognition, ASR) etc.; when the multimedia data is text data, a content description text of text data generally includes at least one item or combination of the following: the title of the text data, the label added by the publisher to the text data, and the content of the text data itself.
[0168] For example, see Figure 5 As shown, it is a schematic diagram of the process of constructing description content for video data in an embodiment of the present application. Figure 5As shown in the illustrated content, in the process of using the big prediction model to generate description content for video data, the subtitle content extracted from the video data, K video screenshots captured from the video data, and the question content constructed for the video data are used as inputs of the big language model, and then the description content output by the big language model for the video data can be obtained.
[0169] When obtaining N groups of comment texts for target multimedia data, a random sampling method without replacement can be adopted, and N random samplings without replacement are performed in the comment text set associated with the target multimedia data, wherein the number of comment texts extracted in each random sampling without replacement can be the same or different, and this application does not impose specific restrictions on this; the comment text set associated with the target multimedia data contains all comment texts published corresponding to the target multimedia data.
[0170] Step 302: The processing device performs the following operations for each group of comment texts: obtains each comment text contained in a group of comment texts, and obtains the comment relevance corresponding to the group of comment texts based on the semantic similarity between each comment text, and obtains the content matching degree between a group of comment texts and the target multimedia data based on the semantic similarity between each comment text and the content description text.
[0171] In an embodiment of the present application, after the processing device obtains the content description text and N groups of comment texts for the target multimedia data, it can calculate the content relevance between the comment texts contained in each group of comment texts, and calculate the content relevance between each comment text and the content description text, wherein the content relevance between the comment texts is measured by the comment relevance; the content relevance between the comment texts and the content description text is measured by the content matching degree.
[0172] The following description only uses the example of obtaining comment relevance and content matching for a set of comment texts to illustrate the relevant processing process:
[0173] In an embodiment of the present application, when obtaining the relevance of comment texts, the processing device obtains the comment relevance corresponding to the group of comment texts based on the semantic similarity between the comment texts after obtaining each comment text contained in the group of comment texts.
[0174] Specifically, the processing device can input each text feature obtained corresponding to each comment text into a trained first evaluation model, and obtain a first prediction value corresponding to the heat cheating behavior after the first evaluation model performs content relevance detection on each text feature, wherein the first evaluation model is obtained by performing multiple rounds of iterative training based on the first training sample set, and a first training sample includes: each sample comment associated with a sample multimedia data, and a sample label for identifying whether there is heat cheating behavior; and then the first prediction value is used as the comment relevance corresponding to a group of comment texts.
[0175] It should be noted that in the embodiment of the present application, in the process of obtaining each text feature corresponding to each comment text, the processing device can perform the following operations for each comment text: use a preset word segmentation method to perform word segmentation on the comment text to obtain each word segmentation content corresponding to the comment text, and then use the trained word vector model (word embedding, Word2Vec) to obtain each word vector corresponding to each word segmentation content, and then perform average pooling on each word vector to obtain the corresponding average pooling result, and determine the average pooling result as the text feature of the comment text.
[0176] Among them, the preset word segmentation method can be Jieba word segmentation. During the training process of the Word2Vec model, the model training is performed in an unsupervised training method based on the sample content after word segmentation processing. This application does not elaborate on this.
[0177] In addition, in the embodiment of the present application, for multimedia data with traffic anomalies, since there are illegal groups that provide brushing services among the audience of the multimedia data, there is no correspondence between most of the comment texts, and the overall content quality is low; while the audience of multimedia data without traffic anomalies is normal users, so the correlation between the comment texts is stronger; based on this, the processing device constructs a first initial model and trains the first initial model to obtain a first evaluation model, so that with the help of the first evaluation model, the correlation between the comment texts can be evaluated. Moreover, in the case where 1 is used as the sample prediction result for the presence of heat cheating behavior and 0 is used as the sample prediction result for the absence of heat cheating behavior, the larger the value of the first prediction value output by the first evaluation model, the stronger the content irrelevance of the sample comments input at the same time, and the greater the probability of brushing comments in the corresponding sample multimedia data.
[0178] Among them, the first initial model can be constructed based on the extreme gradient boosting algorithm (eXtreme Gradient Boosting, XGBoost), or based on the random forest algorithm, or based on a model structure that can realize the heat anomaly assessment function. This application does not impose specific restrictions on this.
[0179] In the process of training to obtain the first evaluation model, the first initial model is iteratively trained for multiple rounds with the help of the first training sample set until the preset convergence condition is met, and the first evaluation model obtained by training the first initial model is obtained, wherein the loss function used in the process of training the first evaluation model can be a cross-entropy loss function; the preset convergence condition can be: the number of training rounds reaches a first threshold value, or the number of times the loss function is continuously lower than the second threshold value reaches a preset third threshold value, wherein the values of the first threshold value, the second threshold value, and the third threshold value are set according to actual processing needs, and this application does not impose specific restrictions on this.
[0180] In an embodiment of the present application, in order to construct a training sample set, the processing device may first collect sample multimedia data that is determined to have traffic anomalies within a specified time length, and collect sample multimedia data that does not have traffic anomalies, wherein being determined to have traffic anomalies can be regarded as the presence of heat cheating behavior, and the value of the specified time length is set according to actual processing needs, and the present application does not impose specific restrictions on this; the quantity ratio between the sample multimedia data with traffic anomalies and the sample multimedia data without traffic anomalies is set according to actual processing needs, such as the quantity ratio can be 1:1. Afterwards, the collected sample multimedia data can be divided according to training needs. For example, there are two models that need to be trained in the present application, so the collected sample multimedia data can be divided into two groups, one group for generating the first training sample set and one group for generating the second training sample set.
[0181] For example, if the specified time length is 3 months, the multimedia data that has been determined to have traffic anomalies in the last three months can be obtained as sample multimedia data with heat cheating behavior; in addition, when obtaining sample multimedia data without heat cheating behavior, random sampling can be performed from the maintained multimedia data without traffic anomalies.
[0182] Afterwards, when generating the first training sample set, the processing device performs the following operations for each sample multimedia data: from the comment text set associated with the sample multimedia data, L groups of sample comments are extracted by random sampling without replacement, and the sample labels corresponding to the sample multimedia data for indicating whether there is heat cheating behavior are determined, and the L groups of sample comments and sample labels are combined respectively to obtain L first training samples, wherein the total number of sample comments included in each group of sample comments can be kept the same as the total number of comment texts in each group of comment texts extracted for the target multimedia data.
[0183] For example, the processing device can treat multimedia data that is judged to have traffic anomalies as positive sample multimedia data, and configure the sample label to 1, indicating the presence of heat cheating behavior; at the same time, multimedia data that does not have traffic anomalies can be treated as negative sample multimedia data, and the sample label can be configured to 0, indicating the absence of traffic cheating behavior. Afterwards, when a first training sample is input into the model, the features of each sample comment corresponding to each sample comment are first obtained through word segmentation, word vectorization, and average pooling processing, and then the splicing result obtained by splicing the features of each sample comment is input into the first initial model. The process of obtaining the features of each sample comment corresponding to each sample comment and the process of obtaining the features of each text corresponding to each comment text of the target multimedia data adopt the same processing method, and this application does not impose specific restrictions on this.
[0184] For another example, assuming that the feature length of the sample comment feature is P, and the total number of sample comments included in a group of sample comments is Q, then the feature length of the splicing result is P*Q.
[0185] It should be noted that in the embodiments of the present application, according to actual processing needs, after sampling to obtain each sample comment group, each sample comment in each sample comment group can be vectorized separately to obtain a splicing result corresponding to a sample comment group; or, when reading the sample comment group for training, each sample comment feature can be obtained corresponding to each sample comment included, and the sample comment features can be spliced to obtain a splicing result. The present application does not impose specific restrictions on the timing and method of text vectorization.
[0186] In this way, the first evaluation model can be used to evaluate the correlation between the input sample comments, and the output first prediction value of the popularity cheating behavior can be used to express the content quality of each sample comment input at the same time as a whole.
[0187] In an embodiment of the present application, when obtaining the content matching degree, the processing device obtains each comment text contained in a group of comment texts, and then obtains the content matching degree between a group of comment texts and the target multimedia data based on the semantic similarity between each comment text and the content description text.
[0188] Specifically, the processing device can input the text features obtained corresponding to each comment text and content description text into a trained second evaluation model to obtain the second evaluation model to perform content relevance detection on each text feature, and output a second prediction value corresponding to the popularity cheating behavior, wherein the second evaluation model is obtained by performing multiple rounds of iterative training based on the second training sample set, and a second training sample includes: each sample comment and sample description text associated with the sample multimedia data, and a sample label for identifying whether there is popularity cheating behavior; then, the second prediction value is used as the content matching degree corresponding to a group of comment texts.
[0189] It should be noted that in the embodiment of the present application, the processing method used to obtain text features for the corresponding text is the same as the method of obtaining text features for each comment text when the first evaluation model is used for processing. The present application does not impose any specific restrictions on this.
[0190] Furthermore, in the embodiments of the present application, for multimedia data with abnormal traffic, since illegal groups typically do not pay attention to the specific content of the multimedia data when performing traffic manipulation, there is usually no correlation between the associated comment text and the content of the multimedia data, while the correlation between normal comment text and the content of the multimedia data is relatively large. Based on this, a second evaluation model can be obtained through training to describe the correlation between the comment text and the content of the multimedia data.
[0191] Among them, the second evaluation model is specifically obtained by training based on the constructed second initial model. The second initial model can be constructed based on the XGBoost algorithm, or based on the random forest algorithm, or based on a model structure that can realize the heat anomaly evaluation function.
[0192] In the process of training to obtain the second evaluation model, the second initial model is iteratively trained for multiple rounds with the help of the second training sample set until the preset convergence condition is met, and the second evaluation model obtained by training the second initial model is obtained, wherein the loss function used in the process of training the second evaluation model can be a cross-entropy loss function; the preset convergence condition can be: the number of training rounds reaches a fourth threshold value, or the number of times the loss function is continuously lower than the fifth threshold value reaches a preset sixth threshold value, wherein the values of the fourth threshold value, the fifth threshold value, and the sixth threshold value are set according to actual processing needs, and this application does not impose specific restrictions on this.
[0193] In an embodiment of the present application, when constructing a second training sample set, the processing device can perform the following operations for each sample multimedia data used to generate the second training sample set: from the comment text set associated with the sample multimedia data, L groups of sample comments are extracted by random sampling without replacement; the sample label corresponding to the sample multimedia data for indicating whether there is heat cheating behavior is determined, and the sample description text associated with the sample multimedia data is obtained; the L groups of sample comments are respectively combined with the sample label and the sample description text to obtain L second training samples, wherein the total number of sample comments included in each group of sample comments can be kept the same as the total number of comment texts in each group of comment texts extracted for the target multimedia data.
[0194] For example, the processing device can treat multimedia data that is judged to have traffic anomalies as positive sample multimedia data, and configure the sample label to 1, indicating the presence of popularity cheating behavior; at the same time, multimedia data that does not have traffic anomalies can be treated as negative sample multimedia data, and the sample label can be configured to 0, indicating the absence of traffic cheating behavior. Afterwards, assuming that the sample description content includes title text, label text, and description content, before inputting a second training sample into the model, through word segmentation, word vectorization processing, and average pooling processing, each sample comment feature corresponding to each sample comment is obtained, the title feature corresponding to the title text is obtained, the label feature corresponding to the label text is obtained, and the description content feature corresponding to the description content is obtained. Furthermore, the title feature, label feature, description content feature, and each sample comment feature are spliced to obtain a splicing result, and the splicing result is then input into the second initial model to obtain a second prediction result, wherein the process of obtaining text features for the corresponding text and the process of obtaining text features for each comment text corresponding to the target multimedia data adopt the same processing method, and this application does not impose specific restrictions on this.
[0195] For another example, assuming that for sample comment features, title features, tag features, and description content features, the feature lengths of the four features are all P, and the total number of sample comments included in a group of sample comments is Q, then the feature length of the splicing result is (Q+3)*P.
[0196] Similar to the above-mentioned process of training to obtain the first evaluation model, this application does not impose specific restrictions on the timing and method of text vectorization, nor does it impose specific restrictions on the timing of constructing the splicing result.
[0197] In this way, with the help of the second evaluation model, the correlation between the input sample comments and the descriptive content of the multimedia data can be evaluated, and with the help of the output second prediction value of the heat cheating behavior, the content correlation degree between the sample comments input at the same time and the data content of the sample multimedia data can be expressed as a whole.
[0198] Step 303: The processing device calculates the degree of abnormality based on the N comment relevances and the N content matching degrees to obtain an abnormality evaluation value corresponding to the target multimedia data.
[0199] In some feasible implementations of the present application, after obtaining N comment relevances and N content matching degrees, the processing device can perform weighted superposition of the mean of the N comment relevances and the mean of the N content matching degrees to obtain an abnormality assessment value.
[0200] In some other feasible implementations of the present application, after the processing device obtains the content description text and N groups of comment texts associated with the target multimedia data, it can also evaluate the inclusion of each abnormal word in the content description text and N groups of comment texts from the perspective of the hit situation of abnormal words, and obtain the abnormal word hit degree corresponding to the target multimedia data; based on this, when calculating the abnormal evaluation value, data fusion calculation can be performed according to the N comment relevances, N content matching degrees, and the abnormal word hit degree of the target multimedia data to obtain the abnormal evaluation value of the target multimedia data.
[0201] Specifically, the average of the relevance of N comments, the average of the content matching degrees, and the abnormal word hit degree of the target multimedia data can be weighted and superimposed to obtain an abnormal evaluation value, wherein the values of the weights involved are set according to actual processing needs, and the sum of the weights is 1.
[0202] In this way, by additionally adding the influence of the abnormal word hit rate when calculating the abnormal evaluation value, more factors that can be considered can be integrated when determining the abnormal evaluation value.
[0203] When calculating the hit rate of abnormal words, the processing device can obtain various abnormal words constructed in advance for traffic cheating behavior; then determine the number of hits for each abnormal word in the content description text and N groups of comment texts; thereafter, determine the hit rate of the abnormal words corresponding to the target multimedia data based on the number of hits of each abnormal word and the abnormal indicators determined in advance for each abnormal word, wherein the abnormal indicator of an abnormal word is determined by calculating the inclusion of an abnormal word in multimedia data with and without traffic cheating behavior.
[0204] In an embodiment of the present application, when the processing device determines each abnormal word and the corresponding abnormal index, it can obtain a comment text set associated with each sample multimedia data, wherein each sample multimedia data includes: positive sample multimedia data associated with heat cheating behavior, and negative sample multimedia data not associated with heat cheating behavior; then, from each sample comment associated with heat cheating behavior, word segmentation is performed to obtain each candidate word; thereafter, the following operations are performed for each candidate word: based on the sample comment set associated with each positive sample multimedia data and the sample comment set associated with each negative sample multimedia data, the inclusion difference of the candidate word is calculated to obtain the abnormal index of the candidate word; when it is determined that the abnormal index exceeds the set value, the candidate word is determined as an abnormal word.
[0205] It should be noted that the sample multimedia data used to calculate the abnormality indicators can be the sample multimedia data constructed by this application to train the first evaluation model and the second evaluation model, or can be the sample multimedia data collected additionally using the same data collection method.
[0206] When calculating the anomaly index, the processing device may calculate the corresponding target group index (TGI) for each candidate word, and determine the TGI index calculated for the candidate word as the anomaly index of the candidate word. Taking the calculation of the TGI index for candidate word i as an example, the calculation formula involved is as follows:
[0207] TGI index = [the proportion of groups in the target group that contain candidate word i / the proportion of groups in the overall population that contain candidate word i] * 100
[0208] The target group is each sample multimedia data that has cheating behavior, and the overall group is all sample multimedia data.
[0209] It should be understood that the higher the TGI value of candidate word i, the greater the hit probability of candidate word i in the sample multimedia data with heat cheating behavior, the lower the possibility that the comment text containing the candidate word i is a normal comment, and the higher the possibility that it is a fake comment.
[0210] In this way, we can determine the scope of abnormal words by word segmentation from each sample comment associated with traffic cheating behavior, and by calculating the occurrence of candidate words with potential anomalies in positive sample multimedia data and negative sample multimedia data, we can evaluate the abnormality of the candidate words, realize the screening of abnormal words and the determination of abnormal indicators.
[0211] In addition, when the processing device calculates the hit degree of abnormal words, in some feasible implementation methods, the hit degree of abnormal words can be calculated by calculating the number of hit abnormal words and combining the abnormal indicators of each hit abnormal word; in other feasible implementation methods, the hit degree of abnormal words can be calculated by calculating the number of hits of each abnormal word and combining the abnormal indicators of each abnormal word.
[0212] For example, assuming there are 10 abnormal words, we can only count the number of abnormal words that are hit. For example, if three abnormal words are hit, we can add up the abnormal indicators of these three abnormal words to get the abnormal word hit degree.
[0213] For another example, assuming there are 10 abnormal words, the number of hits for each abnormal word can be counted, for example, abnormal word 1 hits 5 times, and abnormal word 2 hits 3 times; then, assuming the abnormal index of abnormal word 1 is v1, and the abnormal index of abnormal word 2 is v2, then the abnormal word hit rate is: 5*v1+3*v2.
[0214] In this way, by calculating the hit rate of abnormal words, it is possible to determine the inclusion of abnormal words in the comment text for the target multimedia data, which is equivalent to providing more considerations for the traffic anomaly assessment of the target multimedia data.
[0215] Step 304: When the processing device determines that the abnormality evaluation value exceeds the set threshold, it is determined that traffic cheating behavior exists in the target multimedia data.
[0216] Specifically, the processing device can determine, based on the relationship between the abnormal evaluation value and a set threshold, that a traffic anomaly exists in the target multimedia data when the abnormal evaluation value exceeds the set threshold. The processing device can then obtain pre-set cheating handling strategies at each level, each of which has an associated evaluation value range. The processing device can then determine a target evaluation value range corresponding to the abnormal evaluation value based on whether the abnormal evaluation value falls within each evaluation value range. The processing device then processes the target multimedia data using the cheating handling strategy corresponding to the target evaluation value range.
[0217] It should be noted that, according to actual processing needs, in order to limit the value range of the abnormal evaluation value, after obtaining the abnormal word hit degree, the abnormal word hit degree can be normalized to a value less than 1. In addition, optionally, after determining that there is a traffic anomaly based on the abnormal evaluation value, you can jump to manual verification to increase the flexibility of processing. In the specific manual verification process, based on the manual verification results, the number of target multimedia data that are manually verified as having traffic anomalies can be determined in each target multimedia data under different evaluation value intervals, and the proportion of the total number of target multimedia data under the corresponding evaluation value interval. The accuracy of the recognition results involved in this application can be evaluated based on the obtained proportion results, and when it is determined that the recognition accuracy of this application is insufficient based on the obtained proportion results, the model used to obtain comment relevance and content matching can be retrained, and abnormal words can be re-screened.
[0218] For example, see Figure 6 As shown, it is a schematic diagram of the process of determining the cheating processing strategy in the embodiment of the present application. Figure 6As can be seen from the schematic, the processing device can predefine a correspondence between anomaly evaluation value intervals and cheating handling strategies, such that larger anomaly evaluation values correspond to more stringent cheating handling strategies. Based on this, after obtaining the anomaly evaluation value of the target multimedia data, the matching anomaly handling strategy can be determined based on the correspondence between the evaluation value intervals and the cheating handling strategies.
[0219] Optionally, in a feasible implementation, for the processing of the publishing object, the total number of multimedia data with traffic abnormalities can be counted among the multimedia data published by the publishing object, and different degrees of processing can be performed on the publishing object based on different value ranges of the total number of multimedia data, and after the total number of multimedia data reaches the maximum set value, the publishing object's account can be banned.
[0220] In this way, from the perspective of content, target multimedia data with traffic anomalies can be detected from massive multimedia data, and different degrees of restrictions can be imposed on the target multimedia data according to the degree of anomaly, thereby avoiding the dissemination of some low-quality multimedia data, improving the overall multimedia data publishing environment, and helping to maintain the fairness of multimedia data dissemination.
[0221] The following, combined with the accompanying figures, uses the example of traffic anomaly assessment for published video works to schematically illustrate the feasible processing process:
[0222] See Figure 7A As shown, it is a schematic diagram of the timing of evaluating flow anomaly in the embodiment of the present application. Figure 7A As shown in the following content, the current process of traffic anomaly is as follows:
[0223] Step 1: The publisher publishes the video work;
[0224] Step 2: The publisher purchases a fraudulent traffic-boosting service on an illegal fraudulent traffic-boosting platform in order to increase the popularity of the video.
[0225] Step 3: The publisher provides the link of the video work to be inflated to the inflated volume platform;
[0226] Step 4: The brushing team performs batch operations on video works;
[0227] Step 5: The brushing team completes the brushing, and the popularity of the video work is increased.
[0228] It should be understood that for works that are inflated, the normal audience of the video works is relatively small, and most of the "audience" is the inflated group. Therefore, usually in the process of batch generation of video works by the inflated team described in step 4, the video audience generates a large amount of operation data for the video works; based on this, when detecting the inflated behavior of video works, it is possible to identify the inflated operation after step 4 occurs, and realize the evaluation of traffic anomalies. Among them, for a video work, when the video work is subjected to the inflated operation, the video work simultaneously has popularity cheating behavior, traffic cheating behavior, traffic anomaly, and inflated behavior.
[0229] See Figure 7B As shown in the figure, it is a schematic diagram of the process of evaluating traffic anomalies for video works in the embodiment of the present application. Figure 7B , describe the feasible evaluation process:
[0230] Step 701: The processing device collects various sample video works, wherein the sample video works include positive sample video works with brushing behavior and negative sample video works without brushing behavior.
[0231] Among them, taking into account the changes in the means used by the inflating groups, when collecting positive sample video works and negative sample video works, priority can be given to video works in the most recent period, such as collecting samples from video works in the past three months; the numerical ratio between the total number of positive sample videos and the total number of negative sample works in each sample video work conforms to the preset ratio value.
[0232] Step 702: The processing device obtains corresponding content data for each sample video work.
[0233] Among them, content data includes: content generated by publishing object operations, namely the title of the video work, added tags, and the video work itself; content generated by video terminal operations, namely comments on the video work; and descriptive content generated based on the video work.
[0234] It should be noted that the description content can be the content output by the large language model after the subtitles of the video work, video stream screenshots, and prompts generated based on experience are input into the large language model.
[0235] Step 703: The processing device uses the content data of each sample video work to train a word vector model, and obtains a feature set associated with each sample video work based on the word vector training results, wherein the feature set includes title features, tag features, various comment features, and descriptive content features.
[0236] In the process of training the word vector model, the content data of each sample video work can be used to perform unsupervised training on the word vector model until a preset number of iterations is reached, so that the word vector training results corresponding to each word segmentation of the word vector model can be obtained.
[0237] It should be understood that in the embodiment of the present application, the text content in the content data is spliced and combined to obtain text content sentences in order to obtain good training results when training the word vector model in an unsupervised training manner.
[0238] See Figure 7C As shown in the figure, it is a schematic diagram of the process of training the word vector model in the embodiment of the present application. Figure 7C As shown in the content, when the content data includes work titles, work labels, work comments, and work description content, the work title, work labels, work comments, and work description content of a sample video work are spliced to obtain text content sentences, and then the word vector model is unsupervisedly trained with the help of the text content sentences after word segmentation to obtain the word vector training results under each word segmentation.
[0239] Furthermore, when finally generating the sentence vector, the title features of the work title are obtained based on the various segmentations covered by the work title content and the word vector training results of the covered segmentations; similarly, the label features of the work label are obtained based on the various segmentations covered by the work label and the word vector training results of the covered segmentations; similarly, the comment features of each comment sentence are obtained based on the various segmentations covered by each comment sentence and the word vector training results of the covered segmentations; similarly, the description content features of the work description content are generated based on the various segmentations covered by the work description content and the word vector training results of the covered segmentations.
[0240] For example, see Figure 7D As shown, it is a schematic diagram of the process of determining the features of each sentence in the embodiment of the present application, combined with the attached Figure 7D As can be seen from the schematic, when determining the title features for the title sentence, the title features are obtained after average pooling processing based on the two word vectors covered by the title sentence; similarly, the label features are obtained for the corresponding label sentences, the description content features can be obtained for the corresponding work description content, and for each comment sentence, the comment features after average pooling processing can be obtained based on the word vector training results corresponding to each word segment contained in the comment sentence.
[0241] Optionally, the processing device can take the positive sample video works as the target group in each sample video work, and count the corresponding TGI index for each comment segmentation associated with the target group, and based on the TGI index of each comment segmentation, screen out abnormal words in each comment segmentation, and determine the TGI index corresponding to the abnormal word as the associated abnormal indicator.
[0242] For example, see Figure 7E As shown, it is a schematic diagram of the process of determining the TGI index of the comment segmentation in the embodiment of the present application. Figure 7E As can be seen from the schematic content, the normal work comments and the fake work comments can be spliced together to obtain text content sentences, and then the text content sentences can be segmented, as well as the segmented comments included in the comments of the target group to calculate the corresponding TGI index.
[0243] Step 704: The processing device constructs training sample sets corresponding to the first initial model and the second initial model based on each sample video work and the associated feature set, trains the first initial model to obtain a first evaluation model, and trains the second initial model to obtain a second evaluation model.
[0244] Among them, when constructing the first training sample set corresponding to the first initial model, taking the processing of a sample video work as an example, L comment feature groups are extracted from the various comment features of a sample video work by random sampling without replacement, and then the L comment feature groups are respectively combined with the sample labels of the sample video work to obtain L first training samples, wherein the sample labels are used to indicate whether the sample video work has popularity cheating behavior.
[0245] When constructing the second training sample set corresponding to the second initial model, taking the processing of a sample video work as an example, Y comment feature groups are extracted from the various comment features of a sample video work by random sampling without replacement, and then the Y comment feature groups are respectively combined with the sample label, title feature, label feature, and descriptive content feature of the sample video work to obtain Y second training samples, where the sample label is used to indicate whether the sample video work has popularity cheating behavior.
[0246] Further, when evaluating traffic anomalies for specific video works, refer to Figure 7F As shown in the figure, it is a schematic diagram of the flow anomaly evaluation process in the embodiment of the present application. Figure 7F , the relevant processing process is explained:
[0247] Step 705: The processing device calculates the popularity change value and popularity comparison value of the massive video works in the publishing state, and determines the video works whose popularity change value and popularity comparison value meet the popularity abnormality screening conditions as target video works.
[0248] Step 706: The processing device determines the abnormal word hit degree for each target video work based on each abnormal word and abnormal index, and uses the first evaluation model to obtain N comment relevances, and uses the second evaluation model to obtain N content matching degrees.
[0249] Step 707: The processing device obtains a corresponding abnormality evaluation value based on the abnormal word hit degree, N comment relevances, and N content matching degrees of each target video work, and determines a corresponding handling strategy based on the abnormality evaluation value of each target video work.
[0250] Taking the processing process for a target video work as an example, the severity of the abnormality of the target video work can be determined according to the value of the abnormality assessment value, and matching processing measures can be adopted according to the severity of the abnormality.
[0251] In summary, the technical solution proposed in this application, starting from the perspective of video content, by constructing the first evaluation model and the second evaluation model for realizing brush volume detection, can determine whether it is a brush volume work by accumulating only dozens of comment data without accumulating a large amount of historical data; in addition, when faced with a video work that suddenly becomes popular, since the popularity of the video work comes from normal traffic, the related comments and other content are also normal, so the recognition process involved in this application will not accidentally harm the normal popular video work; furthermore, since it is difficult to forge the comment content that is relevant to the video, it is difficult for the brush volume group to bypass the recognition strategy of this application. Therefore, the recognition method proposed in this application can effectively detect brush volume works and the publishers who purchase brush volume services.
[0252] Moreover, compared with the traffic anomaly assessment based solely on popularity changes, the assessment method adopted in this application can more quickly identify works with traffic anomalies. Moreover, since it is more difficult to forge comments that are consistent with the content of the work, when evaluating traffic anomalies by detecting false comment content, a good assessment effect can be obtained, which is difficult to be circumvented by brushing groups.
[0253] Based on the same inventive concept, see Figure 8 As shown, it is a schematic diagram of the logical structure of the flow anomaly assessment device in an embodiment of the present application. The flow anomaly assessment device 800 includes a response unit 801, an execution unit 802, an acquisition unit 803, and a determination unit 804, wherein,
[0254] A response unit 801 is configured to, in response to an evaluation indication triggered for target multimedia data, obtain content description text and N groups of comment texts associated with the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from a set of comment texts associated with the target multimedia data;
[0255] Execution unit 802 is configured to perform the following operations for each group of comment texts: obtaining each comment text contained in the group of comment texts; performing content relevance detection on each comment text to obtain comment relevance corresponding to the group of comment texts; performing content relevance detection on each comment text and the content description text to obtain content matching between the group of comment texts and the target multimedia data;
[0256] An obtaining unit 803 is configured to calculate an abnormality degree based on the N comment relevances and the N content matching degrees to obtain an abnormality evaluation value corresponding to the target multimedia data;
[0257] The determination unit 804 is configured to determine that traffic cheating behavior exists in the target multimedia data when the abnormality evaluation value exceeds a set threshold.
[0258] Optionally, before responding to the evaluation instruction triggered for the target multimedia data, the apparatus further includes a screening unit 805, the screening unit 805 being configured to:
[0259] Obtain each candidate multimedia data, and for each candidate multimedia data, perform the following operations:
[0260] Based on the historical operation data of the candidate multimedia data in two adjacent time windows, a quantitative comparison of the operation heat is performed to obtain the heat change value of the candidate multimedia data;
[0261] Based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, a quantitative comparison of the operation heat is performed to obtain a heat comparison value between the candidate multimedia data and each historical multimedia data; each historical multimedia data and the candidate multimedia data are associated with the same publishing object;
[0262] When it is determined that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data is used as the target multimedia data and an evaluation indication is triggered.
[0263] Optionally, when obtaining corresponding popularity change values according to changes in historical operation data in two time windows, the screening unit 805 is configured to:
[0264] For each time window, the following operations are performed: Count the historical occurrences of each target operation in the time window; Perform weighted fusion processing on each historical occurrence according to the preset weight values for each target operation to obtain the operation heat value corresponding to the time window;
[0265] Calculating a heat value increment and a heat value ratio between two operation heat values; and determining the heat value increment and the heat value ratio as a heat change value of the candidate multimedia data.
[0266] Optionally, when obtaining a heat comparison value between the candidate multimedia data and each historical multimedia data based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, the screening unit 805 is configured to:
[0267] According to the historical operation data of the candidate multimedia data, the cumulative occurrence times of various target operations are counted, and the operation heat value obtained by weighted fusion of the cumulative occurrence times is obtained;
[0268] Based on the historical operation data of each historical multimedia data, the cumulative number of occurrences of each target operation is counted to obtain an operation heat value obtained by weighted fusion of the cumulative number of occurrences; and based on the operation heat values, the average operation heat value of each historical multimedia data is calculated;
[0269] Calculate the heat value increment and heat value ratio between the operation heat value of the candidate multimedia data and the operation heat average; and use the heat value increment and heat value ratio as the heat comparison value between the candidate multimedia data and each historical multimedia data.
[0270] Optionally, when obtaining a heat comparison value between the candidate multimedia data and each historical multimedia data based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, the screening unit 805 is configured to:
[0271] Corresponding to the candidate multimedia data and each historical multimedia data, determining a corresponding time interval;
[0272] For the candidate multimedia data and each historical multimedia data, perform the following operations respectively:
[0273] Count the number of occurrences of each target operation in the corresponding concurrent time interval to obtain the statistical results of each operation; obtain the corresponding concurrent operation heat based on the statistical results of each operation; obtain the historical average heat based on the concurrent operation heat of each historical multimedia data; determine the heat value increment and heat value ratio between the concurrent operation heat of the candidate multimedia data and the historical average heat as the corresponding heat comparison value.
[0274] Optionally, after obtaining the content description text and N groups of comment texts associated with the target multimedia data, and before calculating the abnormality degree according to the N comment relevances and the N content matching degrees, the execution unit 802 is further configured to:
[0275] Obtain various abnormal words pre-constructed for traffic fraud behavior;
[0276] In the content description text and N groups of comment texts, determine the number of hits for each abnormal word;
[0277] Based on the number of hits for each abnormal word and the abnormal indicators pre-determined for each abnormal word, the abnormal word hit degree corresponding to the target multimedia data is determined. The abnormal indicator of an abnormal word is determined by calculating the inclusion of an abnormal word in multimedia data with and without traffic cheating behavior.
[0278] Optionally, each abnormal word and the corresponding abnormal indicator are determined by the execution unit 802 in the following manner:
[0279] Obtaining a set of comment texts associated with each sample multimedia data, wherein each sample multimedia data includes: positive sample multimedia data associated with popularity cheating behavior, and negative sample multimedia data not associated with popularity cheating behavior;
[0280] From each sample comment related to traffic fraud, word segmentation is performed to obtain candidate words;
[0281] For each candidate word, the following operations are performed: based on the sample comment sets associated with each positive sample multimedia data and the sample comment sets associated with each negative sample multimedia data, the inclusion difference of the candidate word is calculated to obtain the abnormality index of the candidate word; when it is determined that the abnormality index exceeds the set value, the candidate word is determined as an abnormal word.
[0282] Optionally, when calculating the abnormality degree based on the N comment relevances and the N content matching degrees to obtain the abnormality evaluation value corresponding to the target multimedia data, the obtaining unit 803 is used to:
[0283] According to the relevance of N comments, the matching degree of N contents, and the abnormal word hit degree of the target multimedia data, data fusion calculation is performed to obtain the abnormal evaluation value of the target multimedia data.
[0284] Optionally, after determining that traffic cheating occurs in the target multimedia data, the determining unit 804 is further configured to:
[0285] Obtaining pre-set cheating handling strategies at each level, wherein each cheating handling strategy has an associated evaluation value range;
[0286] Determine the target evaluation value interval corresponding to the abnormal evaluation value based on the inclusion of the abnormal evaluation value in each evaluation value interval;
[0287] The target multimedia data is processed using a cheating processing strategy corresponding to the target evaluation value interval.
[0288] Optionally, when performing content relevance detection on each comment text to obtain comment relevance corresponding to a group of comment texts, the execution unit 802 is configured to:
[0289] Inputting each text feature obtained for each comment text into a trained first evaluation model to obtain a first prediction value corresponding to the popularity cheating behavior output by the first evaluation model after performing content relevance detection on each text feature. The first evaluation model is obtained by performing multiple rounds of iterative training based on a first training sample set. Each first training sample includes: each sample comment associated with a sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs;
[0290] The first prediction value is used as the comment relevance corresponding to a group of comment texts.
[0291] Optionally, when performing content relevance detection based on each comment text and the content description text to obtain a content matching degree between a group of comment texts and the target multimedia data, the execution unit 802 is configured to:
[0292] Inputting the text features obtained for each comment text and content description text into the trained second evaluation model to obtain a second prediction value corresponding to the popularity cheating behavior after the second evaluation model performs content relevance detection on each text feature. The second evaluation model is obtained by performing multiple rounds of iterative training based on a second training sample set. Each second training sample includes: each sample comment and sample description text associated with the sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs;
[0293] The second prediction value is used as the content matching degree corresponding to a group of comment texts.
[0294] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0295] After introducing the flow anomaly assessment method and apparatus according to an exemplary embodiment of the present application, an electronic device according to another exemplary embodiment of the present application will be introduced next.
[0296] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0297] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. Figure 9 FIG. 1 is a schematic diagram of the hardware structure of an electronic device using an embodiment of the present application. In one embodiment, the electronic device may be Figure 2 The processing device 220 shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Figure 9 As shown, it includes a memory 901 , a communication module 903 and one or more processors 902 .
[0298] Memory 901 is used to store computer programs executed by processor 902. Memory 901 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.
[0299] Memory 901 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 901 may be a combination of the above memories.
[0300] The processor 902 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 902 is configured to implement the above-mentioned traffic anomaly assessment method when calling the computer program stored in the memory 901 .
[0301] The communication module 903 is used to communicate with the terminal device and the server.
[0302] The specific connection medium between the memory 901, the communication module 903 and the processor 902 is not limited in the embodiment of the present application. Figure 9 In the embodiment, the memory 901 and the processor 902 are connected via a bus 904. Figure 9 The connections between the other components are shown in bold lines for illustration only and are not intended to be limiting. The bus 904 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 9 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.
[0303] The memory 901 stores a computer storage medium, which stores computer executable instructions. The computer executable instructions are used to implement the flow anomaly assessment method of the embodiment of the present application. The processor 902 is used to execute the above-mentioned flow anomaly assessment method, such as Figure 3A shown.
[0304] In another embodiment, the electronic device may also be other electronic devices, see Figure 10 FIG. 1 is a schematic diagram of the hardware structure of another electronic device using an embodiment of the present application. Specifically, the electronic device may be Figure 2 The terminal device 210 shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Figure 10 As shown, it includes: a communication component 1010, a memory 1020, a display unit 1030, a camera 1040, a sensor 1050, an audio circuit 1060, a Bluetooth module 1070, a processor 1080 and other components.
[0305] The communication component 1010 is used to communicate with the server. In some embodiments, it may include a wireless fidelity (WiFi) module. The WiFi module is a short-range wireless transmission technology. Electronic devices can help users send and receive information through the WiFi module.
[0306] Memory 1020 can be used to store software programs and data. Processor 1080 executes the various functions and data processing of terminal device 210 by running the software programs or data stored in memory 1020. In this application, memory 1020 can store an operating system and various application programs, and can also store computer programs that execute the traffic anomaly assessment method of the embodiment of this application.
[0307] The display unit 1030 may also be used to display information input by a user or provided to a user, as well as a graphical user interface (GUI) of various menus of the terminal device 210. Specifically, the display unit 1030 may include a display screen 1032 disposed on the front of the terminal device 210. The display unit 1030 may be used to display multimedia data, etc., as described in the embodiments of the present application.
[0308] The display unit 1030 can also be used to receive input digital or character information and generate signal input related to user settings and function control of the terminal device 210. Specifically, the display unit 1030 may include a touch screen 1031 set on the front of the terminal device 210, which can collect user touch operations on or near it.
[0309] The touch screen 1031 can be covered on the display screen 1032, or the touch screen 1031 and the display screen 1032 can be integrated to realize the input and output functions of the terminal device 210. The integrated touch screen can be simply called a touch display screen. In this application, the display unit 1030 can display applications and corresponding operation steps.
[0310] Camera 1040 can be used to capture still images, and users can post comments on images captured by camera 1040 through the app. The lens generates an optical image of an object and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to processor 1080 for conversion into a digital image signal.
[0311] The client device may further include at least one sensor 1050, such as an accelerometer 1051, a distance sensor 1052, a fingerprint sensor 1053, and a temperature sensor 1054. The client device may also be equipped with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0312] The audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the terminal device 210. The audio circuit 1060 can convert received audio data into electrical signals and transmit them to the speaker 1061, which then converts the signals into sound signals for output. Meanwhile, the microphone 1062 converts the collected sound signals into electrical signals, which are then received by the audio circuit 1060 and converted into audio data. The audio data is then output to the communication component 1010 for transmission to, for example, another terminal device 210, or to the memory 1020 for further processing.
[0313] The Bluetooth module 1070 is used to exchange information with other Bluetooth devices having a Bluetooth module through the Bluetooth protocol.
[0314] The processor 1080 is the control center of the client device. It uses various interfaces and lines to connect various parts of the entire terminal. It executes various functions of the client device and processes data by running or executing software programs stored in the memory 1020 and calling data stored in the memory 1020. In some embodiments, the processor 1080 may include at least one processing unit; the processor 1080 may also integrate an application processor and a baseband processor. In the present application, the processor 1080 can run an operating system, application programs, user interface display and touch response, as well as the traffic anomaly assessment method of the embodiment of the present application. In addition, the processor 1080 is coupled to the display unit 1030.
[0315] In some possible implementations, various aspects of the flow anomaly assessment method provided in the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to enable the electronic device to perform the steps of the flow anomaly assessment method according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device may perform the following steps: Figure 3A Follow the steps shown in .
[0316] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0317] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, apparatus, or device.
[0318] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.
[0319] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0320] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The computer program can be executed entirely on the user electronic device, partially on the user electronic device, as a separate software package, partially on the user electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (for example, using an Internet service provider to connect through the Internet).
[0321] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.
[0322] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0323] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain a computer-usable computer program.
[0324] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program commands. These computer program commands can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the commands executed by the processor of the computer or other programmable data processing device generate commands for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0325] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0326] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for evaluating flow anomaly, characterized in that: include: In response to an evaluation indication triggered for target multimedia data, obtaining content description text and N groups of comment texts associated with the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from a set of comment texts associated with the target multimedia data; For each group of comment texts, the following operations are performed: obtaining each comment text contained in the group of comment texts; performing content relevance detection based on each comment text to obtain comment relevance corresponding to the group of comment texts; performing content relevance detection based on each comment text and the content description text to obtain content matching between the group of comment texts and the target multimedia data; Calculating the degree of abnormality based on the N comment relevances and the N content matching degrees to obtain an abnormality evaluation value corresponding to the target multimedia data; When it is determined that the abnormality evaluation value exceeds a set threshold, it is determined that a traffic abnormality exists in the target multimedia data.
2. The method according to claim 1, wherein Before the evaluation instruction triggered in response to the target multimedia data, the method further includes: Obtain each candidate multimedia data, and for each candidate multimedia data, perform the following operations: Performing a quantitative comparison of the operation heat according to the historical operation data of the candidate multimedia data in two adjacent time windows to obtain a heat change value of the candidate multimedia data; performing a quantitative comparison of operation heats based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data to obtain a heat comparison value between the candidate multimedia data and each historical multimedia data; each historical multimedia data and the candidate multimedia data are associated with the same publishing object; When it is determined that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data is used as the target multimedia data, and an evaluation indication is triggered.
3. The method according to claim 2, wherein The step of performing a quantitative comparison of the operation heat according to the historical operation data of the candidate multimedia data in two adjacent time windows to obtain the heat change value of the candidate multimedia data includes: For each time window, the following operations are performed: the historical occurrence counts of various target operations in the time window are counted; according to the preset weight values for the various target operations, the historical occurrence counts are weighted and fused to obtain the operation heat value corresponding to the time window; Calculating a heat value increment and a heat value ratio between two operation heat values; and determining the heat value increment and the heat value ratio as a heat change value of the candidate multimedia data.
4. The method according to claim 2, wherein The step of performing a quantitative comparison of operation heats based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data to obtain a heat comparison value between the candidate multimedia data and each historical multimedia data includes: According to the historical operation data of the candidate multimedia data, the cumulative occurrence times of various target operations are counted to obtain the operation heat value obtained by weighted fusion of the cumulative occurrence times; Based on the historical operation data of each historical multimedia data, the cumulative number of occurrences of each type of target operation is counted to obtain an operation heat value obtained by weighted fusion of the cumulative number of occurrences; and based on the operation heat values, the average operation heat value of each historical multimedia data is calculated; Calculate the heat value increment and heat value ratio between the operation heat value of the candidate multimedia data and the operation heat mean; and use the heat value increment and the heat value ratio as the heat comparison value between the candidate multimedia data and each historical multimedia data.
5. The method according to claim 2, wherein The step of performing a quantitative comparison of operation heats based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data to obtain a heat comparison value between the candidate multimedia data and each historical multimedia data includes: Determining corresponding time intervals for the candidate multimedia data and each historical multimedia data; For the candidate multimedia data and each historical multimedia data, the following operations are performed respectively: Count the number of occurrences of each target operation in the corresponding concurrent time interval to obtain statistical results of each operation; obtain the corresponding concurrent operation heat based on the statistical results of each operation; obtain the historical average heat based on the concurrent operation heat of each historical multimedia data; determine the heat value increment and heat value ratio between the concurrent operation heat of the candidate multimedia data and the historical average heat as the corresponding heat comparison value.
6. The method according to claim 1, wherein After obtaining the content description text and N groups of comment texts associated with the target multimedia data, and before calculating the abnormality degree according to the N comment relevances and N content matching degrees, the method further includes: Obtain various abnormal words pre-constructed for traffic fraud behavior; Determine the number of hits for each abnormal word in the content description text and the N groups of comment texts; According to the number of hits of each abnormal word and the abnormal indicators determined in advance for each abnormal word, the abnormal word hit degree corresponding to the target multimedia data is determined, wherein the abnormal indicator of an abnormal word is determined by calculating the inclusion of the abnormal word in the multimedia data with and without traffic cheating behavior.
7. The method according to claim 6, wherein The abnormal words and corresponding abnormal indicators are determined in the following way: Obtaining a comment text set associated with each sample multimedia data, wherein each sample multimedia data includes: positive sample multimedia data associated with popularity cheating behavior, and negative sample multimedia data not associated with popularity cheating behavior; From each sample comment related to traffic fraud, word segmentation is performed to obtain candidate words; For each candidate word, the following operations are performed: based on the sample comment sets associated with each positive sample multimedia data and the sample comment sets associated with each negative sample multimedia data, the inclusion difference of the candidate word is calculated to obtain the abnormality index of the candidate word; when it is determined that the abnormality index exceeds the set value, the candidate word is determined as an abnormal word.
8. The method according to claim 6, wherein The calculating of the abnormality degree according to the N comment relevances and the N content matching degrees to obtain the abnormality evaluation value corresponding to the target multimedia data includes: According to the N comment relevances, the N content matching degrees, and the abnormal word hit degree of the target multimedia data, a data fusion calculation is performed to obtain an abnormal evaluation value of the target multimedia data.
9. The method according to claim 1, wherein After determining that the target multimedia data has a traffic anomaly, the method further includes: Obtaining pre-set cheating handling strategies at each level, wherein each cheating handling strategy has an associated evaluation value range; Determine the target evaluation value interval corresponding to the abnormal evaluation value according to the inclusion of the abnormal evaluation value in each evaluation value interval; The target multimedia data is processed using a cheating processing strategy corresponding to the target evaluation value interval.
10. The method according to any one of claims 1 to 9, wherein The performing content relevance detection on the review texts to obtain the review relevance corresponding to the group of review texts includes: Inputting each text feature obtained for each comment text into a trained first evaluation model to obtain a first prediction value corresponding to the popularity cheating behavior output by the first evaluation model after performing content relevance detection on each text feature, wherein the first evaluation model is obtained by performing multiple rounds of iterative training based on a first training sample set, wherein a first training sample includes: each sample comment associated with a sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs; The first prediction value is used as the comment relevance corresponding to the group of comment texts.
11. The method according to any one of claims 1 to 9, wherein: The performing content relevance detection based on the comment texts and the content description texts to obtain a content matching degree between the set of comment texts and the target multimedia data includes: Inputting the text features obtained for each of the comment texts and the content description text into a trained second evaluation model to obtain a second prediction value corresponding to the popularity cheating behavior output by the second evaluation model after performing content relevance detection on each of the text features, wherein the second evaluation model is obtained by performing multiple rounds of iterative training based on a second training sample set, wherein a second training sample includes: each sample comment and sample description text associated with the sample multimedia data, and a sample label for identifying whether the popularity cheating behavior occurs; The second prediction value is used as the content matching degree corresponding to the group of comment texts.
12. A flow anomaly assessment device, characterized in that: include: a response unit, configured to, in response to an evaluation indication triggered for target multimedia data, obtain content description text and N groups of comment texts associated with the target multimedia data, wherein the N groups of comment texts are obtained by performing N times of random sampling without replacement from a set of comment texts associated with the target multimedia data; An execution unit is configured to perform the following operations for each group of comment texts: obtaining each comment text contained in the group of comment texts; performing content relevance detection based on each comment text to obtain comment relevance corresponding to the group of comment texts; performing content relevance detection based on each comment text and the content description text to obtain a content matching degree between the group of comment texts and the target multimedia data; an obtaining unit, configured to calculate an abnormality degree based on the N comment relevances and the N content matching degrees, and obtain an abnormality evaluation value corresponding to the target multimedia data; The determination unit is configured to determine that traffic cheating behavior exists in the target multimedia data when the abnormal evaluation value exceeds a set threshold.
13. The device according to claim 12, wherein Before responding to the evaluation instruction triggered for the target multimedia data, the apparatus further includes a screening unit, the screening unit being configured to: Obtain each candidate multimedia data, and for each candidate multimedia data, perform the following operations: Performing a quantitative comparison of the operation heat according to the historical operation data of the candidate multimedia data in two adjacent time windows to obtain a heat change value of the candidate multimedia data; Performing a quantitative comparison of operation heats based on the historical operation data of the candidate multimedia data and the historical operation data of each historical multimedia data, and obtaining a heat comparison value between the candidate multimedia data and each historical multimedia data; Each of the historical multimedia data and the candidate multimedia data is associated with the same publishing object; When it is determined that the heat change value and the heat comparison value meet the heat anomaly screening condition, the candidate multimedia data is used as the target multimedia data, and an evaluation indication is triggered.
14. The device according to claim 12, wherein After obtaining the content description text and N groups of comment texts associated with the target multimedia data, and before calculating the abnormality degree based on the N comment relevances and N content matching degrees, the execution unit is further configured to: Obtain various abnormal words pre-constructed for traffic fraud behavior; Determine the number of hits for each abnormal word in the content description text and the N groups of comment texts; According to the number of hits of each abnormal word and the abnormal indicators determined in advance for each abnormal word, the abnormal word hit degree corresponding to the target multimedia data is determined, wherein the abnormal indicator of an abnormal word is determined by calculating the inclusion of the abnormal word in the multimedia data with and without traffic cheating behavior.
15. The device according to claim 12, wherein After determining that the target multimedia data has a traffic anomaly, the determining unit is further configured to: Obtaining pre-set cheating handling strategies at each level, wherein each cheating handling strategy has an associated evaluation value range; Determine the target evaluation value interval corresponding to the abnormal evaluation value according to the inclusion of the abnormal evaluation value in each evaluation value interval; The target multimedia data is processed using a cheating processing strategy corresponding to the target evaluation value interval.
16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 11 is implemented.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Cited By
AR content prediction loading method and system based on visual attention trajectory
CN121277351A