Short video value evaluation method based on deep learning and related device

CN115700754BActive Publication Date: 2026-09-18CHINESE ACADEMY OF PRESS & PUBLICATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211341727.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2026-09-18
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

然而已有的图像评估技术并不能很好的评估图片的价值,也就无法对视频价值进行有效评估

Benefits of technology

[0036] This invention provides a deep learning-based method for short video value assessment. First, short videos from at least one video platform are acquired and structured to obtain standard video data. Multi-dimensional value tags are then obtained and labeled based on this standard video data to create a multi-dimensional value training dataset. Next, a deep neural network model is trained on this training dataset to obtain a value assessment model that outputs multi-dimensional value scores based on input short videos. Finally, the short video to be assessed is input into the value assessment model to obtain a multi-dimensional value score as the value assessment result. This method constructs a deep learning model and leverages deep learning algorithms to complete short video value assessment. It offers broader assessment dimensions, is not limited by the knowledge required for human assessment, and is applicable to different short video platforms and fields, significantly improving the efficiency and fairness of short video value assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700754B_ABST
    Figure CN115700754B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a short video value evaluation method and related device based on deep learning, which comprises the following steps: obtaining short videos from at least one video platform, and performing structured processing on the short videos to obtain standard video data; obtaining multi-dimensional value labels based on the standard video data and labeling to obtain a multi-dimensional value training data set; training a deep neural network model based on the multi-dimensional value training data set to obtain a value evaluation model for outputting multi-dimensional value scores according to input short videos; and inputting a short video to be evaluated into the value evaluation model to obtain multi-dimensional value scores as value evaluation results. The embodiment of the application completes short video value evaluation by constructing a deep learning model with the aid of a deep learning algorithm, which has a wider evaluation dimension, is not limited by knowledge during human evaluation, and can be applied to different short video platforms and fields, thereby greatly improving the efficiency and fairness of short video value evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of video processing technology, and in particular to a method and related apparatus for evaluating the value of short videos based on deep learning. Background Technology

[0002] Short videos, typically referring to videos no longer than 10 minutes, are a form of internet content dissemination. With the widespread adoption of mobile devices and faster network speeds, short videos, offering quick, fast, and high-traffic content, have gradually gained favor with major platforms, fans, and investors. Furthermore, with the emergence of the influencer economy, a group of high-quality UGC content creators have gradually risen in the short video industry. Weibo, Miaopai, Kuaishou, and Toutiao have all entered the short video industry, recruiting a number of excellent content production teams to join, which has further promoted the development of the short video industry.

[0003] As the short video industry develops, the quality of short videos on the internet varies greatly. Often, after a short video goes viral, a large number of creators blindly follow suit, resulting in poor-quality imitations. This is detrimental to the long-term development of short videos. Therefore, a short video value assessment method is needed to control the quality of short videos and improve their overall quality.

[0004] The successful application of deep learning technology in computer vision and multimedia has greatly improved video and image analysis techniques. For example, Google research team proposed a deep convolutional neural network, NIMA, which can structure video content into machine-processable text data and predict the distribution of human evaluation opinions on images based on direct perception (technical perspective) and attractiveness (aesthetic perspective). However, existing image evaluation techniques cannot effectively assess the value of images, and therefore cannot effectively evaluate the value of videos. Furthermore, to improve the evaluation of short video value, an intuitive solution is to use user behavior data from short video platforms as the basis for model training. However, over-reliance on this data can lead to evaluation results biased towards specific categories of content, failing to adapt to the ever-evolving market environment and the increasing diversity of video data. Summary of the Invention

[0005] This invention provides a method and apparatus for evaluating the value of short videos based on deep learning. It improves the efficiency and fairness of short video value evaluation by utilizing deep learning, and can be applied to a wider range of fields without being limited by the knowledge of the personnel.

[0006] In a first aspect, embodiments of the present invention provide a method for evaluating the value of short videos based on deep learning, comprising:

[0007] Acquire short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data;

[0008] Based on the standard video data, multi-dimensional value labels are obtained and labeled to obtain a multi-dimensional value training dataset.

[0009] A deep neural network model is trained based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video.

[0010] The short video to be evaluated is input into the value assessment model to obtain a multi-dimensional value score, which is used as the value assessment result.

[0011] In some embodiments, before inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result, the method further includes:

[0012] Based on the short videos to be evaluated, the guidance evaluation results are determined according to the preset guidance evaluation indicators, and short videos that fail the guidance evaluation are eliminated.

[0013] In some embodiments, training a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video includes:

[0014] Video feature labels for different video platforms are extracted and labeled from the standard video data. A deep learning model is trained based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library.

[0015] A value assessment model applicable to different video platforms is obtained by training the multi-dimensional value training dataset and the pre-trained model.

[0016] In some embodiments, the multi-dimensional value tags include content evaluation tags, platform evaluation tags, and business evaluation tags, and the process of obtaining and labeling multi-dimensional value tags based on the standard video data includes:

[0017] Based on the value attributes of short video content, the content value is evaluated and determined as the content evaluation label according to the preset content value indicators.

[0018] The platform value obtained by extracting user behavior within the video platform corresponding to the standard video data, performing time dimension and numerical normalization, is used as the platform evaluation label.

[0019] The commercial revenue of short videos corresponding to the standard video data within a fixed online time range is obtained, and normalized to obtain the commercial value as a commercial evaluation label.

[0020] In some embodiments, after inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result, the method further includes:

[0021] Based on the pre-trained model, the first video feature of the short video to be evaluated and the second video feature of the standard video data are extracted respectively.

[0022] The approximate nearest neighbor algorithm is used to calculate the feature similarity between the first video feature and the second video feature. The multi-dimensional value score is adjusted based on the feature similarity to obtain the value assessment result.

[0023] In a second aspect, embodiments of the present invention provide a short video value assessment device based on deep learning, comprising:

[0024] A data processing module is used to acquire short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data.

[0025] The annotation module is used to obtain and annotate multi-dimensional value labels based on the standard video data to obtain a multi-dimensional value training dataset.

[0026] The model training module is used to train a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video.

[0027] The evaluation module is used to input the short video to be evaluated into the value evaluation model to obtain a multi-dimensional value score as the value evaluation result.

[0028] In some embodiments, the model training module is used for:

[0029] Video feature labels for different video platforms are extracted and labeled from the standard video data. A deep learning model is trained based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library.

[0030] A value assessment model applicable to different video platforms is obtained by training the multi-dimensional value training dataset and the pre-trained model.

[0031] In some embodiments, the evaluation module is further configured to:

[0032] Based on the pre-trained model, the first video feature of the short video to be evaluated and the second video feature of the standard video data are extracted respectively.

[0033] The approximate nearest neighbor algorithm is used to calculate the feature similarity between the first video feature and the second video feature. The multi-dimensional value score is adjusted based on the feature similarity to obtain the value assessment result.

[0034] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the deep learning-based short video value assessment method provided in the first aspect.

[0035] In a fourth aspect, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executed, implements the deep learning-based short video value assessment method provided in the first aspect.

[0036] This invention provides a deep learning-based method for short video value assessment. First, short videos from at least one video platform are acquired and structured to obtain standard video data. Multi-dimensional value tags are then obtained and labeled based on this standard video data to create a multi-dimensional value training dataset. Next, a deep neural network model is trained on this training dataset to obtain a value assessment model that outputs multi-dimensional value scores based on input short videos. Finally, the short video to be assessed is input into the value assessment model to obtain a multi-dimensional value score as the value assessment result. This method constructs a deep learning model and leverages deep learning algorithms to complete short video value assessment. It offers broader assessment dimensions, is not limited by the knowledge required for human assessment, and is applicable to different short video platforms and fields, significantly improving the efficiency and fairness of short video value assessment. Attached Figure Description

[0037] Figure 1 This is a flowchart of a short video value assessment method based on deep learning provided in an embodiment of the present invention;

[0038] Figure 2 This is a flowchart of another short video value assessment method based on deep learning provided in an embodiment of the present invention;

[0039] Figure 3 This is a flowchart of another short video value assessment method based on deep learning provided in an embodiment of the present invention;

[0040] Figure 4 This is a flowchart of another short video value assessment method based on deep learning provided in an embodiment of the present invention;

[0041] Figure 5 This is a flowchart of another short video value assessment method based on deep learning provided in an embodiment of the present invention;

[0042] Figure 6 This is a schematic diagram of the structure of a short video value assessment device based on deep learning provided in an embodiment of the present invention;

[0043] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining the present invention and not for limiting the present invention. It should also be noted that, for ease of description, only the parts relevant to the present invention are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0045] Figure 1 A flowchart of a short video value assessment method based on deep learning provided by an embodiment of the present invention is given. The method of this embodiment can be executed by a short video value assessment device based on deep learning. The device can be implemented by hardware and / or software and can be set inside the electronic device as part of the electronic device.

[0046] like Figure 1 As shown, the deep learning-based short video value assessment method provided in this embodiment includes the following steps:

[0047] Step 101: Obtain short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data.

[0048] Video refers to a continuous stream of moving images presented and disseminated electronically or digitally. Short video refers to a video with a duration not exceeding a certain time, typically no more than 10 minutes. In this embodiment, data structuring refers to the process of cleaning short videos from multiple channels and transferring them according to a preset data format. For example, short videos from platforms such as Douyin, Kuaishou, and Bilibili may have inconsistent data formats and content information, thus requiring conversion to a pre-defined data format. More specifically, in some embodiments, the standard video data obtained after structuring should include some or all of the following short video content information: work title, creator, publishing platform, publishing date, type, work duration, work description, and copyright statement. It should be understood that the requirements for short video content information in data structuring can be changed according to actual needs, and no specific restrictions are imposed here.

[0049] Step 102: Obtain and label multi-dimensional value tags based on the standard video data to obtain a multi-dimensional value training dataset.

[0050] Specifically, in this embodiment, the value labeling of standard video data is based on value assessment indicators. The labeling content mainly includes content assessment labels, platform assessment labels, and commercial assessment labels. The specific process includes:

[0051] Based on the value attributes of short video content, the content value is evaluated and determined as the content evaluation label according to the preset content value indicators.

[0052] The platform value obtained by extracting user behavior within the video platform corresponding to the standard video data, performing time dimension and numerical normalization, is used as the platform evaluation label.

[0053] The commercial revenue of short videos corresponding to the standard video data within a fixed online time range is obtained, and normalized to obtain the commercial value as a commercial evaluation label.

[0054] More specifically, the three labels mentioned above represent three primary indicators in the value assessment metrics: content value, brand value, and commercial value. When determining the final multi-dimensional value labels, content value, brand value, and commercial value can each occupy different weights. That is, determining the multi-dimensional value labels according to the preset value assessment metrics includes: summarizing the content value, brand value, and commercial value according to the preset weights to obtain the multi-dimensional value labels. For example, in some embodiments, content value adopts a percentage-based indicator with a weight of 50%, brand value adopts a percentage-based indicator with a weight of 20%, and commercial value adopts a percentage-based indicator with a weight of 30%.

[0055] More specifically, in some embodiments, the three primary indicators—content value, brand value, and commercial value—can be further subdivided:

[0056] For content value, it further includes secondary indicators of style and artistry (each with corresponding weights, such as style 20% and artistry 30%, where the weights refer to their weights in the overall value assessment indicators, not their weights in the overall content value); for brand value, it further includes secondary indicators of brand awareness, reputation, and activity (each with corresponding weights, such as brand awareness 6%, reputation 6%, and activity 8%, where the weights refer to their weights in the overall value assessment indicators); for commercial value, it further includes secondary indicators of marketability, development potential, and profitability (each with corresponding weights, such as marketability 10%, development potential 10%, and profitability 10%, where the weights refer to their weights in the overall value assessment indicators).

[0057] More specifically, based on the above embodiments, some embodiments can further subdivide each secondary indicator into multiple tertiary indicators:

[0058] Regarding style, this further includes requirements for advocating positive and uplifting themes, maintaining social order and stability, possessing refined taste (avoiding vulgar, negative, or uncivilized themes), and upholding social morality and promoting the inheritance of excellent traditional Chinese culture (each with corresponding weights, such as 5% for positive themes, 5% for stability, 5% for refinement, and 5% for inheritance). Regarding artistry, this further includes requirements for artistic images to be typical and concrete, reflecting the essence of social life and possessing universal social significance; requirements for artistic plots to be vivid and intricate, with diverse plot changes and natural, coherent transitions; and requirements for artistic structure to be rigorous and complete, with clear narration, reasonable plot settings, and logical flow. The requirements for the content include: a smooth and rigorous narrative; accurate and vivid artistic language; fluent writing; language use matching the characters and plot conflicts; absence of clickbait titles, vulgar language, or lowbrow humor; diverse and appropriate expressive means; full expression of content; originality in artistic expression, possessing positive realistic significance or innovative spirit; and innovative use of cutting-edge technology to achieve unique expressive power (each with corresponding weight, e.g., 5%). Regarding brand awareness, further requirements include the frequency and amount of media exposure or mentions in other content products; positive and negative media reviews; the influence of media's emotional perception of the work; and international influence. The criteria for shaping a national image include: the quality of media mentions (emphasizing the social influence and credibility of media outlets publishing evaluations); broad channel coverage (including coverage on platforms like Weibo, WeChat, Xiaohongshu, and short video platforms, each with corresponding weights, such as 2% mention rate, 1% influence, 1% dissemination power, 1% credibility, and 1% coverage); positive reputation (including user recommendation and feedback satisfaction, academic influence and professional institutional discussion and dissemination, and awards / awards recognition, each with corresponding weights, such as 2% satisfaction, 2% attention, and 2% recognition); and activity level (including the number and growth of the creator's followers). The evaluation criteria for speed include user engagement (requiring likes and favorites), number of comments in the comment section, interactiveness of the emotional attributes of the comments, user clicks and shares, number of external sharing, search engine mentions, and timeliness (also weighted accordingly, e.g., user engagement 2%, interactiveness 2%, activity 3%, and timeliness 1%). For marketability, further criteria include mention rate across the entire internet search engine, imagery of the emotional attributes in online comments, search engine mention rate, number of comments on previous works, emotional attributes of comments on previous works, and traffic-generating aspects of user feedback on previous works (also weighted accordingly, e.g., imagery 2% and traffic 8%).Regarding development, this further includes requirements for ranking based on the frequency of a work's appearance on the list, requirements for full-industry chain operation with continuous value-added operations from upstream to downstream, and requirements for capital intervention and investment from industry capital (each with corresponding weights, such as ranking 3%, value-added 4%, and investment 3%). Regarding profitability, this further includes requirements for revenue sharing through the platform, revenue generated from advertising and operating income, and revenue generated from copyright sales (each with corresponding weights, such as profitability 6% and copyright 4%). It should be noted that the weights in this embodiment refer to their weights within the overall value assessment indicators, not their weights within the secondary indicators.

[0059] Step 103: Train a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video.

[0060] Specifically, in some embodiments, step 103 is as follows: Figure 2 The specific steps shown include 1031-1033:

[0061] Step 1031: Extract and label video feature tags for different video platforms from the standard video data.

[0062] Step 1032: Train a deep learning model based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library.

[0063] Step 1033: Based on the multi-dimensional value training dataset and the pre-trained model, a value assessment model applicable to different video platforms is trained.

[0064] In this embodiment, a visual recognition model and an NLP semantic analysis model can be used to construct structured video data and video shooting scripts that include elements such as scene, shot size, angle, camera movement, actors, costumes, props, subtitles, and music, and to train a multi-label model.

[0065] Step 104: Input the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result.

[0066] Specifically, the multi-dimensional value score is the evaluation value corresponding to the multi-dimensional value label. It is a result obtained by evaluating based on multi-dimensional value evaluation indicators through a value evaluation model. Optionally, in some specific embodiments, step 104 needs to consider the impact of innovation standards while evaluating value. Therefore, after obtaining the initial value evaluation result, such as Figure 3 As shown, it also includes:

[0067] Step 105: Extract the first video features of the short video to be evaluated and the second video features of the standard video data based on the pre-trained model.

[0068] Step 106: Calculate the feature similarity between the first video feature and the second video feature using the approximate nearest neighbor algorithm, and adjust the multi-dimensional value score according to the feature similarity to obtain the value assessment result.

[0069] The higher the feature similarity, the higher the likelihood of plagiarism and the lower the innovation of the short video to be evaluated. Accordingly, its multi-dimensional value score should be reduced. Conversely, the multi-dimensional value score should be increased. The adjusted multi-dimensional value score is the final value assessment result.

[0070] More specifically, in some embodiments, such as Figure 4 As shown, before step 104, there is also step 107:

[0071] Step 107: Based on the short video to be evaluated, determine the guidance evaluation result according to the preset guidance evaluation indicators, and remove the short videos to be evaluated that do not meet the guidance evaluation results.

[0072] The guidance assessment indicators are used to evaluate whether the content of short videos conforms to the positive guidance required by morality and law. Specifically, in this embodiment, the guidance assessment indicators include at least one of the following: legality, uniformity, ethnicity, policy, and compliance. Legality refers to compliance with the Constitution, protection of the legitimate rights and interests of others, and the absence of content that insults, defames, or infringes upon the legitimate rights and interests of others. Uniformity refers to upholding national unity, sovereignty, and territorial integrity. Ethnicity refers to maintaining national unity and customs. Policy refers to compliance with national religious policies, ideological policies, and other policies and guidelines. Compliance refers to compliance with relevant laws and regulations, including not promoting obscenity, pornography, violence, gambling, or inciting crime. During the guidance assessment, the assessment result is considered qualified only if the video data to be assessed meets all of the aforementioned guidance assessment indicators. If the video data to be assessed does not meet any one of the following criteria, the guidance assessment result is deemed unqualified.

[0073] It should be further explained that the evaluation of the short videos to be evaluated based on the guidance evaluation indicators in this embodiment can be completed manually or by algorithms. For example, in some specific embodiments, the video data to be evaluated can be displayed to designated evaluators, and the evaluation results of the video data to be evaluated based on legality, uniformity, ethnicity, policy, and compliance can be obtained through a preset evaluation interface. (The evaluation results can be expressed in numerical form to show whether the evaluation results meet the indicator requirements, such as legality 56 points, uniformity 78 points, ethnicity 67 points, policy compliance 77 points, and compliance 99 points, where legality is below the passing score of 60 points and does not meet the legality requirement). The reviewers referred to here are usually professionals in the short video industry, such as the reviewers of various short video platforms and top internet celebrities. In some other embodiments, a guidance evaluation model can be trained using deep learning, and the short videos to be evaluated can be input into the guidance evaluation model to output guidance evaluation results that meet the guidance evaluation indicators. Of course, this is only for illustrative purposes and is not a specific limitation. It should be understood that in this embodiment, only short videos whose guidance assessment results are qualified will continue to undergo value assessment, i.e., be input into the value assessment model.

[0074] The method provided in this embodiment first acquires short videos from at least one video platform and performs structured processing on the short videos to obtain standard video data. Based on the standard video data, multi-dimensional value tags are obtained and labeled to obtain a multi-dimensional value training dataset. Then, a deep neural network model is trained based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short videos. Finally, the short video to be evaluated is input into the value assessment model to obtain a multi-dimensional value score as the value assessment result. This method completes the short video value assessment by constructing a deep learning model and using deep learning algorithms. Its assessment dimensions are broader, it is not limited by the knowledge of human assessment, and it can be applied to different short video platforms and fields, greatly improving the efficiency and fairness of short video value assessment.

[0075] Optionally, in some alternative embodiments, such as Figure 5 As shown, after step 104, step 201 is also included:

[0076] Step 201: Based on the preset short video types and the value assessment result report, classify and rank the videos according to the multi-dimensional value scores to obtain the short video classification ranking results.

[0077] See Table 1 for details:

[0078] Table 1

[0079]

[0080] In this embodiment, the short video types include at least one of the following: narrative (performance-based short videos with a storyline), documentary (short videos that capture and record real life), editing (short videos that primarily feature derivative film and television works), creative (short videos showcasing creative works using special shooting and software techniques), knowledge (short videos that primarily share a certain type of knowledge), and art (short videos that primarily feature talent performances). More specifically, in some embodiments, each type of short video can be further subdivided. For example, as shown in Table 1, the overall code for narrative is A, which is further divided into: emotional (code A1); humorous (code A2); ethical (code A3); historical (code A4); and slice-of-life (code A5) (not listed here).

[0081] The method provided in this embodiment, after obtaining the value assessment results, reflects the value of different short videos in the corresponding video fields through short video classification and ranking. The classification and ranking avoids the problem of inaccurate value comparison between different types of short videos due to different specific emphases of assessment indicators. Moreover, the fairness, objectivity and scientific nature of the evaluation can be better guaranteed when using manual or intelligent evaluation for short videos of the same type.

[0082] Figure 6 This is a schematic diagram of a short video value assessment device based on deep learning, provided as an embodiment of the present invention. This device can be implemented by software and / or hardware and integrated into an electronic device. Figure 6 As shown, the device includes a data processing module 51, a labeling module 52, a model training module 53, and an evaluation module 54.

[0083] The data processing module 51 is used to acquire short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data.

[0084] The annotation module 52 is used to obtain and annotate multi-dimensional value labels based on the standard video data to obtain a multi-dimensional value training dataset;

[0085] The model training module 53 is used to train a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video.

[0086] The evaluation module 54 is used to input the short video to be evaluated into the value evaluation model to obtain a multi-dimensional value score as the value evaluation result.

[0087] The apparatus provided in this embodiment first acquires short videos from at least one video platform and performs structured processing on the short videos to obtain standard video data. Based on the standard video data, it acquires and labels multi-dimensional value tags to obtain a multi-dimensional value training dataset. Then, it trains a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short videos. Finally, it inputs the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result. This method completes short video value assessment by constructing a deep learning model and using deep learning algorithms. Its assessment dimensions are broader, it is not limited by the knowledge of human assessment, and it can be applied to different short video platforms and fields, greatly improving the efficiency and fairness of short video value assessment.

[0088] Based on the above embodiments, the model training module is characterized in that:

[0089] Video feature labels for different video platforms are extracted and labeled from the standard video data. A deep learning model is trained based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library.

[0090] A value assessment model applicable to different video platforms is obtained by training the multi-dimensional value training dataset and the pre-trained model.

[0091] Based on the above embodiments, the evaluation module is further used for:

[0092] Based on the pre-trained model, the first video feature of the short video to be evaluated and the second video feature of the standard video data are extracted respectively.

[0093] The approximate nearest neighbor algorithm is used to calculate the feature similarity between the first video feature and the second video feature. The multi-dimensional value score is adjusted based on the feature similarity to obtain the value assessment result.

[0094] Based on the above embodiments, before inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result, the method further includes:

[0095] Based on the short videos to be evaluated, the guidance evaluation results are determined according to the preset guidance evaluation indicators, and short videos that fail the guidance evaluation are eliminated.

[0096] Based on the above embodiments, the multi-dimensional value tags include content evaluation tags, platform evaluation tags, and business evaluation tags. The process of obtaining and labeling multi-dimensional value tags based on the standard video data includes:

[0097] Based on the value attributes of short video content, the content value is evaluated and determined as the content evaluation label according to the preset content value indicators.

[0098] The platform value obtained by extracting user behavior within the video platform corresponding to the standard video data, performing time dimension and numerical normalization, is used as the platform evaluation label.

[0099] The commercial revenue of short videos corresponding to the standard video data within a fixed online time range is obtained, and normalized to obtain the commercial value as a commercial evaluation label.

[0100] This invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a deep learning-based short video value assessment method. The method includes: acquiring short videos from at least one video platform and performing structured processing on the short videos to obtain standard video data; acquiring and labeling multi-dimensional value tags based on the standard video data to obtain a multi-dimensional value training dataset; training a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on input short videos; and inputting the short video to be assessed into the value assessment model to obtain a multi-dimensional value score as the value assessment result.

[0101] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. Furthermore, storage media may reside in a first computer system in which a program is executed, or may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media that may reside in different locations (e.g., in different computer systems connected via a network). Storage media may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.

[0102] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the deep learning-based short video value assessment operation described above, but can also execute related operations in the deep learning-based short video value assessment method provided in any embodiment of the present invention.

[0103] This invention provides an electronic device that may include the deep learning-based short video value assessment device provided in any embodiment of this invention. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 7 As shown, the electronic device may include: a memory 601, a central processing unit (CPU) 602 (also known as a processor, hereinafter referred to as CPU), and the memory 601 for storing executable program code; the processor 602 reads the executable program code stored in the memory 601 to run a program corresponding to the executable program code, for executing: acquiring short videos from at least one video platform and performing structured processing on the short videos to obtain standard video data; acquiring and labeling multi-dimensional value tags based on the standard video data to obtain a multi-dimensional value training dataset; training a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on input short videos; and inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result.

[0104] The electronic device also includes: a peripheral interface 603, an RF (Radio Frequency) circuit 605, an audio circuit 606, a speaker 611, a power management chip 608, an input / output (I / O) subsystem 609, a touch screen 612, other input / control devices 610, and an external port 604. These components communicate via one or more communication buses or signal lines 607.

[0105] It should be understood that the illustrated electronic device 600 is merely an example of an electronic device, and the electronic device 600 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0106] The following is a detailed description of the electronic device for handling application exceptions provided in this embodiment, which is an example of an electronic device.

[0107] The memory 601 can be accessed by the CPU 602, peripheral interface 603, etc. The memory 601 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other volatile solid-state storage devices.

[0108] Peripheral interface 603 can connect the device's input and output peripherals to CPU 502 and memory 601.

[0109] I / O subsystem 609 connects input / output peripherals on the device, such as touchscreen 612 and other input / control devices 610, to peripheral interface 603. I / O subsystem 609 may include display controller 6091 and one or more input controllers 6092 for controlling other input / control devices 610. The one or more input controllers 6092 receive or send electrical signals to other input / control devices 610, which may include physical buttons (press buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, and click wheels. It is worth noting that input controller 6092 can be connected to any of the following: keyboard, infrared port, USB interface, and pointing device such as a mouse.

[0110] The touch screen 612 is an input and output interface between the user terminal and the user, and displays visual output to the user. The visual output may include graphics, text, icons, videos, etc.

[0111] The display controller 6091 in the I / O subsystem 609 receives or sends electrical signals to the touchscreen 612. The touchscreen 612 detects touches, and the display controller 6091 converts these touches into interactions with user interface objects displayed on the touchscreen 612, thus achieving human-computer interaction. These user interface objects can be icons for running games, connecting to a network, etc. It is worth noting that the device may also include an optical mouse, which is a touch-sensitive surface that does not display visual output, or an extension of the touch-sensitive surface formed by the touchscreen.

[0112] RF circuit 605 is primarily used to establish communication between electronic devices and wireless networks (i.e., the network side), enabling data reception and transmission between the electronic devices and the wireless network. Examples include sending and receiving SMS messages and emails. Specifically, RF circuit 605 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 605 converts electrical signals into electromagnetic signals or vice versa, and uses these electromagnetic signals to communicate with the communication network and other devices. RF circuit 605 may include known circuits for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, CODEC (Coder-Coder) chipsets, Subscriber Identity Modules (SIMs), etc.

[0113] The audio circuit 606 is mainly used to receive audio data from the peripheral interface 603, convert the audio data into an electrical signal, and send the electrical signal to the speaker 611.

[0114] Speaker 611 is used to convert voice signals received by electronic devices from wireless networks via RF circuit 605 into sound and play the sound to the user.

[0115] The power management chip 608 is used to provide power and manage the power supply for the CPU 602, the I / O subsystem, and the hardware connected to the peripheral interface 603.

[0116] The aforementioned electronic device can execute the method provided in any embodiment of the present invention, and has corresponding functional modules for executing the method. The electronic device provided in the embodiment of the present invention first acquires short videos from at least one video platform, and performs structured processing on the short videos to obtain standard video data. Based on the standard video data, it acquires and labels multi-dimensional value tags to obtain a multi-dimensional value training dataset. Then, it trains a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short videos. Finally, it inputs the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result. This method completes the short video value assessment by constructing a deep learning model and using deep learning algorithms. Its assessment dimensions are broader, it is not limited by the knowledge of human assessment, and it can be applied to different short video platforms and fields, greatly improving the efficiency and fairness of short video value assessment.

[0117] The above description is merely a preferred embodiment of the present invention and the technical principles employed. The present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the claims.

Claims

1. A method for evaluating the value of short videos based on deep learning, characterized in that, include: Acquire short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data; Based on the standard video data, multi-dimensional value labels are obtained and labeled to obtain a multi-dimensional value training dataset. A deep neural network model is trained based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video. This includes extracting and labeling video feature tags for different video platforms from the standard video data, training a deep learning model based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library, and training a value assessment model applicable to different video platforms based on the multi-dimensional value training dataset and the pre-trained model. The short video to be evaluated is input into the value assessment model to obtain a multi-dimensional value score, which is used as the value assessment result.

2. The method according to claim 1, characterized in that, Before inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result, the process also includes: Based on the short videos to be evaluated, the guidance evaluation results are determined according to the preset guidance evaluation indicators, and short videos that fail the guidance evaluation are eliminated.

3. The method according to claim 1, characterized in that, The multi-dimensional value tags include content evaluation tags, platform evaluation tags, and business evaluation tags. The process of obtaining and labeling multi-dimensional value tags based on the standard video data includes: Based on the value attributes of short video content, the content value is evaluated and determined as the content evaluation label according to the preset content value indicators. The platform value obtained by extracting user behavior within the video platform corresponding to the standard video data, performing time dimension and numerical normalization, is used as the platform evaluation label. The commercial revenue of short videos corresponding to the standard video data within a fixed online time range is obtained, and normalized to obtain the commercial value as a commercial evaluation label.

4. The method according to any one of claims 1-3, characterized in that, After inputting the short video to be evaluated into the value assessment model to obtain a multi-dimensional value score as the value assessment result, the following steps are also included: Based on the pre-trained model, the first video feature of the short video to be evaluated and the second video feature of the standard video data are extracted respectively. The approximate nearest neighbor algorithm is used to calculate the feature similarity between the first video feature and the second video feature. The multi-dimensional value score is adjusted according to the feature similarity to obtain the value assessment result.

5. A short video value assessment device based on deep learning, characterized in that, include: A data processing module is used to acquire short videos from at least one video platform and perform structured processing on the short videos to obtain standard video data. The annotation module is used to obtain and annotate multi-dimensional value labels based on the standard video data to obtain a multi-dimensional value training dataset. The model training module is used to train a deep neural network model based on the multi-dimensional value training dataset to obtain a value assessment model for outputting multi-dimensional value scores based on the input short video. This includes extracting and labeling video feature tags for different video platforms from the standard video data, training a deep learning model based on the labeled standard video data to obtain a pre-trained model for describing the features of the video library, and training a value assessment model applicable to different video platforms based on the multi-dimensional value training dataset and the pre-trained model. The evaluation module is used to input the short video to be evaluated into the value evaluation model to obtain a multi-dimensional value score as the value evaluation result.

6. The short video value assessment device based on deep learning according to claim 5, characterized in that, The evaluation module is also used for: Based on the pre-trained model, the first video feature of the short video to be evaluated and the second video feature of the standard video data are extracted respectively. The approximate nearest neighbor algorithm is used to calculate the feature similarity between the first video feature and the second video feature. The multi-dimensional value score is adjusted according to the feature similarity to obtain the value assessment result.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the deep learning-based short video value assessment method as described in any one of claims 1-4.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the deep learning-based short video value assessment method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Network article evaluation method and system, computer equipment and readable storage medium

    CN111260197A

  • Video content evaluation method and device, storage medium and computer equipment

    CN111741330A