Artificial Intelligence-Based Video Processing Method, Apparatus, and Electronic Device

By dividing the video into segments, extracting and fusing features, and using recursive updates to predict recommended information clips, the problem of wasted computing resources in the video is solved, and the actual utilization rate and user experience of electronic devices are improved.

CN112287799BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011148760.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-23
Publication Date
2025-07-29
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

In video scenarios, the computational resources used by electronic devices to present recommended information are wasted, resulting in low actual utilization rates and the prior art has not been effectively solved.

Method used

Divide the video into multiple video clips, extract content features and operation features, perform fusion processing, and recursively update the prediction video clips including recommended information to improve prediction accuracy.

Benefits of technology

Accurately identify and delete recommended information clips in videos, improve the computing resource utilization rate of electronic devices, improve user experience and save storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112287799B_ABST
    Figure CN112287799B_ABST
Patent Text Reader

Abstract

The present application provides a video processing method, apparatus, electronic device, and computer-readable storage medium based on artificial intelligence; related to artificial intelligence technology and big data technology; the method includes: dividing a video into multiple video segments; performing feature extraction processing on the video segments to obtain content features of the video segments; obtaining operation data for the video segments, and performing feature mapping processing on the operation data to obtain operation features; performing fusion processing on the content features and operation features of the video segments to obtain fusion features; performing recursive update processing on the fusion features of multiple video segments, and predicting, according to the obtained updated fusion features, the video segments including recommendation information among the multiple video segments. Through the present application, it is possible to accurately predict the video segments including recommendation information in the video, thereby helping to improve the actual utilization rate of the computing resources of the electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence and big data technologies, and in particular to a video processing method, apparatus, electronic device, and computer-readable storage medium based on artificial intelligence. Background Art

[0002] In video scenarios, many videos often include embedded recommendation information. Taking a short video platform as an example, in order to increase their own income, some uploaders will add advertisements to the short videos they shoot and upload them to the short video platform, so that users of the short video platform are forced to watch product advertisements during the process of watching the short video.

[0003] During the presentation of the video, it will consume the computing resources of the electronic device (such as the background server of the short video platform). However, for videos including recommendation information, the computing resources used by the electronic device to send or present the recommendation information will be wasted, that is, the actual utilization rate of the computing resources of the electronic device is low.

[0004] In response to this, the related art does not provide an effective solution. Summary of the Invention

[0005] Embodiments of this application provide a video processing method, apparatus, electronic device, and computer-readable storage medium based on artificial intelligence, which can accurately identify video segments including recommendation information in a video, and help improve the actual utilization rate of the computing resources of the electronic device.

[0006] The technical solution of the embodiments of this application is implemented as follows:

[0007] Embodiments of this application provide a video processing method based on artificial intelligence, including:

[0008] Dividing the video into multiple video segments;

[0009] Performing feature extraction processing on the video segment to obtain the content feature of the video segment;

[0010] Obtaining operation data for the video segment, and performing feature mapping processing on the operation data to obtain operation features;

[0011] Performing fusion processing on the content feature and operation feature of the video segment to obtain a fusion feature;

[0012] Performing recursive update processing on the fusion features of multiple video segments, and predicting video segments including recommendation information in the multiple video segments according to the obtained updated fusion features.

[0013] Embodiments of this application provide a video processing apparatus based on artificial intelligence, including:

[0014] A partitioning module, configured to partition a video into multiple video segments;

[0015] A feature extraction module, configured to perform feature extraction processing on the video segments to obtain content features of the video segments;

[0016] A feature mapping module, configured to obtain operation data for the video segments and perform feature mapping processing on the operation data to obtain operation features;

[0017] A fusion module, configured to perform fusion processing on the content features and operation features of the video segments to obtain fusion features;

[0018] A recursive update module, configured to perform recursive update processing on the fusion features of multiple video segments, and predict, according to the obtained updated fusion features, the video segments including recommendation information among the multiple video segments.

[0019] An embodiment of the present application provides an electronic device, including:

[0020] A memory, configured to store executable instructions;

[0021] A processor, configured to implement the video processing method based on artificial intelligence provided by the embodiment of the present application when executing the executable instructions stored in the memory.

[0022] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, configured to cause a processor to implement the video processing method based on artificial intelligence provided by the embodiment of the present application when executed.

[0023] The embodiment of the present application has the following beneficial effects:

[0024] The video is partitioned into multiple video segments, the content features and operation features of the video segments are determined, and fusion processing is performed to obtain fusion features. In this way, the fusion features include information in multiple dimensions of the video segments. On this basis, recursive update processing is performed on the fusion features of multiple video segments, so that the updated fusion features can accurately and effectively represent the video segments. When predicting the video segments including recommendation information according to the updated fusion features, the prediction accuracy can be improved, and further, it helps to improve the actual utilization rate of the computing resources of the electronic device. Description of the Drawings

[0025] Figure 1 is a schematic architecture diagram of a video processing system based on artificial intelligence provided by an embodiment of the present application;

[0026] Figure 2 is a schematic architecture diagram of a terminal device provided by an embodiment of the present application;

[0027] Figure 3A It is a schematic flowchart of a video processing method based on artificial intelligence provided by an embodiment of the present application;

[0028] Figure 3B It is a schematic flowchart of a video processing method based on artificial intelligence provided by an embodiment of the present application;

[0029] Figure 3C It is a schematic flowchart of a video processing method based on artificial intelligence provided by an embodiment of the present application;

[0030] Figure 3D It is a schematic flowchart of a video processing method based on artificial intelligence provided by an embodiment of the present application;

[0031] Figure 4 It is a schematic diagram for determining multiple types of objects provided by an embodiment of the present application;

[0032] Figure 5 It is a schematic diagram for recursively updating and processing the fusion features of multiple video segments provided by an embodiment of the present application;

[0033] Figure 6 It is a schematic diagram for recursively updating and processing the content features and operation features of a video segment provided by an embodiment of the present application;

[0034] Figure 7 It is a schematic diagram for determining content features provided by an embodiment of the present application;

[0035] Figure 8 It is a schematic architecture diagram of an advertisement prediction model provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0037] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first", "second", and "third" are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In the following description, the term "plurality" means at least two.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0040] In the embodiments of this application, when collecting and processing relevant data in practical applications, the informed consent or separate consent of the personal information subject should be obtained strictly in accordance with the requirements of relevant laws and regulations, and subsequent data use and processing behaviors should be carried out within the scope authorized by laws, regulations, and the personal information subject.

[0041] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0042] 1) Artificial Intelligence (AI): Using a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. Computer Vision Technology (CV) is an important branch of artificial intelligence, mainly studying relevant theories and technologies, and attempting to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. In the embodiments of this application, computer vision technology can be used to identify video segments including recommendation information in a video.

[0043] 2) Content feature: Used to represent the content included in a video segment. The embodiments of this application do not limit the method for determining the content feature. For example, a neural network model can be used to perform feature extraction processing on the video segment to obtain the content feature.

[0044] 3) Object: Represents the subject that can perform specific operations on a video. For example, the object can be a user account or a terminal device. Among them, the operations that can be performed on the video can be preset, such as play operations and skip operations, etc.

[0045] 4) Operation data: The operation data of a video segment, that is, it includes the operations performed on the video segment by several objects.

[0046] 5) Recursive update processing: A processing method for sequence data. For each data in the sequence data, the data is updated by combining it with the adjacent data. In the embodiments of the present application, a Recurrent Neural Network (RNN) model can be used for recursive update processing. A recurrent neural network is also known as a cyclic neural network.

[0047] 6) Recommendation information: Used to recommend specific content, such as products, service content, or cultural and entertainment sports programs, etc. The recommendation information embedded in the video often has a certain difference from the original content of the video. Here, the form of the recommendation information is not limited. For example, it can be an advertisement.

[0048] 7) Negative correlation: A numerical relationship. For example, there are variables A and B, and when A is smaller, B is larger; when A is larger, B is smaller. Then it is said that there is a negative correlation between A and B.

[0049] 8) Fully Connected Layer: Essentially a matrix, used to perform matrix multiplication operations (i.e., linear transformation) on the input vector to extract and integrate the information in the input vector to obtain the output vector. Among them, the matrix multiplication operation is equivalent to weighted processing.

[0050] 9) Database: Similar to an electronic filing cabinet, that is, a place for storing electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. A database can also be understood as a data set stored together in a certain way, shared by multiple users, with as little redundancy as possible, and independent of the application program. In the embodiments of the present application, the database is used to store videos, such as videos on a video platform.

[0051] 10) Big Data: Refers to a data set that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth-rate, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems. In the embodiments of the present application, big data technology can be used to implement the processing of a large number of videos and the training of models, etc.

[0052] The embodiments of the present application provide an artificial intelligence-based video processing method, apparatus, electronic device, and computer-readable storage medium, which can accurately predict video segments including recommendation information in a video, and help improve the actual utilization rate of the computing resources of the electronic device. The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of terminal devices or as a server.

[0053] See Figure 1 , Figure 1 FIG. 1 is a schematic architecture diagram of an artificial intelligence-based video processing system 100 provided by the embodiments of the present application. The terminal device 400 is connected to the server 200 through the network 300, and the server 200 is connected to the database 500. Among them, the network 300 can be a wide area network, a local area network, or a combination of the two.

[0054] In some embodiments, taking the electronic device as a terminal device as an example, the artificial intelligence-based video processing method provided by the embodiments of the present application can be implemented by the terminal device. For example, the terminal device 400 runs the client 410. The client 410 divides the locally stored video into multiple video segments. For each video segment, its content feature and operation feature are determined, and then the content feature and operation feature are fused to obtain a fusion feature. Among them, the client 410 can be a video application, such as a video player. Then, the client 410 performs recursive update processing on the fusion features of multiple video segments, and predicts the video segments including recommendation information in the multiple video segments according to the obtained updated fusion features. The client 410 can delete the video segments including recommendation information predicted in the video, so as to improve the user experience when playing the video. At the same time, it can also save local storage resources and improve the actual utilization rate of the computing resources consumed by the client 410 when playing the video.

[0055] It should be noted that the embodiments of the present application do not limit the timing for the client 410 to predict and delete the video segments including recommendation information in the video. For example, it can be performed while storing the video, or it can be performed when detecting a play operation on the video.

[0056] In some embodiments, taking the electronic device as a server as an example, the video processing method based on artificial intelligence provided by the embodiments of the present application can also be implemented by the server. For example, the server 200 can be a server that provides video services (such as the background server of a short video platform). For the videos stored in the database 500 (such as the videos uploaded by uploaders to the short video platform), the server 200 predicts the video segments that include recommendation information and deletes the video segments that include recommendation information. Since the storage space occupied by the video will be reduced after deleting the video segments that include recommendation information in the predicted video, the storage resources of the server 200 can be saved. The server 200 can also provide video services to the client 410 (such as the foreground client of the short video platform) when receiving a video request from the client 410, that is, send the video from which the predicted video segments that include recommendation information have been deleted to the client 410 for presentation (playback) on the client 410. Figure 1 Taking video A as an example. In this way, the actual utilization rate of the computing resources consumed by the server 200 to send the video and the computing resources consumed by the client 410 to present the video can be improved.

[0057] It should be noted that the process of predicting the video segments that include recommendation information in the video can be integrated in the recommendation information prediction model, and the electronic device (such as the terminal device 400 or the server 200) realizes the prediction by calling the recommendation information prediction model. Before using the recommendation information prediction model, in order to improve the prediction accuracy, the recommendation information prediction model can be trained. For example, the server 200 trains the recommendation information prediction model according to the data set and calls the trained recommendation information prediction model to realize the prediction of the video segments that include recommendation information. Or, the server can also send the trained recommendation information prediction model to the client 410 for local deployment on the client 410, so that the client 410 can call the trained recommendation information prediction model to realize the prediction of the video segments that include recommendation information. Among them, the database includes sample videos and the video segments that include recommendation information marked in each sample video, such as obtained by manual annotation.

[0058] In some embodiments, the terminal device 400 or the server 200 may implement the AI-based video processing method provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or software module in the operating system; it may be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a video application (corresponding to the client 410 above), specifically, a video player, etc.; it may also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it may also be a small program that can be embedded in any APP, such as a small program component embedded in a video application. Among them, the small program component can be controlled by the user to run or close. In short, the above computer program may be any form of application, module or plug-in.

[0059] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and AI platforms. Among them, the cloud service may be a video processing service for the terminal device 400 to call to predict video segments including recommendation information in the video. The terminal device 400 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0060] Taking the electronic device provided in the embodiments of the present application as the terminal device as an example, it can be understood that for the case where the electronic device is a server, Figure 2 Some parts in the shown structure (such as the user interface, the presentation module, and the input processing module) may be omitted. Refer to Figure 2 Figure 2 is a schematic structural diagram of the terminal device 400 provided in the embodiments of the present application. Figure 2 The terminal device 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0061] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0062] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0063] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0064] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0065] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0066] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0067] A network communication module 452 for reaching other computing devices via one or more (wired or wireless) network interfaces 420 , exemplary network interfaces 420 including Bluetooth, WiFi, and USB;

[0068] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0069] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0070] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 An AI-based video processing device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a segmentation module 4551, a feature extraction module 4552, a feature mapping module 4553, a fusion module 4554, and a recursive update module 4555. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0071] The artificial intelligence-based video processing method provided in the embodiments of the present application will be explained in combination with the exemplary application and implementation of the electronic device provided in the embodiments of the present application.

[0072] See also Figure 3A , Figure 3A This is a flow chart of the video processing method based on artificial intelligence provided by the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0073] In step 101, a video is divided into a plurality of video segments.

[0074] Here, the video to be processed is divided into multiple video segments. The embodiment of the present application does not limit the division method. For example, the video can be divided evenly according to the set division time or the set number of video segments. For example, if the duration of the video to be processed is 1 minute and the division time is 5 seconds, the video can be divided evenly into 12 video segments according to the division time, and the duration of each video segment is 5 seconds; for another example, if the set number of video segments is 10, the video can be divided evenly into 10 video segments according to the number of video segments, and the duration of each video segment is 6 seconds.

[0075] In addition, the embodiment of the present application does not limit the source of the video to be processed. For example, it can be a video obtained from the outside world or a video stored locally.

[0076] In step 102, feature extraction processing is performed on the video clip to obtain content features of the video clip.

[0077] For each divided video segment, perform feature extraction processing on the video segment to obtain content features representing the content of the video segment, and the content features may be in vector form. Here, a model constructed based on the principles of artificial intelligence, such as a neural network model, can be used to implement the feature extraction processing. It should be noted that after obtaining the content features, further linear transformation can be performed on the content features to extract and integrate useful information in the content features.

[0078] In some embodiments, the above-mentioned feature extraction processing of the video segment to obtain the content features of the video segment can be implemented in the following way: perform any one of the following processes: perform feature extraction processing on each image in the video segment to obtain image features, and fuse the image features of multiple images into the content features of the video segment; perform feature extraction processing on multiple consecutive images in the video segment, and use the extracted image features as the content features of the video segment.

[0079] The embodiments of the present application provide two ways to determine content features. The first way is to perform feature extraction processing on each image in the video segment separately to obtain image features, and fuse the image features corresponding to all the images in the video segment into the content features of the video segment. Among them, there is no limitation on the way of fusing to obtain the content features. For example, it can be splicing processing, summation processing or weighted summation, etc. In the first way, a specific model, such as a two-dimensional convolutional neural network (2D Convolutional Neural Network, 2D CNN) model, can be used to implement the feature extraction processing of a single image.

[0080] The second way is to perform feature extraction processing on multiple consecutive images (such as all images) in the video segment at the same time, and use the extracted image features as the content features of the video segment. Similarly, a specific model, such as a three-dimensional convolutional neural network (3D Convolutional Neural Network, 3D CNN) model, can be used to implement the feature extraction processing of multiple consecutive images. Through the above methods, the flexibility of determining content features can be improved, and applicable methods can be selected according to the actual application scenario.

[0081] In step 103, obtain operation data for the video segment, and perform feature mapping processing on the operation data to obtain operation features.

[0082] For each video segment in the video to be processed, in addition to determining the content features, the operations performed by a specific object on the video segment are also obtained as operation data. Among them, the operations that can be performed on the video segment can be preset, such as including play operation, pause operation, skip operation, etc. For the skip operation, the start position and end position of the skip operation for the video to be processed can be obtained, and then it can be determined which video segments in the video the skip operation is performed on. For example, a 1-minute video is evenly divided into 10 video segments. The start position of a certain skip operation for the video is the 5th second, and the end position is the 10th second. Then it can be determined that the skip operation is performed on the 1st and 2nd video segments in the video. Among them, the position of the video segment (i.e., which video segment) is counted in the time axis order of the video, and the time axis order is the order from the start of the video to the end of the video.

[0083] The object in the embodiments of the present application refers to an object that can perform the above-set operations on the video. The object can be a terminal device or a user account of a video platform, etc., and no limitation is made thereto. When obtaining the operation data for the video segment, the user account can be used as a unit for obtaining, or the terminal device can be used as a unit for obtaining. In the latter case, it is default that the users of the same terminal device are fixed. Among them, different terminal devices can be distinguished by device identifiers, such as International Mobile Equipment Identity (IMEI), etc. The object that needs to obtain the operation data can be preset according to the actual application scenario. For example, in a video platform, the operation data of all registered user accounts for the video segment can be obtained, or the operation data of a specific registered user account for the video segment can be obtained.

[0084] After obtaining the operation data for the video segment, for the convenience of subsequent processing by the electronic device, the operation data is subjected to feature mapping processing to obtain numerical operation features. In the actual application scenario, most users will perform a skip operation on the recommended information when watching a video including recommended information, that is, the skip operation can largely reflect the existence of the recommended information. Therefore, in the process of feature mapping processing for the operation performed by a certain object on the video segment in the operation data, it can be determined whether to map the performed operation to a first set value or a second set value according to whether the performed operation includes a skip operation. Of course, this does not constitute a limitation on the feature mapping processing. After the feature mapping processing, the obtained operation features can be in vector form (consisting of multiple values) or in numerical form (only including a single value).

[0085] In step 104, the content features and operation features of the video clip are fused to obtain fused features.

[0086] Here, the content features representing the content of the video clip and the operation features representing the operations performed on the video clip are fused to obtain fused features. The embodiments of the present application do not limit the manner of fusion processing. For example, it can be splicing processing or weighted summation, etc.

[0087] In step 105, the fused features of multiple video clips are recursively updated, and based on the obtained updated fused features, the video clips including recommendation information among the multiple video clips are predicted.

[0088] After obtaining the fused features of each video clip in the video, the fused features of all video clips are sorted in the order of the time axis of the video, and the sorted fused features are recursively updated. Among them, for a certain fused feature A, the recursive update processing means updating the fused feature A itself according to the adjacent fused feature (such as the previous fused feature). Through the recursive update processing, the fused features of video clips can be effectively updated according to the sequential relevance existing among multiple video clips themselves, so that the updated fused features can more accurately and effectively represent the video clips. Generally speaking, in a video including recommendation information, the recommendation information only accounts for a small part, and most of the video is normal content, and the relevance between the recommendation information and the normal content is small or non-existent. Through the recursive update processing, the difference between the video clips including recommendation information and the video clips not including recommendation information in the updated fused features can also be increased.

[0089] After the recursive update processing, the updated fused features of each video clip can be obtained. Then, based on these updated fused features, the video clips including recommendation information among all video clips can be predicted, and the prediction method will be elaborated in detail later. The embodiments of the present application do not limit the application of the predicted video clips including recommendation information. For example, the predicted video clips including recommendation information in the video can be directly deleted, or a deletion option can be provided, and when a confirmation operation for the deletion option is received, the predicted video clips including recommendation information in the video are deleted. In this way, the actual utilization rate of the computing resources consumed in the subsequent video presentation process can be improved, and at the same time, the experience of the user watching the video can be improved.

[0090] It should be noted that the above steps can be integrated into the recommendation information prediction model, and the electronic device executes the corresponding steps by invoking the recommendation information prediction model. For example, step 105 can be integrated into the recommendation information prediction model, and step 104 and step 105 can also be integrated into the recommendation information prediction model, and there is no limitation in this regard. To improve the prediction accuracy of the recommendation information prediction model, before actually invoking the recommendation information prediction model, the recommendation information prediction model can be supervised learning, that is, trained.

[0091] For example, the dataset for training can include multiple sample videos and video segments with labeled recommendation information in each sample video. Among them, the video segments with recommendation information in the sample videos can be obtained through manual annotation. Of course, other annotation methods can also be applied. After processing the sample videos through the recommendation information prediction model, the difference (i.e., the loss value) between the predicted video segments with recommendation information and the labeled video segments with recommendation information is determined. According to this difference, backpropagation is performed in the recommendation information prediction model, and during the backpropagation process, the weight parameters of the recommendation information prediction model are updated. Among them, there is no limitation on the loss function used to calculate the loss value. For example, it can be a cross-entropy loss function. After training the recommendation information prediction model with the dataset, the prediction accuracy of the recommendation information prediction model for video segments with recommendation information can be improved.

[0092] As Figure 3A shown, the embodiment of the present application divides the video into multiple video segments and determines the fusion features according to the information of multiple dimensions of the video segments. On this basis, recursive update processing is performed on the fusion features of multiple video segments, which can further improve the accuracy of the fusion features. In this way, when predicting the video segments with recommendation information according to the updated fusion features, the prediction accuracy can be improved, which in turn helps to improve the actual utilization rate of the computing resources of the electronic device.

[0093] In some embodiments, referring to Figure 3B , Figure 3B is a schematic flowchart of a video processing method based on artificial intelligence provided by an embodiment of the present application. Figure 3A The steps shown in step 103 can be implemented through steps 201 to 206, and each step will be described in combination.

[0094] In step 201, the objects other than the target recommendation object of the video are used as the first objects.

[0095] In the embodiments of the present application, operation data of a specific object for a video segment can be obtained. The specific object can include three types: a first object, a second object, and a third object, which will be described separately below. In the process of determining the first object, first, the target recommended object of the video is determined. Among them, the target recommended object can be set manually or obtained through other means.

[0096] Taking the case where the object is a terminal device as an example, when the electronic device is a server, the server can, when detecting a video request sent by a certain terminal device, use the terminal device as the target recommended object of the video requested by the video request; when the electronic device is a terminal device, it can use itself as the target recommended object of the video to be presented.

[0097] Taking the case where the object is a user account as an example, when the electronic device is a server, the server can, when detecting a video request sent by a certain user account, use the user account as the target recommended object of the video requested by the video request; when the electronic device is a terminal device, the terminal device can use the user account in the logged-in state in the running video application as the target recommended object of the video to be presented by the application. Of course, the method of determining the target recommended object is not limited to this.

[0098] After determining the target recommended object, the objects other than the target recommended object are used as the first object. As an example, the embodiments of the present application provide a schematic diagram for determining multiple types of objects as shown in Figure 4 In Figure 4 , taking the case where the object is a user account as an example, all registered user accounts (such as registered in a video application) are shown, specifically including user accounts A, B, C, D, and E. When the target recommended object is determined to be user account A, the remaining user accounts B, C, D, and E are all used as the first object. The operations performed by the first object on the video segment can be used as a global standard and have reference value.

[0099] In step 202, the objects among the multiple first objects that are similar to the target recommended object are used as the second object.

[0100] After determining the first object through step 201, the objects among all the first objects that are similar to the target recommended object are used as the second object. The embodiments of the present application do not limit the method for determining whether objects are similar. For example, by traversing all the first objects, when the information of the traversed first object (such as registration information in a video application) is the same as the information of the target recommended object, the traversed first object is used as the second object, where the information can include at least one of gender, city where located, and hobbies, but is not limited thereto.

[0101] InFigure 4 Taking the case where user account B is similar to user account A and user account C is also similar to user account A as an example, both user accounts B and C are taken as the second object.

[0102] In some embodiments, the object similar to the target recommended object among the multiple first objects can be taken as the second object in the following way: constructing an initial vector with the same dimension as the number of videos in the database; where each value in the initial vector corresponds to a video in the database; for any object, determining the value in the initial vector corresponding to the video on which any object has performed a playback operation, and updating the determined value in the initial vector to a set value to obtain a video playback vector; among the multiple first objects, determining the object whose similarity with the target recommended object on the video playback vector is greater than the similarity threshold as the second object.

[0103] In the embodiments of the present application, starting from the playback operation on the video, the second object can be determined. First, construct an initial vector with the same dimension as the number of videos in the database, where each value in the initial vector corresponds to a video in the database. For the convenience of processing, all values in the initial vector can be initialized to zero. The database is, for example, the database of a video application (video platform), which is used to store all videos related to this application. Of course, all videos related to this application can also be stored in other ways, such as in a distributed file system of a server, or stored locally on a terminal device, etc. Here, only the case of storing in the database is used for illustrative purposes.

[0104] For each object, determine the value in the initial vector corresponding to the video on which the object has performed a playback operation, and update the determined value to a set value (such as the value 1) to obtain the video playback vector of the object. The obtained video playback vector can directly reflect the video preferences of the object. Then, traverse all the first objects. When the similarity between the traversed first object and the target recommended object on the video playback vector is greater than the similarity threshold, take the traversed first object as the object similar to the target recommended object, that is, the second object. Among them, the similarity can be cosine similarity or other similarities, and the similarity threshold can be set according to the actual application scenario. Through the above method, the second object with video preferences similar to those of the target recommended object can be determined.

[0105] In step 203, the object having a follow-up relationship with the target recommended object among the multiple first objects is taken as the third object.

[0106] Here, the objects among all the first objects that have a following relationship with the target recommended object can be used as the third objects. Among them, the following relationship can be a one-way following relationship, such as following the target recommended object or being followed by the target recommended object, or a mutual following relationship.

[0107] In Figure 4 For example, in the case where user account D and user account A have a mutual following relationship, and at the same time user account E also has a mutual following relationship with user account A, both user account D and E are used as the third objects.

[0108] It should be noted that the first object, the second object, and the third object in the embodiments of the present application are not isolated from each other. For example, a certain object may be the first object, the second object, and the third object at the same time.

[0109] In step 204, the operations performed by the first object, the second object, and the third object on the video segment are used as the operation data for the video segment.

[0110] After determining the objects at different levels, that is, the first object, the second object, and the third object, the operations performed by all the first objects on the video segment, the operations performed by all the second objects on the video segment, and the operations performed by all the third objects on the video segment are jointly used as the operation data for the video segment. In this way, the operation data includes feedback information at different levels.

[0111] In step 205, the operations performed by the sample object on the video segment are mapped to operation values, and sub-operation features are constructed based on the operation values of multiple sample objects; where the sample object is any one of the first object, the second object, and the third object.

[0112] In the process of performing feature mapping processing on the operation data, the operations performed by each type of object in the operation data on the video segment are processed separately. For the sake of understanding, the process of processing the operations performed by the sample object on the video segment is described, where the sample object is any one of the first object, the second object, and the third object.

[0113] For the operations performed by each sample object in the operation data on the video segment, they are mapped to specific values. For the sake of distinction, the values obtained here are named operation values. In the embodiments of the present application, the skip operation can largely reflect the existence of recommended information. Therefore, the operations performed by the sample object on the video segment can be mapped to different operation values according to whether the operations include the skip operation.

[0114] After obtaining the operation values of each sample object, sub-operation features are constructed based on the operation values of all sample objects, that is, sub-operation features corresponding to the first object are constructed according to the operation values of all first objects, sub-operation features corresponding to the second object are constructed according to the operation values of all second objects, and sub-operation features corresponding to the third object are constructed according to the operation values of all third objects.

[0115] In some embodiments, the above mapping of the operations performed by the sample object on the video segment to operation values can be achieved in the following way: when the operation performed by the sample object on the video segment includes a skip operation, the first set value is used as the operation value of the sample object; when the operation performed by the sample object on the video segment does not include a skip operation, the second set value is used as the operation value of the sample object.

[0116] Here, the first set value and the second set value are different and can be set according to the actual application scenario. For example, the first set value is 1 and the second set value is 0. When the operation performed by a certain sample object on the video segment includes a skip operation, the first set value is used as the operation value of the sample object; when the operation performed by a certain sample object on the video segment does not include a skip operation, the second set value is used as the operation value of the sample object. In this way, a differential mapping of values can be achieved according to the result of whether the performed operation includes a skip operation, that is, the obtained operation value can directly reflect whether the corresponding operation includes a skip operation.

[0117] In some embodiments, the above construction of sub-operation features based on the operation values of multiple sample objects can be achieved in the following way: perform any one of the following processes: perform an accumulation process on the operation values of multiple sample objects, and use the value obtained from the accumulation process as the sub-operation feature; construct a vector with the same dimension as the number of sample objects based on the operation values of multiple sample objects as the sub-operation feature.

[0118] The embodiments of the present application provide two ways to construct sub-operation features. For the convenience of understanding, the case where the sample object is the first object is taken as an example. The first way is to perform an accumulation process on the operation values of all first objects, and use the value obtained from the accumulation process as the sub-operation feature corresponding to the first object, that is, the obtained sub-operation feature is in the form of a value. Among them, the accumulation process can be a summation process or a weighted summation, etc. The amount of information of the sub-operation feature obtained by this method is small, and the subsequent calculation amount is also small, which is suitable for scenarios where the computing power of the electronic device is weak or the efficiency requirement for video processing is high (such as real-time video processing).

[0119] The second way is to construct a vector with the same dimension as the number of the first objects according to the operation values of all the first objects, so as to serve as the sub-operation feature corresponding to the first object. That is, the sub-operation feature obtained by the second way is in the form of a vector, and each value therein is the same as the operation value of a first object. Compared with the first way, the sub-operation feature obtained by the second way has a larger amount of information, and the subsequent calculation amount is also larger, which is applicable to the scenario where the computing power of the electronic device is relatively strong or the timeliness requirement for video processing is not high (such as offline video processing).

[0120] In some embodiments, the above cumulative processing of the operation values of multiple sample objects can be implemented in such a way that the value obtained by the cumulative processing is used as the sub-operation feature: dividing the number of videos for which the sample object has performed a skip operation by the video base number to obtain the skip video ratio; determining a weight negatively correlated with the skip video ratio; performing weighted processing on the operation values of multiple sample objects according to the weight of the sample object to obtain the sub-operation feature; wherein, the video base number includes any one of the number of videos in the database and the number of videos for which the sample object has performed a play operation.

[0121] Here, the cumulative processing can be weighted summation. For the sake of easy understanding, the case where the sample object is the first object is taken as an example to illustrate the process of the cumulative processing. For each first object, dividing the number of videos for which the first object has performed a skip operation by the video base number to obtain the skip video ratio of the first object. The skip video ratio can directly reflect the skip habit of the user corresponding to the first object. The smaller the skip video ratio is, the greater the possibility that the videos for which the first object has performed a skip operation include recommended information (that is, the user corresponding to the first object doesn't like to skip very much). Among them, the video base number can be the number of videos in the database or the number of videos for which the first object has performed a play operation. Since there may be a large difference in the number of videos for which different first objects have performed a play operation, taking the number of videos for which the first object has performed a play operation as the video base number can further improve the accuracy of the determined skip video ratio.

[0122] After determining the skip video ratio of the first object, a weight negatively correlated with the skip video ratio is determined. The negative correlation between the skip video ratio and the weight can be set according to the actual application scenario. For example, the weight can be the reciprocal of the skip video ratio. Then, according to the weight of each first object, the operation values of all first objects are weighted and summed to obtain the sub-operation feature corresponding to the first object. For example, the first objects include A, B, and C, the operation values are 1, 0, and 1 respectively, and the weights are 0.2, 0.5, and 0.3 respectively. Then the sub-operation feature corresponding to the first object is 0.2×1 + 0.5×0 + 0.3×1 = 0.5. In the above way, the weight of the operation value of the sample object can be adaptively adjusted according to the skip video ratio of the sample object, so that the weighted operation value is more in line with the actual situation and the accuracy of the finally obtained sub-operation feature is improved.

[0123] In some embodiments, the above-mentioned weight negatively correlated with the skip video ratio can be determined in the following way: Sort multiple sample objects according to the first numerical order of the skip video ratio; Uniformly sample the set weight range according to the number of sample objects to obtain multiple weights to be assigned; According to the second numerical order, assign the multiple weights to be assigned to the multiple sorted sample objects in sequence; where the first numerical order is opposite to the second numerical order.

[0124] An example of the negative correlation is provided in the embodiments of the present application. For the convenience of understanding, the case where the sample object is the first object is used as an example. First, all first objects are sorted according to the first numerical order of the skip video ratio, where the first numerical order can be from largest to smallest or from smallest to largest. At the same time, the set weight range is uniformly sampled (i.e., equally spaced sampling) according to the number of first objects to obtain multiple weights to be assigned. For example, the set weight range is [0, 1] and the number of first objects is 3. After uniform sampling, the weights to be assigned can include 0, 0.5, and 1.

[0125] Then, according to the second numerical order opposite to the first numerical order, all the weights to be assigned are sequentially assigned to all the sorted first objects, where each first object is assigned a weight. For example, the first numerical order is from largest to smallest, and the second numerical order is from smallest to largest. Then, when assigning, the smallest weight to be assigned is assigned to the first object ranked first, that is, the first object with the largest skip video ratio, and so on. The above way provides an example of determining the weight of the sample object, but this does not limit the embodiments of the present application.

[0126] In some embodiments, the above-mentioned operation values of multiple sample objects can be used to construct a vector with the same dimension as the number of sample objects in the following way: sort the operation values of multiple sample objects according to the object parameters of the sample objects; construct a vector with the same dimension as the number of sample objects according to the sorted multiple operation values; wherein, the object parameters include any one of the registration time, the time of the last execution of the playback operation, and the number of videos for which the playback operation has been executed.

[0127] In addition to performing cumulative processing to obtain sub-operation features in numerical form, sub-operation features in vector form can also be constructed. For the sake of easy understanding, the case where the sample object is the first object is taken as an example. Since a vector has a direction and the order of the values in the vector will affect subsequent processing, in the embodiments of the present application, the operation values of all first objects can be sorted according to the object parameters of each first object, that is, the order of the operation values is regularized. Among them, the object parameters can include any one of the registration time (such as the registration time in a video application), the time of the last execution of the playback operation, and the number of videos for which the playback operation has been executed, and of course, it is not limited thereto. In addition, the order followed during sorting is not limited. For example, when the object parameter is the registration time, the operation values of all first objects can be sorted in ascending order of the registration time of the first objects.

[0128] Then, according to all the sorted operation values, construct a vector with the same dimension as the number of first objects as the sub-operation feature corresponding to the first object. For example, if the sorted operation values are 0, 0, 1, and 1 in sequence, the constructed vector is [0, 0, 1, 1]. It should be noted that in the process of constructing the sub-operation feature in vector form, the skip video ratio of the first object can also be determined, and the operation values of the first object can be weighted according to the weight negatively correlated with the skip video ratio (multiply the operation values of the first object by the weight of the first object), and then a vector is constructed according to the weighted operation values.

[0129] Through the above method, it can be ensured that the sorting principles of the operation values in the vectors corresponding to the first object, the second object, and the third object are the same, which is convenient for learning effective rules from these vectors.

[0130] In step 206, the sub-operation features corresponding to the first object, the sub-operation features corresponding to the second object, and the sub-operation features corresponding to the third object are concatenated into the operation feature of the video segment.

[0131] After obtaining the sub-operation features corresponding to the first object, the sub-operation features corresponding to the second object, and the sub-operation features corresponding to the third object, these sub-operation features at different levels are fused together, for example, by splicing, to obtain the operation features of the video segment. Of course, other methods can also be used here, such as weighted summation, etc. Figure 3B Only splicing is used as an example in

[0132] such as Figure 3B As shown, in the embodiments of the present application, by determining objects at different levels and then fusing the sub-operation features at different levels into the operation features of the video segment, the comprehensiveness and hierarchy of the information contained in the operation features can be improved, which is convenient for learning useful rules from the operation features subsequently.

[0133] In some embodiments, referring to Figure 3C , Figure 3C is a schematic flowchart of a video processing method based on artificial intelligence provided by the embodiments of the present application. Figure 3A The step 104 shown can be implemented through steps 301 to 302, and will be described in conjunction with each step.

[0134] In step 301, the content features and operation features of the video segment are recursively updated to obtain the updated content features and the updated operation features.

[0135] The embodiments of the present application provide an example of fusing the content features and operation features of a video segment. First, the content features and operation features of the video segment are sorted and recursively updated to obtain the updated content features and the updated operation features, that is, to make the features more accurate. Here, the sorting order is not limited and can be set according to the actual application scenario. For example, when the operation features are obtained by splicing the sub-operation features corresponding to the first object, the sub-operation features corresponding to the second object, and the sub-operation features corresponding to the third object, the sorting can be performed in the order of content features - sub-operation features corresponding to the first object - sub-operation features corresponding to the second object - sub-operation features corresponding to the third object.

[0136] In step 302, the updated content features and the updated operation features are spliced to obtain the fusion features.

[0137] Here, the updated content features and the updated operation features of the video segment are spliced to obtain the fusion features. Of course, other methods can also be used here to obtain the fusion features, such as weighted summation, etc. Figure 3C Only splicing is used as an example in

[0138] In Figure 3C , Figure 3AThe shown step 105 can be implemented through steps 303 to 305, which will be described in conjunction with each step.

[0139] In step 303, the fusion features of multiple video segments are recursively updated to obtain the updated fusion features.

[0140] In step 304, the updated fusion features of the video segments are weighted to obtain the probability that the video segments include recommendation information.

[0141] Here, the updated fusion features of the video segments can be linearly transformed, that is, weighted, and finally a numerical value is obtained. This numerical value is the probability that the video segment includes recommendation information. Among them, this numerical value can also be the probability that the video segment does not include recommendation information, depending on the actual application scenario. Here, only the probability that the video segment includes recommendation information is used as an example for illustration.

[0142] In step 305, the video segments in the video whose probability of including recommendation information is greater than the probability threshold are used as the predicted video segments including recommendation information.

[0143] Here, the probability threshold can be preset, such as set to 0.5. After obtaining the probability that each video segment includes recommendation information, the video segments whose probability of including recommendation information is greater than the probability threshold are used as the predicted video segments including recommendation information. For a video, the number of predicted video segments including recommendation information may be 0, 1, or multiple.

[0144] In Figure 3C After step 305, in step 306, the predicted video segments including recommendation information can also be deleted.

[0145] If the video segments including recommendation information have been predicted, the predicted video segments including recommendation information can be deleted from the video. In order to reduce the adverse effects (such as picture fragmentation) on the original content of the video (that is, the normal content except the recommendation information) after deletion, in step 101, the duration of the video segments can be minimized as much as possible, for example, the video is divided according to a division duration of 5 seconds.

[0146] In an actual application scenario, taking the case where the electronic device is a server as an example, the server can process the video uploaded by the video uploader or the video requested by the video request sent by the terminal device, that is, predict the video segments including recommendation information in the video and delete them. Taking the case where the electronic device is a terminal device as an example, the terminal device can process the video stored locally or the video to be played soon, that is, predict the video segments including recommendation information in the video and delete them. Of course, the video processing scenario is not limited to this.

[0147] As Figure 3C shown, by deleting the video segments predicted to include recommendation information in the video in the embodiments of the present application, the storage space occupied by the video can be reduced, and the storage resources of the electronic device can be saved. When subsequent operations such as sending or playing the video are performed, the actual utilization rate of the computing resources consumed by the electronic device can also be improved.

[0148] In some embodiments, referring to Figure 3D , Figure 3D FIG. is a schematic flowchart of an artificial intelligence-based video processing method provided by the embodiments of the present application. Figure 3A The step 105 shown can be implemented through steps 401 to 404, and will be described in conjunction with each step.

[0149] In step 401, according to the time axis order of the video, the fusion features of multiple video segments are sorted.

[0150] Here, an example of the recursive update processing method is described. After obtaining the fusion features of all video segments in the video, the fusion features of all video segments can be sorted according to the time axis order of the video, that is, the order from the start to the end of the video.

[0151] In step 402, according to the fusion feature of any video segment and the intermediate feature of the previous video segment, the intermediate feature of any video segment is fused.

[0152] As an example, the embodiments of the present application provide a schematic diagram of recursively updating the fusion features of multiple video segments as shown in Figure 5 . In Figure 5 , the sorted fusion features are successively the fusion feature of video segment 1, the fusion feature of video segment 2, and the fusion feature of video segment 3. For the sake of easy understanding, taking video segment 2 as an example, a method of fusing to obtain the intermediate feature is described. Among them, the intermediate feature is also called the hidden layer state in the RNN structure. First, the fusion feature of video segment 2 and the intermediate feature of the previous video segment (i.e., video segment 1) are weighted and summed, the result of the weighted sum is added with a bias term, and the obtained result is activated to obtain the intermediate feature of video segment 2. Among them, the activation processing can be implemented by an activation function such as the hyperbolic tangent function. For different video segments, the weight parameters corresponding to the fusion features in the weighted sum process can be the same, and the weight parameters and bias terms corresponding to the intermediate features of the previous video segment are the same, that is, different video segments share the weight parameters and bias terms.

[0153] It should be noted that for the first video segment after sorting (such as Figure 5For the video segment 1) in [ ], since there is no previous video segment, the intermediate features of the previous video segment can be set to zero during the fusion process.

[0154] In step 403, the intermediate features are weighted to obtain the updated fusion features of any video segment.

[0155] Here, taking Figure 5 the video segment 2 in [ ] as an example, the intermediate features of the video segment 2 are weighted (a bias term can also be added after the weighting process, and the meaning of the bias term here is different from that in step 402), to obtain the updated fusion features of the video segment 2. After updating the fusion features according to the sequential relevance between video segments, the updated fusion features can more accurately and effectively represent the corresponding video segments.

[0156] It should be noted that the recursive update processing method shown in steps 402 to 403 is also applicable to step 301. As an example, the embodiments of the present application provide a schematic diagram of the recursive update processing of the content features and operation features of video segments as shown in Figure 6 [ ]. In Figure 6 [ ], four features, namely, the content features obtained after sorting of a certain video segment, the sub-operation features corresponding to the first object, the sub-operation features corresponding to the second object, and the sub-operation features corresponding to the third object, are shown. For ease of description, they are named feature 1, feature 2, feature 3, and feature 4 respectively. Taking feature 2 as an example, during the recursive update processing, according to feature 2 and the intermediate features corresponding to the previous feature (i.e., feature 1), the intermediate features corresponding to feature 2 are fused, and then the intermediate features corresponding to feature 2 are weighted to obtain the updated feature 2, and so on.

[0157] In step 404, based on the obtained updated fusion features, the video segments including recommendation information among multiple video segments are predicted.

[0158] Based on the updated fusion features of each video segment in the video, the video segments including recommendation information can be predicted.

[0159] It should be noted that the above steps (such as steps 402 to 404) can be integrated into an RNN model, taking the sorted multiple fusion features as the input of the RNN model, and implementing the recursive update processing through the RNN structure in the RNN model. During the training of the RNN model, the weight parameters and bias terms used in the weighting processes in steps 402 and 403 can be updated to make the effect of the recursive update processing better. Among them, the purpose of adding the bias term is to help the RNN model converge quickly and improve the training effect.

[0160] Such asFigure 3D As shown in the figure, based on the principle of RNN, the embodiment of the present application recursively updates the fusion features of all video segments in the video, so that the updated fusion features can more accurately and effectively represent the corresponding video segments, which helps to improve the prediction accuracy of video segments including recommendation information.

[0161] Next, an exemplary application of the embodiment of the present application in an actual application scenario will be described. For the sake of easy understanding, an example will be given where the recommendation information is an advertisement and the object is a user account. In a video scenario, in order to increase their own revenue, video uploaders on a video platform may add advertisements to the video, forcing users to watch the embedded advertisements in the video when watching the video. Such videos with embedded advertisements will result in a poor viewing experience for users. At the same time, since the video uploader recommends advertisements through illegal channels (i.e., adding advertisements to the video), the video platform cannot obtain the due revenue.

[0162] The embodiment of the present application provides an artificial intelligence-based video processing method, which can be applied to video applications, such as the background server of an application, or the front-end client of an application, such as various video players. Specifically, video processing can be implemented through a video content representation module, a user behavior collection and analysis module, and an advertisement prediction module, which will be described separately below.

[0163] 1) Video content representation module.

[0164] For a video to be processed, it is evenly divided into multiple video segments, and the division standard can be preset. For example, the video to be processed is divided according to a duration of 5 seconds to obtain multiple video segments. Then, each obtained video segment is modeled to obtain content features for representing the content of the video segment. As an example, the embodiment of the present application provides a schematic diagram for determining content features as shown in Figure 7 In Figure 7 , the continuous multiple images in the video segment can be subjected to feature extraction processing through a 3D CNN model, and the extracted image features (expressed in vector form) are processed through two fully connected layers to obtain the content features of the video segment. Among them, the fully connected layer is the fully connected layer, and the purpose is to perform a linear transformation on the vector, that is, to perform a weighted process. In a deep learning model, usually, the more times of linear transformation (that is, the deeper the model), the better the effect, and the greater the computational complexity. Of course, the above process of determining content features does not limit the embodiment of the present application. For example, the image features output by the 3D CNN model can be directly used as content features.

[0165] 2) User behavior collection and analysis module.

[0166] In this module, the following information is collected:

[0167] ① For all user accounts except the current user account of the video to be processed, determine whether each user account has performed a skip operation on the video to be processed, as well as the start position and end position corresponding to the skip operation. Based on the start position and end position corresponding to the skip operation, it is possible to determine which video segments in the video to be processed have been skipped. Here, the current user account of the video to be processed corresponds to the target recommendation object above, and the user accounts except the current user account correspond to the first object above. The current user account of the video to be processed can be set in advance, or the user account in the logged-in state in the video application can be used as the current user account, and there is no limitation on this.

[0168] ② For all user accounts except the current user account of the video to be processed, determine the ratio between the number of videos for which each user account has performed a skip operation and the number of videos for which each user account has performed a play operation. For the convenience of description, this ratio will be named the skipped video ratio in the following text.

[0169] ③ Determine the user accounts (corresponding to the third object above) that have a mutual follow relationship with the current user account of the video to be processed.

[0170] Based on the above information collected, three different user behavior characteristics of the video segments (corresponding to the sub-operation characteristics above) can be obtained:

[0171] ① Global user skip behavior characteristic (i.e., the sub-operation characteristic corresponding to the first object above).

[0172] Here, for each video segment of the video to be processed, among all user accounts except the current user account, count the user accounts that have performed a skip operation on the video segment. For each user account counted, set the operation value of this user account to 1 (corresponding to the first set value above). Finally, accumulate the operation values of all user accounts except the current user account, and use the value obtained from the accumulation process as the global user skip behavior characteristic of this video segment. Among them, the accumulation process is such as the summation process.

[0173] When the proportion of skipped videos of a user account is smaller (i.e., the user holding the user account doesn't like to skip much), if the user account performs a skip operation on a certain video segment, then this video segment is more likely to be a video segment including an advertisement. Therefore, during the accumulation process, the operation values corresponding to different user accounts can be weighted according to the proportion of skipped videos. For example, all user accounts except the current user account can be sorted according to the proportion of skipped videos, and the weight of the user account with the largest proportion of skipped videos is assigned 0, and the weight of the user account with the smallest proportion of skipped videos is assigned 1. For other user accounts, they are evenly distributed according to the ranking. In this way, each user account can obtain an assigned weight.

[0174] During the accumulation process, the weight of the user account is multiplied by the operation value corresponding to the user account to achieve weighting. Finally, the weighted operation values of all user accounts except the current user account are summed to obtain the global user skip behavior feature. For example, if the weight of a user account is 0.5 and the operation value is 1, then the weighted operation value is 0.5×1 = 0.5.

[0175] ② Skip behavior feature of like-minded friends (i.e., the sub-operation feature corresponding to the second object above).

[0176] Here, user accounts with video preferences similar to those of the current user account are determined as like-minded friends, corresponding to the second object above. For example, if the current user account is account A and another user account is account B, then the calculation can be performed according to the following formula:

[0177] Account similarity = cos(video play vector of account A, video play vector of account B).

[0178] Among them, cos refers to the cosine function, which is used to calculate the cosine similarity between account A and account B in the video play vector as the account similarity between the two accounts. The video play vector is a one-hot vector, and the dimension is the number of videos in the database. Each dimension corresponds to a video. If the user account has performed a play operation on a certain video, then the value corresponding to this video in the video play vector of the user account is assigned 1; conversely, if the user account has not performed a play operation on this video, then the value corresponding to this video in the video play vector of the user account is assigned 0.

[0179] After determining the account similarity between the current user account and each other user account, the user accounts with an account similarity greater than the similarity threshold are taken as the like-minded friends of the current user account. Then, for all the like-minded friends of the current user account, the like-minded friend skip behavior features of each video segment in the video to be processed are calculated in the same way as calculating the global user skip behavior features described above.

[0180] ③ Attention friend skip behavior features (i.e., the sub-operation features corresponding to the third object above).

[0181] Here, the user accounts that have a mutual follow relationship with the current user account are taken as the attention friends, corresponding to the third object above. Then, for all the attention friends of the current user account, the attention friend skip behavior features of each video segment in the video to be processed are calculated in the same way as calculating the global user skip behavior features described above.

[0182] So far, in the user behavior collection and analysis module, for each video segment in the video to be processed, three features can be obtained, namely the global user skip behavior features, the like-minded friend skip behavior features, and the attention friend skip behavior features.

[0183] 3) Advertising prediction module.

[0184] Here, for each video segment in the video to be processed, the content features of the video segment, the global user skip behavior features, the like-minded friend skip behavior features, and the attention friend skip behavior features are concatenated into a sequence feature, and the sequence features of multiple video segments are used as the input of the advertising prediction model (corresponding to the recommendation information prediction model above). The advertising prediction model will perform sequence annotation on the input sequence features and finally output the probability of each video segment including an advertisement. Based on these probabilities, the video segments including advertisements in the video to be processed can be predicted. For example, the video segments with a probability of including an advertisement greater than the probability threshold are taken as the video segments including advertisements.

[0185] In the training stage of the advertising prediction model, it can be trained on a supervised data set, which includes multiple sample videos and the video segments including advertisements in each sample video that are marked. After processing the sequence features of the sample videos through the advertising prediction model, the difference (i.e., the loss value) between the predicted video segments including advertisements and the actually marked video segments including advertisements is determined, and backpropagation is performed in the advertising prediction model according to this difference. During the backpropagation process, the weight parameters of the advertising prediction model are updated. Among them, the loss function used to calculate the loss value is not limited. For example, it can be a cross-entropy loss function.

[0186] As an example, the embodiments of the present application provide as Figure 8Schematic diagram of the architecture of the advertisement prediction model shown. The advertisement prediction model adopts the architecture of hierarchical RNN. In Figure 8 , the shown eigenvalue represents the above four features of the video segment. Among them, the content feature is in vector form, and the global user skip behavior feature, the skip behavior feature of like-minded friends, and the skip behavior feature of followed friends are all in numerical form. The numerical value can be regarded as a vector with a dimension of 1. After linearly transforming the four features of the video segment through the fully connected layer, the first layer of RNN is used to fuse the four linearly transformed features to generate a vector. For the sake of easy understanding, the generated vector is named the fusion feature.

[0187] The fusion features of all video segments are used as a new sequence feature and processed through the second layer of RNN to obtain the updated fusion feature. For the updated fusion feature of each video segment, through the processing of the MultiLayer Perceptron (MLP) and the fully connected layer (i.e., performing two linear transformations), a score is generated, and this score is the probability that the video segment includes an advertisement (of course, it can also be the probability of not including an advertisement, depending on the actual annotation situation).

[0188] Through the embodiments of the present application, at least the following technical effects can be achieved: 1) Obtain the content feature and the information of user feedback, and consider the original content of the video (i.e., the normal content other than the advertisement) during the processing of the advertisement prediction model, which can effectively improve the recognition accuracy of the video segment including the advertisement; 2) Can automatically delete the advertisement in the video, improve the user experience when the user watches the video, and avoid bringing advertisement interference to the user; 3) Prompt the video uploader to publish the advertisement through the channels specified by the video platform, which can increase the legitimate income of the video platform.

[0189] Next, continue to describe the exemplary structure in which the video processing device 455 based on artificial intelligence provided by the embodiments of the present application is implemented as a software module. In some embodiments, as Figure 2 shown, the software module stored in the video processing device 455 based on artificial intelligence in the memory 450 may include: a division module 4551 for dividing the video into multiple video segments; a feature extraction module 4552 for performing feature extraction processing on the video segment to obtain the content feature of the video segment; a feature mapping module 4553 for obtaining the operation data for the video segment and performing feature mapping processing on the operation data to obtain the operation feature; a fusion module 4554 for fusing the content feature and the operation feature of the video segment to obtain the fusion feature; a recursive update module 4555 for recursively updating the fusion features of multiple video segments and predicting the video segments including the recommendation information among the multiple video segments according to the obtained updated fusion features.

[0190] In some embodiments, the feature mapping module 4553 is further configured to: use an object other than the target recommended object of the video as a first object; use an object similar to the target recommended object among the multiple first objects as a second object; use an object having a follow-up relationship with the target recommended object among the multiple first objects as a third object; and use the operations performed by the first object, the second object, and the third object on the video segment as operation data for the video segment.

[0191] In some embodiments, the feature mapping module 4553 is further configured to: map the operations performed by a sample object on a video segment to operation values, and construct a sub-operation feature based on the operation values of multiple sample objects; where the sample object is any one of the first object, the second object, and the third object; and splice the sub-operation feature corresponding to the first object, the sub-operation feature corresponding to the second object, and the sub-operation feature corresponding to the third object into an operation feature of the video segment.

[0192] In some embodiments, the feature mapping module 4553 is further configured to perform any one of the following processes: perform an accumulation process on the operation values of multiple sample objects, and use the value obtained from the accumulation process as a sub-operation feature; or construct a vector having the same dimension as the number of sample objects based on the operation values of multiple sample objects, and use it as a sub-operation feature.

[0193] In some embodiments, the feature mapping module 4553 is further configured to: divide the number of videos on which the sample object has performed a skip operation by the video base number to obtain a skip video ratio; determine a weight negatively correlated with the skip video ratio; and perform a weighted process on the operation values of multiple sample objects according to the weight of the sample object to obtain a sub-operation feature; where the video base number is either the number of videos in the database or the number of videos on which the sample object has performed a play operation.

[0194] In some embodiments, the feature mapping module 4553 is further configured to: sort multiple sample objects according to the first numerical order of the skip video ratio; uniformly sample a set weight range according to the number of sample objects to obtain multiple weights to be assigned; and sequentially assign the multiple weights to be assigned to the sorted multiple sample objects according to the second numerical order; where the first numerical order is opposite to the second numerical order.

[0195] In some embodiments, the feature mapping module 4553 is further configured to: sort the operation values of multiple sample objects according to the object parameters of the sample object; and construct a vector having the same dimension as the number of sample objects based on the sorted multiple operation values; where the object parameters include any one of the registration time, the time of the most recent play operation, and the number of videos on which a play operation has been performed.

[0196] In some embodiments, the feature mapping module 4553 is further configured to: construct an initial vector with the same dimension as the number of videos in the database; wherein each value in the initial vector corresponds to a video in the database; for any object, determine the value in the initial vector corresponding to the video on which any object has performed a playback operation, and update the determined value in the initial vector to a set value to obtain a video playback vector; among multiple first objects, determine an object whose similarity to the target recommended object on the video playback vector is greater than the similarity threshold as the second object.

[0197] In some embodiments, the recursive update module 4555 is further configured to: sort the fusion features of multiple video segments according to the time axis order of the videos; for any video segment in the video, perform the following processing: fuse the fusion feature of any video segment and the intermediate feature of the previous video segment to obtain the intermediate feature of any video segment; perform a weighting process on the intermediate feature to obtain the updated fusion feature of any video segment.

[0198] In some embodiments, the fusion module 4554 is further configured to: perform a recursive update process on the content feature and the operation feature of the video segment to obtain the updated content feature and the updated operation feature; splice the updated content feature and the updated operation feature to obtain a fusion feature.

[0199] In some embodiments, the feature extraction module 4552 is further configured to perform any one of the following processes: perform feature extraction processing on each image in the video segment to obtain image features, and fuse the image features of multiple images into the content feature of the video segment; perform feature extraction processing on multiple consecutive images in the video segment, and use the extracted image features as the content feature of the video segment.

[0200] In some embodiments, the recursive update module 4555 is further configured to: perform a weighting process on the updated fusion feature of the video segment to obtain the probability that the video segment includes recommendation information; use the video segment in the video whose probability of including recommendation information is greater than the probability threshold as the predicted video segment including recommendation information.

[0201] In some embodiments, the video processing device 455 based on artificial intelligence further includes: a deletion module, configured to delete the predicted video segment including recommendation information.

[0202] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the artificial intelligence-based video processing method described above in the present invention.

[0203] The embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the embodiment of the present application, for example, Figure 3A 、 Figure 3B 、 Figure 3C and Figure 3D The video processing method based on artificial intelligence is shown.

[0204] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0205] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0206] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0207] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0208] In summary, the following technical effects can be achieved through the embodiments of the present application:

[0209] 1) Determine the fusion feature based on the content feature and operation feature of the video segment, so that the fusion feature includes information in multiple dimensions in the video segment. On this basis, perform recursive update processing on the fusion features of multiple video segments. By referring to the normal content in the reference video, the updated fusion feature can more accurately and effectively represent the video segment. In this way, when predicting the video segment including the recommendation information based on the updated fusion feature, the prediction accuracy can be improved.

[0210] 2) By deleting the video segment including the recommendation information predicted in the video, the storage space occupied by the video can be reduced, and the storage resources of the electronic device can be saved. When the electronic device subsequently performs operations such as sending or playing the video, the actual utilization rate of the computing resources consumed by the electronic device can be improved.

[0211] 3) In the process of determining the content feature, feature extraction processing can be performed on a single image, or on multiple consecutive images together, which improves the flexibility of feature extraction.

[0212] 4) By determining objects at different levels (global objects, similar objects, and concerned objects), and then fusing the sub-operation features at different levels into the operation feature of the video segment, the comprehensiveness and hierarchy of the information contained in the operation feature can be improved, which is convenient for subsequently learning effective data rules from the operation feature.

[0213] 5) The embodiments of the present application can construct sub-operation features in numerical form or vector form, which are applicable to different scenarios. For example, the sub-operation feature in numerical form can be applicable to the scenario of real-time video processing, and the sub-operation feature in vector form can be applicable to the scenario of offline video processing.

[0214] 6) The embodiments of the present application can be implemented by being integrated into the recommendation information prediction model. The recommendation information prediction model can be deployed in the electronic device in a componentized form, which is simple to deploy and has a wide range of applications. For example, it can be deployed in the video player of the terminal device as a function of filtering recommendation information, and it also supports the user to decide whether to enable this function.

[0215] The above are only the embodiments of the present application, and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A video processing method based on artificial intelligence, characterized in that, The method includes: Dividing the video into multiple video segments; Performing feature extraction processing on the video segments to obtain the content features of the video segments; Obtaining operation data for the video segments and performing feature mapping processing on the operation data to obtain operation features, where the operation data is an operation performed on the video segments, and the operation includes at least one of the following: play operation, pause operation, skip operation; Performing recursive update processing on the content features and operation features of the video segments to obtain updated content features and updated operation features; splicing the updated content features and the updated operation features to obtain fused features; Performing recursive update processing on the fused features of multiple video segments, and predicting the video segments including recommendation information among the multiple video segments according to the obtained updated fused features.

2. The method according to claim 1, characterized in that, The obtaining of the operation data for the video segments includes: Regarding the objects other than the target recommended object of the video as the first objects; Regarding the objects similar to the target recommended object among the multiple first objects as the second objects; Regarding the objects having a follow-up relationship with the target recommended object among the multiple first objects as the third objects; Regarding the operations performed by the first objects, the second objects, and the third objects on the video segments as the operation data for the video segments.

3. The method according to claim 2, wherein The performing of feature mapping processing on the operation data to obtain operation features includes: Mapping the operations performed by the sample objects on the video segments to operation values, and constructing sub-operation features according to the operation values of the multiple sample objects; where the sample objects are any one of the first objects, the second objects, and the third objects; Splicing the sub-operation features corresponding to the first objects, the sub-operation features corresponding to the second objects, and the sub-operation features corresponding to the third objects into the operation features of the video segments.

4. The method according to claim 3, characterized in that, The constructing of sub-operation features according to the operation values of the multiple sample objects includes: Performing any one of the following processes: Performing cumulative processing on the operation values of the multiple sample objects, and using the value obtained by the cumulative processing as the sub-operation feature; Constructing a vector with the same dimension as the number of the sample objects according to the operation values of the multiple sample objects to be used as the sub-operation feature.

5. The method according to claim 4, wherein The performing of cumulative processing on the operation values of the multiple sample objects and using the value obtained by the cumulative processing as the sub-operation feature includes: Dividing the number of videos for which the sample objects have performed skip operations by the video base number to obtain the skip video ratio; Determining a weight negatively correlated with the skip video ratio; Performing weighted processing on the operation values of the multiple sample objects according to the weights of the sample objects to obtain the sub-operation feature; where the video base number includes any one of the number of videos in the database and the number of videos for which the sample objects have performed play operations.

6. The method according to claim 5, wherein The determining of the weight negatively correlated with the skip video ratio includes: Sorting the multiple sample objects according to the first numerical order of the skip video ratio; Uniformly sample the set weight range according to the number of the sample objects to obtain multiple weights to be assigned; According to the second numerical order, sequentially assign the multiple weights to be assigned to the multiple sorted sample objects; Wherein, the first numerical order is opposite to the second numerical order.

7. The method according to claim 4, characterized in that The constructing a vector with the same dimension as the number of the sample objects according to the operation values of the multiple sample objects includes: Sort the operation values of the multiple sample objects according to the object parameters of the sample objects; Construct a vector with the same dimension as the number of the sample objects according to the multiple sorted operation values; Wherein, the object parameters include any one of the registration time, the time of the most recent execution of the play operation, and the number of videos for which the play operation has been executed.

8. The method according to claim 2, wherein The taking, as the second object, the object similar to the target recommended object among the multiple first objects includes: Construct an initial vector with the same dimension as the number of videos in the database; wherein, each value in the initial vector corresponds to a video in the database; For any one object, determine the value in the initial vector corresponding to the video for which the play operation has been executed by the any one object, and update the determined value in the initial vector to a set value to obtain a video play vector; Among the multiple first objects, determine the object whose similarity with the target recommended object on the video play vector is greater than the similarity threshold as the second object.

9. The method according to any one of claims 1 to 8, characterized in that The recursively updating and processing the fusion features of the multiple video segments includes: Sort the fusion features of the multiple video segments according to the time axis order of the video; For any one video segment in the video, perform the following processing: Fuse the fusion feature of the any one video segment and the intermediate feature of the previous video segment to obtain the intermediate feature of the any one video segment; Perform a weighting process on the intermediate feature to obtain the updated fusion feature of the any one video segment.

10. The method according to any one of claims 1 to 8, characterized in that The extracting the content feature of the video segment by performing a feature extraction process on the video segment includes: Perform any one of the following processes: Perform a feature extraction process on each image in the video segment to obtain an image feature, and fuse the image features of the multiple images into the content feature of the video segment; Perform a feature extraction process on multiple consecutive images in the video segment, and use the extracted image feature as the content feature of the video segment.

11. The method according to any one of claims 1 to 8, characterized in that, The predicting the video segment including the recommendation information among the multiple video segments according to the obtained updated fusion feature includes: Perform a weighting process on the updated fusion feature of the video segment to obtain the probability that the video segment includes the recommendation information; Take the video segment whose probability of including the recommendation information in the video is greater than the probability threshold as the predicted video segment including the recommendation information; The method further includes: Delete the predicted video segment including the recommendation information.

12. A video processing device based on artificial intelligence, characterized in that, The apparatus includes: A dividing module, configured to divide a video into multiple video segments; A feature extraction module, configured to perform feature extraction processing on the video clip to obtain the content features of the video clip; A feature mapping module, configured to obtain operation data for the video clip and perform feature mapping processing on the operation data to obtain operation features, where the operation data is an operation performed on the video clip, and the operation includes at least one of the following: a play operation, a pause operation, a skip operation; A fusion module, configured to perform recursive update processing on the content features and operation features of the video clip to obtain updated content features and updated operation features; perform splicing processing on the updated content features and the updated operation features to obtain fusion features; A recursive update module, configured to perform recursive update processing on the fusion features of multiple video clips, and predict the video clips including recommendation information among the multiple video clips according to the obtained updated fusion features.

13. An electronic device, characterized in that, Comprising: A memory, configured to store executable instructions; A processor, configured to implement the artificial intelligence-based video processing method according to any one of claims 1 to 11 when executing the executable instructions stored in the memory.

14. A computer-readable storage medium, characterized in that, Stored with executable instructions, configured to implement the artificial intelligence-based video processing method according to any one of claims 1 to 11 when being executed by a processor.

15. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer executable instructions or the computer program are executed by a processor, the artificial intelligence-based video processing method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Video playing method and device and storage medium

    CN111209440A

  • Multimedia file detection method and device, multimedia file playing method and device, equipment and storage medium

    CN111416996A