Method, device and medium for generating emotional frame indication information of video

By generating emotional frame indication information of the video, the problem of low efficiency in selecting emotional state change segments in the existing technology is solved, and rapid positioning and improved video editing efficiency are achieved.

CN119697427BActive Publication Date: 2025-09-05SUZHOU XIAOMIANAO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411955970.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-28
Publication Date
2025-09-05
Estimated Expiration
2044-12-28

AI Technical Summary

Technical Problem

The existing technology is inefficient in selecting emotion state change segments from the original video. A method is needed to generate emotion frame indication information of the video to quickly locate the desired segment.

Method used

By obtaining the emotional information of each video content information in the target video information, performing weighted fusion, calculating the comprehensive emotional value of the video frame, and generating emotional frame indication information, it can help users quickly understand the emotional change trend of the video information.

Benefits of technology

It can quickly locate the segments with changing emotional states in the video, thus improving the efficiency of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697427B_ABST
    Figure CN119697427B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video processing technology. In particular, it relates to a method, device and medium for generating emotion frame indication information of a video. The method includes: obtaining the emotion information corresponding to each video content information in the target video information; for each video content information, weighted fusion is performed based on one or more emotion tags corresponding to the video content information, and the probability information and first weight information corresponding to each emotion tag to obtain the emotion value of each video content information; for each video frame content information, weighted fusion is performed based on the emotion value corresponding to each video content information and the second weight information corresponding to each video content information to obtain the comprehensive emotion value of each video frame content information; based on the comprehensive emotion value corresponding to each video frame content information, emotion frame indication information of the target video information is generated. By generating emotion frame indication information, users can quickly understand the emotion change trend of video information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a method, device, and medium for generating emotion frame indication information of a video. Background Art

[0002] With the popularity of short videos and the booming development of streaming platforms, more and more people are beginning to participate in video editing. Video editing includes but is not limited to: selecting one or more desired clips from the original video for video editing and creation, or selecting multiple desired clips from one or more original videos for video editing and creation. The original video often shows changes in emotional state (for example, in the live broadcast of a short video account, there may be plots with large emotional fluctuations). The existing technology usually uses manual selection to find the desired emotional state change clips from the original video, which is very inefficient. Therefore, the existing technology urgently needs a method that can generate emotional frame indication information of the video to help users quickly locate the desired clips and greatly improve efficiency. Summary of the Invention

[0003] In order to solve the above problems, the present application provides a method, device and medium for generating emotion frame indication information of a video.

[0004] According to one aspect of the present application, a method for generating emotion frame indication information of a video is provided, the method comprising:

[0005] Obtaining emotion information corresponding to each video content information in the target video information, wherein the target video information includes a plurality of video frame content information, the plurality of video frame content information is arranged in time frame order, each video frame content information includes a plurality of video content information, and the emotion information includes one or more emotion labels corresponding to each video content information, and probability information corresponding to each emotion label;

[0006] For each video content information, weighted fusion is performed based on one or more emotion tags corresponding to the video content information, probability information corresponding to each emotion tag, and first weight information corresponding to each emotion tag to obtain an emotion value for each video content information;

[0007] For each video frame content information, performing weighted fusion based on the emotion value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain a comprehensive emotion value of each video frame content information;

[0008] According to the comprehensive emotion value corresponding to each video frame content information, the emotion frame indication information of the target video information is generated, wherein the emotion frame indication information includes the comprehensive emotion value corresponding to each frame sequence number in the target video information.

[0009] According to one aspect of the present application, a computer device for generating emotion frame indication information of a video is provided, the device comprising:

[0010] processor; and

[0011] A memory arranged to store computer executable instructions which, when executed, cause the processor to perform the operations of any of the methods described above.

[0012] According to one aspect of the present application, a computer-readable medium storing instructions is provided. When the instructions are executed, the system performs the operations of any of the methods described above.

[0013] Compared with the existing methods, the present application obtains the emotional information corresponding to each video content information in the target video information, and calculates the emotional value of the video content information based on the emotional information of each video content information; for each video frame content information, the comprehensive emotional value of the video frame content information is calculated according to the emotional value of each video content information corresponding to the video frame content information, so as to obtain the comprehensive emotional value of multiple video frame content information in the target video information; based on the comprehensive emotional value of each video frame content information in the multiple video frame content information, the emotional indication information of the target video information is obtained, so as to help users quickly understand the emotional change trend of the video information. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0015] Figure 1 A flow chart of a method for generating emotion frame indication information of a video according to one embodiment of the present application is shown;

[0016] Figure 2 A schematic diagram showing the structure of a device for generating emotion frame indication information of a video according to one embodiment of the present application is shown;

[0017] Figure 3 An exemplary system is shown that can be used to implement the various embodiments described in this application. DETAILED DESCRIPTION

[0018] The present application is described in further detail below with reference to the accompanying drawings.

[0019] In a typical configuration of the present application, the terminal, the device of the service network, and the trusted party each include one or more processors (eg, a central processing unit (CPU)), an input / output interface, a network interface, and a memory.

[0020] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of a computer-readable medium.

[0021] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or means. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory methods, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0022] The devices referred to in this application include, but are not limited to, terminals, network devices, or devices formed by integrating terminals and network devices via a network. The terminals include, but are not limited to, any mobile electronic product that can interact with a user (e.g., through a touchpad), such as a smartphone, a tablet computer, etc. The mobile electronic product can use any operating system, such as the Android operating system, the iOS operating system, etc. The network device includes an electronic device that can automatically perform numerical calculations and information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc. The network device includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud consisting of multiple servers; herein, a cloud is composed of a large number of computers or network servers based on cloud computing. Cloud computing is a type of distributed computing, a virtual supercomputer composed of a group of loosely coupled computers. The network includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless self-organizing network (Ad Hoc network), etc. Preferably, the device may also be a program running on the terminal, the network device, or a device formed by integrating the terminal and the network device, the network device, the touch terminal, or the network device and the touch terminal via a network.

[0023] Of course, those skilled in the art should understand that the above-mentioned devices are only examples, and other existing or future devices that are applicable to this application should also be included in the scope of protection of this application and are included here by reference.

[0024] In the description of the present application, “plurality” means two or more, unless otherwise clearly defined.

[0025] Figure 1A method for generating emotional frame indication information of a video according to an embodiment of the present application is shown. The method includes steps S11, S12, and S13. In step S11, the emotional information corresponding to each video content information in the target video information is obtained, wherein the target video information includes multiple video frame content information, the multiple video frame content information is arranged in time frame order, each video frame content information includes multiple video content information, and the emotional information includes one or more emotional tags corresponding to each video content information, and probability information corresponding to each emotional tag; in step S12, for each video content information, weighted fusion is performed according to the one or more emotional tags corresponding to the video content information, the probability information corresponding to each emotional tag, and the first weight information corresponding to each emotional tag to obtain the emotional value of each video content information; in step S13, for each video frame content information, weighted fusion is performed according to the emotional value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain the comprehensive emotional value of each video frame content information; in step S14, emotional frame indication information of the target video information is generated according to the comprehensive emotional value corresponding to each video frame content information, wherein the emotional frame indication information includes the comprehensive emotional value corresponding to each frame sequence number in the target video information.

[0026] Specifically, in step S11, emotional information corresponding to each piece of video content information in the target video information is obtained. The target video information includes multiple video frame content information, which is arranged in time frame order. Each video frame content information includes multiple pieces of video content information. The emotional information includes one or more emotion tags corresponding to each piece of video content information, as well as probability information corresponding to each emotion tag. In some embodiments, the target video information includes, but is not limited to, the original video. For example, recorded videos of live streams from short video accounts, movie videos, etc. In some embodiments, the video content information includes, but is not limited to, image information, audio information, etc. Those skilled in the art will appreciate that video information includes multiple frames of images and audio arranged in a continuous sequence. For example, multiple pieces of video content information (e.g., image information, audio information, etc.) in the same time frame are considered as one piece of video frame content information. In some embodiments, emotion tags include, but are not limited to, happy, sad, neutral, etc. For example, each piece of video content information corresponds to one or more emotion tags, as well as probability information for each emotion tag. For example, image information 1 corresponds to three emotional labels: happy, sad, and neutral, where the probability of happy is 0.6, the probability of sad is 0.3, and the probability of neutral is 0.1; audio 1 corresponds to three emotional labels: happy, sad, and neutral, where the probability of happy is 0.7, the probability of sad is 0.2, and the probability of neutral is 0.1. In some embodiments, different video content information has different methods for obtaining emotional labels and probability information. For example, image information can obtain emotional labels and probability information by inputting image information into a classification model, or by extracting emotional features from the image information and determining emotional labels and probability information based on the emotional features in the image information; audio information can obtain emotional labels and probability information based on the audio features by extracting audio features. For specific instructions on determining emotional labels and probability information, please refer to the corresponding embodiments below and will not be repeated here. For example, target video information 1 includes video frame content information 1 and video frame content information 2. Video frame content information 1 includes image information 1 and audio information 1, and video frame content information 2 includes image information 2 and audio information 2. Image information 1 has three emotion labels: happy, sad, and neutral. The probability of happy is 0.6, the probability of sad is 0.3, and the probability of neutral is 0.1. Audio information 1 has three emotion labels: happy, sad, and neutral. The probability of happy is 0.7, the probability of sad is 0.2, and the probability of neutral is 0.1. Image information 2 has three emotion labels: happy, sad, and neutral. The probability of happy is 0.5, the probability of sad is 0.3, and the probability of neutral is 0.2. Audio information 2 has three emotion labels: happy, sad, and neutral. The probability of happy is 0.6, the probability of sad is 0.2, and the probability of neutral is 0.2.

[0027] In step S12, for each video content information, weighted fusion is performed based on one or more emotion tags corresponding to the video content information, the probability information corresponding to each emotion tag, and the first weight information corresponding to each emotion tag to obtain the emotion value of each video content information. In some embodiments, different emotion tags correspond to first weight information. For each video content information, the system performs weighted fusion based on one or more emotion tags corresponding to the video content information, the probability information corresponding to each emotion tag, and the first weight information corresponding to each emotion tag to obtain the emotion value of each video content information. For example, the first weight information of happiness is 0.8, the first weight information of sadness is 1.0, and the first weight information of neutrality is 0. In the above embodiment, the emotion value of image information 1 is k1=0.6*0.8+0.3*1+0.1*0=0.78; the emotion value of audio information 1 is h1=0.7*0.8+0.2*1+0.1*0=0.76; the emotion value of image information 2 is k2=0.5*0.8+0.3*1+0.2*0=0.7; and the emotion value of audio information 2 is h2=0.6*0.8+0.2*1+0.2*0=0.68. In some embodiments, different emotion tags correspond to different first weight information. In some embodiments, in order to filter out video frames with large emotion changes, step S12 includes: for each video content information, based on one or more emotion tags corresponding to the video content information, and the probability information corresponding to each emotion tag, and the first weight information corresponding to each emotion tag, weighted fusion is performed to obtain the emotion value of each video content information, wherein the one or more emotion tags include at least one target emotion tag, and the first weight information corresponding to the target emotion tag is higher than that of other emotion tags in the one or more emotion tags. For example, the target emotion tags include but are not limited to emotions with large fluctuations, such as happiness, excitement, sadness, quarrels, etc. In some embodiments, the first weight information of each target emotion tag in at least one target emotion tag can be randomly assigned, but the first weight information of the target emotion tag is higher than the first weight information of other emotion tags. In other embodiments, the first weight information of each target emotion tag in at least one target emotion tag can also be pre-set. For example, at least one target emotion tag includes excitement and happiness. Generally speaking, the emotional fluctuation of excitement is higher than the emotional fluctuation of happiness. Therefore, the first weight information of excitement can be higher than the first weight information of happiness.

[0028] In step S13, for each video frame content information, a weighted fusion is performed based on the emotion values ​​corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information, to obtain a comprehensive emotion value for each video frame content information. In some embodiments, different video content information corresponds to second weight information, and the second weight information corresponding to different video content information may vary. For example, in some embodiments, the video content information may also include text information (e.g., letters in video image information, text information in images). Compared to audio information, text information generally expresses emotions more effectively through audio information. Therefore, the second weight information corresponding to audio information may be higher than the second weight information corresponding to text information. For example, after obtaining the emotion value corresponding to each video content information, a comprehensive emotion value for each video frame content information is then obtained based on the emotion value of each video content information. For example, the second weight information corresponding to image information is 0.55, and the second weight information corresponding to audio information is 0.45. Continuing with the above embodiment as an example, the emotion value k1 of image information 1 is 0.78, and the emotion value h1 of audio information 1 is 0.76. Then, the comprehensive emotion value S of video frame content information 1 is 0.78*0.55+0.76*0.45=0.771. In some embodiments, when the image information includes a face, the face often expresses emotion more accurately. Therefore, whether the image information includes a face has an important impact on the allocation of the second weight. In some embodiments, step S13 includes: for each video frame content information, if the video content information included in the video frame content information includes facial image information, determining to use the second target weight combination information; otherwise, determining to use the second candidate weight combination information; performing weighted fusion based on the emotion value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain a comprehensive emotion value for each video frame content information, wherein if the video content information includes facial image information, the second weight information belongs to the second target weight combination information; otherwise, the second weight information belongs to the second candidate weight combination information. For example, the system may preset two second weight information allocation combinations. For each video frame content information, the system uses facial recognition technology to identify whether the image information includes a face. If a face is included, the second target weight combination information is determined to be used. If the image information does not include a face, the second candidate weight combination information is used. In some embodiments, the second target weight combination information includes second weight information corresponding to different video content information, and the second candidate weight combination information includes second weight information corresponding to different video content information.In some embodiments, in the second target weight combination information, the second weight corresponding to the image information is higher than the weight of other video content information; in the second candidate weight combination information, the second weight corresponding to the image information is lower than the weight of other video content information. For example, the video content information includes image information, audio information, and text information. The second target weight combination information includes: the second weight information corresponding to the image information is 0.65, the second weight information corresponding to the audio information is 0.25, and the second weight information corresponding to the text information is 0.10; the second candidate weight combination information includes: the second weight information corresponding to the image information is 0.15, the second weight information corresponding to the audio information is 0.65, and the second weight information corresponding to the text information is 0.15. For each video frame content information, the system uses face recognition technology to detect whether the image information includes a face. If the image information includes a face, it determines to use the second target weight combination information, that is, the second weight information corresponding to the image information is 0.65, the second weight information corresponding to the audio information is 0.25, and the second weight information corresponding to the text information is 0.10. If the image information does not include a face, the second candidate weight combination information is determined to be used, that is, the second weight information corresponding to the image information is 0.15, the second weight information corresponding to the audio information is 0.65, and the second weight information corresponding to the text information is 0.15. Of course, those skilled in the art will understand that the second target weight combination information and the second candidate weight combination information described above are only examples, and other existing or future second target weight combination information and second candidate weight combination information, if applicable to this embodiment, are also within the scope of protection of this application and are included herein by reference.

[0029] In step S14, emotion frame indication information of the target video information is generated according to the comprehensive emotion value corresponding to the content information of each video frame, wherein the emotion frame indication information includes the comprehensive emotion value corresponding to each frame number in the target video information. For example, after determining the comprehensive emotion value of each frame, emotion indication information about the target video information is generated according to the comprehensive emotion value corresponding to each frame number. In some embodiments, the emotion indication information can be a curve change on a timeline (e.g., a time frame) (e.g., a curve generated based on the comprehensive emotion value of each frame), which can help users distinguish between high and low emotions. In other embodiments, the emotion indication information can also be the comprehensive emotion value of each marked frame. Of course, those skilled in the art will understand that the expression form of the emotion indication information described above is only an example, and other existing or future possible expression forms, if applicable to this embodiment, are also within the scope of protection of this application and are included herein by reference.

[0030] In some embodiments, step S11 includes step S111 (not shown), step S112 (not shown), step S113 (not shown), and step S114 (not shown). In step S111, target video information is acquired. In step S112, the target video information is segmented into multiple video frame content information using a video decomposition method, where the multiple video frame content information is arranged in time frame order. In step S113, for each video frame content information, multiple video content information is extracted from the video frame content information using a content extraction method, where the multiple video content information includes image information, text information, and audio information. In step S114, for each video content information, one or more emotion tags corresponding to the video content information and probability information corresponding to each emotion tag are acquired using a classification method corresponding to the video content information. In some embodiments, the content extraction method includes, but is not limited to, OCR (Optical Character Recognition) technology and audio extraction technology. For example, using video decomposition technology, the input video is segmented into several static image frames. Then, using OCR technology, text information such as subtitles and text annotations is identified from the static image frames. Meanwhile, audio clips are extracted from the target video information to obtain multiple video frame content information. Each video frame content information includes multiple video content information (e.g., static images, subtitles, text annotations, and audio clips). In some embodiments, classification methods include, but are not limited to, classification models and feature extraction methods. For detailed descriptions of these methods, please refer to the corresponding embodiments below and are not elaborated here.

[0031] In some embodiments, step S114 includes: for image information, extracting features from the image information to obtain emotional feature information in the image information; determining one or more emotional tags corresponding to the image information and probability information corresponding to each emotional tag based on the emotional feature information; for text information, inputting the text information into a text emotion classification model, and outputting one or more emotional tags corresponding to the text information and probability information corresponding to each emotional tag through the text emotion classification model; for audio information, using multiple consecutive time frames of audio information as a processing unit, extracting audio features from the audio information of the multiple time frames, inputting the audio features into an audio emotion classification model, and outputting one or more emotional tags corresponding to the audio information of the multiple time frames and probability information corresponding to each emotional tag through the audio emotion classification model, wherein the emotional information corresponding to each time frame of the audio information of the multiple time frames is the same. In some embodiments, different video content information requires different classification methods to obtain one or more emotional tags corresponding to the video content information and probability information corresponding to each emotional tag. In some embodiments, emotional feature information includes but is not limited to facial feature information of a person's face, scene information, object information, hue information, texture features, composition features, etc. For example, based on the emotional feature information in the image information, one or more emotional tags corresponding to the image information and the probability information corresponding to each emotional tag are determined. For a specific description of the image information, please refer to the corresponding embodiment below, which will not be repeated here. In some embodiments, the text emotion classification model includes but is not limited to a Transformer model trained using BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pre-trained Transformer). Here, those skilled in the art will understand that a text emotion classification model can be obtained based on a large amount of training text information and one or more emotional tags corresponding to each training text information, as well as the probability information of each emotional tag, so that the text information can be input into the text emotion classification model in the future, and the one or more emotional tags corresponding to the text information and the probability information of each emotional tag are output. In some embodiments, for audio information, audio features can be obtained through short-time Fourier transform (STFT) or Mel-frequency cepstral coefficients (MFCC), and then a pre-trained audio emotion classification model (such as a support vector machine SVM or a deep neural network DNN) is used to perform emotion recognition on the audio features to obtain one or more emotion labels (such as anger, happiness, negativity, etc.), as well as probability information of each emotion label.Here, similar to the above-mentioned text emotion classification model, an audio emotion classification model can be obtained based on a large number of audio feature training with emotion labels and corresponding probability information. In some embodiments, audio information of multiple consecutive time frames is used as a processing unit. For example, audio features are obtained from a segment of audio information through short-time Fourier transform (STFT) or Mel-frequency cepstral coefficient (MFCC), and one or more emotion labels of the audio information (audio information of multiple consecutive time frames) and probability information of each emotion label are obtained. The emotion information corresponding to the audio information of each frame in the multiple consecutive time frames of audio information is the same. For example, video frame content information 1 includes audio information 1, video frame content information 2 includes audio information 2, and video frame content information 3 includes audio information 3. The audio information of the three consecutive frames is taken as a processing unit, the audio features of the audio information are extracted, and the emotional information of the three consecutive frames is obtained through the audio emotion classification model (for example, happy, sad, and neutral, where the probability of happiness is 0.6, the probability of sadness is 0.2, and the probability of neutral is 0.5). Then the emotional information of audio information 1, audio information 2, and audio information 3 are happy, sad, and neutral, respectively, where the probability of happiness is 0.6, the probability of sadness is 0.2, and the probability of neutral is 0.5.

[0032] In some embodiments, feature extraction is performed on image information to obtain emotional features in the image information; based on the emotional features, one or more emotional tags corresponding to the image information and probability information corresponding to each emotional tag are determined, including: if the image information includes a face, facial feature information of the face is extracted using face detection technology, the facial feature information is input into an emotion recognition model as emotional feature information, and the emotion recognition model outputs one or more emotional tags for the image information and probability information for each emotional tag; otherwise, multiple groups of emotional feature information are extracted from the image information using multiple feature extraction methods, wherein each group of emotional feature information includes one or more categories of emotional feature information, each category of emotional feature information representing an emotional tag; for each group of emotional feature information, the probability information of each emotional tag is determined based on the proportion of each category of emotional feature information in the group of emotional feature information. For example, image information may include a face or not. In some embodiments, the system can detect and mark whether the image information includes a face using face detection technology (e.g., Google FaceMesh technology). If the image information includes a face, facial feature information is extracted using face detection technology. This facial feature information is then input into an emotion recognition model, which outputs one or more emotion labels for the image information, along with probability information for each emotion label. Those skilled in the art will appreciate that the emotion recognition model can be trained using a large amount of facial feature information with emotion labels and probability information. The model architecture can be a deep neural network (DNN), recurrent neural network (RNN), or variations thereof. If the image information does not contain a face, multiple sets of emotion feature information are extracted from the image information using various feature extraction methods. In some embodiments, emotion feature information includes, but is not limited to, scene information, object information, hue information, texture features, and composition features. In some embodiments, the system extracts multiple sets of emotional feature information from image information using multiple feature extraction methods. Each set of emotional feature information includes one or more emotional feature information. For each set of emotional feature information, the system classifies the one or more emotional feature information included in the set according to a preset mapping relationship, resulting in one or more categories of emotional feature information. (For example, if the set of emotional feature information includes two primary colors, orange and blue, with orange accounting for 60% and blue accounting for 40%, and the emotional label represented by orange is comfort and the emotional label represented by blue is depression, then the probability of comfort is 0.6 and the probability of depression is 0.4.) For example, emotional feature information includes scenes and objects. Different scenes are often associated with specific emotions, such as natural scenery, which is often associated with tranquility and relaxation. A convolutional neural network (CNN) in deep learning, such as the ResNet model, is used to perform scene classification and object detection on the image information, outputting one or more scene information and objects included in the image information. Scene classification can determine the environment of the image, such as natural scenery, city, or indoors.Specific objects, such as blood, ice and snow, bonfires, cats and dogs, can also be used to determine the emotional orientation of an image. For example, a ResNet model outputs scene information and objects, including natural scenery, and then classifies the detected scenes and objects based on a preset mapping relationship to obtain one or more categories of emotional feature information. Each category of emotional feature information represents an emotional label, and the probability of each emotional label is determined based on the proportion of each category of emotional feature information. Another example of emotional feature information includes color and lighting. Different colors and lighting convey different emotions. First, image color histogram analysis techniques are used to count the dominant tones in the image and then label their emotional categories. For example, warm tones (such as orange and red) convey comfort or enthusiasm, while cool tones (such as blue or gray) convey tranquility or melancholy. Another example of emotional feature information includes composition and content. Edge detection and shape analysis techniques are used to identify structural and layout features within the image. Compositional elements such as lines, shapes, and framing are analyzed to identify their emotional orientation. For example, diagonal lines or irregular shapes often convey movement or instability, while horizontal lines often convey calmness. For another example, emotional feature information includes texture and details, and texture analysis technology (Gabor filter) is applied to evaluate the detailed features of the image. The richness and type of texture details can also affect emotions. For example, rough textures generally convey complexity or tension, while smooth textures generally convey peace. Of course, those skilled in the art will understand that the feature extraction methods and emotional feature information described above are merely examples, and other feature extraction methods are currently available or may appear in the future. If emotional feature information can be applied to this embodiment, it is also within the scope of protection of this application and is incorporated herein by reference.

[0033] In some embodiments, the method further includes step S15 (not shown). In step S15, an emotional peak frame is determined based on the emotional frame indication information of the target video information to mark the emotional peak frame, wherein the emotional peak frame includes the frame number in the target video information whose comprehensive emotional value meets the target emotional value condition. In some embodiments, in order to quickly help users locate the emotional peak position for video editing, the system determines the emotional peak frame based on the emotional frame indication information of the target video information. In some embodiments, the comprehensive emotional value meeting the target emotional value condition includes but is not limited to the comprehensive emotional value being equal to or higher than the target threshold, or the comprehensive emotional value being ranked before the target among all comprehensive emotional values ​​in the target video information (for example, in the top 10%). In some embodiments, after the system determines the emotional peak frame, it can mark the emotional peak frame to quickly locate the emotional peak frame. For example, the emotional peak frame can be marked by setting a marking symbol on the time frame or the playback progress bar.

[0034] Figure 2The schematic diagram of the structure according to an embodiment of the present application is shown, and the device includes a first module, a second module, a third module, and a fourth module. The first module is used to obtain the emotion information corresponding to each video content information in the target video information, wherein the target video information includes a plurality of video frame content information, the plurality of video frame content information is arranged in time frame order, each video frame content information includes a plurality of video content information, and the emotion information includes one or more emotion tags corresponding to each video content information, and probability information corresponding to each emotion tag; the first and second modules are used for obtaining the emotion information corresponding to each video content information according to the one or more emotion tags corresponding to the video content information, and the probability information corresponding to each emotion tag; The first module is used to perform weighted fusion based on the probability information corresponding to the emotion tag and the first weight information corresponding to each emotion tag to obtain the emotion value of each video content information; the first module is used to perform weighted fusion based on the emotion value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain the comprehensive emotion value of each video frame content information; the first module is used to generate emotion frame indication information of the target video information based on the comprehensive emotion value corresponding to each video frame content information, wherein the emotion frame indication information includes the comprehensive emotion value corresponding to each frame number in the target video information.

[0035] Here, the specific implementations corresponding to module 11, module 12, and module 13 are the same as or similar to the specific embodiments of step S11, step S12, and step S13, and are therefore not repeated here and are included here by reference.

[0036] In addition to the methods and devices described in the above embodiments, the present application also provides a computer-readable storage medium, which stores computer code. When the computer code is executed, the method described in any of the above items is executed.

[0037] The present application also provides a computer program product. When the computer program product is executed by a computer device, the method described in any one of the preceding items is executed.

[0038] The present application also provides a computer device, comprising:

[0039] one or more processors;

[0040] a memory for storing one or more computer programs;

[0041] When the one or more computer programs are executed by the one or more processors, the one or more processors are caused to implement the method as described in any one of the preceding items.

[0042] Figure 3shows an exemplary system that can be used to implement the various embodiments described in this application;

[0043] like Figure 3 In some embodiments, the system 300 can function as any of the devices described in the various embodiments. In some embodiments, the system 300 can include one or more computer-readable media (e.g., system memory or NVM / storage device 320) having instructions and one or more processors (e.g., processor(s) 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the modules and thereby perform the actions described herein.

[0044] For one embodiment, system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of processor(s) 305 and / or any suitable device or component in communication with system control module 310 .

[0045] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0046] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. For one embodiment, system memory 315 can include any suitable volatile memory, such as a suitable DRAM. In some embodiments, system memory 315 can include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0047] For one embodiment, system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to NVM / storage device 320 and communication interface(s) 325 .

[0048] For example, NVM / storage 320 may be used to store data and / or instructions. NVM / storage 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0049] NVM / storage device 320 may include storage resources that are physically part of the device on which system 300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 320 may be accessed over a network via communication interface(s) 325.

[0050] Communication interface(s) 325 may provide an interface for system 300 to communicate over one or more networks and / or with any other suitable devices. System 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0051] For one embodiment, at least one of the processor(s) 305 may be packaged together with the logic of one or more controllers of the system control module 310 (e.g., the memory controller module 330). For one embodiment, at least one of the processor(s) 305 may be packaged together with the logic of one or more controllers of the system control module 310 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 305 may be integrated on the same die with the logic of one or more controllers of the system control module 310. For one embodiment, at least one of the processor(s) 305 may be integrated on the same die with the logic of one or more controllers of the system control module 310 to form a system-on-chip (SoC).

[0052] In various embodiments, system 300 may be, but is not limited to, a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or a different architecture. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0053] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, and the like. In addition, some steps or functions of the present application can be implemented in hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.

[0054] In addition, a part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or method scheme according to the present application through the operation of the computer. Method personnel in this field should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0055] Communication media include media by which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media may include guided transmission media such as cables and wires (e.g., fiber optic, coaxial, etc.) and wireless (unguided transmission) media capable of propagating energy waves, such as acoustic, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data may be embodied as, for example, a modulated data signal in a wireless medium such as a carrier wave or similar mechanism such as that embodied as part of a spread spectrum method. The term "modulated data signal" refers to a signal that has one or more characteristics changed or set in such a manner as to encode information in the signal. Modulation may be analog, digital, or a hybrid modulation method.

[0056] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or approach for storage of information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memory such as random access memory (RAM, DRAM, SRAM); and non-volatile memory such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or later developed that can store computer-readable information / data for use by a computer system.

[0057] Here, according to one embodiment of the present application, a device is included, which includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to run the method and / or method scheme based on the aforementioned multiple embodiments according to the present application.

[0058] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or essential features of the present application.

Claims

1. A method for generating emotional frame indication information of a video, characterized in that: The method comprises: Acquire emotional information corresponding to each video content information in target video information, wherein the target video information includes a plurality of video frame content information, the plurality of video frame content information is arranged in time frame order, each video frame content information includes a plurality of video content information, and the emotional information includes one or more emotional labels corresponding to each video content information, and probability information corresponding to each emotional label; For each video content information, performing weighted fusion based on one or more emotion tags corresponding to the video content information, probability information corresponding to each emotion tag, and first weight information corresponding to each emotion tag to obtain an emotion value for each video content information; For each video frame content information, performing weighted fusion based on the emotion value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain a comprehensive emotion value of each video frame content information; Generate emotional frame indication information of the target video information according to the comprehensive emotional value corresponding to each video frame content information, wherein the emotional frame indication information includes the comprehensive emotional value corresponding to each frame sequence number in the target video information; The step of obtaining the emotion information corresponding to each video content information in the target video information includes: Get target video information; By using a video decomposition method, the target video information is divided into a plurality of video frame content information, wherein the plurality of video frame content information is arranged in a time frame order; For each video frame content information, extracting multiple video content information from the video frame content information by using a content extraction method, wherein the multiple video content information includes image information, text information, and audio information; For each video content information, obtaining one or more emotion tags corresponding to the video content information and probability information corresponding to each emotion tag through a classification method corresponding to the video content information; For each video content information, performing weighted fusion according to one or more emotion tags corresponding to the video content information, probability information corresponding to each emotion tag, and first weight information corresponding to each emotion tag to obtain the emotion value of each video content information includes: For each piece of video content information, weighted fusion is performed based on one or more emotion tags corresponding to the video content information, probability information corresponding to each emotion tag, and first weight information corresponding to each emotion tag to obtain an emotion value for each piece of video content information, wherein the one or more emotion tags include at least one target emotion tag, and the first weight information corresponding to the target emotion tag is higher than that of other emotion tags in the one or more emotion tags; The method of performing weighted fusion on each video frame content information according to the emotion value corresponding to each video content information in the plurality of video content information included in the video frame content information and the second weight information corresponding to each video content information to obtain a comprehensive emotion value of each video frame content information includes: For each video frame content information, if the video content information included in the video frame content information includes facial image information, determine to use the second target weight combination information; otherwise, determine to use the second candidate weight combination information; According to the emotional value corresponding to each video content information in the multiple video content information included in the video frame content information and the second weight information corresponding to each video content information, weighted fusion is performed to obtain a comprehensive emotional value of each video frame content information, wherein, if the video content information includes facial image information, the second weight information belongs to the second target weight combination information; otherwise, the second weight information belongs to the second candidate weight combination information.

2. The method according to claim 1, characterized in that For each video content information, obtaining one or more emotion tags corresponding to the video content information and probability information corresponding to each emotion tag through a classification method corresponding to the video content information includes: For image information, extracting features from the image information to obtain emotional feature information in the image information; determining one or more emotional tags corresponding to the image information and probability information corresponding to each emotional tag based on the emotional feature information; For text information, the text information is input into a text sentiment classification model, and the text sentiment classification model outputs one or more sentiment tags corresponding to the text information, as well as probability information corresponding to each sentiment tag; For audio information, the audio information of multiple consecutive time frames is used as a processing unit, and audio features are extracted from the audio information of the multiple time frames to input the audio features into an audio emotion classification model. The audio emotion classification model outputs one or more emotion labels corresponding to the audio information of the multiple time frames, as well as probability information corresponding to each emotion label, wherein the emotion information corresponding to the audio information of each time frame in the audio information of the multiple time frames is the same.

3. The method according to claim 2, characterized in that The emotional features in the image information are obtained by extracting features from the image information; Determining one or more emotion tags corresponding to the image information and probability information corresponding to each emotion tag based on the emotion feature includes: If the image information includes a face, facial feature information of the face is extracted through face detection technology, and the facial feature information is input into an emotion recognition model as emotion feature information, and the emotion recognition model outputs one or more emotion labels of the image information, as well as probability information of each emotion label; otherwise, multiple groups of emotion feature information are extracted from the image information through multiple feature extraction methods, wherein each group of emotion feature information includes one or more categories of emotion feature information, and each category of emotion feature information represents an emotion label; for each group of emotion feature information, the probability information of each emotion label is determined according to the proportion of each category of emotion feature information in the group of emotion feature information.

4. The method according to claim 1, wherein In the second target weight combination information, the second weight corresponding to the image information is higher than the weight of other video content information; in the second candidate weight combination information, the second weight corresponding to the image information is lower than the weight of other video content information.

5. The method according to claim 1, characterized in that The method further comprises: An emotion peak frame is determined according to the emotion frame indication information of the target video information to mark the emotion peak frame, wherein the emotion peak frame includes a frame number in the target video information whose comprehensive emotion value meets the target emotion value condition.

6. A computer device for generating emotional frame indication information of a video, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Emotion analysis system and method based on probability emotion dictionary

    CN111859925A

  • Landing cylindrical atmosphere lamp illumination control method and system based on nested situation recognition

    CN116916497A