A personalized video special effect matching method, device and computer-readable storage medium based on artificial intelligence

Through the YOLO-based face emotion recognition model, the emotions of video characters are identified and personalized video special effects are matched, and the problem of low efficiency and accuracy of video special effects matching in the existing technology is solved, and efficient and accurate real-time special effects matching and emotional reflection are achieved.

CN119048965BActive Publication Date: 2025-05-27BEIJING MIAOYIN ANIMATION CULTURE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411456365.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-05-27
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

In the prior art, the efficiency and accuracy of video special effects matching are not high, especially in real-time video processing, real-time matching cannot be achieved, and the selected special effects fail to truly reflect the emotions of the video characters.

Method used

A face emotion recognition model based on YOLO is used to recognize the face emotions of the video characters and match the corresponding personalized video special effects based on the recognition results. The model includes the Backbone layer, the Neck layer and the Head layer, improves feature extraction capabilities through the improved CBS module and the C3AM module, and uses improved loss functions to improve the detection accuracy of the model.

Benefits of technology

It significantly improves the efficiency and accuracy of video special effects matching, realizes real-time special effects matching without manual participation, and better reflects the emotions of the video characters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048965B_ABST
    Figure CN119048965B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a personalized video special effect matching method, device, and computer-readable storage medium based on artificial intelligence, including obtaining video training sample data; constructing a face emotion recognition model based on YOLO, and inputting the video training sample data into the face emotion recognition model for training; obtaining a video to be processed, and inputting the video to be processed into the face emotion recognition model to obtain a corresponding face emotion recognition result; and matching corresponding personalized video special effects for the characters in the video according to the face emotion recognition result. The present invention recognizes the face emotions of video characters through a face emotion recognition model based on YOLO, and matches corresponding video special effects according to the recognition results. The accuracy of detection in this process is high, and no manual participation is required, significantly improving the efficiency and accuracy of video special effect matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a personalized video special effects matching method, device and computer-readable storage medium based on artificial intelligence. Background Art

[0002] Video special effects refer to a series of post-processing technologies and creative means used in the video production process to enhance the visual impact, improve the artistic effect or tell the story, so as to enhance the audience's viewing experience. These special effects can include color adjustment, picture deformation, dynamic graphic overlay, light and shadow effects, etc. Video special effects are difficult to achieve in the actual video shooting process. The application range of video special effects is very wide, including movies, television, advertising, music videos, games, short videos, virtual reality, etc. Studies have shown that adding special effects to videos will greatly increase users' interest in browsing or watching.

[0003] The method of matching video special effects in the prior art is mainly achieved through manual selection, and the specific process includes: when the user locates the position where the special effects are added to the video, the corresponding video special effects are selected according to the video special effects library provided by the system. However, the above process has the following disadvantages: First, the process of manually selecting special effects requires manual participation, which is time-consuming and labor-intensive, and inefficient. Especially when matching video special effects for videos with high real-time requirements such as live broadcasts, special effects matching cannot be achieved in real time, and it is difficult to meet the real-time requirements of video special effects matching; second, when manually selecting special effects, due to the large number of special effects types provided by video special effects, users mainly choose special effects based on personal preferences. The selected special effects do not truly reflect the real emotions of the video characters, resulting in the video special effects and video characters not achieving the best match. Therefore, in the process of matching video special effects in the prior art, there are technical problems such as low matching efficiency and matching accuracy. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a personalized video special effects matching method based on artificial intelligence, which can solve the technical problems of low efficiency and accuracy of video special effects matching in the prior art.

[0005] The first aspect of the present invention provides a personalized video special effects matching method based on artificial intelligence, comprising the following steps:

[0006] S1: Obtain a video data set, and annotate facial emotions in the video data set as video training sample data;

[0007] S2: constructing a YOLO-based facial emotion recognition model, inputting the video training sample data into the facial emotion recognition model for training, and updating the parameters of the facial emotion recognition model using a loss function;

[0008] S3: Obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result;

[0009] S4: Matching corresponding personalized video effects for the characters in the video according to the facial emotion recognition results.

[0010] As a further improvement of the present invention, the facial emotion recognition model includes a Backbone layer, a Neck layer and a Head layer; the Backbone layer includes a Focus layer, a first ICBS module, a first C3AM module, a second ICBS module, a second C3AM module, a third ICBS module, a third C3AM module, a fourth ICBS module, a fourth C3AM module, and an SPPF module connected in sequence; the Neck layer includes a second CDR module, a second upsampling module, a second concat module, a first C3 module, a first CDR module, a first upsampling module, a first concat module, a second C3 module, a fifth ICBS module, a third C3AM module, and a SPPF module connected in sequence. three concat modules, a third C3 module, a sixth ICBS module, a fourth concat module, and a fourth C3 module; the Head layer includes a first feature map P1, a second feature map P2 and a third feature map P3; the output of the second C3AM module is input to the first concat module, the output of the fourth C3AM module is input to the second concat module, the output of the SPPF module is input to the second CDR module, the output of the second CDR module is input to the fourth concat module, the output of the first CDR module is input to the third concat module, the second C3 module, the third C3 module, and the fourth C3 module are connected to P1, P2, and P3 respectively.

[0011] As a further improvement of the present invention, the structures of the first ICBS module to the sixth ICBS module are the same; the ICBS module is an improved CBS module, including a CDR module and an SD module; the CDR module includes a CBS module, a deep convolution module (DSConv), a concat module and a recombine module connected in sequence, the output of the CBS module is input into the concat module, and the recombine module is used to realize random reordering of different channels; the SD module includes 4 parallel Slice modules and a concat module, the output of the CDR module is respectively input into the parallel Slice modules, and the output of the Slice module is respectively input into the concat module.

[0012] As a further improvement of the present invention, the structures of the first C3AM module to the fourth C3AM module are the same; the C3AM module is an attention mechanism module (AM) embedded in the original C3 module, and the type of the attention mechanism module is CA attention mechanism or ECA attention mechanism.

[0013] As a further improvement of the present invention, the loss function is:

[0014]

[0015] Among them, Loss is the loss function, IoU is the area intersection ratio between the predicted box and the real box, b is the center point of the predicted box, and b gt is the center point of the real box, ρ(b,b gt ) represents the distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closed area containing the predicted box and the real box, α represents the weight coefficient, v represents the similarity coefficient of the aspect ratio of the predicted box and the real box, w gt is the width of the real box, h gt is the height of the real box, w is the width of the predicted box, h is the height of the predicted box, and k 1 and k 2 is the preset proportional coefficient, d 1 is the distance between the midpoint of the bottom border of the predicted box and the real box, d 2 is the distance between the midpoint of the left border of the predicted box and the real box, d 3 is the distance between the minimum edge vertex coordinates of the predicted box and the real box, d 4 is the distance between the maximum edge vertex coordinates of the predicted box and the true box.

[0016] As a further improvement of the present invention, the facial emotion recognition result corresponds to multiple personalized video effects, and the multiple personalized video effects have different priorities.

[0017] As a further improvement of the present invention, the facial emotion recognition results include happiness, sadness, embarrassment, and sadness.

[0018] As a further improvement of the present invention, the personalized video special effects include color correction, text animation, and ray tracing.

[0019] The second aspect of the present invention provides a personalized video special effects matching device based on artificial intelligence, including a training sample acquisition module, a model training module, a face emotion recognition module, and a personalized special effects matching module;

[0020] The training sample acquisition module is used to acquire a video data set and annotate facial emotions in the video data set as video training sample data;

[0021] The model training module is used to build a YOLO-based facial emotion recognition model, input the video training sample data into the facial emotion recognition model for training, and use the loss function to update the parameters of the facial emotion recognition model;

[0022] The facial emotion recognition module is used to obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result;

[0023] The personalized special effects matching module is used to match corresponding personalized video special effects for the characters in the video according to the facial emotion recognition results.

[0024] A third aspect of the present invention provides a computer-readable storage medium, which stores a computer program. The computer program can be executed by a processor to implement the aforementioned personalized video special effects matching method based on artificial intelligence.

[0025] Compared with the prior art, the present invention has at least the following beneficial effects:

[0026] (1) The present invention recognizes the facial emotions of video characters through a YOLO-based facial emotion recognition model, and matches the corresponding video special effects according to the recognition results. The detection accuracy of this process is high and no human intervention is required, which significantly improves the efficiency and accuracy of video special effects matching. At the same time, the present invention improves the original YOLO model, improves the feature extraction capability, reduces the loss of feature information, and can better ensure the accuracy of video special effects matching.

[0027] (2) The loss function provided by the present invention improves the parameters of the original CIOU damage function and adds a distance difference penalty term, which can improve the convergence speed of the loss function on the one hand and prevent the degradation of the CIOU loss function on the other hand. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0029] Figure 1 It is a flowchart of a personalized video special effects matching method based on artificial intelligence provided by an embodiment of the present invention.

[0030] Figure 2 It is a YOLO model structure diagram provided by the prior art.

[0031] Figure 3 It is a structural block diagram of a YOLO-based facial emotion recognition model provided by an embodiment of the present invention.

[0032] Figure 4 It is a structural block diagram of an improved CBS (ICBS) provided by an embodiment of the present invention.

[0033] Figure 5 It is a structural block diagram of a C3AM provided by an embodiment of the present invention.

[0034] Figure 6 It is a schematic diagram of a distance penalty item of an improved CIOU provided by an embodiment of the present invention.

[0035] Figure 7 It is a structural block diagram of a personalized video special effects matching device based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0037] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0038] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0039] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0040] Embodiment 1:

[0041] See also Figure 1 As shown, the first aspect of the present invention provides a personalized video special effects matching method based on artificial intelligence, comprising the following steps:

[0042] S1: Obtain a video data set, and annotate facial emotions in the video data set as video training sample data;

[0043] Step S1 is to obtain training sample data. Those skilled in the art may also select a commonly used facial expression data set, such as the Fer2013 facial expression data set, which has 35,886 facial expression pictures, including 7 expressions, namely: Anger, Disgust, Fear, Happy, Neurtal, Sad, Surprise (anger, disgust, fear, happiness, natural, sadness, surprise). The number of expressions in the data set is unbalanced and contains a small number of erroneous samples, which can better simulate facial expression data.

[0044] S2: constructing a YOLO-based facial emotion recognition model, inputting the video training sample data into the facial emotion recognition model for training, and updating the parameters of the facial emotion recognition model using a loss function;

[0045] In step S2, a facial emotion recognition model is established based on YOLO, and the model is trained using training samples. The YOLO model is a regression method based on deep learning and is widely used in image classification. Figure 2 It is a YOLO model architecture provided by the existing technology. The CSP structure (Cross Stage Partial) is used in the Backbone part, and the path aggregation network (PAN) + feature pyramid network (FPN) structure is used in the Neck network part. The FPN structure outputs the feature maps of different layers in the Backbone network to obtain the features between different layers. The PAN structure is based on the low-sampling features obtained by FPN, and then adds upsampling for feature fusion. In YOLOv5, the PAN+FPN structure can make full use of and fuse the features extracted from the Backbone network to obtain better detection performance.

[0046] The CBS module, namely the Conv BatchNorm SiLU module, is a basic convolutional neural network building block. It consists of three main parts: a convolutional layer (Conv), a batch normalization layer (BatchNorm), and a SiLU activation function. The C3 module combines multiple bottleneck structures and residual connections to form an efficient feature extraction module. The C3 module is usually composed of an initial 1x1 convolutional layer, multiple bottleneck structures, a residual connection, and a final 1x1 convolutional layer. In the C3 module, the feature map is extracted through multiple bottleneck structures, and the initial features and the extracted features are added through the residual connection to form the final output feature map. The above modules are all conventional component structures in the YOLO model, and the embodiments of the present invention will not be described in detail.

[0047] However, the detection accuracy of the above model is relatively poor, mainly due to the poor feature extraction capability. Therefore, the present invention improves the above model to reduce the loss of effective features during feature extraction. It can extract features more comprehensively and control the surge in model parameters.

[0048] See also Figure 3 As shown, the facial emotion recognition model includes a Backbone layer, a Neck layer and a Head layer; the Backbone layer includes a Focus layer, a first ICBS module, a first C3AM module, a second ICBS module, a second C3AM module, a third ICBS module, a third C3AM module, a fourth ICBS module, a fourth C3AM module, and an SPPF module connected in sequence; the Neck layer includes a second CDR module, a second upsampling module, a second concat module, a first C3 module, a first CDR module, a first upsampling module, a first concat module, a second C3 module, a fifth ICBS module, a third concat module, and a SPPF module connected in sequence. at module, third C3 module, sixth ICBS module, fourth concat module, fourth C3 module; the Head layer includes a first feature map P1, a second feature map P2 and a third feature map P3; the output of the second C3AM module is input to the first concat module, the output of the fourth C3AM module is input to the second concat module, the output of the SPPF module is input to the second CDR module, the output of the second CDR module is input to the fourth concat module, the output of the first CDR module is input to the third concat module, the second C3 module, the third C3 module, and the fourth C3 module are connected to P1, P2, and P3 respectively.

[0049] The structures of the first ICBS module to the sixth ICBS module are the same; the ICBS module is an improved CBS module, including a CDR module and an SD module; the CDR module includes a CBS module, a deep convolution module (DSConv), a concat module and a recombine module connected in sequence, the output of the CBS is input to the concat module, and the recombine module is used to realize random reordering of different channels; the SD module includes 4 parallel Slice modules and a concat module, the output of the CDR module is respectively input to the parallel Slice modules, and the output of the Slice module is respectively input to the concat module.

[0050] The structures of the first C3AM module to the fourth C3AM module are the same; the C3AM module is an original C3 module in which an attention mechanism module (AM) is embedded, and the type of the attention mechanism module is a CA attention mechanism or an ECA attention mechanism.

[0051] The present invention replaces the original CBS module before each C3 module in the Backbone network of the YOLO framework in the prior art with an ICBS module. In the Neck network, the first two CBS modules are replaced by CDR modules, and the last two CBS modules are replaced by ICBS modules. Among them, the ICBS module is obtained by improving the original CBS module, which consists of a CDR module (the abbreviation of its substructures CBS module, DSConv module, and recombine module) and an SD module (Space Depth). The function of the CDR module is to reduce information loss during feature extraction and suppress the growth of model parameters. The structure of the CDR module is as follows Figure 4 As shown in Figure 2. The function of this module is to fuse different convolutional features, reduce the probability of local feature information loss, and improve the information correlation between channels. The function of the SD module is to compress the image size while retaining the feature information.

[0052] S3: Obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result;

[0053] S4: Matching corresponding personalized video effects for the characters in the video according to the facial emotion recognition results.

[0054] The present invention recognizes the facial emotions of video characters through a YOLO-based facial emotion recognition model, and matches corresponding video special effects according to the recognition results. The process has high detection accuracy and does not require human intervention, which significantly improves the efficiency and accuracy of video special effects matching. At the same time, the present invention improves the original YOLO model, improves the feature extraction capability, reduces the loss of feature information, and can better ensure the accuracy of video special effects matching.

[0055] As a further improvement of the present invention, the loss function is:

[0056]

[0057] Among them, Loss is the loss function, IoU is the area intersection ratio between the predicted box and the real box, b is the center point of the predicted box, and b gt is the center point of the real box, ρ(b,b gt ) represents the distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closed area containing the predicted box and the real box, α represents the weight coefficient, v represents the similarity coefficient of the aspect ratio of the predicted box and the real box, w gt is the width of the real box, h gt is the height of the real box, w is the width of the predicted box, h is the height of the predicted box, and k 1and k 2 is the preset proportional coefficient, d 1 is the distance between the midpoint of the bottom border of the predicted box and the real box, d 2 is the distance between the midpoint of the left border of the predicted box and the real box, d 3 is the distance between the minimum edge vertex coordinates of the predicted box and the real box, d 4 is the distance between the maximum edge vertex coordinates of the predicted box and the true box.

[0058] In object detection based on YOLO, the commonly used loss function is CIOU (Complete Intersection over Union), which takes into account three important geometric factors: overlapping area, center point distance and aspect ratio. CIoU measures the overlapping area of ​​the target and the real box through IoU, which is measured by the angle of distance and corresponding aspect ratio. The definition of CIOU is as follows:

[0059]

[0060] where ρ(b,b gt ) represents the distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closed area containing the predicted box and the real box, α represents the weight coefficient, and v represents the similarity coefficient of the aspect ratio of the predicted box and the real box. gt is the width of the real box, h gt is the height of the real box, w is the width of the predicted box, and h is the height of the predicted box. However, the traditional CIoU loss function has the following disadvantages: 1. v is less robust and is greatly affected by outliers, resulting in dramatic changes in the loss function value and slow convergence; 2. When the aspect ratio of the predicted box and the real box is the same, the CIOU loss function will degenerate into the DIOU loss function, affecting the recognition accuracy.

[0061] First, in order to improve the robustness of the parameters and increase the convergence speed of the model, the present invention improves the parameter v of the CIOU loss function, specifically:

[0062]

[0063] The improved parameter v is more robust and smoother than the original parameter v, which enables the regression loss function to achieve faster convergence speed, better positioning results and model performance.

[0064] Secondly, in order to solve the degradation problem of CIOU loss function, the present invention adds a distance difference penalty term to the loss function. Figure 6As shown, when the aspect ratio of the predicted box and the real box is the same, αv in the original CIOU loss function will become 0. At this time, the CIOU loss function will deteriorate and become a DIOU loss function, resulting in inaccurate calculation of the loss function, further affecting the recognition result of the model. Therefore, in order to avoid the above problem, the present invention adds a distance difference penalty term. Even if αv in the original CIOU loss function becomes 0 due to the same aspect ratio of the predicted box and the real box, due to the existence of the distance difference penalty term, the CIOU loss function will not deteriorate, thereby improving the accuracy of the CIOU loss function, thereby further ensuring the accuracy of model recognition.

[0065] As a further improvement of the present invention, the facial emotion recognition result corresponds to multiple personalized video effects, and the multiple personalized video effects have different priorities.

[0066] In order to enhance the user experience, the present invention sets a plurality of corresponding video special effects for each emotion recognition result. For example, the special effects corresponding to happiness include enlarging the mouth or increasing the sound of laughter. In addition, different priorities are set for different special effects, so that after the user sets it, the system can automatically select the corresponding special effect type according to the user's preference.

[0067] As a further improvement of the present invention, the facial emotion recognition results include happiness, sadness, embarrassment, and sadness.

[0068] As a further improvement of the present invention, the personalized video special effects include color correction, text animation, and ray tracing.

[0069] As a further improvement of the present invention, the parameters are mapped to the whale positions of the whale optimization algorithm, and the whale optimization algorithm specifically includes:

[0070] S21: Initialize parameters, including the whale position X(t), the maximum number of iterations N max , current number of iterations N, fitness function f, iteration parameters Pace coefficient A = 2ar-a, weight coefficient C = 2a, r, r 3 is a random number between 0 and 1, l is a random number between -1 and 1, h is a random number between 0 and 1, b is the parameter of the spiral path, and β is a preset weight factor;

[0071] S22: Calculate the fitness value of each whale, let X * is the location of the whale with the best fitness value;

[0072] S23: Start looping and update all parameters; if h<1 / 2, execute step S24, otherwise execute step S25;

[0073] S24: If |A|≥1, enter the global search phase and update the whale position using the following formula: Where X(t+1) is the updated whale position, X rand is the randomly selected whale position, sign is the sign function, Σ(Y(t)) is the cumulative sum of each element of Y(t), and ||·|| represents the L2 norm; if |A|<1, the shrinking and encircling stage is entered, and the whale position is updated using the following formula: X(t+1)=X * -A|CX * -X(t)|;

[0074] S25: Enter the spiral update phase and use the following formula to update the whale position:

[0075] S26: recalculate the fitness value of the whale and repeat steps S23-S25;

[0076] S27: Determine whether convergence has been achieved or the maximum number of iterations has been reached. If so, output the result.

[0077] Whale Optimization Algorithm (WOA) is a new type of swarm intelligence optimization search method proposed by Mirjalili et al. of Griffith University, Australia in 2016. This algorithm simulates the hunting behavior of humpback whale groups in nature. Its core idea is to find the optimal solution by simulating the self-organization and adaptability of whale groups. However, the traditional whale optimization algorithm is prone to fall into the local optimal solution. The embodiment of the present invention improves the traditional whale optimization algorithm by adding adaptive weights to the position update formula of each stage, which can improve the search ability of the algorithm, ensure that the algorithm will not fall into the local optimal solution, accelerate the convergence speed of the algorithm, and better ensure that the model parameters are optimized.

[0078] Embodiment 2:

[0079] The second aspect of the present invention provides a personalized video special effects matching device based on artificial intelligence, including a training sample acquisition module, a model training module, a face emotion recognition module, and a personalized special effects matching module;

[0080] The training sample acquisition module is used to acquire a video data set and annotate facial emotions in the video data set as video training sample data;

[0081] The model training module is used to build a YOLO-based facial emotion recognition model, input the video training sample data into the facial emotion recognition model for training, and use the loss function to update the parameters of the facial emotion recognition model;

[0082] The facial emotion recognition module is used to obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result;

[0083] The personalized special effects matching module is used to match corresponding personalized video special effects for the characters in the video according to the facial emotion recognition results.

[0084] Embodiment 3:

[0085] A third aspect of the present invention provides a computer-readable storage medium, which stores a computer program. The computer program can be executed by a processor to implement the aforementioned personalized video special effects matching method based on artificial intelligence.

[0086] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A personalized video special effects matching method based on artificial intelligence, characterized by: The steps include: S1: Obtain a video data set, and annotate facial emotions in the video data set as video training sample data; S2: constructing a facial emotion recognition model based on YOLOv5, inputting the video training sample data into the facial emotion recognition model for training, and updating the parameters of the facial emotion recognition model using a loss function; S3: Obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result; S4: matching corresponding personalized video effects for the characters in the video according to the facial emotion recognition results; The original CBS module before each C3 module in the Backbone layer of the facial emotion recognition model is replaced by an ICBS module, the first two CBS modules in the Neck layer are replaced by a CDR module, and the last two CBS modules are replaced by an ICBS module; the ICBS module is an improved CBS module, including a CDR module and an SD module; the CDR module includes a CBS module, a deep convolution module (DSConv), a concat module and a recombine module connected in sequence, the output of the CBS module is input to the concat module, and the recombine module is used to realize random reordering of different channels; the SD module includes 4 parallel Slice modules and a concat module, the output of the CDR module is respectively input to the parallel Slice modules, and the output of the Slice module is respectively input to the concat module; the CBS module includes a convolution layer (Conv), a batch normalization layer (BatchNorm), and a SiLU activation function.

2. The method according to claim 1, characterized in that: The facial emotion recognition model includes a Backbone layer, a Neck layer and a Head layer; the Backbone layer includes a Focus layer, a first ICBS module, a first C3AM module, a second ICBS module, a second C3AM module, a third ICBS module, a third C3AM module, a fourth ICBS module, a fourth C3AM module, and an SPPF module connected in sequence; the Neck layer includes a second CDR module, a second upsampling module, a second concat module, a first C3 module, a first CDR module, a first upsampling module, a first concat module, a second C3 module, a fifth ICBS module, a third conca t module, third C3 module, sixth ICBS module, fourth concat module, fourth C3 module; the Head layer includes a first feature map P1, a second feature map P2 and a third feature map P3; the output of the second C3AM module is input to the first concat module, the output of the fourth C3AM module is input to the second concat module, the output of the SPPF module is input to the second CDR module, the output of the second CDR module is input to the fourth concat module, the output of the first CDR module is input to the third concat module, the second C3 module, the third C3 module, and the fourth C3 module are connected to P1, P2, and P3 respectively.

3. The method according to claim 2, characterized in that The structures of the first C3AM module to the fourth C3AM module are the same; the C3AM module is an original C3 module in which an attention mechanism module (AM) is embedded, and the type of the attention mechanism module is a CA attention mechanism or an ECA attention mechanism.

4. The method according to claim 1, characterized in that: The loss function is: Among them, Loss is the loss function, IoU is the area intersection ratio between the predicted box and the real box, b is the center point of the predicted box, and b gt is the center point of the real box, ρ(b,b gt ) represents the distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closed area containing the predicted box and the real box, α represents the weight coefficient, v represents the similarity coefficient of the aspect ratio of the predicted box and the real box, w gt is the width of the real box, h gt is the height of the real box, w is the width of the predicted box, h is the height of the predicted box, k1 and k2 are preset scale factors, d1 is the distance between the midpoint of the bottom border of the predicted box and the real box, d2 is the distance between the midpoint of the left border of the predicted box and the real box, d3 is the distance between the minimum edge vertex coordinates of the predicted box and the real box, and d4 is the distance between the maximum edge vertex coordinates of the predicted box and the real box.

5. The method according to claim 1, characterized in that: The facial emotion recognition result corresponds to a plurality of personalized video effects, and the plurality of personalized video effects have different priorities.

6. The method according to claim 1, characterized in that: The facial emotion recognition results include happiness, sadness, embarrassment, and sadness.

7. The method according to claim 1, characterized in that: The personalized video special effects include color correction, text animation, and ray tracing.

8. A personalized video special effects matching device based on artificial intelligence, comprising a training sample acquisition module, a model training module, a face emotion recognition module, and a personalized special effects matching module; The training sample acquisition module is used to acquire a video data set and annotate facial emotions in the video data set as video training sample data; The model training module is used to build a facial emotion recognition model based on YOLOv5, input the video training sample data into the facial emotion recognition model for training, and use the loss function to update the parameters of the facial emotion recognition model; The facial emotion recognition module is used to obtain a video to be processed, and input the video to be processed into the facial emotion recognition model to obtain a corresponding facial emotion recognition result; The personalized special effects matching module is used to match corresponding personalized video special effects for the characters in the video according to the facial emotion recognition results; The original CBS module before each C3 module in the Backbone layer of the facial emotion recognition model is replaced by an ICBS module, the first two CBS modules in the Neck layer are replaced by a CDR module, and the last two CBS modules are replaced by an ICBS module; the ICBS module is an improved CBS module, including a CDR module and an SD module; the CDR module includes a CBS module, a deep convolution module (DSConv), a concat module and a recombine module connected in sequence, the output of the CBS module is input to the concat module, and the recombine module is used to realize random reordering of different channels; the SD module includes 4 parallel Slice modules and a concat module, the output of the CDR module is respectively input to the parallel Slice modules, and the output of the Slice module is respectively input to the concat module; the CBS module includes a convolution layer (Conv), a batch normalization layer (BatchNorm), and a SiLU activation function.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the artificial intelligence-based personalized video special effects matching method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Real-time expression recognition method based on YOLOv5l and attention mechanism

    CN115497140A

  • Classroom facial expression recognition method and device based on YOLOv4

    CN116453178A

  • Image processing method and device, computer equipment, storage medium and program product

    CN116645706A

  • Facial expression recognition system based on artificial intelligence

    CN118135638A