Enhanced human video extinction model design method, equipment and medium

Through optimized data acquisition and processing, automatic extinction and feature encoding technology, combined with EGRUAN decoding module, the limitations of the existing technology in handling dynamic backgrounds and camera viewing angle changes are solved, and efficient video extinction effect is achieved.

CN119991726APending Publication Date: 2025-05-13山东浪潮智慧建筑科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057376.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing human video extinction techniques have limitations in dealing with dynamic backgrounds or camera perspective changes, requiring high cost of manual participation and insufficient use of time domain information.

Method used

Multi-category video data is collected through preset data acquisition rules, automatic extinction and data verification processing is performed, data sets are integrated and the encoding calculation is performed using the feedforward network structure based on the feature encoding module, and combined with EGRUAN decoding to replace RVM decoding, the extinction model parameters are optimized.

Benefits of technology

Deep semantic estimation and fine pixel segmentation of human prospects in video sequences are realized, extinction effect is improved, manual participation costs are reduced, and time-domain information is fully utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991726A_ABST
    Figure CN119991726A_ABST
Patent Text Reader

Abstract

The invention discloses an enhanced human video extinction model design method, equipment and a medium, belongs to the technical field of artificial intelligence, and is used for solving the problems that the existing human video extinction technology has limitation when processing a dynamic background or camera visual angle change, needs certain artificial participation cost, and is inconvenient to use. And the time domain information cannot be fully utilized. The method comprises the following steps: carrying out data acquisition processing on multi-category video data containing character features to obtain a multi-dimensional data set; performing feedforward network structure coding calculation processing based on a feature coding module on an input image in the model training data set to obtain a convolutional deep feature map; carrying out a collocation operation on the convolutional deep feature maps, carrying out input and output control on a feature decoding module, and carrying out a channel splitting and collocation operation on other deep feature maps to obtain a final deep feature map; and completing content information extraction of the to-be-processed character feature video data through the human video extinction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device and medium for designing an enhanced human video extinction model. Background Art

[0002] Alpha Matting, also known as alpha matting, is a technique for separating the foreground and background from an image. In the field of portrait matting, it focuses on extracting the foreground of a portrait from an image. Video matting is further extended to the video field, where it separates video content by predicting the transparency mask of consecutive frames. This technology has important applications in many fields, including film, advertising, and augmented reality.

[0003] Current video matting techniques can be divided into two main categories. The first category is techniques that rely on auxiliary information, which may include a ternary map (Trimap) and a static background image. In the ternary map-based method, the foreground, background, and transition areas in the image need to be manually marked to create a mask map for model input. This method requires human participation and is therefore not suitable for a fully automated matting process. It is mainly used in image and video editing scenarios that require user interaction, such as the foreground, background, and alpha matting model (FBA). Background-based techniques rely on pre-captured scene images. Although they can automatically obtain background information, they have limitations when dealing with dynamic backgrounds or changes in camera perspectives, such as the background matting model (BGMV2). The second category of techniques does not rely on any auxiliary information, such as the real-time portrait matting model without a ternary map (MODNet). MODNet reduces jitter in video sequences by comparing adjacent frames, but fails to make full use of temporal information. Summary of the invention

[0004] The embodiments of the present application provide an enhanced human video extinction model design method, device and medium for solving the following technical problems: the existing human video extinction technology has limitations when processing dynamic backgrounds or changes in camera perspectives, requires certain manual participation costs, and fails to fully utilize time domain information.

[0005] The present application embodiment adopts the following technical solutions:

[0006] On the one hand, an embodiment of the present application provides an enhanced human video extinction model design method, including: according to preset data collection rules, data collection processing is performed on multi-category video data containing human features to obtain a multidimensional data set; automatic extinction and data verification processing is performed on the multidimensional data set, and the optimized multidimensional data set is integrated to obtain a model training data set for the human video extinction model; the input image in the model training data set is subjected to feedforward network structure encoding calculation processing based on a feature encoding module to obtain a convolutional deep feature map; the convolutional deep feature map is juxtaposed, and the eighth deep feature map is determined based on a bilinear interpolation module; according to the eighth deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain a final deep feature map; according to the training process corresponding to the final deep feature map, the original extinction model is parameter optimized to obtain the human video extinction model; and through the human video extinction model, the content information extraction of the processed human feature video data is completed.

[0007] The embodiment of the present application is guided by the gradient path design strategy, comprehensively utilizes structural reparameterization, cross-stage local links and efficient layer aggregation technology, constructs EGRUAN decoding, and uses EGRUAN decoding to replace the RVM (Robust Video Matting, enhanced video matting) decoder module. Through the use of EGRUAN decoding, the human video matting model (RVMPP) can obtain rich gradient flows and realize deep semantic estimation and fine pixel segmentation of human foreground in video sequences.

[0008] In a feasible implementation, according to preset data collection rules, data collection and processing are performed on multi-category video data containing human features to obtain a multidimensional data set, which specifically includes: determining a human with specific features as foreground data, and determining several scenes as background data; synthesizing the foreground data and the background data through foreground and background synthesis technology, and creating the data collection rules; according to the data collection rules, data collection and processing are performed on the foreground and background videos containing human features under relevant random index values ​​to determine them as a first data set; data collection is performed on the dynamic background video that does not contain human features to determine a second data set; data collection and processing are performed on the green screen and foreground videos containing human features under relevant random index values ​​to determine them as a third data set; wherein the multi-category video data includes: the foreground and background videos, the dynamic background video, and the green screen and foreground videos; the multidimensional data set includes: the first data set, the second data set, and the third data set.

[0009] In a feasible implementation manner, the multidimensional dataset is subjected to automatic extinction and data verification processing, and the optimized multidimensional dataset is subjected to dataset integration to obtain a model training dataset for a human video extinction model, specifically including: performing automatic extinction processing on the first dataset in the multidimensional dataset through the MODNet extinction model to obtain basic Alpha extinction data; performing automatic extinction processing on the third dataset in the multidimensional dataset through the MODNet extinction model to obtain foreground and Alpha extinction data; performing data verification processing on the basic Alpha extinction data and the foreground and Alpha extinction data through the labelme annotation tool to obtain an optimized first optimized dataset and a third optimized dataset; synthesizing the first optimized dataset with the second dataset in the multidimensional dataset, and merging the synthesized dataset with the third optimized dataset and the PPM open source data to obtain the model training dataset.

[0010] In a feasible implementation, the input image in the model training data set is subjected to a feedforward network structure coding calculation based on a feature coding module to obtain a convolutional deep feature map, which specifically includes: performing time domain generation processing on a randomly selected time period video in the model training data set to obtain an image sequence; and recursively selecting the image sequence according to an index value to obtain a model input image; inputting the model input image into a replication module and a standard convolution downsampling module in sequence to output a high-resolution feature map; and inputting the high-resolution feature map into a replication module and a standard convolution downsampling module in sequence. After inputting the second feature encoding module, the third feature encoding module, the fourth feature encoding module and the fifth feature encoding module, the second deep feature map, the third deep feature map, the fourth deep feature map and the fifth deep feature map are obtained respectively; the fifth deep feature map is input into the atrous spatial convolution pooling pyramid module to obtain the sixth deep feature map; wherein the convolution deep feature map includes: the second deep feature map, the third deep feature map, the fourth deep feature map, the fifth deep feature map and the sixth deep feature map.

[0011] In a feasible implementation, after the high-resolution feature map is sequentially input into the second feature coding module, the third feature coding module, the fourth feature coding module and the fifth feature coding module, the second deep feature map, the third deep feature map, the fourth deep feature map and the fifth deep feature map are correspondingly obtained, specifically including: the second feature coding module, the third feature coding module, the fourth feature coding module and the fifth feature coding module are all composed of BNeck modules in series; and the network structure of each of the BNeck modules is consistent; the high-resolution feature map is input into the first BNeck module, and the 2-1 deep feature maps are obtained. The 2-1st deep feature map is calculated and processed in sequence through deep separable convolution, global pooling of squeeze-stimulate attention module, fully connected network, input of activation function and convolution to reduce channel dimension, so as to obtain the 2-6th deep feature map of the first BNeck module; the 2-6th deep feature map is input into the second BNeck module; based on channel number expansion processing, jump link to activation function layer, input of deep separable convolution, global average pooling in spatial dimension, spectrum product processing and convolution to reduce spatial and channel dimension, the 2nd deep feature map of the second BNeck module is obtained; the 2nd deep feature map is input into the second BNeck module; Input into the third BNeck module; based on the inverse residual structure with a linear bottleneck, calculate and obtain the 3rd-6th deep feature map of the third BNeck module; input the 3rd-6th deep feature map into the fourth BNeck module, calculate and obtain the 3rd deep feature map of the fourth BNeck module; input the 3rd deep feature map into the fifth BNeck module, calculate and obtain the 4th-6th deep feature map of the fifth BNeck module; input the 4th-6th deep feature map into the sixth BNeck module, calculate and obtain the 4th-12th deep feature map of the sixth BNeck module; The 4th-12th deep feature map is input into the seventh BNeck module, and the 4th deep feature map of the seventh BNeck module is calculated and obtained; based on the 4th deep feature map, the 5th-6th deep feature map of the eighth BNeck module is calculated and obtained; based on the 5th-6th deep feature map, the 5th-12th deep feature map of the ninth BNeck module is calculated and obtained; based on the 5th-12th deep feature map, the 5th-18th deep feature map of the tenth BNeck module is calculated and obtained; based on the 5th-18th deep feature map, the 5th deep feature map of the eleventh BNeck module is calculated and obtained.

[0012] In a feasible implementation, the convolutional deep feature maps are juxtaposed, and based on a bilinear interpolation module, an eighth deep feature map is determined, specifically including: performing short-term memory information storage processing on the sixth deep feature map in the convolutional deep feature map through a convolutional gated recurrent unit to obtain the seventh deep feature map; juxtaposing the sixth deep feature map with the seventh deep feature map to form an identity mapping residual module, and outputting a deep feature map with short-term memory information through the identity mapping residual module; performing a 2-fold upsampling processing on the deep feature map with short-term memory information through a bilinear interpolation module to obtain the eighth deep feature map.

[0013] In a feasible implementation manner, according to the 8th deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain the final deep feature map, specifically including: performing channel splitting processing on the 8th deep feature map to obtain the 8th-1st deep feature map and the 8th-2nd deep feature map; juxtaposing the 8th-1st deep feature map with the 5th deep feature map, and inputting the juxtaposed deep feature map into the tenth feature decoding module to obtain the 10th deep feature map; wherein the feature decoding module includes the tenth feature decoding module, the twelfth feature decoding module and the fourteenth feature decoding module; performing channel splitting processing on the 10th deep feature map to obtain the 10th-1st deep feature map and the 10th-2nd deep feature map; inputting the 8th-2nd deep feature map into the cross-stage local link module of the reparameterized convolutional network to obtain the 9th deep feature map; wherein the reparameterized convolutional network The network cross-stage local link module package includes: a ninth local link module, an eleventh local link module and a thirteenth local link module; according to the feature decoding module and the re-parameterized convolutional network cross-stage local link module, the calculated 8-2nd deep feature map, the 10-2nd deep feature map, the 12-2nd deep feature map, the 14-2nd deep feature map, the 9th deep feature map, the 11th deep feature map and the 13th deep feature map are juxtaposed, Obtain the 17-2th intermediate deep feature map; input the 14-1th deep feature map into the convolution batch normalization rectified linear unit activation function module in sequence to obtain the 17-1th intermediate deep feature map; juxtapose the 17-1th intermediate deep feature map and the 17-2th intermediate deep feature map, and use the juxtaposed deep feature maps as the input of the standard convolution module to obtain the 17th deep feature map; define the 17th deep feature map as the final deep feature map.

[0014] In a feasible implementation manner, according to the training process corresponding to the final deep feature map, the parameters of the original extinction model are optimized to obtain the human video extinction model, specifically including: performing model training processing on the original extinction model through a preset training data set and the training process corresponding to the final deep feature map, and performing reverse chain derivation of the parameters of the original extinction model through a loss function and a back propagation algorithm; and performing parameter optimization processing on the original extinction model in a convergence state through an Adam optimizer to obtain the optimized human video extinction model.

[0015] In a second aspect, an embodiment of the present application also provides an enhanced human video extinction model design device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor so that the at least one processor can execute an enhanced human video extinction model design method described in any of the above-mentioned embodiments.

[0016] In the third aspect, an embodiment of the present application also provides a non-volatile computer storage medium, characterized in that the storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each of which includes instructions, and when the instructions are executed by a terminal, the terminal executes an enhanced human video extinction model design method described in any of the above-mentioned embodiments.

[0017] The present application provides an enhanced human video extinction model design method, device and medium. Compared with the prior art, the embodiments of the present application have the following beneficial technical effects:

[0018] 1. Data collection and processing optimization: Through the preset data collection rules, multi-category video data containing human characteristics can be efficiently obtained to ensure the diversity and accuracy of the data set. The processing after data collection can generate a multidimensional data set, providing richer information for model training.

[0019] 2. Automatic extinction and data verification: Automatic extinction processing can reduce noise and interference in the data and improve data quality. Data verification can ensure the integrity and reliability of the data set and avoid errors in the model training process.

[0020] 3. Dataset integration: By integrating the optimized multidimensional data sets, we obtain the data sets used for model training, which improves the utilization rate of the data sets and the training efficiency of the model.

[0021] 4. Feature encoding and calculation: Using a feedforward network structure based on the feature encoding module for encoding calculation can extract the deep features of the image, which is crucial to the accuracy and robustness of the extinction model.

[0022] 5. Feature map optimization: Through the juxtaposition operation of the convolutional deep feature map and the bilinear interpolation module, the quality of the feature map can be optimized and the accuracy of the extinction processing can be improved.

[0023] 6. Feature decoding and control: The input and output control of the feature decoding module, as well as the channel splitting and juxtaposition operations of the deep feature map, help generate more effective feature representations, thereby improving the performance of the extinction model.

[0024] 7. Parameter optimization: Through the training process corresponding to the final deep feature map, the parameters of the original extinction model are optimized, which can significantly improve the extinction effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0026] Figure 1 A flow chart of a method for designing an enhanced human video extinction model provided in an embodiment of the present application;

[0027] Figure 2 A schematic diagram of the overall network structure of a human video extinction model provided in an embodiment of the present application;

[0028] Figure 3 A schematic diagram of the structure of an enhanced human video extinction model design device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0030] The present application embodiment provides an enhanced human video extinction model design method, such as Figure 1As shown, the enhanced human video extinction model design method specifically includes steps S101-S106:

[0031] S101. Perform data collection and processing on multi-category video data containing human features according to preset data collection rules to obtain a multi-dimensional data set.

[0032] Specifically, firstly, a person with specific features is determined as foreground data, and several scenes are determined as background data. Through the foreground and background synthesis technology, the foreground data and the background data are synthesized and processed, and data collection rules are created.

[0033] Furthermore, according to the data collection rules, data collection processing is performed on the foreground and background videos containing human features under the relevant random index values ​​to determine them as the first data set.

[0034] Furthermore, data collection is performed on a dynamic background video that does not contain human features to determine a second data set.

[0035] Furthermore, data collection and processing under the random index values ​​are performed on the green screen and foreground videos containing human features to determine them as the third data set.

[0036] The multi-category video data includes: foreground and background video, dynamic background video, and green screen and foreground video. The multi-dimensional data set includes: a first data set, a second data set, and a third data set.

[0037] In one embodiment, the open source RVM model (Robust Video Matting) is an advanced video matting algorithm that is specifically optimized for character video matting and can maintain excellent matting quality while processing high-resolution videos in real time. The core innovation of RVM is that it uses a recurrent neural network to process video sequences, making full use of the timing information of the video, so that RVM can achieve better results than traditional frame-by-frame processing methods, especially when dealing with scenes such as complex backgrounds and fast motion. Training is performed by combining synthetic data of the foreground and background, but this method may lead to inconsistencies in brightness and lighting conditions. Although semantic segmentation tasks are introduced and multi-task training is performed using real data sets, the model may still over-adapt to synthetic scenes. In order to overcome the limitations of RVM in training data sets, this application independently collects and annotates a set of high-definition, high-quality private data sets, called VMPP Datasets, whose data distribution matches the actual application scenarios.

[0038] In one embodiment, data collection rules are formulated: Considering the huge workload of manually annotating the Video Matting dataset, it is chosen to continue to use the foreground and background synthesis method to create the application dataset. The goal is to collect a large number of background videos and a small number of videos containing people (foreground). In view of the fact that actual application data mostly appear in exhibitions and exhibition halls with dense and dynamically changing crowds, 10 such scenes were selected for video collection to ensure the consistency of data distribution. At the same time, 20 employees were selected as foreground materials, and there were no special requirements for their clothing and hair accessories.

[0039] In one embodiment, the collection of foreground and background videos: the combination of 1 to 19 people selected from 20 employees is calculated, and after these combinations are shuffled, 10 combinations are selected as foreground materials according to the random index value. These foreground materials are combined with 10 background videos, and a total of 100 videos containing foreground and background are required, each video is 10 to 20 seconds long, forming the FBSets data set (the first data set).

[0040] In one embodiment, dynamic background video collection: 20 dynamic background videos without foreground were shot in 10 exhibition and exhibition hall environments, each video was also 10 to 20 seconds long, and the DBGVSets dataset (second dataset) was constructed.

[0041] In one embodiment, the collection of green screen and foreground videos: In the green screen environment, the combination of 1 to 19 people is calculated from 20 employees. After these combinations are shuffled, 20 combinations are selected as foreground materials according to the random index value. These foreground materials are combined with the green screen background. A total of 20 videos containing foreground and background are required, each video lasting 10 to 20 seconds, forming the GFSets data set (the third data set).

[0042] S102, performing automatic extinction and data verification processing on the multidimensional data set, and integrating the optimized multidimensional data set to obtain a model training data set for a human video extinction model.

[0043] Specifically, it is necessary to first perform automatic extinction processing on the first data set in the multidimensional data set through the MODNet extinction model to obtain basic Alpha extinction data.

[0044] Furthermore, the third data set in the multidimensional data set is automatically extincted by using the MODNet extinction model to obtain foreground and alpha extinction data.

[0045] Furthermore, the labelme annotation tool is used to perform data verification processing on the basic Alpha extinction data and the foreground and Alpha extinction data to obtain the optimized first optimized data set and the third optimized data set.

[0046] Furthermore, the first optimized data set is synthesized with the second data set in the multidimensional data set, and the synthesized data set is merged with the third optimized data set and the PPM open source data to obtain a model training data set.

[0047] In one embodiment, although the open source RVM model uses motion and time data enhancement techniques, these enhancement strategies are still insufficient to cope with the data distribution differences between background images and actual scenes. To solve this problem, we first use the foreground and alpha extinction data of open source datasets such as VideoMatte240K, Distinctions-646, Adobe Image Matting, and background datasets such as DVM and BGMV2 to generate training datasets through synthesis processing. Subsequently, based on these processed datasets and the PPM open source dataset, a fine-tuning dataset was created to reduce the overfitting of the model to the synthetic scene. The specific steps are as follows:

[0048] S2-1) Automatic extinction processing: The open source video extinction model MODNet is used to perform automatic extinction processing on the FBSets dataset (the first dataset) to obtain Alpha extinction data.

[0049] S2-2), further automatic extinction: Also using the MODNet model, automatic extinction is performed on the GFSets dataset (the third dataset) to obtain foreground and alpha extinction data.

[0050] S2-3) Data verification: Use the open source annotation tool labelme to manually verify the manually generated extinction data to ensure the accuracy of the data. This step generates the verified datasets vFBSets (the first optimized dataset) and vGFSets (the second optimized dataset).

[0051] S2-4) Data synthesis and integration: By writing Python scripts, the vFBSets and DBGVSets (second dataset) datasets are synthesized, and the synthesized dataset is merged with the vGFSets dataset and the PPM open source dataset, finally forming the VMPPDatasets dataset (model training dataset). The construction of this dataset aims to provide more realistic and diverse data support for model training.

[0052] S103, performing feedforward network structure encoding calculation processing based on the feature encoding module on the input image in the model training data set to obtain a convolutional deep feature map.

[0053] Specifically, it is necessary to first perform time domain generation processing on the randomly selected time period videos in the model training data set to obtain an image sequence, and then recursively select the image sequence according to the index value to obtain the model input image.

[0054] Furthermore, after the model input image is sequentially input into the replication module and the standard convolution downsampling module, a high-resolution feature map is output.

[0055] Further, Figure 2 A schematic diagram of the overall network structure of a human video extinction model provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, after the high-resolution feature map is sequentially input into the second feature encoding module, the third feature encoding module, the fourth feature encoding module and the fifth feature encoding module, the second deep feature map, the third deep feature map, the fourth deep feature map and the fifth deep feature map are obtained accordingly.

[0056] Furthermore, the fifth deep feature map is input into the atrous spatial convolutional pooling pyramid module to obtain the sixth deep feature map.

[0057] Among them, the convolutional deep feature map includes: the second deep feature map, the third deep feature map, the fourth deep feature map, the fifth deep feature map and the sixth deep feature map.

[0058] As a feasible implementation method, Figure 2As shown, the second feature encoding module, the third feature encoding module, the fourth feature encoding module and the fifth feature encoding module are all composed of BNeck modules (wherein the BNECK module refers to the "bottleneck" structure in computer science, which represents the lowest performance part of the system and limits the performance of the entire system. BNECK modules usually refer to bottleneck points in the system, which may be limitations of hardware resources (such as CPU, memory, disk), or limitations of program algorithms or network transmission) connected in series. And the network structure of each BNeck module is consistent. The high-resolution feature map is input into the first BNeck module to obtain the 2-1 deep feature map. Through deep separable convolution, global pooling of the squeeze-excited attention module, fully connected network, input of activation function and convolution to reduce the channel dimension, the 2-1 deep feature map is calculated and processed in turn to obtain the 2-6 deep feature maps of the first BNeck module. The 2-6 deep feature maps are input into the second BNeck module. Based on the channel number expansion processing, jump link to the activation function layer, input of the depth-separable convolution, global average pooling in the spatial dimension, spectral product processing and convolution to reduce the spatial and channel dimensions, the second deep feature map of the second BNeck module is obtained. The second deep feature map is input into the third BNeck module. Based on the inverse residual structure with a linear bottleneck, the 3rd-6th deep feature maps of the third BNeck module are calculated and obtained. The 3rd-6th deep feature maps are input into the fourth BNeck module, and the 3rd deep feature map of the fourth BNeck module is calculated and obtained; the 3rd deep feature map is input into the fifth BNeck module, and the 4th-6th deep feature maps of the fifth BNeck module are calculated and obtained. The 4th-6th deep feature maps are input into the sixth BNeck module, and the 4th-12th deep feature maps of the sixth BNeck module are calculated and obtained. The 4th to 12th deep feature maps are input into the 7th BNeck module, and the 4th deep feature map of the 7th BNeck module is calculated and obtained. Based on the 4th deep feature map, the 5th to 6th deep feature maps of the 8th BNeck module are calculated and obtained. Based on the 5th to 6th deep feature maps, the 5th to 12th deep feature maps of the 9th BNeck module are calculated and obtained. Based on the 5th to 12th deep feature maps, the 5th to 18th deep feature maps of the 10th BNeck module are calculated and obtained. Based on the 5th to 18th deep feature maps, the 5th deep feature map of the 11th BNeck module is calculated and obtained.

[0059] It should be noted that if Figure 2 As shown, the 2-1 deep feature map corresponds to FM 2-1 , the second deep feature map corresponds to FM2, and the 5th to 18th deep feature maps correspond to FM 5-18 , and so on, the rest of the deep feature maps can be obtained.

[0060] In one embodiment, Figure 2 As shown in Figure 1, a video with a duration of 10 seconds and a frame rate of 25 FPS is randomly selected from the VMPPDatasets dataset (model training dataset) and an image sequence I is generated according to the time domain. i ∈R h×w×3 , where i∈{0, 1, ..., 248, 249}, and recursively select I according to the index value 0 to 249 i , scale it to 1024×1024 resolution as the model input image I∈R 1024×1024×3 .

[0061] S3-1), the input image I passes through the copy module Silence numbered 0 and the standard convolution downsampling module DownSample numbered 1 in turn to obtain a high-resolution feature map Img∈R 512×512×16 .

[0062] S3-2), Img passes through the feature encoding modules EncoderBlock numbered 2, 3, 4, and 5 in sequence to obtain the feature maps FM2∈R from shallow to deep 256×256×24 FM3∈R 128×128×40 FM4∈R 64×64×80 FM5∈R 32×32×160 , shallow high-resolution features are rich in detail information, and deep low-resolution features are rich in semantic information.

[0063] Specifically, the EncoderBlocks numbered 2, 3, 4, and 5 in step S3-2 are composed of 2, 2, 3, and 4 BNeck modules connected in series, respectively. The network structure of each BNeck module is consistent, and is composed of an inverse residual structure with a linear bottleneck, a depth-separable convolution, a squeeze-stimulated attention, and an activation function connected in series. The overall feedforward process of step 2 includes:

[0064] S3-2-1), Img is used as the input of the first BNeck module. On the inverse residual structure with a linear bottleneck, a 1×1 convolution is used, and the number of channels remains unchanged at 16 to obtain FM 2-1 ∈R 512×512×16 , with residual edges and jump links to the activation function layer. FM 2-1 As the input of the depthwise separable convolution, a 1x1 convolution is performed to keep the channel dimension unchanged, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 2-2 ∈R 512×512×16 FM 2-2 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 2-3 ∈R 1×1×16, and then bring it into the 2-layer fully connected network to get FM 2-4 ∈R 1×1×16 , calculate FM 2-4 and FM 2-2 The product of FM 2-5 ∈R 512×512×16 , FM 2-5 As the input of h-swish activation function, FM is obtained 2-6-1 ∈R 512 ×512×16 , juxtaposed FM 2-6-1 And Img and bring it into 1x1 convolution to reduce the channel dimension, and calculate the output FM of the first BNeck module 2-6 ∈R 512×512×16 .

[0065] S3-2-2), FM 2-6 As the input of the second BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 16 to 64, obtaining FM 2-7 ∈R 512×512×64 , with residual edges and jump links to the activation function layer. FM 2-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 24, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 2-8 ∈R 512×512×24 FM 2-8 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 2-9 ∈R 1×1×24 , and then bring it into the 2-layer fully connected network to get FM 2-10 ∈R 1 ×1×24 , calculate FM 2-10 and FM 2-8 The product of FM 2-11 ∈R 512×512×24 , FM 2-11 As the input of h-swish activation function, FM is obtained 2-12 ∈R 512×512×24 , juxtaposed FM 2-12 and FM 2-6 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output FM2∈R of the second BNeck module 256×256×24 .

[0066] S3-2-3), FM2 is used as the input of the third BNeck module. On the inverse residual structure with a linear bottleneck, 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 24 to 72 to obtain FM 3-1 ∈R 256×256×72, with residual edges and jump links to the activation function layer. FM 3-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 24, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 3-2 ∈R 256×256×24 FM 3-2 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 3-3 ∈R 1×1×24 , and then bring it into the 2-layer fully connected network to get FM 3-4 ∈R 1 ×1×24 , calculate FM 3-4 and FM 3-2 The product of FM 3-5 ∈R 256×256×24 , FM 3-5 As the input of h-swish activation function, FM is obtained 3-6-1 ∈R 256×256×24 , juxtaposed FM 3-6-1 And FM2 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the third BNeck module 3-6 ∈R 256×256×24 .

[0067] S3-2-4), FM 3-6 As the input of the fourth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 16 to 72, and the FM 3-7 ∈R 256×256×72 , with residual edges and jump links to the activation function layer. FM 3-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, and the FM is calculated 3-8 ∈R 256×256×40 FM 3-8 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 3-9 ∈R 1×1×40 , and then bring it into the 2-layer fully connected network to get FM 3-10 ∈R 1×1×40 , calculate FM 3-10 and FM 3-8 The product of FM 3-11 ∈R 256×256×40 , FM 3-11 As the input of h-swish activation function, FM is obtained 3-12 ∈R 256×256×40 , juxtaposed FM 3-12 and FM 3-6And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output FM3∈R of the fourth BNeck module 128×128×40 .

[0068] S3-2-5), FM3 is used as the input of the fifth BNeck module. On the inverse residual structure with a linear bottleneck, 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 120, and FM 4-1 ∈R 128×128×120 , with residual edges and jump links to the activation function layer. FM 4-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 4-2 ∈R 128×128×40 FM 4-2 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 4-3 ∈R 1×1×40 , and then bring it into the 2-layer fully connected network to get FM 4-4 ∈R 1 ×1×40 , calculate FM 4-4 and FM 4-2 The product of FM 4-5 ∈R 128×128×40 , FM 4-5 As the input of h-swish activation function, FM is obtained 4-6-1 ∈R 128×128×40 , juxtaposed FM 4-6-1 And FM3 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the fifth BNeck module 4-6 ∈R 128×128×40 .

[0069] S3-2-6), FM 4-6 As the input of the sixth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 120, and the FM 4-7 ∈R 128×128×120 , with residual edges and jump links to the activation function layer. FM 4-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 4-8 ∈R 128×128×40 FM 4-8 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 4-9 ∈R 1×1×40 , and then bring it into the 2-layer fully connected network to get FM 4-10 ∈R1 ×1×40 , calculate FM 4-10 and FM 4-8 The product of FM 4-11 ∈R 128×128×40 , FM 4-11 As the input of h-swish activation function, FM is obtained 4-12-1 ∈R 128×128×40 , juxtaposed FM 4-12-1 and FM 4-6 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the sixth BNeck module 4-12 ∈R 128×128×40 .

[0070] S3-2-7), FM 4-12 As the input of the seventh BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 240, and the FM 4-13 ∈R 128×128×240 , with residual edges and jump links to the activation function layer. FM 4-13 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 4-14 ∈R 128×128×80 FM 4-14 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 4-15 ∈R 1×1×80 , and then bring it into the 2-layer fully connected network to get FM 4-16 ∈R 1×1×80 , calculate FM 4-16 and FM 4-14 The product of FM 4-17 ∈R 128×128×80 , FM 4-17 As the input of h-swish activation function, FM is obtained 4-18 ∈R 128×128×80 , juxtaposed FM 4-18 and FM 4-12 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output FM4∈R of the seventh BNeck module 64×64×80 .

[0071] S3-2-8), FM4 is used as the input of the eighth BNeck module. On the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 200, and FM 5-1 ∈R 64×64×200 , with residual edges and jump links to the activation function layer. FM 5-1As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 5-2 ∈R 64×64×80 FM 5-2 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 5-3 ∈R 1×1×80 , and then bring it into the 2-layer fully connected network to get FM 5-4 ∈R 1 ×1×80 , calculate FM 5-4 and FM 5-2 The product of FM 5-5 ∈R 64×64×80 , FM 5-5 As the input of h-swish activation function, FM is obtained 5-6-1 ∈R 64×64×80 , juxtaposed FM 5-6-1 And FM4 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the eighth BNeck module 5-6 ∈R 64×64×80 .

[0072] S3-2-9), FM 5-6 As the input of the ninth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 184, obtaining FM 4-7 ∈R 64×64×184 , with residual edges and jump links to the activation function layer. FM 5-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 5-8 ∈R 64×64×80 FM 4-8 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 5-9 ∈R 1×1×80 , and then bring it into the 2-layer fully connected network to get FM 5-10 ∈R 1 ×1×80 , calculate FM 5-10 and FM 5-8 The product of FM 5-11 ∈R 64×64×80 , FM 5-11 As the input of h-swish activation function, FM is obtained 5-12-1 ∈R 64×64×80 , juxtaposed FM 5-12-1 and FM 5-6And bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the ninth BNeck module 5-12 ∈R 64×64×80 .

[0073] S3-2-10), FM 5-12 As the input of the tenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 184, and the FM 5-13 ∈R 64×64×184 , with residual edges and jump links to the activation function layer. FM 5-13 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 5-14 ∈R 64×64×80 FM 5-14 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 5-15 ∈R 1×1×80 , and then bring it into the 2-layer fully connected network to get FM 5-16 ∈R 1×1×80 , calculate FM 5-16 and FM 5-14 The product of FM 5-17 ∈R 64×64×80 , FM 5-17 As the input of h-swish activation function, FM is obtained 5-18-1 ∈R 64×64×80 , juxtaposed FM 5-18-1 and FM 5-12 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output FM of the tenth BNeck module 5-18 ∈R 64×64×80 .

[0074] S3-2-11), FM 5-18 As the input of the eleventh BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 480, and the FM 5-19 ∈R 64×64×480 , with residual edges and jump links to the activation function layer. FM 5-19 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 112, and then a 3x3 convolution is performed to keep the spatial dimension unchanged to obtain FM 5-20 ∈R 64×64×112 FM 5-20 As the input of the squeeze-stimulated attention module, we first perform global average pooling in the spatial dimension to obtain FM 5-21 ∈R 1×1×112, and then bring it into the 2-layer fully connected network to get FM 5-22 ∈R 1×1×112 , calculate FM 5-22 and FM 5-20 The product of FM 5-23 ∈R 64×64×112 , FM 5-23 As the input of h-swish activation function, FM is obtained 5-24-1 ∈R 64×64×112 , juxtaposed FM 5-24-1 and FM 5-18 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output FM5∈R of the eleventh BNeck module 32×32×112 .

[0075] S3-3), FM5 passes through the hole spatial convolution pooling pyramid module ASPP numbered 6 to obtain the deep feature map FM6∈R 16×16×160 .

[0076] S104, performing a juxtaposition operation on the convolutional deep feature maps, and determining an eighth deep feature map based on a bilinear interpolation module.

[0077] Specifically, first, a convolutional gated recurrent unit is used to perform short-term memory information storage processing on the 6th deep feature map in the convolutional deep feature map to obtain the 7th deep feature map.

[0078] Furthermore, the sixth deep feature map and the seventh deep feature map are juxtaposed to form an identity mapping residual module, and a deep feature map with short-term memory information is output through the identity mapping residual module.

[0079] Furthermore, through the bilinear interpolation module, the deep feature map with short-term memory information is upsampled by 2 times to obtain the 8th deep feature map.

[0080] In one embodiment, Figure 2 As shown, in a feasible specific implementation:

[0081] S4-1), FM6 (the sixth deep feature map) passes through the convolutional gated recurrent unit ConvGRU numbered 7 to store short-term memory information, and obtain the deep feature map FM7∈R 16×16×160 , then FM6 and FM7 are juxtaposed to form an identity mapping residual module to avoid gradient disappearance and enable the network to learn more useful information, and obtain a deep feature map FM with short-term memory information input8 ∈R 16×16×320 .

[0082] S4-2)、FM input8After the bilinear interpolation module Bilinear2X numbered 8, the deep feature map FM8∈R is obtained by 2 times upsampling. 32×32×320 . Perform channel splitting operation on FM8 to obtain two branch feature maps FM 8-1 ∈R 32 ×32×192 and FM 8-2 ∈R 32×32×128 .

[0083] S105. According to the 8th deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain the final deep feature map.

[0084] Specifically, the 8th deep feature map is subjected to channel splitting processing to obtain the 8-1st deep feature map and the 8-2nd deep feature map.

[0085] Further, the 8-1st deep feature map and the 5th deep feature map are concatenated, and the concatenated deep feature map is input into the tenth feature decoding module to obtain the tenth deep feature map. The feature decoding module includes the tenth feature decoding module, the twelfth feature decoding module and the fourteenth feature decoding module.

[0086] Furthermore, the 10th deep feature map is subjected to channel splitting processing to obtain the 10-1st deep feature map and the 10-2nd deep feature map.

[0087] Further, the 8-2 deep feature map is input into the re-parameterized convolutional network cross-stage local connection module to obtain the 9th deep feature map. The re-parameterized convolutional network cross-stage local connection module package includes: a ninth local connection module, an eleventh local connection module and a thirteenth local connection module.

[0088] Furthermore, according to the feature decoding module and the reparameterized convolutional network cross-stage local connection module, the calculated 8-2nd deep feature map, 10-2nd deep feature map, 12-2nd deep feature map, 14-2nd deep feature map, 9th deep feature map, 11th deep feature map and 13th deep feature map are juxtaposed to obtain the 17-2nd intermediate deep feature map.

[0089] Furthermore, the 14-1th deep feature map is sequentially input into the convolutional batch normalization rectified linear unit activation function module to obtain the 17-1th intermediate deep feature map.

[0090] Further, the 17-1st intermediate deep feature map and the 17-2nd intermediate deep feature map are juxtaposed, and the juxtaposed deep feature maps are used as inputs of the standard convolution module to obtain the 17th deep feature map. The 17th deep feature map is defined as the final deep feature map.

[0091] In one embodiment, Figure 2 As shown, in combination with step S104, the specific real-time method also includes:

[0092] S4-3) and FM 8-1 Concatenate with FM5 to get the feature map FM input10 ∈R 32×32×1152 , and input it into the feature decoding module DecoderBlock numbered 10 to obtain the feature map FM 10 ∈R 64×64×320 . 10 Perform channel splitting operation to obtain two branch feature maps FM respectively 10-1 ∈R 64×64×192 and FM 10-2 ∈R 64×64×128 , FM 10-1 As one of the inputs of DecoderBlock (feature decoding module) numbered 12.

[0093] S4-4) The network structures of the decoder blocks numbered 10, 12, and 14 are consistent, and their input and output processing steps are consistent with step S4-3. The output of the decoder block numbered n∈{10, 12, 14} is the feature map FM n , perform channel splitting operation on it, and obtain the characteristic spectra FM of the two molecules respectively n-1 and FM n-2 .

[0094] S4-5)、FM 8-2 After the reparameterized convolutional network cross-stage local connection module RepNCSP numbered 9, the feature map FM9∈R is obtained. 64×64×32 .

[0095] S4-6) The network structures of RepNCSPs numbered 9, 11, and 13 are consistent. The output of RepNCSP numbered n∈{9, 11, 13} is the feature map FM n .

[0096] S4-7) and FM 8-2 ,FM 10-2 ,FM 12-2 ,FM 14-2 , FM9, FM 11,FM 13 Perform the juxtaposition operation to obtain the feature map FM input17-2 ∈R 256×256×640 (17-2 intermediate deep feature map). The network structure constructed by steps S4-3 and S4-4 is a standard efficient layer aggregation network ELAN. Guided by the gradient path design strategy, it can obtain rich gradient flows, allowing the network to learn more useful fine-grained features. Steps S4-5 and S4-6 add a shortcut structure with RepNCSP (RepNCSPELAN4 is the feature extraction-fusion module in YOLOv9) to each layer, forming a nested ELAN, which further enriches the gradient flow and enables ConvGRU (convolutional gated recurrent unit) to memorize more accurate and rich short-term and long-term information.

[0097] S4-8)、FM 14-1 After passing through the convolution batch normalization corrected linear unit activation function modules CBR numbered 15 and 16, the feature map FM is obtained. input17-1 ∈R 256×256×160 (17-1 intermediate deep feature map), for FM input17-1 and FM input17-2 Perform juxtaposition processing to obtain the feature map FM input17 ∈R 256×256×800 , which will be used as the input of the standard convolution module Conv, and the feature map FM will be obtained after processing 17 ∈R 512×512×8 (Final deep feature map).

[0098] Meanwhile, the overall structure of DecoderBlock in step S4-3 is as follows Figure 1 As shown in the example in the lower left corner, it includes:

[0099] S4-3-1) Input feature map FM input10 After passing through the convolution layer Conv, the batch normalization layer BatchNorm, and the rectified linear unit activation function ReLU in sequence, the feature map FM is obtained. d1 .

[0100] S4-3-2) for the characteristic map FM d1 Perform channel splitting operation to obtain two feature maps FM respectively d1-1 and FM d1-2 , FM d1-1 After the convolution gated recurrent unit ConvGRU stores short-term and long-term memory information, the feature map FM is obtained d2-1 . Then FM d2-1 With FM d1-2 Perform the juxtaposition operation to obtain the feature map FM d2, the entire structure is a residual unit to avoid gradient disappearance.

[0101] S4-3-3) Feature Spectrum FM d2 After the bilinear interpolation module Bilinear2X, the feature map FM is obtained by upsampling by 2 times 10 .

[0102] S106, according to the training process corresponding to the final deep feature map, the original extinction model is optimized to obtain a human video extinction model (RVMPP). And through the human video extinction model, the content information of the character feature video data to be processed is extracted.

[0103] Specifically, the original extinction model is trained through the preset training data set and the training process corresponding to the final deep feature map, and the parameters of the original extinction model are reversely chain-derived through the loss function and the back-propagation algorithm. The parameters of the original extinction model are optimized in the convergence state through the Adam optimizer to obtain the optimized human video extinction model.

[0104] In one embodiment, according to the VMPPDatasets dataset (model training dataset), the training set, validation set, and test set are divided into 80%, 10%, and 10% respectively. Using the training set, the following steps are followed to perform forward reasoning of the model:

[0105] 1. Model training: Use the training set data to perform forward reasoning based on the method steps included in the training process corresponding to the final deep feature map.

[0106] 2. Parameter optimization: Through the loss function and back propagation (BP) algorithm, the model parameters are reverse chained and derived.

[0107] 3. Adam optimizer: Use Adam optimizer to optimize the model parameters until the model reaches convergence. The criterion for model convergence is that the loss value on the validation set changes less than 0.002 in 10 consecutive iterations.

[0108] As a feasible implementation method, after the model training is completed, it will be evaluated on the test set to verify the performance of the model. In addition, the model will also be deployed in the production environment to further verify its effect in actual application scenarios. Specific indicators for evaluating model performance include:

[0109] 1. SAD (Sum of Absolute Differences): The sum of absolute errors is used to measure the difference between the model prediction value and the true value.

[0110] 2. MAD (Mean Absolute Difference): Mean absolute difference, reflecting the average error of model prediction.

[0111] 3. MSE (Mean Squared Error): Mean square error, which measures the average of the sum of squares of the differences between the predicted value and the true value.

[0112] 4. GradientError: Gradient error, which evaluates the prediction accuracy of the model in the edge area.

[0113] 5. ConnectivityError: Connectivity error, which measures the performance of the model in handling image connectivity.

[0114] Through these comprehensive evaluation indicators, we can fully understand the performance of the model in practical applications and make further optimization and adjustments accordingly.

[0115] In addition, the present application embodiment also provides an enhanced human video extinction model design device, such as Figure 3 As shown, the enhanced human video extinction model design device 300 specifically includes:

[0116] At least one processor 301; and a memory 302 in communication with the at least one processor 301; wherein the memory 302 stores instructions executable by the at least one processor 301, so that the at least one processor 301 can execute:

[0117] According to the preset data collection rules, data collection and processing are performed on multi-category video data containing human features to obtain a multidimensional data set;

[0118] Performing automatic extinction and data verification processing on the multidimensional data set, and integrating the optimized multidimensional data set to obtain a model training data set for a human video extinction model;

[0119] The input image in the model training data set is encoded by a feedforward network structure based on the feature encoding module to obtain a convolutional deep feature map.

[0120] The convolution deep feature maps are juxtaposed, and based on the bilinear interpolation module, the eighth deep feature map is determined;

[0121] According to the 8th deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain the final deep feature map;

[0122] According to the training process corresponding to the final deep feature map, the parameters of the original extinction model are optimized to obtain the human video extinction model; and through the human video extinction model, the content information extraction of the character feature video data to be processed is completed.

[0123] The embodiment of the present application uses the gradient path design strategy as a guide, comprehensively utilizes structural reparameterization, cross-stage local links and efficient layer aggregation technology to construct EGRUAN, and uses EGRUAN to replace the decoder module of RVM. Through the use of EGRUAN, RVMPP can obtain rich gradient flows and realize deep semantic estimation and fine pixel segmentation of human foreground in video sequences.

[0124] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0125] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0126] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0127] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0128] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0130] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0131] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0132] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0133] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0134] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for designing an enhanced human video extinction model, characterized in that: The method comprises: According to the preset data collection rules, data collection and processing are performed on multi-category video data containing human features to obtain a multi-dimensional data set; Performing automatic extinction and data verification processing on the multidimensional data set, and integrating the optimized multidimensional data set to obtain a model training data set for a human video extinction model; Performing feedforward network structure encoding calculation processing based on a feature encoding module on the input image in the model training data set to obtain a convolutional deep feature map; The convolution deep feature maps are juxtaposed, and based on a bilinear interpolation module, an eighth deep feature map is determined; According to the eighth deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain a final deep feature map; According to the training process corresponding to the final deep feature map, the parameters of the original extinction model are optimized to obtain the human video extinction model; and through the human video extinction model, the content information extraction of the character feature video data to be processed is completed.

2. The enhanced human video extinction model design method according to claim 1, characterized in that: According to the preset data collection rules, data collection and processing are performed on multi-category video data containing human features to obtain a multidimensional data set, which specifically includes: A person with specific features is determined as foreground data, and a number of scenes are determined as background data; By using foreground and background synthesis technology, the foreground data and the background data are synthesized and processed, and the data collection rules are created; According to the data collection rule, data collection processing is performed on the foreground and background videos containing human features under the relevant random index values ​​to determine them as the first data set; Collecting data of a dynamic background video that does not contain human features to determine a second data set; Performing data collection and processing on the green screen and foreground videos containing human features under the relevant random index values ​​to determine them as the third data set; The multi-category video data includes: the foreground and background videos, the dynamic background video, and the green screen and foreground videos; the multi-dimensional data set includes: the first data set, the second data set, and the third data set.

3. The enhanced human video extinction model design method according to claim 1, characterized in that: The multidimensional data set is subjected to automatic extinction and data verification processing, and the optimized multidimensional data set is subjected to data set integration to obtain a model training data set for a human video extinction model, specifically including: Automatically extinct the first data set in the multidimensional data set by using the MODNet extinction model to obtain basic Alpha extinction data; Automatically extinction processing is performed on the third data set in the multidimensional data set by using the MODNet extinction model to obtain foreground and alpha extinction data; Using the labelme annotation tool, performing data verification processing on the basic Alpha extinction data and the foreground and Alpha extinction data to obtain an optimized first optimized data set and a third optimized data set; The first optimized data set is synthesized with the second data set in the multidimensional data set, and the synthesized data set is merged with the third optimized data set and the PPM open source data to obtain the model training data set.

4. The enhanced human video extinction model design method according to claim 1, characterized in that: The input image in the model training data set is subjected to feedforward network structure encoding calculation processing based on the feature encoding module to obtain a convolutional deep feature map, specifically including: Performing time domain generation processing on randomly selected time period videos in the model training data set to obtain an image sequence; and recursively selecting the image sequence according to the index value to obtain a model input image; After the model input image is sequentially input into the replication module and the standard convolution downsampling module, a high-resolution feature map is output; After the high-resolution feature map is sequentially input into the second feature coding module, the third feature coding module, the fourth feature coding module and the fifth feature coding module, a second deep feature map, a third deep feature map, a fourth deep feature map and a fifth deep feature map are obtained correspondingly; Inputting the fifth deep feature map into the dilated spatial convolutional pooling pyramid module to obtain a sixth deep feature map; Among them, the convolutional deep feature map includes: the second deep feature map, the third deep feature map, the fourth deep feature map, the fifth deep feature map and the sixth deep feature map.

5. The enhanced human video extinction model design method according to claim 4, characterized in that: After the high-resolution feature map is sequentially input into the second feature coding module, the third feature coding module, the fourth feature coding module and the fifth feature coding module, the second deep feature map, the third deep feature map, the fourth deep feature map and the fifth deep feature map are obtained correspondingly, specifically including: The second feature encoding module, the third feature encoding module, the fourth feature encoding module and the fifth feature encoding module are all composed of BNeck modules in series; and the network structure of each BNeck module is consistent; Input the high-resolution feature map into the first BNeck module to obtain the 2-1 deep feature map; through the depth-separable convolution, the global pooling of the squeeze-stimulate attention module, the fully connected network, the input of the activation function and the convolution to reduce the channel dimension, the 2-1 deep feature map is sequentially calculated and processed to obtain the 2-6 deep feature map of the first BNeck module; Input the 2nd to 6th deep feature maps into the second BNeck module; based on channel number expansion processing, jump link to the activation function layer, input of depthwise separable convolution, global average pooling in spatial dimension, map product processing and convolution to reduce spatial and channel dimensions, obtain the 2nd deep feature map of the second BNeck module; Inputting the second deep feature map into the third BNeck module; calculating and obtaining the 3rd to 6th deep feature maps of the third BNeck module based on the inverse residual structure with a linear bottleneck; Inputting the 3rd to 6th deep feature maps into the fourth BNeck module, calculating and obtaining the 3rd deep feature map of the fourth BNeck module; Inputting the third deep feature map into the fifth BNeck module, calculating and obtaining the 4th to 6th deep feature maps of the fifth BNeck module; Inputting the 4th to 6th deep feature maps into the 6th BNeck module, calculating and obtaining the 4th to 12th deep feature maps of the 6th BNeck module; Inputting the 4th to 12th deep feature maps into the 7th BNeck module, calculating and obtaining the 4th deep feature map of the 7th BNeck module; Based on the fourth deep feature map, calculate and obtain the fifth and sixth deep feature maps of the eighth BNeck module; Based on the 5th to 6th deep feature maps, the 5th to 12th deep feature maps of the ninth BNeck module are calculated and obtained; Based on the 5th to 12th deep feature maps, the 5th to 18th deep feature maps of the tenth BNeck module are calculated and obtained; Based on the 5th to 18th deep feature maps, the 5th deep feature map of the eleventh BNeck module is calculated and obtained.

6. The enhanced human video extinction model design method according to claim 1, characterized in that: The convolution deep feature maps are juxtaposed, and based on a bilinear interpolation module, an eighth deep feature map is determined, specifically including: Performing short-term memory information storage processing on the sixth deep feature map in the convolutional deep feature map through a convolutional gated recurrent unit to obtain a seventh deep feature map; The sixth deep feature map and the seventh deep feature map are juxtaposed to form an identity mapping residual module, and a deep feature map with short-term memory information is output through the identity mapping residual module; The deep feature map with short-term memory information is upsampled by a factor of 2 through a bilinear interpolation module to obtain the eighth deep feature map.

7. The enhanced human video extinction model design method according to claim 1, characterized in that: According to the eighth deep feature map, the feature decoding module is input and output controlled, and the remaining deep feature maps are channel split and juxtaposed to obtain the final deep feature map, which specifically includes: Perform channel splitting processing on the 8th deep feature map to obtain an 8-1st deep feature map and an 8-2nd deep feature map; Concatenate the 8-1st deep feature map with the 5th deep feature map, and input the concatenated deep feature map into the tenth feature decoding module to obtain the tenth deep feature map; wherein the feature decoding module includes the tenth feature decoding module, the twelfth feature decoding module and the fourteenth feature decoding module; Performing channel splitting processing on the 10th deep feature map to obtain a 10-1th deep feature map and a 10-2th deep feature map; Input the 8-2 deep feature map into the re-parameterized convolutional network cross-stage local link module to obtain the 9th deep feature map; wherein the re-parameterized convolutional network cross-stage local link module package includes: a 9th local link module, an 11th local link module and a 13th local link module; According to the feature decoding module and the re-parameterized convolutional network cross-stage local link module, the calculated 8-2nd deep feature map, the 10-2nd deep feature map, the 12-2nd deep feature map, the 14-2nd deep feature map, the 9th deep feature map, the 11th deep feature map and the 13th deep feature map are juxtaposed to obtain the 17-2nd intermediate deep feature map; The 14-1 deep feature maps are sequentially input into the convolution batch normalization rectified linear unit activation function module to obtain the 17-1 intermediate deep feature maps; The 17-1 intermediate deep feature map and the 17-2 intermediate deep feature map are juxtaposed, and the juxtaposed deep feature maps are used as input of a standard convolution module to obtain a 17th deep feature map; The 17th deep feature map is defined as the final deep feature map.

8. The enhanced human video extinction model design method according to claim 1, characterized in that: According to the training process corresponding to the final deep feature map, the original extinction model is optimized to obtain the human video extinction model, which specifically includes: The original extinction model is subjected to model training processing through a preset training data set and a training process corresponding to the final deep feature map, and the parameters of the original extinction model are reversely chain-derived through a loss function and a back-propagation algorithm; The Adam optimizer is used to perform parameter optimization processing on the original extinction model in a converged state to obtain the optimized human video extinction model.

9. An enhanced human video extinction model design device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the enhanced human video extinction model design method according to any one of claims 1-8.

10. A non-volatile computer storage medium, characterized in that: The storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each of which includes instructions, and when the instructions are executed by the terminal, the terminal executes an enhanced human video extinction model design method according to any one of claims 1-8.