Role recognition method, device, computer equipment and storage medium

By aligning and resizing image regions based on key point identification, the method improves the accuracy of object recognition by preserving the aspect ratio, addressing the distortion issues in traditional target recognition methods.

CN113822142BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110857929.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-28
Publication Date
2025-07-15
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

Traditional object recognition methods can easily lead to distortion of length and width ratio when adjusting the image size, affecting the recognition accuracy of target objects in the image.

Method used

By obtaining the pending video frames, performing object detection and key point positioning, and using alignment registration processing to map the character area into a target image of a preset size, avoiding direct image size adjustment and maintaining coordination of length and width ratios.

Benefits of technology

Improve the accuracy of character recognition, ensure that the length and width ratio of character objects in the image are coordinated, and the role information can be accurately obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822142B_ABST
    Figure CN113822142B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for role recognition. The method includes: obtaining a video to be processed, and extracting target video frames from the video to be processed; performing object detection on each of the target video frames to determine the role regions where the role objects appearing in each of the target video frames are located respectively; determining the key position information of the target key points of each role object in the corresponding target video frame; based on the key position information, performing alignment and registration processing on the images formed by the role regions where each of the role objects is located respectively, to obtain target images with a preset size corresponding to each role object respectively; performing role recognition based on the target images to obtain the role information corresponding to each of the role objects in the video to be processed. Using this method can improve the recognition accuracy of role objects in images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a role recognition method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of computer technology, object recognition technology has emerged. Through object recognition, target objects in images or videos can be located and recognized to obtain information about the target objects in the images or videos. For example, by using a recognition model to recognize the people in an image or video to determine information such as the roles played by the people in the image or video.

[0003] However, different images may have different sizes. Traditional object recognition methods often directly change the size of the image to make the image meet the processing requirements of the recognition model. However, this method is likely to cause distortion of the aspect ratio of the image content, affecting the recognition accuracy of the target objects in the image. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a role recognition method, apparatus, computer device, and storage medium that can improve the recognition accuracy of role objects in images.

[0005] A role recognition method, the method comprising:

[0006] Obtaining a video to be processed, and extracting target video frames from the video to be processed;

[0007] Performing object detection on each of the target video frames to determine the role regions where the role objects appearing in each of the target video frames are located;

[0008] Determining the key position information of the target key points of each role object in the corresponding target video frame;

[0009] Based on the key position information, performing alignment and registration processing on the images formed by the role regions where each role object is located, respectively, to obtain target images of a preset size corresponding to each role object;

[0010] Performing role recognition based on the target images to obtain role information corresponding to each role object in the video to be processed.

[0011] A role recognition apparatus, the apparatus comprising:

[0012] An obtaining module, configured to obtain a video to be processed, and extract target video frames from the video to be processed;

[0013] A detection module, configured to perform object detection on each of the target video frames respectively to determine the role regions where the role objects appearing in each of the target video frames are located;

[0014] A determination module, configured to determine the key position information of the target key points of each role object in the corresponding target video frame;

[0015] A registration module, configured to perform alignment and registration processing on the images formed by the role regions where each role object is located respectively based on the key position information to obtain target images with a preset size corresponding to each role object respectively;

[0016] An identification module, configured to perform role identification based on the target images to obtain the role information corresponding to each role object in the video to be processed.

[0017] In one embodiment, the detection module is further configured to perform convolution processing on each of the target video frames respectively to obtain corresponding video frame features; slide preset detection frames on each of the target video frames respectively to obtain each candidate frame corresponding to each of the target video frames respectively; determine the role objects appearing in each of the target video frames based on each candidate frame corresponding to each of the target video frames respectively, and determine the role regions where the role objects are located in the corresponding target video frames.

[0018] In one embodiment, the detection module is further configured to, for each of the target video frames, determine the role objects appearing in the corresponding target video frame according to each candidate frame corresponding to the corresponding target video frame respectively; for each role object, enlarge the size of the candidate frame where the corresponding role object is located, and use the region included in the enlarged candidate frame as the role region where the role object is located in the corresponding target video frame.

[0019] In one embodiment, the determination module is further configured to, when there are target parts for the role object in the corresponding target video frame, extract feature points from the target parts of each role object respectively to obtain the target key points corresponding to each target part respectively, and determine the key position information of the target key points in the corresponding target video frame; when there are no target parts for the role object in the corresponding target video frame, predict the target parts of the role object in the corresponding video frame, and predict the key position information of the target key points of the target parts in the corresponding video frame.

[0020] In one embodiment, the registration module is further configured to obtain the preset position information corresponding to each preset feature point in the preset template; determine the mapping relationship between each role object and the preset template according to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point; and based on the mapping relationship corresponding to each role object, adjust the image formed by the role area where the corresponding role object is located to a target image with a preset size.

[0021] In one embodiment, for each role object, when the target part exists in the corresponding target video frame of the role object, the registration module is further configured to directly map the image formed by the role area where the corresponding role object is located to a target image with a preset size based on the mapping relationship corresponding to the role object; when the target part does not exist in the corresponding target video frame of the role object, the registration module is further configured to map the image formed by the role area where the corresponding role object is located to a part of the target image with a preset size based on the mapping relationship corresponding to the role object, and fill the remaining pixels in the target image with zero pixels to obtain a target image with a preset size.

[0022] In one embodiment, the recognition module is further configured to, for the target image with a preset size corresponding to each role object, perform convolution processing on the target image through a feature extraction network to obtain corresponding image features; perform pooling processing on the image features to obtain corresponding pooling features; perform residual processing on the pooling features, and fuse the features obtained by the residual processing with the corresponding image features to obtain target feature vectors corresponding to each role object; and determine the role information corresponding to each role object based on the target feature vectors corresponding to each role object.

[0023] In one embodiment, the recognition module is further configured to, for the target feature vector corresponding to each role object, calculate the feature similarity between the target feature vector and each preset feature vector respectively to obtain at least one feature similarity corresponding to the corresponding role object; for each role object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among the corresponding at least one feature similarity, and use the role information corresponding to the target preset feature vector as the role information corresponding to the corresponding role object.

[0024] In one embodiment, the recognition module is further configured to calculate the Euclidean distance between the target feature vector corresponding to each role object and each preset feature vector respectively, to obtain at least one Euclidean distance corresponding to the corresponding role object; for each role object, determine the minimum Euclidean distance among the corresponding at least one Euclidean distance; when the minimum Euclidean distance is less than the distance threshold, use the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector.

[0025] In one embodiment, the role information includes at least one of the movie name and movie link corresponding to the role object; the apparatus further includes: a push module; the push module is configured to obtain each user account that has browsed the to-be-processed video, and calculate the relevance between each user account and the to-be-processed video respectively; push at least one of the movie name and movie link corresponding to the role object to the user account corresponding to the relevance that meets the relevance condition.

[0026] In one embodiment, the apparatus further includes: a processing module; the processing module is configured to obtain the playback channel of the to-be-processed video; when it is determined based on the role information that the playback channel does not have the playback permission, delete the to-be-processed video under the playback channel, and process the user account that spreads the to-be-processed video.

[0027] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0028] Obtain a to-be-processed video, and extract a target video frame from the to-be-processed video;

[0029] Perform target detection on each of the target video frames respectively, to determine the role areas where the role objects appearing in each of the target video frames are located;

[0030] Determine the key position information of the target key points of each role object in the corresponding target video frame;

[0031] Based on the key position information, perform alignment and registration processing on the images composed of the role areas where each role object is located respectively, to obtain target images with a preset size corresponding to each role object respectively;

[0032] Perform role recognition based on the target images, to obtain the role information corresponding to each role object in the to-be-processed video.

[0033] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0034] Obtain the video to be processed, and extract the target video frames from the video to be processed;

[0035] Perform object detection on each of the target video frames to determine the role regions where the role objects appearing in each of the target video frames are located respectively;

[0036] Determine the key position information of the target key points of each role object in the corresponding target video frame;

[0037] Based on the key position information, perform alignment and registration processing on the images formed by the role regions where each role object is located respectively, to obtain target images of a preset size corresponding to each role object respectively;

[0038] Perform role recognition based on the target images to obtain the role information corresponding to each role object in the video to be processed respectively.

[0039] The above-mentioned role recognition method, device, computer device and storage medium obtain the video to be processed, and extract the target video frames from the video to be processed, and perform object detection on each target video frame respectively to accurately determine the role regions where the role objects appearing in each target video frame are located respectively. Determine the key position information of the target key points of each role object in the corresponding target video frame. Based on the key position information, the images formed by the role regions where the role objects are located can be accurately mapped into target images of a preset size, so as to avoid the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size. The aspect ratios of the role objects in the target images of the preset size are coordinated. Based on the target images, role recognition can be performed to accurately obtain the role information corresponding to each role object in the video to be processed respectively. Description of the Drawings

[0040] Figure 1 It is an application environment diagram of the image recognition method in an embodiment;

[0041] Figure 2 It is a flowchart of the image recognition method in an embodiment;

[0042] Figure 3 It is a flowchart of performing object detection on each target video frame respectively to determine the role regions where the role objects appearing in each target video frame are located respectively in an embodiment;

[0043] Figure 4 It is a flowchart of performing role recognition based on the target images to obtain the role information corresponding to each role object in the video to be processed respectively in another embodiment;

[0044] Figure 5A flowchart showing the steps of calculating the feature similarity between a target feature vector and each preset feature vector in an embodiment;

[0045] Figure 6 A flowchart showing the application of the image recognition method in a video push scenario in an embodiment;

[0046] Figure 7 A flowchart showing the image recognition method in another embodiment;

[0047] Figure 8 A structural block diagram of an image recognition device in an embodiment;

[0048] Figure 9 A structural block diagram of an image recognition device in another embodiment;

[0049] Figure 10 An internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] The present application relates to the field of artificial intelligence (AI) technology. Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. The solutions provided in the embodiments of the present application relate to the role recognition method of artificial intelligence, which will be specifically described through the following embodiments.

[0052] The role recognition method provided by the present application can be applied to a role recognition system as shown in Figure 1 As shown in Figure 1As shown in the figure, the role recognition system includes a terminal 110 and a server 120. In one embodiment, both the terminal 110 and the server 120 can independently execute the role recognition method provided in the embodiments of the present application. The terminal 110 and the server 120 can also be used in cooperation to execute the role recognition method provided in the embodiments of the present application. When the terminal 110 and the server 120 are used in cooperation to execute the role recognition method provided in the embodiments of the present application, the terminal 110 acquires a video to be processed and sends the video to be processed to the server 120. The server 120 extracts target video frames from the video to be processed and performs target detection on each target video frame respectively to determine the role regions where the role objects appearing in each target video frame are located. The server 120 determines the key position information of the target key points of each role object in the corresponding target video frame. Based on the key position information, the server 120 performs alignment and registration processing on the images composed of the role regions where each role object is located respectively to obtain target images of a preset size corresponding to each role object respectively. The server 120 performs role recognition based on the target images to obtain role information corresponding to each role object in the video to be processed, and returns the role information to the terminal 110.

[0053] Among them, the server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services or a cloud server cluster composed of multiple cloud servers. The terminal 110 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV, etc., but is not limited thereto. An application program can be installed on the terminal 110, and the application program can be a communication application, a mail application, a video application, a music application, etc., without limitation. The terminal 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0054] In one embodiment, multiple servers can form a blockchain, and the server is a node on the blockchain.

[0055] In one embodiment, the data related to the role recognition method can be stored on the blockchain. For example, data such as the video to be processed, the target video frames, the target key points and key position information of the role objects, the target images of the preset size, and the role information corresponding to the role objects can be stored on the blockchain.

[0056] In one embodiment, as Figure 2 shown, a role recognition method is provided. Taking the method applied to the Figure 1 computer device (the computer device can specifically be a terminal or a server) as an example for illustration, the method includes the following steps:

[0057] Step S202: Obtain the video to be processed and extract the target video frames from the video to be processed.

[0058] The video to be processed refers to a video that needs to perform character recognition, and may include at least one of movies, TV dramas, programs, and animations. The video to be processed can be a legitimate video, a pirated video, or a mixed video, and can also be a video clip in a legitimate video or a video clip in a pirated video. For example, the video to be processed can be a mixed video obtained by cutting and splicing at least one of movies, TV dramas, programs, and animations.

[0059] Specifically, the computer device can obtain the video to be processed, perform video frame extraction on the video to be processed, and obtain the target video frames. Further, the computer device performs video frame extraction on the video to be processed to obtain a preset number of target video frames. Video frame extraction refers to sampling the video frames in the video to be processed.

[0060] In one embodiment, the computer device performs video frame extraction on the video to be processed at a preset interval duration to obtain each target video frame. For example, if one frame is extracted per second for the video to be processed, a 30-second video can obtain 30 target video frames.

[0061] In one embodiment, the terminal can perform video frame extraction on the video to be processed to obtain each candidate video frame, and select the video frames with character objects from the candidate video frames as the target video frames. For example, among the 30 candidate video frames obtained by video frame extraction, 15 candidate video frames have character objects, then these 15 candidate video frames with character objects are used as the target video frames.

[0062] Step S204: Perform object detection on each target video frame respectively to determine the role regions where the character objects appearing in each target video frame are located.

[0063] Object detection refers to automatically processing the region of interest and selectively ignoring the region of no interest (region of interest, abbreviated as ROI) when facing a scene. The region of interest is called the main region, that is, the region where the main body is located. For example, the character object in the target video frame is the main body, and the role region is the region of interest.

[0064] Specifically, the computer device performs object detection on each target video frame respectively to obtain the role regions where each character object in each target video frame is located.

[0065] In one embodiment, the computer device adjusts each target video frame to a fixed size in equal proportion, and performs object detection on each target video frame of the fixed size respectively to obtain the role regions where each character object in each target video frame is located.

[0066] For a target video frame that cannot be directly scaled to a fixed size in equal proportion, the computer device scales the length of the target video frame to a fixed length in equal proportion and scales the height of the target video frame to the corresponding height to obtain an intermediate video frame. The height in the intermediate target frame is padded with zero pixels to the fixed width, thereby obtaining a target video frame with a fixed length and a fixed height.

[0067] For example, the role regions where the role objects appear in each target video frame are determined through Mask RCNN, and the fixed size of this Mask RCNN is 800 * 800 pixels. The target video frame is scaled to a length of 800 pixels in equal proportion, and the width is padded with zeros to 800 pixels to obtain a target video frame of 800 * 800. For example, if the target video frame is 400 * 200, it is scaled to an intermediate video frame of 800 * 400 in equal proportion, and then the width is padded from 400 pixels to 800 pixels using zero pixels to obtain a target video frame of 800 * 800.

[0068] Step S206, determining the key position information of the target key points of each role object in the corresponding target video frame.

[0069] There may be at least one role object in a video frame. When there is one role object in the target video frame, the key position information of this role object in the target video frame is determined. When there are multiple role objects in the target video frame, the key position information of each role object in the target video frame is determined respectively.

[0070] The key position information may specifically be the coordinate information corresponding to the target key points in the target video frame.

[0071] Specifically, for each target key point of each role object, the computer device determines the key position information corresponding to the target key point in the corresponding target video frame. For example, the computer device determines the coordinate information of each target key point in the corresponding target video frame.

[0072] Step S208, based on the key position information, performing alignment and registration processing on the images formed by the role regions where each role object is located respectively to obtain target images with a preset size corresponding to each role object respectively.

[0073] Among them, the alignment and registration processing refers to the alignment of two or more images in spatial positions to match and superimpose two or more images obtained at different times, by different imaging devices, or under different conditions. The target images may be upper body images, full body images, etc., but are not limited thereto.

[0074] Specifically, for each character object in the target video frame, determine the image formed by the character area where each character object is located in the corresponding target video frame. Based on the key position information of the target key points corresponding to the character object, perform alignment and registration processing on the image formed by the character area where the character object is located, so as to map the image to a target image with a preset size. According to the same processing method, target images with a preset size corresponding to each character object can be obtained.

[0075] Step S210, perform character recognition based on the target image to obtain the character information corresponding to each character object in the video to be processed.

[0076] Among them, the character information includes at least one of, but is not limited to, the character identifier corresponding to the character object, the movie and television name, the movie and television link, and the preset playback channel of the movie and television video corresponding to the character object. The character identifier can specifically be the character name.

[0077] Specifically, the computer device performs character recognition based on the target image with a preset size to obtain the character information corresponding to each character object in each target image.

[0078] For example, when the character information is the character name, the computer device performs character recognition based on the target image with a preset size to obtain the character name of each character object in the corresponding movie and television video in each target image.

[0079] In this embodiment, obtain the video to be processed, extract the target video frames from the video to be processed, and perform object detection on each target video frame respectively to accurately determine the character areas where the character objects appearing in each target video frame are located. Determine the key position information of the target key points of each character object in the corresponding target video frame. Based on the key position information, the image formed by the character area where the character object is located can be accurately mapped to a target image with a preset size, so as to avoid the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size. The aspect ratios of the character objects in the target image with a preset size are coordinated. Based on this target image for character recognition, the character information corresponding to each character object in the video to be processed can be accurately obtained.

[0080] In one embodiment, as Figure 3 shown, performing object detection on each target video frame respectively to determine the character areas where the character objects appearing in each target video frame are located includes:

[0081] Step S302, perform convolution processing on each target video frame respectively to obtain the corresponding video frame features.

[0082] Specifically, the computer device can perform convolution processing on each target video frame respectively to obtain the video frame features corresponding to each target video frame respectively.

[0083] In one embodiment, the computer device may input the target video frame into a trained neural network, such as ResNeXt, to obtain the video frame features corresponding to each target video frame output by ResNeXt.

[0084] Step S304, slide the preset detection box on each target video frame respectively to obtain each candidate box corresponding to each target video frame respectively.

[0085] Wherein, the candidate box is an image area in the target video frame where a role object may exist.

[0086] Specifically, the computer device may use the preset detection box to slide on the target video frame. Each time it slides, a candidate box can be obtained. The size of the candidate box is the same as the size of the preset detection box. When the preset detection box traverses the target video frame, multiple candidate boxes corresponding to the target video frame can be obtained. According to the same processing method, multiple candidate boxes corresponding to each target video frame can be obtained.

[0087] Step S306, based on each candidate box corresponding to each target video frame respectively, determine the role object that appears in each target video frame, and determine the role area where the role object is located in the corresponding target video frame.

[0088] Specifically, the computer device normalizes the pixels in the image area of the candidate box to determine whether the image area in the candidate box is a foreground area or a background area. The foreground area is the area where the role object is located.

[0089] Furthermore, the computer device can determine the probability that the image area in each candidate box is a foreground area. When the probability that the image area in the candidate box is a foreground area is greater than the probability threshold, it is determined that there is a role object in the candidate box, and the image area in the candidate box is used as the role area where the role object is located in the corresponding target video frame.

[0090] In this embodiment, convolution processing is performed on each target video frame respectively to obtain the corresponding video frame features. The preset detection box is slid on each target video frame respectively to obtain each candidate box where a role object may exist. Based on each candidate box corresponding to each target video frame respectively, the role object in the target video frame and the role area where the role object is located in the target video frame can be accurately identified.

[0091] In one embodiment, the computer device can also determine the role areas where the role objects that appear in each target video frame are located respectively through Faster RCNN, CenterNet, etc.

[0092] In one embodiment, the computer device determines the role regions where the role objects appear in each target video frame through Mask RCNN. Specifically, the computer device inputs the target video frame into a pre-trained neural network, such as the ResNeXt network, etc., to obtain the corresponding video frame features. For each key point in the video frame features, a preset number of ROIs are set with the key point as the center, so as to obtain multiple candidate ROIs. The scales of the preset number of ROIs are different. For example, one is a 7*7 ROI and one is a 9*9 ROI, but not limited to this. Each candidate ROI is sent into the RPN (Region Proposal Network) network for binary classification (foreground or background) and bounding-box regression (abbreviated as BB regression) to filter out some candidate ROIs. Then, ROIAlign operations are performed on these remaining ROIs to map and obtain ROIs with a fixed size. Full convolutional operations are performed on each fixed-size ROI to obtain the role region where each role object in the target video frame is located.

[0093] In one embodiment, based on each candidate box corresponding to each target video frame, determining the role objects that appear in each target video frame and determining the role regions where the role objects are located in the corresponding target video frame includes:

[0094] For each target video frame, according to each candidate box corresponding to the corresponding target video frame, determining the role objects that appear in the corresponding target video frame; for each role object, enlarging the size of the candidate box where the corresponding role object is located, and taking the area included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame.

[0095] The computer device normalizes the pixels in the image region of the candidate box to determine whether the image region in the candidate box is a foreground region or a background region. The foreground region is the region where there is a role object. After the computer device determines each candidate box with a role object, it enlarges the size of the candidate box with the role object in the corresponding target video frame, and takes the area included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame. In the same processing manner, the role regions where each role object in the target video frame is located can be obtained.

[0096] In one embodiment, for each role object, enlarging the size of the candidate box where the corresponding role object is located and taking the area included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame includes:

[0097] For each character object, determine the length and width of the candidate box where the character object is located, increase the length and width of the candidate box to a preset length and a preset width respectively, and use the area enclosed by the candidate box with the preset length and width as the character area where the character object is located in the corresponding target video frame.

[0098] In this embodiment, based on each candidate box corresponding to each target video frame, the character objects in the target video frame can be accurately recognized. For each character object, enlarge the size of the candidate box where the corresponding character object is located, and use the area enclosed by the enlarged candidate box as the character area where the character object is located in the corresponding target video frame, so as to accurately determine the area where the complete character object is located and avoid the situation that the detected character object is partially missing due to the candidate box being too small.

[0099] In one embodiment, determining the key position information of the target key points of each character object in the corresponding target video frame includes:

[0100] When there are target parts of the character object in the corresponding target video frame, perform feature point extraction on each target part of each character object to obtain the target key points corresponding to each target part, and determine the key position information of the target key points in the corresponding target video frame; when there are no target parts of the character object in the corresponding target video frame, predict the target parts of the character object in the corresponding video frame, and predict the key position information of the target key points of the target parts in the corresponding video frame.

[0101] Among them, the target parts can be parts such as the head, neck, shoulders, chest, arms of the character object, but are not limited thereto, and can be set as needed.

[0102] Specifically, when the computer device determines that there are target parts of the character object in the corresponding target video frame, for each target part corresponding to the same character object, perform feature point extraction on each target part respectively to obtain the target key points corresponding to each target part. The computer device determines the key position information of each target key point in the corresponding target video frame.

[0103] When the computer device determines that there are no target parts of the character object in the corresponding target video frame, predict the target parts of the character object in the corresponding video frame, and predict the target key points of the target parts. The computer device predicts the key position information of each target key point in the corresponding video frame.

[0104] In one embodiment, the key position information is the coordinates of the target key points in the target video frame. When the target part exists on the corresponding target video frame for the character object, feature points are extracted for each target part of each character object to obtain the target key points corresponding to each target part, and the coordinates of the target key points in the corresponding target video frame are determined; when the target part does not exist on the corresponding target video frame for the character object, the target part of the character object in the corresponding video frame is predicted, and the coordinates corresponding to the target key points of the target part in the corresponding video frame are predicted.

[0105] In one embodiment, when a preset number of target parts exist on the corresponding target video frame for the character object, feature points are extracted for each target part of each character object to obtain the target key points corresponding to each target part, and the key position information of the target key points in the corresponding target video frame is determined; the preset number of target parts are different parts; when the preset number of target parts do not exist on the corresponding target video frame for the character object, the target part of the character object in the corresponding video frame is predicted, and the key position information corresponding to the target key points of the target part in the corresponding video frame is predicted.

[0106] For example, the preset number of target parts includes the head, neck, shoulders, and chest. The key position information is coordinates. When the head, neck, shoulders, and chest exist on the target video frame for the character object, the feature points of the head, neck, shoulders, and chest are extracted respectively to obtain the corresponding target key points, and the coordinates of each target key point in the target video frame are determined. When the head and neck exist on the target video frame for the character object, but the shoulders and chest do not exist, the feature points of the head and neck are extracted respectively to obtain the corresponding target key points, and the coordinates of each target key point in the target video frame are determined. Also, the key point positions of the shoulders and chest in the target video frame are predicted, and the coordinates of each key point position in the target video frame are determined.

[0107] In one embodiment, the SDM algorithm (Supervised Descent Method) is used to extract feature points for each target part of the character object to obtain the target key points corresponding to each target part, and the key position information of the target key points is obtained. For the character object without a target part, each target part, the target key points of each target part, and the key position information of each target key point are predicted.

[0108] In other embodiments, the determination of the key position information of the target key points can also be achieved through the AlphaPose algorithm.

[0109] In this embodiment, for a role object with a target part, feature points are extracted from the target part respectively to obtain target key points corresponding to each target part, so as to accurately determine the key position information of the target key points in the corresponding target video frame. For a role object without a target part, the target part of the role object, the target key points of the target part, and the key position information of the target key points are predicted, so that each target part of the role object can be supplemented for subsequent registration processing of the role object.

[0110] In one embodiment, based on the key position information, image registration processing is respectively performed on the images formed by the role areas where each role object is located to obtain target images with a preset size corresponding to each role object, including:

[0111] Obtain the preset position information corresponding to each preset feature point in the preset template; determine the mapping relationship between each role object and the preset template according to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point; based on the mapping relationship corresponding to each role object, adjust the image formed by the role area where the corresponding role object is located to a target image with a preset size.

[0112] Specifically, the computer device can preset a preset template, and the preset template has a preset size. For example, the preset size is 128*256. The preset template may include preset feature points and the preset position information corresponding to each preset feature point. The computer device performs matching processing on the target key points of the role object and each preset feature point, and calculates the mapping relationship between the role object and the preset template according to the key position information of the matched target key points and the preset position information of the matched preset feature points. Based on the mapping relationship between the role object and the preset template, the computer device maps each pixel in the image formed by the role area where the role object is located to the same spatial layout as the preset template, so as to obtain a target image with a preset size. For each role object, the same processing is performed to obtain the mapping relationship between each role object and the preset template, and then according to the mapping relationship corresponding to each role object, the image formed by the role area where the corresponding role object is located can be mapped to a target image with a preset size.

[0113] In one embodiment, each preset feature point in the preset template corresponds to its own key position information and has no corresponding pixels. In other embodiments, each preset feature point may correspond to pixels, and each pixel is 0.

[0114] In one embodiment, for the image formed by the role area where the role object is located, the matching process of each target key point and each preset feature point of the role object includes: for the image formed by the role area where the role object is located, the matching process of each target key point corresponding to the target part of the role object and each preset feature point corresponding to the target part in the preset template is performed. For example, the key points of the head of the role object are matched with the head feature points in the preset template, and the key points of the neck of the role object are matched with the neck feature points in the preset template.

[0115] In one embodiment, the preset template can be a preset image of a preset size, and the preset image contains a preset object. The preset feature points are the key feature points extracted from the target part of the preset object. The computer device can match the target key points corresponding to each target part of the role object with the preset feature points corresponding to the corresponding target part of the preset object, and calculate the mapping relationship between the role object and the preset object according to the coordinate information of the matched target key points and the coordinate information of the preset feature points. Based on the mapping relationship, each pixel in the image formed by the role area where the role object is located is mapped to the same image space as the preset image to obtain a target image of the preset size.

[0116] In this embodiment, by obtaining the preset position information corresponding to each preset feature point in the preset template, and according to the key position information corresponding to the target key point of each role object and the preset position information of each preset feature point, the mapping relationship between each role object and the preset template can be accurately determined. Based on the mapping relationship corresponding to each role object, the image formed by the role area where the corresponding role object is located is accurately adjusted to a target image of the preset size, so that each role object can be accurately mapped to a target image of the preset size, thereby avoiding the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size.

[0117] In one embodiment, based on the mapping relationship corresponding to each role object, adjusting the image formed by the role area where the corresponding role object is located to a target image of the preset size includes:

[0118] For each role object, when there is a target part of the role object in the corresponding target video frame, based on the mapping relationship corresponding to the role object, the image formed by the role area where the corresponding role object is located is directly mapped to a target image of the preset size; when there is no target part of the role object in the corresponding target video frame, based on the mapping relationship corresponding to the role object, the image formed by the role area where the corresponding role object is located is mapped to a part of the target image of the preset size, and the remaining pixels in the target image are filled with zero pixels to obtain a target image of the preset size.

[0119] Specifically, after determining the mapping relationship between each character object and the preset template, it is determined whether there is a target part in the corresponding target video frame for each character object. When it is determined that there is a target part in the corresponding target video frame for a character object, based on the mapping relationship corresponding to the character object, the image formed by the character area where the character object is located is directly mapped to a target image of a preset size. When there is no target part in the corresponding target video frame for the character object, based on the mapping relationship corresponding to the character object, the image formed by the character area where the corresponding character object is located is mapped to a part of the target image of a preset size, and the remaining pixels in the target image are filled with zero pixels to obtain a target image of a preset size. A zero pixel refers to a pixel with a pixel value of zero.

[0120] In one embodiment, the preset size can be 128*256. When there are preset feature points corresponding to the head, neck, shoulders, and chest of the preset object in the preset template, and the character object has a head, neck, shoulders, and chest in the target video frame, based on the mapping relationship between the character object and the preset object, the pixels of the head, neck, shoulders, and chest of the character object are directly mapped to the pixels of the head, neck, shoulders, and chest in the target image to obtain a target image of 128*256. When the character object has a head and neck in the target video frame but no shoulders and chest, based on the mapping relationship between the character object and the preset object, the pixels of the head and neck of the character object are mapped to the pixels of the head and neck in the 128*256 target image, and the remaining pixels in the 128*256 target image are filled with zero pixels, that is, the pixels of parts such as the shoulders and chest in the 128*256 target image are filled with 0 to obtain a 128*256 target image.

[0121] In this embodiment, for a character object with a target part, based on the mapping relationship corresponding to the character object, the image formed by the character area where the corresponding character object is located is directly mapped to a target image of a preset size, so that the character object can be directly mapped to a target image of a fixed size, and the aspect ratio of the character object in the obtained target image of the preset size can be kept coordinated. For a character object without a target part, when there is no target part in the corresponding target video frame for the character object, based on the mapping relationship corresponding to the character object, the image formed by the character area where the corresponding character object is located is mapped to a part of the target image of a preset size, and the remaining pixels in the target image are filled with zero pixels, so that the aspect ratio of the character object in the target image of the preset size is kept coordinated, ensuring the clarity of the character object in the target image of the preset size, and improving the accuracy of identifying the character object.

[0122] In one embodiment, as Figure 4, perform role recognition based on the target image to obtain the role information corresponding to each role object in the video to be processed, including:

[0123] Step S402, for the target image with a preset size corresponding to each role object, perform convolutional processing on the target image through a feature extraction network to obtain the corresponding image features.

[0124] Specifically, the computer device inputs the target image with a preset size corresponding to each role object into the feature extraction network. The feature extraction network respectively performs convolutional processing on each target image with a preset size to obtain the image features corresponding to each target image.

[0125] In one embodiment, the feature extraction network can be a classification network with Resnet50 as the backbone, or it can also be Vgg, Resnet101, but not limited to this. The classification network with Resnet50 as the backbone uses the softmax cross-entropy loss function and the triplet loss function as the target loss function to train the feature extraction network.

[0126] Step S404, perform pooling processing on the image features to obtain the corresponding pooling features.

[0127] Specifically, for the image features corresponding to each target image, the feature extraction network respectively performs pooling processing on the image features corresponding to the same target image to obtain the pooling features corresponding to the target image. Further, a target image can be convolved to obtain multiple image features. The feature extraction network performs pooling processing on multiple image features corresponding to a target image to obtain at least one pooling feature corresponding to the target image.

[0128] In one embodiment, the feature extraction network performs average pooling processing on the image features of the target image to obtain the pooling features corresponding to the target image.

[0129] Step S406, perform residual processing on the pooling features, and fuse the features obtained by the residual processing with the corresponding image features to obtain the target feature vectors corresponding to each role object.

[0130] Specifically, the feature extraction network performs residual processing on the pooling features corresponding to the same target image to obtain the features after residual processing. Fuse the features obtained by the residual processing with the image features corresponding to the same target image to obtain the target feature vector corresponding to the same target image. The target feature vector is the target feature vector corresponding to the role object in the target image. According to the same processing method, the target feature vectors corresponding to each role object can be obtained.

[0131] Step S408: Determine the role information corresponding to each role object based on the target feature vectors corresponding to each role object.

[0132] Specifically, the computer device determines the role information corresponding to each role object based on the feature similarity between the target feature vector corresponding to each role object and the preset feature vector.

[0133] In this embodiment, the input image of the feature extraction network is of a preset size. Before the role object is input into the feature extraction network, the image formed by the role area where the role object is located in the corresponding video frame is mapped into a target image of the preset size, so that the size of the target image containing the role object meets the size of the input image of the feature extraction network, avoiding the problem that after inputting an image that does not meet the size requirements, the feature recognition network directly stretches the length and width of the image to the preset size, resulting in distortion of the aspect ratio of the image, and thus resulting in a relatively poor recognition effect of the feature extraction network.

[0134] In one embodiment, determining the role information corresponding to each role object based on the target feature vectors corresponding to each role object includes:

[0135] For the target feature vector corresponding to each role object, calculate the feature similarity between the target feature vector and each preset feature vector respectively, to obtain at least one feature similarity corresponding to the corresponding role object; for each role object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among the at least one corresponding feature similarity, and use the role information corresponding to the target preset feature vector as the role information corresponding to the corresponding role object.

[0136] Among them, the similarity condition can be that there is a maximum value among multiple feature similarities, or there is a feature similarity greater than the similarity threshold.

[0137] Specifically, the computer device obtains the preset feature vector, which is pre-extracted from each role object, and the preset feature vectors of each role object are associated and stored with the corresponding role information. For the target feature vector corresponding to each role object extracted by the feature extraction network, the computer device calculates the feature similarity between the target feature vector of the role object and each preset feature vector respectively, to obtain at least one feature similarity corresponding to the role object. The computer device can determine whether there is a feature similarity that meets the similarity condition among the at least one feature similarity corresponding to the role object. If so, use the role information corresponding to the target preset feature vector corresponding to the feature similarity that meets the similarity condition as the role information corresponding to the role object.

[0138] According to the same processing method, the feature similarity corresponding to each role object can be obtained, so as to obtain the role information corresponding to each role object.

[0139] In one embodiment, when there is no feature similarity that meets the similarity condition among at least one feature similarity corresponding to the role object, it is determined that the recognition fails, that is, the role information of the role object cannot be recognized.

[0140] In this embodiment, for the target feature vector corresponding to each role object, the feature similarity between the target feature vector and each preset feature vector is calculated respectively, and at least one feature similarity corresponding to the corresponding role object is obtained, so that the role information corresponding to each role object can be accurately recognized based on the feature similarity that meets the similarity condition.

[0141] In one embodiment, as Figure 5 shown, for the target feature vector corresponding to each role object, the feature similarity between the target feature vector and each preset feature vector is calculated respectively, and at least one feature similarity corresponding to the corresponding role object is obtained, including step S502.

[0142] Step S502, for the target feature vector corresponding to each role object, calculate the Euclidean distance between the target feature vector and each preset feature vector respectively, and obtain at least one Euclidean distance corresponding to the corresponding role object.

[0143] Specifically, the feature similarity can be characterized by the Euclidean distance. For the target feature vector corresponding to each role object extracted by the feature extraction network, the computer device calculates the Euclidean distance between the target feature vector of the role object and each preset feature vector respectively, and obtains at least one Euclidean distance corresponding to the role object.

[0144] For each role object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among the corresponding at least one feature similarity, including step S504 and step S506.

[0145] Among them, step S504, for each role object, determine the minimum Euclidean distance among the corresponding at least one Euclidean distance.

[0146] Step S506, when the minimum Euclidean distance is less than the distance threshold, use the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector.

[0147] Specifically, the computer device determines the minimum Euclidean distance among at least one Euclidean distance corresponding to the role object and obtains a distance threshold. When the minimum Euclidean distance is less than the distance threshold, the preset feature vector corresponding to the minimum Euclidean distance is used as the target preset feature vector, and thus the role information corresponding to the target preset feature vector is used as the role information corresponding to the role object.

[0148] In the same processing manner, the Euclidean distances corresponding to each role object can be obtained, and thus the target preset feature vectors corresponding to each role object can be obtained.

[0149] In this embodiment, the Euclidean distance between the target feature vector of the role object and each preset feature vector is calculated. When the minimum Euclidean distance is less than the distance threshold, the minimum Euclidean distance is used as the target preset feature vector corresponding to the role object, so that the role information corresponding to each role object can be accurately identified according to the Euclidean distance.

[0150] In one embodiment, the feature similarity can be characterized by the cosine similarity. For the target feature vector corresponding to each role object, the feature similarity between the target feature vector and each preset feature vector is calculated respectively, and at least one feature similarity corresponding to the corresponding role object is obtained, including: for the target feature vector corresponding to each role object, the cosine similarity between the target feature vector and each preset feature vector is calculated respectively, and at least one cosine similarity corresponding to the corresponding role object is obtained;

[0151] For each role object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among at least one corresponding feature similarity, including: for each role object, determine the maximum cosine similarity among at least one corresponding cosine similarity; when the maximum cosine similarity is greater than the similarity threshold, the preset feature vector corresponding to the maximum cosine similarity is used as the target preset feature vector.

[0152] In one embodiment, the role information includes at least one of the movie name and movie link corresponding to the role object; the method further includes:

[0153] Obtain each user account that has browsed the video to be processed, and calculate the relevance between each user account and the video to be processed respectively; push at least one of the movie name and movie link corresponding to the role object to the user account corresponding to the relevance that meets the relevance condition.

[0154] Specifically, the character information includes at least one of the movie / TV name and the movie / TV link corresponding to the character object. An application program is installed on the terminal, and the user can log in to the application program through the user account, so as to browse the video to be processed in the application program. The computer device obtains each user account that has browsed the video to be processed, calculates the relevance between each user account and the video to be processed respectively, and filters out the user accounts that meet the relevance condition from each user account. At least one of the movie / TV name and the movie / TV link corresponding to the character object is pushed to the user account corresponding to the relevance that meets the relevance condition.

[0155] In one embodiment, calculating the relevance between the user account and the video to be processed includes: obtaining the interactive comments, browsing time, and browsing times of the user account for the video to be processed, and calculating the relevance between the user account and the video to be processed according to the interactive comments, browsing time, and browsing times.

[0156] In this embodiment, the higher the relevance between the user account and the video to be processed, the more interested the user is in the video to be processed or a certain part of the content in the video to be processed. Then, at least one of the movie / TV name and the movie / TV link corresponding to the character object is pushed to the interested users, so as to achieve accurate video recommendation.

[0157] In one embodiment, as Figure 6 shown, a character recognition method is provided, which is applied to a computer device and specifically applied to a video push scenario, including:

[0158] Step S602, video frame extraction.

[0159] The computer device obtains the mixed video browsed by the user account A, and the mixed video is obtained by splicing the character segments played by the same actor in different TV dramas, movies, and variety shows. The computer device extracts video frames from the mixed video every 2 seconds to obtain a plurality of target video frames.

[0160] Step S604, human body detection and interception.

[0161] The Mask RCNN network is trained using human detection sample data, and its input is an image with a fixed size of 800 * 800 pixels. The computer device scales each target video frame proportionally so that the longest side is 800 pixels and the shorter side is padded with zeros to 800 pixels, obtaining each target video frame of 800 * 800. Each 800 * 800 target video frame is input into the Mask RCNN network, and the Mask RCNN network predicts the detection boxes of the human body rectangular coordinates of possible role objects in each target video frame. Image cropping is performed according to the detection boxes to obtain the final human body image containing the role object. The human body image can be a face image, an image containing the shoulders and above, an upper body image, or a full body image. All human body images are recorded as dataset I.

[0162] Step S606, head and shoulder key point localization.

[0163] The SDM algorithm (Supervised Descent Method) is used to localize the head and shoulder key points in each human body image in dataset I, and the coordinate information of the head, neck, and shoulder key points is located. For human body images with only a head, the coordinate information of the head key point is determined, and the coordinate information of the neck and shoulder key points is predicted. For human body images with a head and a neck, the coordinate information of the head key point and the neck key point is determined, and the coordinate information of the shoulder key point is predicted. Each human body image corresponds to a set of head and shoulder key point coordinate information, and all head and shoulder key point information is recorded as L.

[0164] Step S608, alignment and registration of human body images.

[0165] Using the head and shoulder key point information L in step S606, the human body image I in step S604 is subjected to alignment and registration processing to obtain a new aligned target human body image. Specifically, it is judged whether the human body image I is a full body image through the head and shoulder key point information L. For a half body image, it can be padded with black edges to obtain a target human body image with a height of 256 and a width of 128. The target human body image can be a full body image. All target human body images are recorded as M.

[0166] Step S610, feature extraction of human body images.

[0167] The feature extraction network is a classification network with Resnet50 as the backbone, trained using personal data, and its input is of a fixed size of 128*256. The computer device inputs the target personal image M obtained after alignment and registration in step S608 into the feature extraction network, and can obtain a 512-dimensional feature vector embedding. Each target personal image corresponds to a feature vector embedding, and this vector can be understood as the feature attribute of the corresponding movie or TV drama character of the target personal image. All the embeddings obtained from all target personal images through the feature extraction network are denoted as E.

[0168] Step S612, search for character names.

[0169] For the movie or TV drama character names that need to be recognized in the actual business line, under the precondition of manual participation, collect the movie or TV drama video data corresponding to the character names, and pre-obtain the feature vector embedding of each character object through the corresponding processing of the above steps S602, S604, S608, and S610. Each character name can correspond to one or more feature vector embeddings. All the feature vectors of all character names are denoted as E2 and are called the registration library. The number of special vector embeddings in the registration library E2 is called the size of the registration library, denoted as n.

[0170] After the above-mentioned mixed video passes through step S610, a feature vector set E can be obtained. For a certain feature vector e in E, the Euclidean distance can be calculated with each embedding in the entire feature vector set E2 in the registration library, that is, n distances can be obtained, which are d1, d2,..., dn respectively. Calculate the minimum distance among the n distances, denoted as dk (1≤k≤n); when dk is less than the threshold t, it is considered that the feature vector e corresponds to the kth character name, that is, the character object corresponding to the feature vector e is the kth character. When dk is greater than or equal to the threshold t, it is considered that the feature vector e is an invalid recognition.

[0171] Step S614, video push.

[0172] After obtaining the character name of each character object, the corresponding movie or TV drama name can be determined according to the character name, and the movie or TV drama name or movie or TV drama link is pushed to the user account A. Other videos corresponding to the character name can also be pushed to the user account A to achieve precise video push.

[0173] In this embodiment, frames are extracted from video data. For each frame, the coordinates of the smallest rectangular box containing the human body target are obtained using a human body detection algorithm. Within this region, the head and shoulder key point algorithm is used to predict and output the coordinates of the head and shoulder key points. Then, using the head and shoulder key point coordinates, the human body image is aligned and registered to a unified scale. For human body images that are not full-body in the frame (such as headshot images, head and shoulder images, and half-body images), after compensation by filling black borders, a fixed-size human body image can be obtained. Taking the fixed-size human body image as the input image, referring to the pedestrian re-identification (ReID) technology, with Resnet50 as the model backbone network, character Embedding features are extracted, enabling the rapid identification of character objects in a short video or movie purely from the perspective of the image without referring to information such as video titles, subtitles, and voiceovers. Moreover, in this embodiment, the human body image obtained through the registration process can avoid the problem of image aspect ratio distortion caused by directly using non-full-body images in movies and TV shows as inputs, ensuring the proportional coordination of the character objects in the image, thereby making the extracted Embedding more discriminative for identification and achieving a more robust movie and TV character recognition effect.

[0174] Furthermore, in the embodiment, using a half-body image or a full-body image as the input to the feature extraction network, the same person wearing different costumes or different people wearing similar costumes can be regarded as different characters, and the corresponding feature vectors are pre-stored, which can effectively avoid the problem that the face recognition cannot determine the character name when the same actor appears in multiple movies and TV shows. The pixel ratio of the half-body image or the full-body image is significantly higher than that of the human face, which can also solve the problem of low recall due to blurred human faces to a certain extent.

[0175] In one embodiment, the method further includes: obtaining the playback channel of the video to be processed; when it is determined based on the character information that the playback channel does not have the playback permission, deleting the video to be processed under the playback channel and processing the user account that spreads the video to be processed.

[0176] Specifically, the role information includes the preset playback channels of the video corresponding to the role object, and the preset playback channels have the playback rights of the video. After the computer device determines the preset playback channels of the video corresponding to the role object in the video to be processed, it obtains the playback channels of the video to be processed, compares the playback channels of the video to be processed with the preset playback channels, and determines whether the playback channels of the video to be processed belong to the preset playback channels. When the playback channels of the video to be processed belong to the preset playback channels, it is determined that the playback channels have the playback rights. When the playback channels of the video to be processed do not belong to the preset playback channels, it is determined that the playback channels do not have the playback rights, which means that the playback channels play the video to be processed without permission. Then, the computer device deletes the video to be processed under the playback channels, obtains the user account that spreads the video to be processed, and reports the user account that spreads the video to be processed, or gives a prompt to the user account that spreads the video to be processed.

[0177] In this embodiment, when obtaining the playback channels of the video to be processed and determining that the playback channels do not have the playback rights based on the role information, it means that the video to be processed is an illegally recorded video or the playback channels are irregular playback channels. The video to be processed under the playback channels is deleted, and the user account that spreads the video to be processed is processed, so as to avoid the large-scale spread of illegally recorded videos, thereby effectively protecting the copyright of the video and the playback rights of the video.

[0178] In one embodiment, as Figure 7 shown, a role recognition method is provided, which is applied to a computer device and includes:

[0179] Step S702, the computer device obtains the video to be processed and extracts the target video frames from the video to be processed.

[0180] Step S704, the computer device performs convolution processing on each target video frame respectively to obtain the corresponding video frame features; slides the preset detection frames on each target video frame respectively to obtain the candidate frames corresponding to each target video frame respectively; for each target video frame, determines the role objects that appear in the corresponding target video frame according to the candidate frames corresponding to the corresponding target video frame respectively.

[0181] Step S706, for each role object, the computer device enlarges the size of the candidate frame where the corresponding role object is located, and uses the area included in the enlarged candidate frame as the role area where the role object is located in the corresponding target video frame.

[0182] Step S708, when there are target parts in the corresponding target video frame for the role object, extracts the feature points of each target part of each role object respectively to obtain the target key points corresponding to each target part respectively, and determines the key position information of the target key points in the corresponding target video frame.

[0183] Step S710: When the target part does not exist in the corresponding target video frame of the role object, predict the target part of the role object in the corresponding video frame, and predict the target key point information of the target part at the key position corresponding to the corresponding video frame.

[0184] Step S712: The computer device obtains the preset position information corresponding to each preset feature point in the preset template; according to the key position information corresponding to the target key point of each role object and the preset position information of each preset feature point, determine the mapping relationship between each role object and the preset template.

[0185] Step S714: For each role object, when the target part exists in the corresponding target video frame of the role object, based on the mapping relationship corresponding to the role object, directly map the image formed by the role area where the role object is located into a target image of a preset size.

[0186] Step S716: When the target part does not exist in the corresponding target video frame of the role object, based on the mapping relationship corresponding to the role object, map the image formed by the role area where the role object is located into a part of the target image of a preset size, and fill the remaining pixels in the target image with zero pixels to obtain a target image of a preset size.

[0187] Step S718: For the target image of a preset size corresponding to each role object respectively, perform convolution processing on the target image through a feature extraction network to obtain the corresponding image features; perform pooling processing on the image features to obtain the corresponding pooling features; perform residual processing on the pooling features, and fuse the features obtained by the residual processing with the corresponding image features to obtain the target feature vectors corresponding to each role object respectively.

[0188] Step S720: For the target feature vector corresponding to each role object, calculate the Euclidean distance between the target feature vector and each preset feature vector respectively to obtain at least one Euclidean distance corresponding to the corresponding role object; for each role object, determine the minimum Euclidean distance among the corresponding at least one Euclidean distance.

[0189] Step S722: When the minimum Euclidean distance is less than the distance threshold, the computer device uses the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector, and uses the role information corresponding to the target preset feature vector as the role information corresponding to the corresponding role object.

[0190] Step S724, the role information includes at least one of the movie name and movie link corresponding to the role object; obtain each user account that has viewed the video to be processed, and calculate the relevance between each user account and the video to be processed respectively. The computer device pushes at least one of the movie name and movie link corresponding to the role object to the user accounts corresponding to the relevance that meet the relevance conditions.

[0191] Step S726, the computer device obtains the playback channel of the video to be processed; when it is determined based on the role information that the playback channel does not have the playback permission, delete the video to be processed under the playback channel, and process the user accounts that spread the video to be processed.

[0192] In this embodiment, obtain the video to be processed, extract the target video frames from the video to be processed, perform convolution processing on each target video frame respectively to obtain the corresponding video frame features, and slide the preset detection frames on each target video frame respectively to obtain each candidate frame where a role object may exist. Based on each candidate frame corresponding to each target video frame respectively, the role object in the target video frame and the role area where the role object is located in the target video frame can be accurately identified.

[0193] According to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point, the mapping relationship between each role object and the preset template can be accurately determined. Based on the mapping relationship corresponding to each role object, the image formed by the role area where the corresponding role object is located is accurately adjusted to the target image of the preset size, so that each role object can be accurately mapped to the target image of the preset size, avoiding the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size. The aspect ratio of the role objects in the target image of the preset size is coordinated, and based on this target image for role recognition, the role information corresponding to each role object in the video to be processed can be accurately obtained.

[0194] The higher the relevance between the user account and the video to be processed, the more interested the user is in the video to be processed or a certain part of the content in the video to be processed. Then, at least one of the movie name and movie link corresponding to the role object is pushed to the interested users, so as to achieve accurate recommendation of the video.

[0195] Moreover, when it is determined based on the role information that the playback channel does not have the playback permission, it means that the video to be processed is a pirated video or the playback channel is an informal playback channel. Delete the video to be processed under the playback channel, and process the user accounts that spread the video to be processed, avoiding the large-scale spread of pirated videos, so as to effectively protect the copyright of the video and the playback permission of the video.

[0196] It should be understood that although Figure 2-7The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-7 At least a part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0197] In one embodiment, as Figure 8 shown, a role recognition device 800 is provided. This device can be a software module, a hardware module, or a combination of both to be part of a computer device. Specifically, the device includes: an acquisition module 802, a detection module 804, a determination module 806, a registration module 808, and an identification module 810, where:

[0198] The acquisition module 802 is configured to acquire a video to be processed and extract a target video frame from the video to be processed.

[0199] The detection module 804 is configured to perform object detection on each target video frame respectively to determine the role regions where the role objects appearing in each target video frame are located.

[0200] The determination module 806 is configured to determine the key position information of the target key points of each role object in the corresponding target video frame.

[0201] The registration module 808 is configured to perform alignment and registration processing on the images formed by the role regions where each role object is located respectively based on the key position information to obtain target images with a preset size corresponding to each role object respectively.

[0202] The identification module 810 is configured to perform role recognition based on the target images to obtain role information corresponding to each role object in the video to be processed.

[0203] In this embodiment, a video to be processed is obtained, and target video frames are extracted from the video to be processed. Object detection is performed on each target video frame respectively to accurately determine the role regions where the role objects appearing in each target video frame are located. The key position information of the target key points of each role object in the corresponding target video frame is determined. Based on the key position information, the image formed by the role regions where the role objects are located can be accurately mapped to a target image with a preset size, so as to avoid the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size. The aspect ratios of the role objects in the target image with the preset size are coordinated. Based on this target image, role recognition can be performed to accurately obtain the role information corresponding to each role object in the video to be processed.

[0204] In one embodiment, the detection module 804 is further configured to perform convolution processing on each target video frame respectively to obtain corresponding video frame features; slide a preset detection box on each target video frame respectively to obtain each candidate box corresponding to each target video frame respectively; based on each candidate box corresponding to each target video frame respectively, determine the role objects appearing in each target video frame, and determine the role regions where the role objects are located in the corresponding target video frames.

[0205] In this embodiment, convolution processing is performed on each target video frame respectively to obtain corresponding video frame features. A preset detection box is slid on each target video frame respectively to obtain each candidate box where role objects may exist. Based on each candidate box corresponding to each target video frame respectively, the role objects in the target video frame and the role regions where the role objects are located in the target video frame can be accurately recognized.

[0206] In one embodiment, the detection module 804 is further configured to, for each target video frame, determine the role objects appearing in the corresponding target video frame according to each candidate box corresponding to the corresponding target video frame respectively; for each role object, enlarge the size of the candidate box where the corresponding role object is located, and use the region included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame.

[0207] In this embodiment, based on each candidate box corresponding to each target video frame respectively, the role objects in the target video frame can be accurately recognized. For each role object, enlarge the size of the candidate box where the corresponding role object is located, and use the region included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame, so as to accurately determine the region where the complete role object is located and avoid the situation that the detected role object is partially missing due to the candidate box being too small.

[0208] In one embodiment, the determination module 806 is further configured to, when a target part exists in a corresponding target video frame for a role object, extract feature points from each target part of each role object respectively to obtain target key points corresponding to each target part, and determine the key position information of the target key points in the corresponding target video frame; when a target part does not exist in a corresponding target video frame for a role object, predict the target part of the role object in the corresponding video frame, and predict the key position information of the target key points of the target part in the corresponding video frame.

[0209] In this embodiment, for a role object with a target part, feature points are extracted from the target part respectively to obtain target key points corresponding to each target part, so as to accurately determine the key position information of the target key points in the corresponding target video frame. For a role object without a target part, the target part of the role object, the target key points of the target part, and the key position information of the target key points are predicted, so that each target part of the role object can be supplemented for subsequent registration processing of the role object.

[0210] In one embodiment, the registration module 808 is further configured to obtain the preset position information corresponding to each preset feature point in a preset template; determine the mapping relationship between each role object and the preset template according to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point; and based on the mapping relationship corresponding to each role object, adjust the image formed by the role area where the corresponding role object is located to a target image with a preset size.

[0211] In this embodiment, by obtaining the preset position information corresponding to each preset feature point in the preset template and according to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point, the mapping relationship between each role object and the preset template can be accurately determined. Based on the mapping relationship corresponding to each role object, the image formed by the role area where the corresponding role object is located is accurately adjusted to a target image with a preset size, so that each role object can be accurately mapped to a target image with a preset size, avoiding the problem of aspect ratio distortion caused by directly adjusting the size of the image to a fixed size.

[0212] In one embodiment, the registration module 808 is further configured to, for each role object, when the target part exists in the corresponding target video frame of the role object, directly map the image formed by the role area where the corresponding role object is located to a target image of a preset size based on the mapping relationship corresponding to the role object; when the target part does not exist in the corresponding target video frame of the role object, map the image formed by the role area where the corresponding role object is located to a part of the target image of a preset size based on the mapping relationship corresponding to the role object, and fill the remaining pixels in the target image with zero pixels to obtain a target image of a preset size.

[0213] In this embodiment, for the role object with the target part, based on the mapping relationship corresponding to the role object, the image formed by the role area where the corresponding role object is located is directly mapped to a target image of a preset size, so that the role object can be directly mapped to a target image of a fixed size, and the aspect ratio of the role object in the obtained target image of the preset size can be kept coordinated. For the role object without the target part, when the target part does not exist in the corresponding target video frame of the role object, based on the mapping relationship corresponding to the role object, the image formed by the role area where the corresponding role object is located is mapped to a part of the target image of a preset size, and the remaining pixels in the target image are filled with zero pixels, so that the aspect ratio of the role object in the target image of the preset size is kept coordinated, ensuring the clarity of the role object in the target image of the preset size, and improving the accuracy of role object recognition.

[0214] In one embodiment, the recognition module 810 is further configured to, for the target image of a preset size corresponding to each role object respectively, perform convolution processing on the target image through a feature extraction network to obtain corresponding image features; perform pooling processing on the image features to obtain corresponding pooling features; perform residual processing on the pooling features, and fuse the features obtained by the residual processing with the corresponding image features to obtain a target feature vector corresponding to each role object respectively; determine the role information corresponding to each role object respectively based on the target feature vector corresponding to each role object respectively.

[0215] In this embodiment, since the input image of the feature extraction network is of a preset size, before the role object is input into the feature extraction network, the image formed by the role area where the role object is located in the corresponding video frame is mapped to a target image of a preset size, so that the size of the target image containing the role object meets the size of the input image of the feature extraction network, avoiding the problem that after an image that does not meet the size requirement is input, the feature recognition network directly stretches the length and width of the image to the preset size, resulting in distortion of the aspect ratio of the image and thus a relatively poor recognition effect of the feature extraction network.

[0216] In one embodiment, the recognition module 810 is further configured to calculate the feature similarity between the target feature vector and each preset feature vector for each target feature vector corresponding to a role object, so as to obtain at least one feature similarity corresponding to the corresponding role object; for each role object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among the corresponding at least one feature similarity, and use the role information corresponding to the target preset feature vector as the role information corresponding to the corresponding role object.

[0217] In this embodiment, by calculating the feature similarity between the target feature vector and each preset feature vector for each target feature vector corresponding to a role object, at least one feature similarity corresponding to the corresponding role object is obtained, so that the role information corresponding to each role object can be accurately identified based on the feature similarity that meets the similarity condition.

[0218] In one embodiment, the recognition module 810 is further configured to calculate the Euclidean distance between the target feature vector and each preset feature vector for each target feature vector corresponding to a role object, so as to obtain at least one Euclidean distance corresponding to the corresponding role object; for each role object, determine the minimum Euclidean distance among the corresponding at least one Euclidean distance; when the minimum Euclidean distance is less than the distance threshold, use the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector.

[0219] In this embodiment, by calculating the Euclidean distance between the target feature vector of the role object and each preset feature vector, and using the minimum Euclidean distance as the target preset feature vector corresponding to the role object when the minimum Euclidean distance is less than the distance threshold, the role information corresponding to each role object can be accurately identified according to the Euclidean distance.

[0220] In one embodiment, as Figure 9 shown, a role recognition device 900 is provided, which specifically includes: an acquisition module 902, a detection module 904, a determination module 906, a registration module 908, a recognition module 910, and a push module 912. Among them, for the descriptions of the acquisition module 902, the detection module 904, the determination module 906, the registration module 908, and the recognition module 910, refer to Figure 8 the descriptions of the acquisition module 802, the detection module 804, the determination module 806, the registration module 808, and the recognition module 810 in

[0221] The push module 912 is configured to obtain each user account that has browsed the video to be processed, and calculate the relevance between each user account and the video to be processed respectively; push at least one of the movie name and movie link corresponding to the role object to the user account corresponding to the relevance that meets the relevance condition.

[0222] In this embodiment, the higher the relevance between the user account and the video to be processed, the more interested the user is in the video to be processed or a certain part of the content in the video to be processed. Then, at least one of the movie name and movie link corresponding to the role object is pushed to the interested user, so as to achieve accurate video recommendation.

[0223] In one embodiment, the apparatus 900 further includes: a processing module 914; the processing module 914 is configured to obtain the playback channel of the video to be processed; when it is determined based on the role information that the playback channel does not have the playback permission, delete the video to be processed under the playback channel, and process the user account that spreads the video to be processed.

[0224] In this embodiment, when obtaining the playback channel of the video to be processed and determining based on the role information that the playback channel does not have the playback permission, it means that the video to be processed is an illegally recorded video or the playback channel is an informal playback channel. Delete the video to be processed under the playback channel and process the user account that spreads the video to be processed, so as to avoid the large-scale spread of illegally recorded videos, and thus effectively protect the copyright of the video and the playback permission of the video.

[0225] For the specific limitations of the role recognition device, reference can be made to the limitations of the role recognition method in the above text, which will not be elaborated here. Each module in the above role recognition device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0226] In one embodiment, a computer device is provided, and the computer device can be a terminal or a server. In this embodiment, the computer device is taken as an example of a terminal for illustration, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, operator network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes a role recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0227] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0228] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are realized.

[0229] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are realized.

[0230] In one embodiment, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0231] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0232] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0233] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A role recognition method, characterized in that, The method includes: Obtain a video to be processed, and extract target video frames from the video to be processed; the video to be processed is a mixed video, and the mixed video includes at least two video segments of the same actor, and the same actor plays different role objects in the at least two video segments; Perform object detection on each of the target video frames respectively to determine the role regions where the role objects appearing in each of the target video frames are located; Determine the key position information of the target key points of each role object in the corresponding target video frame; Based on the key position information, perform alignment and registration processing on the images composed of the role regions where each role object is located respectively to obtain target images with a preset size corresponding to each role object respectively; Perform role recognition based on the target images to obtain the role information corresponding to each role object in the video to be processed; the role information includes the film and television video to which the role object belongs, and the role identifier of the role object in the film and television video to which it belongs.

2. The method according to claim 1, wherein, The performing object detection on each of the target video frames respectively to determine the role regions where the role objects appearing in each of the target video frames are located includes: Perform convolution processing on each of the target video frames respectively to obtain corresponding video frame features; Slide a preset detection box on each of the target video frames respectively to obtain each candidate box corresponding to each of the target video frames respectively; Based on each candidate box corresponding to each of the target video frames respectively, determine the role objects appearing in each of the target video frames, and determine the role regions where the role objects are located in the corresponding target video frames.

3. The method according to claim 2, characterized in that, The based on each candidate box corresponding to each of the target video frames respectively, determine the role objects appearing in each of the target video frames, and determine the role regions where the role objects are located in the corresponding target video frames includes: For each of the target video frames, determine the role objects appearing in the corresponding target video frame according to each candidate box corresponding to the corresponding target video frame; For each role object, enlarge the size of the candidate box where the corresponding role object is located, and use the region included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame.

4. The method according to claim 1, characterized in that, The determining the key position information of the target key points of each role object in the corresponding target video frame includes: When there are target parts of the role object in the corresponding target video frame, extract feature points from the target parts of each role object respectively to obtain target key points corresponding to each target part respectively, and determine the key position information of the target key points in the corresponding target video frame; When there are no target parts of the role object in the corresponding target video frame, predict the target parts of the role object in the corresponding video frame, and predict the key position information of the target key points of the target parts in the corresponding video frame.

5. The method according to claim 1, wherein The based on the key position information, perform alignment and registration processing on the images composed of the role regions where each role object is located respectively to obtain target images with a preset size corresponding to each role object respectively includes: Obtain the preset position information corresponding to each preset feature point in the preset template; Determine the mapping relationship between each character object and the preset template according to the key position information corresponding to the target key points of each character object and the preset position information of each preset feature point; Based on the mapping relationship corresponding to each character object, adjust the image formed by the character area where the corresponding character object is located to a target image of a preset size.

6. The method according to claim 5, characterized in that, The adjusting the image formed by the character area where the corresponding character object is located to a target image of a preset size based on the mapping relationship corresponding to each character object includes: For each character object, when there is a target part in the corresponding target video frame, directly map the image formed by the character area where the corresponding character object is located to a target image of a preset size based on the mapping relationship corresponding to the character object; When there is no target part in the corresponding target video frame for the character object, map the image formed by the character area where the corresponding character object is located to a part of the target image of a preset size based on the mapping relationship corresponding to the character object, and fill the remaining pixels in the target image with zero pixels to obtain a target image of a preset size.

7. The method according to claim 1, characterized in that The performing character recognition based on the target image to obtain the character information corresponding to each character object in the video to be processed includes: For the target image of a preset size corresponding to each character object respectively, perform convolution processing on the target image through a feature extraction network to obtain corresponding image features; Perform pooling processing on the image features to obtain corresponding pooling features; Perform residual processing on the pooling features, and fuse the features obtained by the residual processing and the corresponding image features to obtain target feature vectors corresponding to each character object; Based on the target feature vectors corresponding to each character object, determine the character information corresponding to each character object.

8. The method according to claim 7, wherein The determining the character information corresponding to each character object based on the target feature vectors corresponding to each character object includes: For the target feature vector corresponding to each character object, calculate the feature similarity between the target feature vector and each preset feature vector respectively to obtain at least one feature similarity corresponding to the corresponding character object; For each character object, determine the target preset feature vector corresponding to the feature similarity that meets the similarity condition among the at least one corresponding feature similarity, and use the character information corresponding to the target preset feature vector as the character information corresponding to the corresponding character object.

9. The method according to claim 8, wherein The feature similarity is characterized by the Euclidean distance; the calculating the feature similarity between the target feature vector and each preset feature vector respectively for the target feature vector corresponding to each character object to obtain at least one feature similarity corresponding to the corresponding character object includes: For the target feature vector corresponding to each character object, calculate the Euclidean distance between the target feature vector and each preset feature vector respectively to obtain at least one Euclidean distance corresponding to the corresponding character object; For each role object, determining a target preset feature vector corresponding to a feature similarity that meets the similarity condition among at least one determined feature similarity includes: For each role object, determining the minimum Euclidean distance among at least one determined Euclidean distance; When the minimum Euclidean distance is less than a distance threshold, using the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector.

10. The method according to any one of claims 1 to 9, characterized in that, The role information includes at least one of the movie name and movie link corresponding to the role object; the method further includes: Obtaining each user account that has viewed the to-be-processed video, and respectively calculating the relevance between each user account and the to-be-processed video; Pushing at least one of the movie name and movie link corresponding to the role object to the user account corresponding to the relevance that meets the relevance condition.

11. The method according to any one of claims 1 to 9, characterized in that The method further includes: Obtaining the playback channel of the to-be-processed video; When it is determined based on the role information that the playback channel does not have playback permission, deleting the to-be-processed video under the playback channel, and processing the user accounts that have spread the to-be-processed video.

12. A role recognition device, characterized in that, The device includes: An obtaining module, configured to obtain a to-be-processed video, and extract a target video frame from the to-be-processed video; the to-be-processed video is a mixed video, and the mixed video includes at least two video segments of the same actor, and the same actor plays different role objects in the at least two video segments; A detection module, configured to perform object detection on each of the target video frames respectively to determine the role regions where the role objects that appear in each of the target video frames are located; A determination module, configured to determine the key position information of the target key points of each role object in the corresponding target video frame; A registration module, configured to perform alignment and registration processing on the images formed by the role regions where each role object is located respectively based on the key position information, to obtain a target image with a preset size corresponding to each role object respectively; An identification module, configured to perform role identification based on the target image to obtain role information corresponding to each role object in the to-be-processed video; the role information includes the movie video to which the role object belongs, and the role identifier of the role object in the movie video.

13. The device according to claim 12, characterized in that, The detection module is further configured to perform convolution processing on each of the target video frames respectively to obtain corresponding video frame features; slide a preset detection frame on each of the target video frames respectively to obtain each candidate box corresponding to each of the target video frames respectively; based on each candidate box corresponding to each of the target video frames respectively, determine the role objects that appear in each of the target video frames, and determine the role regions where the role objects are located in the corresponding target video frames.

14. The device according to claim 13, characterized in that, The detection module is further configured to, for each of the target video frames, determine the role objects that appear in the corresponding target video frame according to each candidate box corresponding to the corresponding target video frame; for each role object, enlarge the size of the candidate box where the corresponding role object is located, and use the region included in the enlarged candidate box as the role region where the role object is located in the corresponding target video frame.

15. The device according to claim 12, characterized in that The determining module is further configured to, when a target part exists in a corresponding target video frame for a role object, extract feature points for the target parts of each role object respectively, obtain target key points corresponding to the respective target parts, and determine key position information of the target key points in the corresponding target video frame; when a target part does not exist in a corresponding target video frame for a role object, predict the target part of the role object in the corresponding video frame, and predict key position information of the target key points of the target part in the corresponding video frame.

16. The device according to claim 12, characterized in that The registration module is further configured to obtain preset position information corresponding to each preset feature point in a preset template; determine a mapping relationship between each role object and the preset template according to the key position information corresponding to the target key points of each role object and the preset position information of each preset feature point; Based on the mapping relationship corresponding to each role object, adjust an image formed by a role area where the corresponding role object is located to a target image of a preset size.

17. The device according to claim 16, wherein The registration module is further configured to, for each role object, when a target part exists in a corresponding target video frame for the role object, directly map an image formed by a role area where the corresponding role object is located to a target image of a preset size based on the mapping relationship corresponding to the role object; when a target part does not exist in a corresponding target video frame for the role object, map an image formed by a role area where the corresponding role object is located to a part of a target image of a preset size based on the mapping relationship corresponding to the role object, and fill the remaining pixels in the target image with zero pixels to obtain a target image of a preset size.

18. The device according to claim 12, characterized in that, The recognition module is further configured to, for a target image of a preset size corresponding to each role object respectively, perform convolution processing on the target image through a feature extraction network to obtain corresponding image features; perform pooling processing on the image features to obtain corresponding pooling features; perform residual processing on the pooling features, and fuse the features obtained by the residual processing with the corresponding image features to obtain target feature vectors corresponding to the respective role objects; determine role information corresponding to the respective role objects based on the target feature vectors corresponding to the respective role objects.

19. The device according to claim 18, characterized in that, The recognition module is further configured to, for the target feature vector corresponding to each role object, calculate the feature similarity between the target feature vector and each preset feature vector respectively to obtain at least one feature similarity corresponding to the corresponding role object; for each role object, determine a target preset feature vector corresponding to the feature similarity that meets the similarity condition among the corresponding at least one feature similarity, and use the role information corresponding to the target preset feature vector as the role information corresponding to the corresponding role object.

20. The device according to claim 19, characterized in that The feature similarity is characterized by the Euclidean distance; the recognition module is further configured to, for the target feature vector corresponding to each role object, calculate the Euclidean distance between the target feature vector and each preset feature vector respectively to obtain at least one Euclidean distance corresponding to the corresponding role object; For each character object, determine the minimum Euclidean distance among the corresponding at least one Euclidean distance; when the minimum Euclidean distance is less than the distance threshold, use the preset feature vector corresponding to the minimum Euclidean distance as the target preset feature vector.

21. The device according to any one of claims 12 to 20, characterized in that, The character information includes at least one of the movie name and movie link corresponding to the character object; the apparatus further includes: A push module, configured to obtain each user account that has browsed the to-be-processed video, and calculate the relevance between each user account and the to-be-processed video respectively; push at least one of the movie name and movie link corresponding to the character object to the user account corresponding to the relevance that meets the relevance condition.

22. The device according to any one of claims 12 to 20, characterized in that, The apparatus further includes: A processing module, configured to obtain the playback channel of the to-be-processed video; when it is determined based on the character information that the playback channel does not have the playback permission, delete the to-be-processed video under the playback channel, and process the user account that spreads the to-be-processed video.

23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

24. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

25. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Film and television recommendation method and device and storage medium

    CN110399527A

  • Drawn face image recognition method, computer readable storage medium and related equipment

    CN111275005A

  • Object recognition method and device, storage medium and electronic device

    CN111444822A