Live broadcast scene detection method, storage medium and electronic device
Patent Information
- Application Number
- CN202111159815.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The prior art has poor generalization and weakened image background features in live broadcast scene detection.
By segmenting the video frames, the qualities of the moving target in the foreground area and the environmental attribute characteristics in the background area are identified, and multimodal features are fused to obtain video features, and then predict whether the video is a live video.
It improves the accuracy and generalization of live broadcast scene detection and enhances the user's viewing experience.
Smart Images

Figure CN113902989B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and specifically to a live broadcast scene detection method, storage medium and electronic device. Background Art
[0002] In recent years, micro videos, mainly in the form of live broadcast, have been widely disseminated on social networks as a novel media form, adding a lot of fun to people's lives. Due to the large variety and quantity of micro videos, there is an urgent need for an effective method to identify and filter them, so as to obtain meaningful live videos for users to watch, so as to improve the user's viewing experience.
[0003] Traditionally, feature extraction methods, such as scale-invariant feature conversion algorithms or template matching algorithms, can be used to identify local features of objects in videos to achieve the purpose of scene recognition. Among them, local features of objects can be, for example, bedroom corners, door handles, etc., which are used to identify iconic objects in video images. However, this method is too targeted at local areas, resulting in poor generalization and is not conducive to popularization.
[0004] Some scene recognition and detection technologies use various convolutional neural networks to perform image classification operations and directly train end-to-end, such as using classic models such as convolutional neural networks and residual networks. This method does not work well in live broadcast scenarios. This is because in a live broadcast environment, key information such as people exists in the image and occupies most of the image area, weakening the original features of the image background, resulting in poor detection results. Summary of the invention
[0005] Therefore, the embodiments of the present invention intend to provide a live scene detection method and device, as well as related storage media and electronic devices, which can effectively solve the problems of poor generalization when using local features for live scene detection and end-to-end training that weakens the original features of the image background.
[0006] In a first aspect, a live broadcast scene detection method is provided, comprising:
[0007] For a video frame of a video to be detected, segment a foreground area and a background area, wherein the foreground area includes a moving target;
[0008] identifying the posture of the moving target in the foreground area to obtain posture features;
[0009] identifying attributes of the environment in the background area to obtain attribute features; and
[0010] Multimodal feature fusion is performed on the posture features of the moving target and the attribute features of the environment to obtain video features of the video, and based on the video features, it is predicted whether the video is a live video.
[0011] In a second aspect, a live broadcast scene detection device is provided, comprising:
[0012] A segmentation unit, configured to segment a foreground area and a background area for a video frame of a video to be detected, wherein the foreground area includes a moving target;
[0013] a posture recognition unit configured to recognize the posture of the moving target in the foreground area to obtain posture features;
[0014] an environment attribute recognition unit, configured to recognize the attributes of the environment in the background area to obtain attribute features;
[0015] The fusion prediction unit is configured to perform multimodal feature fusion on the posture features of the moving target and the attribute features of the environment to obtain video features of the video, and predict whether the video is a live video based on the video features.
[0016] In a third aspect, a storage medium is provided, storing a computer program, wherein the computer program is configured to execute any live scene detection method of the embodiments of the present invention when executed.
[0017] In a fourth aspect, an electronic device is provided, comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute any live scene detection method of any embodiment of the present invention when running the computer program.
[0018] According to the above technical solution, live scene detection can be performed through multimodal features. This method comprehensively considers the impact of multiple features on the detection results, realizes the complementarity of multiple features, and thus improves the accuracy of the detection results. In addition, the above technical solution extracts multiple features in the video from a global perspective, and live scene detection is performed based on these multiple features, which can enhance the generalization of the live scene detection method.
[0019] Other optional features and technical effects of the embodiments of the present invention are partially described below, and partially can be understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The elements shown are not limited to the proportions shown in the accompanying drawings. The same or similar reference numerals in the accompanying drawings represent the same or similar elements, wherein:
[0021] Figure 1A schematic flow chart of a live broadcast scene detection method according to an embodiment of the present invention is shown;
[0022] Figure 2 A schematic diagram showing a video frame of a video to be detected according to an embodiment of the present invention is shown;
[0023] Figure 3 A schematic flow chart of a method for segmenting a foreground area and a background area according to an embodiment of the present invention is shown;
[0024] Figure 4 The embodiment of the present invention is shown Figure 2 A schematic diagram of a binarized image of a video frame shown;
[0025] Figure 5 A schematic block diagram of segmenting a foreground area and a background area according to an embodiment of the present invention is shown;
[0026] Figure 6 The embodiment of the present invention is shown Figure 4 A schematic diagram of at least a portion of the boundary of a foreground area in a binarized image shown;
[0027] Figure 7 The present invention is based on an embodiment of the present invention. Figure 6 A schematic diagram of the foreground area defined by the boundaries shown;
[0028] Figure 8 A schematic flow chart of a method for estimating a person's posture according to an embodiment of the present invention is shown;
[0029] Fig. 9 A schematic diagram showing key parts of a character according to an embodiment of the present invention is shown;
[0030] Fig.10 A schematic flow chart of a method for determining key points of a person's posture based on key parts according to an embodiment of the present invention is shown;
[0031] Fig.11 The embodiment of the present invention is shown Fig. 9 Schematic diagram of the pose estimation framework for the character shown;
[0032] Fig.12 A schematic flow chart of a method for detecting facial expressions of a person according to an embodiment of the present invention is shown;
[0033] Fig.13 A schematic flow chart of a method for identifying attributes of an environment in a background area according to an embodiment of the present invention is shown;
[0034] Fig.14 A schematic diagram showing grid division of a background area according to an embodiment of the present invention is shown;
[0035] Fig.15 A schematic diagram of template matching according to an embodiment of the present invention is shown;
[0036] Fig.16 A schematic flow chart of a method for performing beat detection on sound of a video according to an embodiment of the present invention is shown;
[0037] Fig.17 A schematic diagram showing a waveform of audio amplitude data of sound in a video according to an embodiment of the present invention;
[0038] Fig.18 A schematic diagram showing a waveform of a frequency domain signal of a difference audio sequence of sound in a video according to an embodiment of the present invention;
[0039] Fig.19 A schematic flow chart of a method for obtaining video features by feature fusion and predicting live video based on the video features according to an embodiment of the present invention is shown;
[0040] Fig. 20 A schematic diagram showing a live broadcast scene detection method according to an embodiment of the present invention is shown;
[0041] Fig.21 A schematic diagram showing the structure of a live broadcast scene detection device according to an embodiment of the present invention is shown; and
[0042] Fig. 22 An exemplary structural diagram of an electronic device capable of implementing the live broadcast scene detection method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific implementation methods and drawings. Here, the exemplary implementation methods of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0044] As described in the background technology, the existing live scene detection method only judges the live scene through a single image, ignoring the implied timing information, the host's demeanor information, etc. As a result, the accuracy and universality of the existing live scene detection method are low. Currently, there is a demand to improve the accuracy of live scene detection and enhance the generalization of live scene detection in response to complex situations in the live environment.
[0045] The embodiment of the present invention provides a live scene detection method. The method can effectively adapt to complex situations in a live broadcast environment, improve the accuracy of scene recognition, and thus improve the user's viewing experience. In addition, the embodiment of the present invention also relates to corresponding devices, computer systems that implement the above methods, and storage media storing programs that can execute the above methods. In some embodiments, the device, component, unit, or model can be implemented by software, hardware, or a combination of software and hardware.
[0046] Figure 1 FIG. 1 is a schematic flow chart of a live broadcast scene detection method 100 according to an embodiment of the present invention. Figure 1 The live broadcast scene detection method 100 may include steps S110 to S170.
[0047] Step S110, for a video frame of a video to be detected, segment a foreground area and a background area, wherein the foreground area includes a moving target.
[0048] The video to be detected can be any video suitable for scene detection. Among them, the scene can be used to describe the activities of the target and the environment in which it is located. For example, a person reading or sleeping in the bedroom, a person dancing in the square or singing indoors, a panda eating bamboo outdoors, etc. can all be regarded as a scene. It can be understood that for recording live videos, the camera is usually facing the target in the live scene, so as to better play the target's demeanor. For scenes where the target is a person, the person can adjust his or her demeanor at any time and can also interact with the audience. In short, live videos usually include moving targets, which can be people or animals. The moving target is usually in motion during the live broadcast. It can be understood that the motion state is not limited to large-scale movements, such as dancing, doing exercises, etc., but can also be small-scale movements, such as singing, reading, etc.
[0049] The video to be detected may include multiple video frames, and the foreground area and background area of the video frames are segmented. Any existing or future developed segmentation method such as the Otsu method-maximum inter-class variance method (OTSU algorithm) can be used. As long as the target is not stationary, it can be considered as a moving target. The foreground area and background area in the video frame can be segmented based on whether the target is moving. The foreground area includes the moving target, and the remaining areas, such as stationary objects or surroundings far away from the lens, can be considered as background areas. Figure 2 FIG. 4 is a schematic diagram showing a video frame of a video to be detected according to an embodiment of the present invention. Figure 2 As shown, the person facing the camera is the foreground area of this video frame, while the person facing away from the camera and the surrounding environment where the two people are located may be the background area of this video frame.
[0050] Step S130, identifying the posture of the moving target in the foreground area to obtain posture features.
[0051] After obtaining the foreground area and the background area, the posture of the moving target in the foreground area can be identified to obtain posture features. The posture of the moving target may include the posture and expression of the moving target. Taking the moving target as a person as an example, different actions can show different postures in the video, such as standing, sitting, walking, dancing, etc. The posture and action of the person can be identified by various feature information, such as the outline of the person, and the person's posture recognition can also be performed based on motion capture technology. Specifically, the motion trajectory of the person is identified by locating the joint points of the person and storing the motion data information of the joint points. The person has different emotions and can show different expressions in the video. The expression of the person can be identified based on the shape and position of the key points of the face. The posture of the identified moving target can be represented by posture features.
[0052] Step S150, identifying the attributes of the environment in the background area to obtain attribute features.
[0053] Based on the background area obtained after segmentation, the attributes of the environment are identified to obtain the attribute characteristics of the environment. For example, the attributes of the environment identify the specific environment, such as a bright and spacious balcony with strong light, or a bedroom with warm colors and romantic style. For example, the attributes of the environment can be determined based on the intensity of light, color matching, landmark buildings and facilities, and decoration style, and the corresponding attribute characteristics can be obtained. For example, based on the identified desks, computers, etc., it can be identified that the environment in the background area is an office.
[0054] Step S170, multi-modal feature fusion is performed on the posture features of the moving target and the attribute features of the environment to obtain video features, and based on the video features, it is predicted whether the video is a live video.
[0055] Multimodal features can be multiple indicators and states that measure a target. The demeanor features of the moving target and the attribute features of the environment in the background area obtained according to the above steps can be regarded as features of different modes obtained based on the video to be detected. These features describe the video from different aspects, with less redundant information. The above multiple features can be multi-modal feature fusion to obtain video features, that is, multiple features are fused into one feature. It can be understood that since the main body of the live video is the "moving target", environmental factors have little effect on the live broadcast process. Therefore, different weights can be assigned to the above features to indicate their importance. For example, the demeanor features of the moving target can be assigned a weight of 70%, while the attribute features of the environment in the background area can be assigned a weight of 20%, and the remaining 10% can be used to represent other features.
[0056] For example, the video features can be input into the neural network so that the neural network can predict whether the corresponding video is a live video. When the prediction result indicates that the video is a live video, it can be retained. If the prediction result indicates that the video is not a live video, it can be deleted. In addition, the live video can be classified according to the prediction result, for example, into singing, dancing, and other categories.
[0057] The feature fusion method is not specifically limited in the present application, and any existing or future method that can achieve multimodal feature fusion is within the protection scope of the present application.
[0058] According to the above technical solution, live scene detection can be performed using video features after multimodal feature fusion. This method comprehensively considers the impact of multiple features on the detection results, realizes the complementarity of multiple features, and thus improves the accuracy of the detection results. In addition, the above technical solution extracts multiple features from the video from a global perspective, and live scene detection is performed based on these multiple features, which can enhance the generalization of the live scene detection method.
[0059] Figure 3 FIG. 4 shows a schematic flow chart of segmenting a foreground area and a background area according to an embodiment of the present invention. Figure 3 In the embodiment of the present invention, step S110 of segmenting the foreground area and the background area can be implemented by the following steps.
[0060] Step S111, calculating a differential image between a current frame and a previous frame of the video.
[0061] Since the target in the foreground area is moving in the live broadcast scene, there are certain differences between adjacent video frames. The inter-frame difference can be used to segment the foreground area and the background area. First, the two adjacent video frames in the video to be detected are denoted as F n and F n-1 , then the grayscale value of the corresponding pixel in the two video frames is recorded as F n (x, y) and F n-1 (x, y). For example, it can be based on the formula:
[0062] D n (x, y) = |F n (x, y)-F n-1 (x, y)|
[0063] Calculate the gray value D of the pixel in the differential image n (x, y), and then obtain the difference image D n .
[0064] Step S112, binarizing the difference image to obtain a binarized image.
[0065] There are many methods for binarizing the difference image. For example, when the gray value of a pixel in the difference image is less than or equal to the threshold T, the pixel corresponding to the gray value is set to 0 (black). On the contrary, when the gray value of a pixel in the difference image is greater than the threshold T, the pixel corresponding to the gray value is set to 255 (white). Then the binary image R can be obtained. n Generally, the threshold T can be set to 127. Specifically, it can be implemented using the following formula:
[0066]
[0067] Figure 4 It is shown that according to an embodiment of the present invention Figure 2 The binarized image is obtained by binarizing the differential image corresponding to the video frame shown.
[0068] Step S113, performing connectivity analysis on the binary image to obtain a foreground area and a background area.
[0069] According to the above steps, a binary image has been obtained, in which the pixel points with a gray value of 255 are points in the foreground area. Connectivity analysis refers to finding each connected area in the binary image and marking it. Connectivity analysis can use the following algorithms: Two-Pass and Seed Filling. Those of ordinary skill in the art can understand the specific implementation steps and results of these two algorithms, and will not be described in detail here.
[0070] To help understand, Figure 5 FIG. 4 shows a schematic block diagram of segmenting a foreground area and a background area according to an embodiment of the present invention. Figure 5 As shown, first determine the current frame F from the video to be detected n and the previous frame F n-1 Then, the difference image between the two is calculated. According to the set threshold, the difference image is binarized. Finally, the foreground area and background area of the video frame are obtained through connectivity analysis.
[0071] Therefore, by making full use of the characteristics of the foreground area in the video including the moving target through the above simple calculation, the foreground area and the background area of the video frame in the video can be segmented. The algorithm is simple and easy to implement, and the amount of calculation is small, which can save the calculation cost.
[0072] In some embodiments, step S113 may include: first, performing a dilation operation and an erosion operation on the binary image to obtain at least a partial boundary of the foreground area; then, determining the foreground area and the background area based on at least a partial boundary.
[0073] For example, Figure 4 The binary image shown in the figure can be obtained by performing dilation and erosion operations respectively. Figure 6 The image shown includes a portion of the foreground region's boundary, which is Figure 6 In some cases, the moving target does not appear completely in the video, such as a person sitting at a desk reading a book. In this case, the upper body of the person usually appears in the lower part of the live video, and the lower body may not appear in the video. Therefore, the same position in the live video is generally taken as the foreground area. Figure 6 For example, the foreground area refers to the connected area contained under the white boundary. In other cases, the moving object in the video frame may appear completely in the video, such as a person dancing in the middle. The entire boundary of the person can be obtained through dilation and erosion operations. Therefore, the area contained within the entire boundary can be regarded as the foreground area, and the area outside the entire boundary can be regarded as the background area.
[0074] The above expansion operation can be understood as expanding Figure 4 The bright white area in the image. Specifically, a structural element can be used to scan each pixel in the image, and each pixel in the structural element is used to perform an "OR" operation with the pixels it covers. If both are 0, the pixel is 0, otherwise it is 1. Conversely, the erosion operation can be used to scan each pixel in the image with a structural element, and each pixel in the structural element is used to perform an "AND" operation with the pixels it covers. If both are 1, the pixel is 1, otherwise it is 0. Usually, these two operations are performed sequentially.
[0075] According to the above technical solution, the dilation operation can merge all the background points that are in contact with the moving target into the moving target, fill the small holes in the area, and make the moving target larger. The erosion operation can eliminate the boundary points of the moving target, making it smaller, and can also eliminate the noise points that are smaller than the structural elements. Therefore, through these simple operations, a smoother and more accurate boundary can be obtained. Thus, the purpose of accurately segmenting the foreground area and the background area with a smaller calculation cost is achieved.
[0076] According to one embodiment of the present invention, determining the foreground area and the background area based on at least a portion of the boundary may also include: first, determining the border of the video frame; then, determining the area enclosed by the lower border of the video frame and at least a portion of the boundary as the foreground area, and determining other areas outside the foreground area as the background area.
[0077] As mentioned above, the upper body of the person usually appears in the lower part of the live video. Figure 6 In this embodiment, the person in the video frame is sitting, and a bust of the person is obtained, so the white border is part of the border of the foreground area (person), and the rest of the border can be the bottom border of the video frame. Figure 7 The present invention is based on an embodiment of the present invention. Figure 6 Schematic diagram of the foreground area determined by the boundary shown. Figure 7 , the area enclosed by the above-mentioned partial border and the lower border is the foreground area ( Figure 7 The rest is the background area.
[0078] In live video, if a moving target, such as a person, does not appear completely in the video frame, it is more likely to appear in the middle and lower positions of the video frame. According to the above technical solution, the boundary of the moving target can be effectively determined, and then the foreground / background area of the image can be accurately segmented based on the complete boundary.
[0079] As mentioned above, in the embodiment of the present invention, the moving target may include a person. Step S130 of identifying the demeanor of the moving target in the foreground area may include: performing posture estimation on the person to obtain the posture characteristics of the person; and / or performing expression detection on the person to obtain the expression characteristics of the person. For example, the posture characteristics of the person may be represented by the position information of the person's trunk or limbs. According to the position information, the amplitude range and intensity of the movement made by the person at this time can be roughly known, so as to predict the activities performed by the person. When the amplitude range of the movement is large, it is predicted that the person may be doing some intense exercise at this time, such as running, dancing, etc. When the amplitude range of the movement is small, it is predicted that the person may be doing some gentle movements at this time, such as reading, sleeping, etc. Similarly, the expression characteristics of the person may be represented by the position information of the facial features of the person. According to the position information, the shape and position of the facial features can be roughly known, so as to predict the emotions of the person at this time. For example, when the positions of the facial features are all within the normal range, it may indicate that the emotional changes of the person at this time are not large, and then it is predicted that the person may be doing activities with high concentration such as reading or sleeping without emotional changes at this time. For example, when the ratio of the open mouth area to the face area exceeds the normal ratio, it may mean that the person's emotions are fluctuating at this time, and then predict that the person may be eating, talking or singing, etc.
[0080] For live videos, many of them are live videos of people. On the one hand, the posture and expression of the person are obvious characteristics of the person; on the other hand, the posture estimation and expression prediction of the person are also relatively easy to achieve. Therefore, this ensures the feasibility and reliability of live scene detection. Finally, these two identify the person from different aspects. If these two are considered comprehensively, it will provide a guarantee for better results in live scene detection.
[0081] Figure 8 FIG. 4 is a schematic flow chart of a method for estimating a person's posture according to an embodiment of the present invention. Figure 8 , the posture estimation of the person may include steps S131 to S133.
[0082] Step S131, semantic segmentation is performed on the foreground area to obtain key parts of the person.
[0083] It is understood that semantic segmentation refers to classifying each pixel in the foreground area. The categories may be the upper arm, forearm, thigh, calf, torso, etc. of a person. Each category is a key part. The specific algorithm for implementing semantic segmentation is not limited in this application, and any existing or future algorithm that can implement semantic segmentation is within the scope of protection of this application. Fig. 9 FIG. 2 shows a schematic diagram of key parts of a character according to an embodiment of the present invention. Fig. 9 As shown, the forearm of the character is obtained through semantic segmentation, see Fig. 9 Medium highlight area.
[0084] Step S132, determining the key points of the character's posture based on the key parts.
[0085] After obtaining the key parts, the corresponding posture key points can be determined. The posture key points are usually the end points of the joints or parts of the human body. Usually, the key points are located at the boundaries or ends of the key parts. Based on this, the key parts can be used to determine the posture key points of the character. See again Fig. 9 The character's forearm is a key part, and there are two key points involved, namely the midpoint of the intersection of the wrist and the elbow.
[0086] Step S133, determining the posture features of the character based on the posture key points.
[0087] According to an embodiment of the present invention, the determined posture key points can be numbered. Then the associated key points are connected, for example, the key points belonging to the same part. Taking the above example as an example, the above step S132 obtains two key points, and the two key points are connected. For the above example, similarly, the elbow and shoulder key points can also be connected. After the associated key points are connected, a figure composed of line segments similar to a human skeleton diagram will be obtained. Based on the figure, the posture of the character can be predicted, thereby determining the posture characteristics of the character. For example, the sitting posture, standing posture, running posture, etc. of the character.
[0088] As one of the most important features in multimodal features, the posture of the person has a great impact on the results of live scene detection. In fact, in live videos, the postures of people are usually some relatively common sitting postures, standing postures, etc. The above steps obtain a relatively accurate estimation result of the person's posture at a relatively low computational cost, which provides stable and accurate information for the subsequent live scene detection, thereby ensuring the accuracy of the live scene detection results.
[0089] Fig.10 FIG. 2 shows a schematic flow chart of a method for determining a key point of a person's posture based on key parts according to an embodiment of the present invention. Fig.10 As shown, step S132 of determining the key points of the character's posture based on the key parts may include the following steps S132a to S132c.
[0090] Step S132a, fitting the key parts with parallelograms.
[0091] As mentioned above, the key parts are extracted based on the body of the character. According to the physiological characteristics, they can be fitted using parallelograms. Fig. 9 As shown, after semantic segmentation of the foreground area, the key part of the forearm is obtained. The forearm can be fitted with a parallelogram. For example, a vertex of the video frame can be used as the origin to establish a rectangular coordinate system. Thus, the position coordinates of each vertex of the parallelogram can be obtained.
[0092] Step S132b, determining the short side of the parallelogram based on the position coordinates of the vertices of the parallelogram.
[0093] After obtaining the position coordinates of each vertex according to the above steps, randomly select one of the vertices as the base point, and then calculate the distances from the two vertices connected to it to the base point. The line connecting the points with the shorter distance between the two points is the short side of the parallelogram.
[0094] Step S132c, determine the midpoint of the short side as the posture key point.
[0095] After determining the short side of the parallelogram, the midpoint of the short side can be determined based on the position and length of the short side. The two midpoints taken from the short side are then used as posture key points. For the same parallelogram, the posture can be estimated by connecting the key points taken on it. It can be understood that after semantic segmentation, multiple key parts can be obtained. For simplicity, only one key part is described in the above scheme. After performing the above operations on multiple key parts, the following can be obtained: Fig.11 Schematic diagram of the character pose estimation framework shown.
[0096] For the detection of live scene, posture estimation is a part of the operation. The above steps have a small amount of calculation on the basis of ensuring the accuracy of posture key points. Therefore, the detection speed is improved on the basis of ensuring the accuracy of the live scene detection results.
[0097] According to an embodiment of the present invention, identifying the posture features of the moving target in the foreground area includes performing facial expression detection on the person. Fig.12 FIG. 1 is a schematic flow chart of a method for detecting facial expressions of a person according to an embodiment of the present invention. Fig.12 The step of detecting facial expressions of a person may include step S134 and step S135.
[0098] Step S134, detecting key points of the person's face.
[0099] First, the facial recognition frame of the person in the video frame can be obtained. Then, the key points of the face are detected in the facial recognition frame. The key points of the face of the person can be represented by the position information of the facial features of the person, for example, the position coordinates of the eyes, eyebrows, and mouth. Exemplarily, 68 key points can be sampled for the face area. Similar to the posture key points, the key points of different parts can be numbered, for example, the facial contour is numbered 1-16, the two eyebrows are numbered 17-21 and 22-26 respectively, the nose is numbered 27-36, the two eyes are numbered 37-42 and 43-48, and the mouth is numbered 49-68.
[0100] Step S135, determining the facial expression features of the person based on the facial key points.
[0101] For example, when the area enclosed by key points numbered 49-68 occupies a larger proportion in the facial recognition frame, it can mean that the person's mouth is opening wider, the person is more excited, and may be very surprised. For another example, when the ratio of the distance between key points numbered 17-21 or 22-26 and the upper boundary of the facial recognition frame to the height of the facial recognition frame is smaller, it can mean that the person's eyebrows are raised and may be very happy. On the contrary, the person may be very sad.
[0102] Based on the above operations, the facial expressions of the characters can be accurately determined, providing more accurate input data for the subsequent live video detection process.
[0103] Fig.13 FIG. 2 is a schematic flow chart showing a method for identifying attributes of an environment in a background area according to an embodiment of the present invention. Fig.13 , step S150 of identifying the attributes of the environment in the background area may include the following steps S151 to S153.
[0104] Step S151, dividing the background area into grids.
[0105] Preferably, the background area is divided into grids so that at least one of the divided grids does not include the foreground area. Because the grid does not include pixels in the foreground area, all of its pixels are from the background area, and thus the grid is more conducive to identifying the attributes of the environment in the background area, avoiding interference of pixels in the foreground area in identifying the attributes of the environment in the background area. Fig.14 FIG. 2 shows a schematic diagram of meshing a background area according to an embodiment of the present invention. Fig.14 As shown, in this embodiment, the background area is evenly divided into 5*5 grids. Among them, the 5 grids in the first column and the three grids in the upper right corner are all grids that do not include the foreground area.
[0106] Step S152: determining images in one or more continuous grids in the divided image as template images.
[0107] The image in any one or more continuous grids in the divided image can be determined as the template image. Preferably, the grid that has no intersection with the foreground area is determined as the template image to avoid the interference of the foreground area on the identification of the background area. Fig.14 For example, the template image may be determined in five grids in the first column, one grid in the first row and fourth column, and two grids in the first two rows of the fifth column.
[0108] Step S153: template matching is performed on the template image and sample images in the image database to determine the attribute characteristics of the environment of the video frame, wherein the sample images respectively include environments with different attributes.
[0109] It can be understood that the image database contains a large number of sample images, each of which may include environments with different attributes. Fig.15 A schematic diagram of template matching according to an embodiment of the present invention is shown. Fig.15The right side of shows some sample images in the image database. As shown in the figure, the sample images include images showing environments with different attributes, such as images of a teenager's bedroom, a romantic bedroom, and a wooden kitchen. Exemplarily, when performing template matching, the sample images can be preliminarily screened according to the contrast, brightness, or hue of the template image and the sample image to reduce the number of sample images, thereby reducing the amount of calculation for subsequent template matching. Then, based on the screened sample images, template matching is performed using the template image. For example, the template image can be regarded as a slider, which is slid pixel by pixel on the sample image, and the similarity between the template image and the area in the sample image covered by it is calculated each time it is slid. The template matching result is determined based on the similarity. The higher the similarity, the greater the probability that the attributes of the environment in the template image are the same as the attributes of the environment in the sample image; otherwise, vice versa. Thus, the attribute characteristics of the environment in the video frame can be determined based on the attributes of the environment in the matched sample image.
[0110] The matching algorithm used in the above matching process has high precision, and the calculation process is simple and easy to implement, thereby ensuring the efficiency and accuracy of live video detection.
[0111] In a specific embodiment, the live scene recognition method 100 may further include: performing beat detection on the sound of the video to generate a beat feature. Step S170 in the method 100 may also be implemented by the following steps, performing multimodal feature fusion of the beat feature with the posture feature of the moving target and the attribute feature of the environment to obtain the video feature of the video. Thus, it is possible to predict whether the video is a live video based on the video feature. In this embodiment, a beat feature is added, and the beat feature of the video can be represented in the form of a vector, for example, Among them, the elements of the vector can represent the probability of containing a beat of a specific period in the sound. The specific period can be 1 second, 2 seconds, etc. For example, the above vector represents that the probability of containing a beat with a period of 1 second, 2 seconds, 3 seconds, 4 seconds, and 5 seconds in the sound is 0.1, 0.5, 0.4, 0.1, and 0.1 respectively. Therefore, according to the above vector, it can be determined that the sound of the video has a high probability of containing a beat with a period of between 2 and 3 seconds. Similar to the posture features of the character, the expression features of the character, and the attribute features of the environment, the beat feature can also be used as one of the multimodal features.
[0112] The beat feature provides relevant information about the sound in the video, which is a powerful supplement to the relevant information of the video frame. Especially for live videos, a large proportion of them are singing videos and dancing videos, and some are daily life videos, but they are also accompanied by background music. Therefore, by adding auditory features on the basis of visual features, the accuracy of live scene detection results is significantly improved. Moreover, the beat feature is a relatively easy-to-detect sound feature, and using this beat feature will not increase the computational complexity of live video detection too much.
[0113] Fig.16 FIG. 2 shows a schematic flow chart of performing beat detection on the sound of a video according to an embodiment of the present invention. Fig.16 As shown, step S160 of performing beat detection on the sound of the video may include the following steps S161 to S164.
[0114] Step S161, extracting audio amplitude data of sound from the video.
[0115] Fig.17 A schematic diagram showing a waveform of audio amplitude data of a sound in a video according to an embodiment of the present invention is shown. The audio amplitude data may represent the volume of the sound, which may be stored in the form of a one-dimensional vector.
[0116] Step S162, calculating the difference between the audio amplitude at the current moment and the audio amplitude at the previous moment to obtain a difference audio sequence.
[0117] It can be understood that before the audio amplitude data is processed, it may contain noise data. Therefore, it can be filtered and noise-reduced first. For example, the current moment differs from the previous moment by 1 millisecond. The difference between the current moment audio amplitude and the previous moment audio amplitude is calculated, and the amplitude difference corresponding to the current moment can be obtained. For multiple moments in the duration of the sound, the above steps are performed respectively to obtain multiple amplitude differences. Using the amplitude difference as the ordinate, the time as the abscissa and a step size of 1 millisecond, a difference audio sequence can be obtained.
[0118] Step S163: Perform Fourier transform on the difference audio sequence to obtain a frequency domain signal.
[0119] Compared with time domain signals, frequency domain signals are easier to analyze the beat of sound. Fig.17 The difference audio sequence of the sound shown in Fourier transform can be obtained as follows Fig.18 The frequency domain signal is shown.
[0120] Step S164, determining the beat feature based on the frequency domain signal. Exemplarily, this step can be implemented using a trained neural network.
[0121] According to the above steps, the information irrelevant to the beat in the sound is effectively removed. It is easier to analyze and obtain the beat features based on the frequency domain information obtained through the above steps, and the algorithm of the above steps is simple and easy to implement.
[0122] It is understood that the above technical solution is only for illustration and does not constitute a limitation of the present invention. For example, the order of step S130 and step S150 can be exchanged, and can even be executed in parallel. In addition, steps S131 to S133 and steps S134 and S135 are only used to distinguish different operations in step S130, and do not represent the order of the steps. In short, the order in the above scheme is only exemplary, and does not limit the order of steps in the actual live scene detection process.
[0123] Fig.19 FIG. 1 is a schematic flow chart showing a method for obtaining video features by feature fusion and predicting live video based on the video features according to an embodiment of the present invention. Fig.19 As shown, step S170 performs multimodal feature fusion of the posture features of the moving target and the attribute features of the environment to obtain video features, which may include the following steps S171 and S172.
[0124] Step S171, obtaining a vector corresponding to the posture feature of the moving target and a vector corresponding to the attribute feature of the environment.
[0125] As mentioned above, the posture features of the characters, the facial expressions of the characters, the attribute features of the environment, etc. extracted from the video can all be used as a modal feature of the video. Each of these features can be represented by a vector. In this step, the vectors corresponding to the posture features of the moving target and the attribute features of the environment can be obtained.
[0126] Step S172, constructing a feature matrix using vectors corresponding to the posture features of the moving target and the attribute features of the environment, wherein the feature matrix is used to represent the video features of the video.
[0127] According to the aforementioned step S171, the vector corresponding to the posture feature of the moving target and the vector corresponding to the attribute feature of the environment can be obtained. If multiple features are directly combined in series, it will lead to serious feature redundancy and ignore the weight difference between multimodal features. According to an embodiment of the present application, the feature matrix x of the i-th video frame of the video to be detected can be constructed based on these vectors. i =(x 1 , x 2 ,,...). Among them, x 1 , x 2,,...represent the vectors corresponding to the posture features of the moving target and the vectors corresponding to the attribute features of the environment. It can be understood that the dimensions of the vectors corresponding to the posture features of the moving target and the attribute features of the environment may be the same or different. When the dimensions of the vectors are different, the largest dimension of these vectors is used as the reference, and the insufficient elements of the remaining vectors are supplemented with "0" to construct the feature matrix.
[0128] Furthermore, the predicting whether a video is a live video based on video features may include steps S173 and S174.
[0129] Step S173, inputting the feature matrix into a trained multi-layer perceptron, so that the multi-layer perceptron outputs a scene classification vector, wherein the elements in the scene classification vector are used to represent a scene of a live video or a scene of a non-live video;
[0130] Step S174, determining whether the video is a live video according to the elements in the scene classification vector.
[0131] Multilayer Perceptron (MLP) is a feedforward artificial neural network model. A trained multilayer perceptron is a model in which the parameters of the neural network are optimized. The multilayer perceptron can output the scene classification vector y i For example, the output scene classification vector can be Each element represents the probability that the video is the scene corresponding to the element. It can be understood that when the value of the element is larger, it can be said that the probability that the scene in the video frame belongs to the scene corresponding to the element is higher. Exemplarily, multiple different scenes may include live video scenes and non-live video scenes. In other words, for a certain scene, it either belongs to a live video scene or a non-live video scene. As described above, the video to be detected may contain multiple video frames, and for each of the video frames, the above steps are performed separately to obtain multiple scene classification vectors. Exemplarily, when more than a certain number of video frames, such as 50% of the video frames, belong to the same scene, the detection result of the video to be detected can be output as belonging to the scene. For example, the video to be detected contains 50 video frames, that is, 50 scene classification vectors can be obtained. Among them, the scene classification vectors obtained based on the first 5 video frames indicate that these video frames are indoor walking scenes, and the scene classification vectors obtained based on the last 45 video frames indicate that these video frames are indoor singing scenes. Obviously, 45 exceeds 50% of 50, so the video to be detected is detected as an indoor singing scene. The indoor singing scene is one of the live video scenes, so the video is a live scene. On the contrary, if the scene classification vectors obtained based on the first 45 video frames indicate that these video frames are indoor reading scenes, and the scene classification vectors obtained based on the last 5 video frames indicate that these video frames are indoor singing scenes, then the video to be detected is detected as an indoor reading scene. The indoor reading scene is one of the non-live video scenes. Therefore, the video is a non-live scene.
[0132] In summary, through the nonlinear transformation of the multi-layer perceptron, x is established i and i The mapping between .
[0133] Existing scene recognition and detection technologies for videos mainly use various convolutional neural networks to perform image classification operations. The neural network can be, for example, a classic model such as the VGG16 model and the residual network (ResNet) model in a deep neural network. This method is very effective for natural scenes and can achieve a high accuracy rate. However, for live videos with moving objects such as people occupying most of the area, the scene recognition effect of this method is not ideal.
[0134] It can be understood that the above steps S171, S172, S173 and S174 take the multimodal feature fusion of the vector corresponding to the posture feature of the moving target and the attribute feature of the environment as an example to describe the specific process of obtaining the video feature of the video. As mentioned above, the beat feature of the sound can also be fused with the posture feature of the moving target and the attribute feature of the environment to obtain the video feature of the video. The process is similar to the above process and will not be repeated here for the sake of brevity.
[0135] Therefore, in the present application, a multi-layer perceptron is used to perform live scene detection on the video to be detected based on the above-mentioned multimodal features, so that non-repetitive features in the video features can be extracted, effectively avoiding the redundancy of features, and thus effectively compensating for the problems of unsatisfactory scene recognition effects in the prior art.
[0136] In an embodiment of the present invention, a multi-layer perceptron can be trained based on a cross entropy loss function using a training video and scene label data corresponding to the training video. In addition, a mini-batch gradient descent algorithm can be used during the training process to optimize the network weights and biases of the multi-layer perceptron.
[0137] Scene label data It can be standardized data of the scene corresponding to each video frame in the training video, which can be obtained by, for example, manual or machine annotation. The scene label data includes the scene label vector corresponding to the real scene in the training video. Exemplarily, the scene label data can be represented in the form of one-hot encoding. For example, for an indoor singing video, the element at the position corresponding to the indoor singing scene in the vector of its video label data is "1", and the remaining elements corresponding to scenes such as dancing and reading are all "0". Similar to the video to be detected, the training video annotated with scene label data is used as input, and the aforementioned steps S110 to S170 are executed to obtain the corresponding scene classification vector. The scene label data can be calculated using the cross entropy loss function. and the scene classification vector y i The prediction loss between . The cross entropy loss function E n (W, B) can be calculated according to the following formula:
[0138]
[0139] Wherein, bn represents the number of samples contained in each batch of data after all scene label data corresponding to all video frames of the training video are divided into N batches; W and B represent the network weights and biases of the multi-layer perceptron, respectively.
[0140] Based on the calculated function values of multiple cross entropy loss functions, the network weights and biases of the multilayer perceptron can be optimized using the minimum batch gradient descent algorithm. For example, after the function values of the loss function of each batch of data are back-propagated, the network weights and biases are continuously adjusted until the minimum batch gradient descent algorithm converges. It can be understood that the convergence of the algorithm can mean that after multiple iterations, the output value tends to a specific value.
[0141] The batch gradient descent algorithm targets the function values of all cross entropy loss functions, that is, the entire data set. The direction of the gradient can be solved by calculating all samples in the data set. In the above algorithm, all sample data must be used in each iteration. For a particularly large amount of data, a lot of computational cost is required. The mini-batch gradient descent algorithm can use some samples instead of all samples in each iteration.
[0142] After the above training process, the network of the multi-layer perceptron can be optimized, and then the live scene detection results output by the optimized multi-layer perceptron are also optimized, which improves the accuracy and reliability of the detection results. In addition, using the small batch gradient descent algorithm to optimize the network of the multi-layer perceptron can significantly reduce the amount of calculation, save a lot of calculation costs, and ensure that the accuracy of the calculation results is not affected in any way.
[0143] Fig. 20 FIG. 2 is a schematic diagram showing a live broadcast scene detection method according to an embodiment of the present invention. Fig. 20 As shown, first, the video frame of the video to be detected is segmented to obtain a foreground area and a background area. Based on the foreground area, the moving target, such as a person, can be estimated and / or the expression detection can be performed to obtain the posture characteristics and / or the expression characteristics of the person. Template matching can be performed based on the background area to determine the attribute characteristics of the environment in the video frame. While performing the above-mentioned regional segmentation, audio data can also be extracted for the video to be detected. The extracted audio data can be subjected to difference processing and Fourier transformation to obtain beat features. The aforementioned posture features of the person, the expression features of the person, the attribute features of the environment, and the beat features can all be represented in the form of vectors as one of the multimodal features. The multimodal features are fused to obtain video features. The video features are input into a multilayer perceptron to predict whether the video to be detected is a live video.
[0144] In some embodiments of the present invention, Fig.21 As shown, a live scene detection device 2100 is also provided. The device 2100 may include a segmentation unit 2101, a posture recognition unit 2102, an environmental attribute recognition unit 2103, and a fusion prediction unit 2104.
[0145] In the illustrated embodiment, the segmentation unit 2101 may be configured to segment a foreground area and a background area for a video frame of a video to be detected. The foreground area includes a moving target. In the illustrated embodiment, the posture recognition unit 2102 may be configured to recognize the posture of the moving target in the foreground area to obtain posture features. In the illustrated embodiment, the environmental attribute recognition unit 2103 may be configured to recognize the attributes of the environment in the background area to obtain attribute features. In the illustrated embodiment, the fusion prediction unit 2104 may be configured to fuse the multimodal features of the video to obtain video features, and predict whether the video is a live video based on the video features. The multimodal features include the posture features of the moving target and the attribute features of the environment.
[0146] Those skilled in the art will understand that, without causing any contradiction, the device of this embodiment can be combined with the method features described in other embodiments, and vice versa.
[0147] In an embodiment of the present invention, a storage medium is provided, wherein the storage medium stores a computer program, and the computer program is configured to execute any live scene detection method of the embodiment of the present invention when executed.
[0148] In an embodiment of the present invention, an electronic device is provided, including: a processor and a memory storing a computer program, wherein the processor is configured to execute any live scene detection method of the embodiment of the present invention when running the computer program.
[0149] Fig. 22 A schematic diagram of an electronic device 2200 that can implement the live scene detection method of an embodiment of the present invention is shown. In some embodiments, more or fewer electronic devices than shown in the figure may be included. In some embodiments, it can be implemented using a single or multiple electronic devices. In some embodiments, it can be implemented using cloud or distributed electronic devices.
[0150] like Fig. 22As shown, the electronic device 2200 includes a central processing unit (CPU) 2201, which can perform various appropriate operations and processes according to the programs and / or data stored in the read-only memory (ROM) 2202 or the programs and / or data loaded from the storage part 2208 to the random access memory (RAM) 2203. CPU 2201 can be a multi-core processor, or it can include multiple processors. In some embodiments, CPU 2201 can include a general main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In RAM 2203, various programs and data required for the operation of the electronic device 2200 are also stored. CPU 2201, ROM 2202 and RAM 2203 are connected to each other via bus 2204. Input / output (I / O) interface 2205 is also connected to bus 2204.
[0151] The processor and the memory are used together to execute the program stored in the memory. When the program is executed by the computer, the steps or functions of the live scene detection method or device described in the above embodiments can be implemented.
[0152] The following components are connected to the I / O interface 2205: an input section 2206 including a keyboard, a mouse, etc.; an output section 2207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 2208 including a hard disk, etc.; and a communication section 2209 including a network interface card such as a LAN card, a modem, etc. The communication section 2209 performs communication processing via a network such as the Internet. A drive 2210 is also connected to the I / O interface 2205 as needed. A removable medium 2211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 2210 as needed, so that a computer program read therefrom is installed into the storage section 2208 as needed. Fig. 22 Only some components are schematically shown in the figure, which does not mean that the computer system 2200 only includes Fig. 22 Components shown.
[0153] The systems, devices, modules or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server or a combination thereof.
[0154] In a preferred embodiment, the live scene detection method can be implemented or realized at least partially or completely on a cloud-based machine learning platform or partially or completely in a self-built machine learning system, such as a GPU array.
[0155] In a preferred embodiment, the live scene detection method and device can be implemented or realized in a server, such as a cloud or distributed server. In a preferred embodiment, the server can also be used to push or send data or content to the interrupt based on the generated result.
[0156] Storage media in embodiments of the present invention include permanent and non-permanent, removable and non-removable items that can be used to store information by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0157] The methods, programs, systems, devices, etc. of the embodiments of the present invention may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.
[0158] Those skilled in the art should understand that the embodiments of the present specification may be provided as methods, systems or computer program products. Therefore, those skilled in the art may imagine that the functional modules / units or controllers and related method steps described in the above embodiments may be implemented in software, hardware or a combination of software / hardware.
[0159] Unless explicitly stated, the actions or steps of the methods, programs, and embodiments of the present invention do not have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel / merged processing of each step is also possible or may be advantageous.
[0160] Herein, “first”, “second”, are used to distinguish different elements in the same embodiment, and do not refer to order or relative importance.
[0161] In this article, multiple embodiments of the present invention are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this article, "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" are meant to be applicable to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. In the absence of contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0162] The exemplary systems and methods of the present invention have been specifically shown and described with reference to the above embodiments, which are merely examples of the best modes for implementing the present systems and methods. It will be appreciated by those skilled in the art that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A live broadcast scene detection method, It is characterized in that include: For a video frame of a video to be detected, segmenting a foreground area and a background area based on whether the target is moving, wherein the foreground area includes a moving target; identifying the posture of the moving target in the foreground area to obtain posture features; Identifying the attributes of the environment in the background area to obtain attribute features; as well as Multimodal feature fusion is performed on the posture features of the moving target and the attribute features of the environment to obtain video features of the video, and based on the video features, it is predicted whether the video is a live video.
2. The method according to claim 1, It is characterized in that The step of performing multimodal feature fusion on the posture features of the moving target and the attribute features of the environment to obtain the video features of the video includes: Obtaining a vector corresponding to the posture feature of the moving target and a vector corresponding to the attribute feature of the environment; A feature matrix is constructed using vectors corresponding to the posture features of the moving target and the attribute features of the environment, wherein the feature matrix is used to represent the video features of the video.
3. The method according to claim 2, It is characterized in that The predicting whether the video is a live video based on the video feature includes: Inputting the feature matrix into a trained multi-layer perceptron so that the multi-layer perceptron outputs a scene classification vector, wherein elements in the scene classification vector are used to represent a scene of a live video or a scene of a non-live video; Determine whether the video is a live video according to the elements in the scene classification vector.
4. The method according to any one of claims 1 to 3, It is characterized in that in, The moving target includes a person, and the step of identifying the posture features of the moving target in the foreground area includes: performing posture estimation on the person to obtain posture features of the person; and / or Perform expression detection on the character to obtain the expression features of the character.
5. The method according to claim 4, It is characterized in that The step of estimating the posture of the person comprises: Performing semantic segmentation on the foreground area to obtain key parts of the person; Determining key points of the character's posture based on the key parts; and The posture features of the character are determined based on the posture key points.
6. The method according to claim 5, It is characterized in that The determining of the key points of the posture of the character based on the key parts includes: Fitting the key parts with a parallelogram; Determine the short side of the parallelogram based on the position coordinates of the vertices of the parallelogram; The midpoint of the short side is determined as the posture key point.
7. The method according to claim 4, It is characterized in that The detecting of facial expressions of the person includes: Detecting facial key points of the person; Determine the facial expression features of the person according to the facial key points.
8. The method according to any one of claims 1 to 3, It is characterized in that The segmenting of the foreground area and the background area comprises: Calculating a differential image between a current frame and a previous frame of the video; Binarizing the difference image to obtain a binary image; and Connectivity analysis is performed on the binary image to obtain the foreground area and the background area.
9. The method according to claim 8, It is characterized in that The performing connectivity analysis on the binary image to obtain the foreground area and the background area includes: Performing a dilation operation and an erosion operation on the binary image to obtain at least a partial boundary of the foreground area; Based on the at least partial boundary, the foreground region and the background region are determined.
10. The method according to claim 9, It is characterized in that The determining the foreground area and the background area based on at least part of the boundary comprises: Determining a border of the video frame; An area enclosed by a lower border of the video frame and at least a portion of the border is determined as the foreground area, and other areas outside the foreground area are determined as the background area.
11. The method according to any one of claims 1 to 3, It is characterized in that The identifying the attributes of the environment in the background area includes: Performing grid division on the background area; Determine, in the divided image, images within one or more continuous grids as template images; Template matching is performed on the template image and sample images in an image database to determine attribute features of the environment of the video frame, wherein the sample images respectively include environments with different attributes.
12. The method according to any one of claims 1 to 3, It is characterized in that Also includes: Performing beat detection on the sound of the video to generate a beat feature; The beat feature is fused with the posture feature of the moving target and the attribute feature of the environment in a multimodal manner to obtain the video feature of the video.
13. The method according to claim 12, It is characterized in that The step of performing beat detection on the sound of the video includes: Extracting audio amplitude data of the sound from the video; Calculate the difference between the current audio amplitude and the previous audio amplitude to obtain a difference audio sequence; Performing Fourier transform on the difference audio sequence to obtain a frequency domain signal; The beat feature is determined based on the frequency domain signal.
14. A storage medium, It is characterized in that The storage medium stores a computer program, and the computer program is configured to execute the live scene detection method according to any one of claims 1 to 13 when executed.
15. An electronic device, It is characterized in that include: A processor and a memory storing a computer program, wherein the processor is configured to execute the live scene detection method according to any one of claims 1 to 13 when running the computer program.
Citation Information
Patent Citations
Video scene recognition method and device, electronic equipment and storage medium
CN111291692A
Label identification method and device
CN113407778A