Target moving object recognition method and device, electronic equipment and storage medium
By performing motion detection and intra-frame feature screening on ball game video frames, identifying target moving objects, the problem of difficulty for online viewers to identify balls for games is solved, the recognition accuracy and efficiency are improved, and the viewing experience is enhanced.
Patent Information
- Application Number
- CN202510568950.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-08-15
AI Technical Summary
In large-scale ball events, it is difficult for online viewers to recognize balls for the game. The existing recognition methods are low in accuracy and efficiency. Especially when using ultra-wide-angle cameras or panoramic cameras to collect global images of the field, the balls for the game account for a small proportion, resulting in a decrease in viewing experience.
By performing motion detection of target video frames, the moving objects are selected, and the recognition rules corresponding to intra-features and number are used to identify target moving objects from the moving objects to be identified, reducing interference and improving recognition accuracy and efficiency.
It improves the accuracy and efficiency of the recognition of target sports objects, eliminates interference from non-target sports objects, and enhances the viewing experience of online audiences.
Smart Images

Figure CN120495339A_ABST
Abstract
Description
[0001] This application is a divisional application with application number 202110219773.9, application date 2021-02-26, and invention name “Target motion object identification method, device, electronic device and storage medium”. Technical Field
[0002] The present application relates to the field of image recognition technology, and more specifically, to a method, device, electronic device and storage medium for identifying a target moving object. Background Art
[0003] In some ball games, in order to obtain the situation of each part of the field, it is usually hoped to capture the overall image of the field. However, some ball games have large fields, and in the captured overall image of the field, the match ball occupies a small proportion of the screen. Online viewers may not be able to identify the match ball well, thereby reducing the audience's online viewing experience.
[0004] In this case, it is usually necessary to identify the match ball in the video to assist online viewers in watching the game. However, using traditional image feature recognition methods to identify the match ball in the video has the problem of low recognition accuracy. Summary of the Invention
[0005] In view of this, the embodiments of the present application propose a target moving object recognition method, device, electronic device and storage medium to improve the above-mentioned problems.
[0006] In a first aspect, an embodiment of the present application provides a method for identifying a target motion object, the method comprising: acquiring a target video frame; performing motion detection on the target video frame to obtain a motion object in the target video frame; screening the motion object to be identified from the motion objects based on the intra-frame features of the motion object; and identifying the target motion object from the motion objects to be identified based on an identification rule corresponding to the number of motion objects to be identified.
[0007] In a second aspect, embodiments of the present application provide a target moving object identification device, comprising: a target video frame acquisition module, a motion detection module, a screening module, and an identification module. The target video frame acquisition module is configured to acquire a target video frame; the motion detection module is configured to perform motion detection on the target video frame to obtain a moving object in the target video frame; the screening module is configured to screen the moving objects to be identified based on their intra-frame features; and the identification module is configured to identify the target moving object from the moving objects to be identified based on an identification rule corresponding to the number of moving objects to be identified.
[0008] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, wherein the above method is executed when the program code is run by a processor.
[0010] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method.
[0011] The embodiments of the present application provide a method, device, electronic device, and storage medium for identifying a target moving object. After acquiring a target video frame, the method first performs motion detection on the target video frame to obtain a moving object in the target video frame. Then, based on the intra-frame features of the moving object, the moving object to be identified is screened from the moving objects. Finally, based on the recognition rules corresponding to the number of moving objects to be identified, the target moving object is identified from the moving objects to be identified. Since only the moving objects to be identified that are screened out from the moving objects are identified, it is equivalent to a preliminary screening of the moving objects, which can eliminate interference from other moving objects to a certain extent, thereby improving the accuracy of target moving object identification. At the same time, since only the moving objects to be identified that are screened out from the moving objects are identified, the number of identifications using preset recognition rules is reduced, thereby improving the efficiency of target moving object identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0013] Figure 1 A flow chart of a target moving object recognition method proposed in one embodiment of the present application is shown;
[0014] Figure 2 Shown Figure 1 A flowchart of an implementation of S110 in a target moving object recognition method proposed in the illustrated embodiment;
[0015] Figure 3 Shown Figure 1 A flowchart of an implementation of S120 in a target moving object recognition method proposed in the illustrated embodiment;
[0016] Figure 4 A schematic diagram of a target video frame in an embodiment of the present application is shown;
[0017] Figure 5 A flowchart of another target moving object recognition method proposed in one embodiment of the present application is shown;
[0018] Figure 6 Shown Figure 5 A schematic diagram of a flow chart of S230 in a target moving object recognition method proposed in the illustrated embodiment;
[0019] Figure 7 Shown Figure 5 A schematic diagram of a flow chart of S230 in another target moving object recognition method proposed in the illustrated embodiment;
[0020] Figure 8 Shown Figure 5 A schematic diagram of a flow chart of S230 in another target moving object recognition method proposed in the illustrated embodiment;
[0021] Figure 9 A flowchart of another target moving object recognition method proposed in one embodiment of the present application is shown;
[0022] Figure 10 A block diagram of a moving object recognition device proposed in one embodiment of the present application is shown;
[0023] Figure 11 A structural block diagram of an electronic device for executing a target moving object recognition method according to an embodiment of the present application is shown;
[0024] Figure 12 A storage unit for storing or carrying program codes for implementing a target moving object recognition method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] In large-scale ball games such as football and basketball, in order to present a global image of the stadium to online viewers so that they can understand information such as the tactical distribution of the game through the global image, an ultra-wide-angle camera or a panoramic camera can be used to capture the global image of the stadium. The global image of the stadium refers to the image covering the entire stadium area.
[0027] However, since the stadiums used for large-scale ball games are generally large, and the match balls are relatively small compared to the stadiums, the match balls occupy a small proportion of the screen in the collected global images of the stadiums. Online viewers may not be able to identify the match balls well in the video images, thereby reducing the audience's online viewing experience.
[0028] In this case, it's often necessary to identify the match ball in the video to assist online viewers in watching the game. For example, the match ball can be identified first and then labeled to help online viewers quickly find the match ball and improve their viewing experience. However, related art methods for identifying match balls in video images suffer from low recognition accuracy.
[0029] For example, in some methods, the game ball in the video screen can be identified by using feature value recognition. This method directly extracts the feature values of the entire video frame, and then identifies the ball based on the feature values. Since the identification is based on the feature values of the entire video frame, there is interference from many similar features, so there is a problem of low recognition accuracy.
[0030] In other methods, recognition can be achieved through neural network deep learning. However, neural network deep learning recognition has certain requirements for the size of the object to be recognized and the input video frame resolution. For example, the object to be recognized cannot be smaller than 20*20 resolution, and the input video frame resolution cannot be larger than 2K. To achieve the desired display effect, the video frame resolution is much larger than the resolution required by neural network deep learning. This requires scaling the video frame resolution to reduce it to meet the input source size requirements of deep learning. As a result, the already small football in the video frame becomes even smaller due to scaling, and the accuracy of deep learning football recognition becomes even lower.
[0031] Therefore, the inventors proposed the target motion object identification method, device, electronic device and storage medium provided in the present application. In this method, after obtaining the target video frame, motion detection is first performed on the target video frame to obtain the motion object in the target video frame, and then based on the intra-frame features of the motion object, the motion object to be identified is screened from the motion objects, and finally based on the identification rules corresponding to the number of motion objects to be identified, the target motion object is identified from the motion objects to be identified.
[0032] In the aforementioned method, since only the moving objects to be identified that have been screened out from the moving objects are identified, it is equivalent to a preliminary screening of the moving objects, which can eliminate the interference of other moving objects to a certain extent, thereby improving the accuracy of target moving object identification. At the same time, since only the moving objects to be identified that have been screened out from the moving objects are identified, the number of moving objects identified using preset recognition rules is reduced, and the recognition efficiency is improved.
[0033] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0034] See also Figure 1 , Figure 1 FIG2 is a flow chart of a method for identifying a moving target object according to an embodiment of the present application, the method comprising the following steps:
[0035] S110, obtaining a target video frame.
[0036] During the video image acquisition process, the video is captured in the form of video image frames (i.e., video frames). Therefore, the video image is composed of multiple video frames. The target video frame is the video frame in the video image that will be used for motion detection. The video image is a full-scale video image of the stadium of a large-scale ball game captured using an ultra-wide-angle camera or a panoramic camera.
[0037] Optionally, the collected global video image may be a real-time video image or a video image collected in advance.
[0038] It is understood that the more video frames presented to the audience per unit time, that is, the higher the video frame rate, the smoother the audience's viewing experience. Therefore, in some embodiments, in order to improve the display effect of the target moving object in the video image, all video frames in the video image can be used as target video frames, that is, obtaining the target video frame is to obtain all video frames in the video image, so as to determine the target moving object from all video frames. In this way, the target moving object can be identified in each video frame in the video image, so that the number of video frames in which the target moving object is identified in the video image finally output to the online audience is greater, and the target moving object viewed by the audience is smoother.
[0039] The target sports object refers to the sports object that the online audience is paying attention to. For example, in ball games, users are paying attention to the game balls, that is, ball objects, such as football, basketball, volleyball, etc.
[0040] However, considering the device performance limitation, or in order to save the device performance, in other embodiments, the target video frame can also be obtained by obtaining a portion of the video frame in the video image, and the obtained portion of the video frame is used as the target video frame. Figure 2As shown, obtaining the target video frame may specifically include the following steps:
[0041] S111, obtaining a target video image and a video frame extraction frame rate.
[0042] The video frame extraction frame rate may be understood as the frame rate of a video image composed of target video frames extracted from a target video image. In this embodiment, the video frame extraction frame rate may be preset as needed.
[0043] S112 , extracting a target video frame from the target video image based on the video frame extraction frame rate.
[0044] To improve the display effect of a target moving object in a video image, as an embodiment, based on the video frame extraction frame rate, the target video frames can be extracted from the target video source uniformly, that is, the target video frames are extracted at a certain frame interval. Specifically, the frame interval can be determined based on the frame rate of the target video image and the video frame extraction frame rate.
[0045] For example, assuming that the frame rate of the target video image is 60 frames / second and the extracted frame rate of the video image is 30 frames / second, the interval frame number can be determined to be 1 frame, that is, 1 frame of video frame is extracted as the target video frame every 1 frame of video frame interval.
[0046] As another embodiment, based on the video frame extraction frame rate, the target video frames can be extracted from the target video source by randomly extracting the target video frames from the target video image. Still taking the video image extraction frame rate as 30 frames per second as an example, in this case, 30 video image frames can be randomly extracted from the target video image every second as the target video frames.
[0047] S120: Perform motion detection on the target video frame to obtain a moving object in the target video frame.
[0048] Moving objects refer to all objects that can move in a video image. For example, in a ball game, moving objects may include a game ball, players, or other objects that may move.
[0049] Research has found that in ball games, the background of the field and other objects do not move, but the game ball, athletes, etc. move. Therefore, in order to reduce the difficulty of subsequent recognition and the amount of recognition data, we can first perform motion detection on the target video frame to obtain the moving object in the target video frame, thereby excluding objects such as the field background, greatly improving the accuracy and efficiency of target moving object recognition.
[0050] Therefore, performing motion detection on a target video frame can be understood as detecting a moving object in the target video frame. There are multiple ways to perform motion detection on a target video frame. Alternatively, motion detection can be performed on the target video frame using a Gaussian mixture model. Alternatively, motion detection can be performed on the target video frame using inter-frame differences.
[0051] In addition, considering that when using ultra-wide-angle cameras or panoramic cameras and other camera equipment to capture global images of the stadium, objects outside the stadium will inevitably be captured, such as spectators outside the stadium, cleaning staff, or players using balls to warm up, and these objects can also move. If these objects are also used for subsequent identification of target moving objects, the recognition accuracy and efficiency will be reduced.
[0052] Therefore, in order to avoid the influence of the moving objects outside the stadium on the recognition accuracy and efficiency, as an implementation method, Figure 3 As shown, performing motion detection on a target video frame to obtain a moving object in the target video frame may specifically include the following steps:
[0053] S121, obtaining a picture of a valid motion area in a target video frame.
[0054] The effective movement area can be understood as the area within the possible movement range of the target moving object. For example, in a ball game, the game ball may typically move within the playing field or within a certain range of the playing field's boundaries. Therefore, the effective movement area can refer to the area within the playing field or within a certain range of the playing field's boundaries.
[0055] Considering that when using an ultra-wide-angle camera or panoramic camera to capture a global image of the stadium, the camera's capture position typically does not change, so the scene area of the captured video image does not change. Therefore, the effective motion area in the acquired target video frame can be pre-calibrated to obtain an image of the effective area. For example, the stadium boundary can be pre-calibrated in the video image, or a certain range outside the stadium boundary can be calibrated.
[0056] Optionally, a polygon may be used to enclose the playing field or a range within a certain range of the playing field boundary, thereby obtaining the image within the polygon as the image of the effective motion area.
[0057] For example, refer to Figure 4 , Figure 4 A schematic diagram of a target video frame is shown. Figure 4 The target video frame shown in shows both the scene inside the stadium and the scene outside the stadium. At this time, the stadium area can be circled with a polygon to obtain the scene of the effective motion area in the target video frame.
[0058] S122: Perform motion detection on the picture in the effective motion area to obtain the moving objects in the picture in the effective motion area.
[0059] By performing motion detection on images in the effective motion area, the influence of spectators, cleaning staff, or players using balls for warm-up on the target moving object is eliminated, further improving the accuracy and efficiency of target moving object recognition.
[0060] S130 , based on the intra-frame features of the moving objects, filter out the moving objects to be identified.
[0061] The moving objects to be identified are the moving objects that are subsequently identified using preset recognition rules. It is understandable that even the moving objects obtained from the effective motion area picture still include all objects that have moved, such as athletes, game balls, flags waving by referees, or other objects that may be in motion. The number of these objects may still be large. If all these moving objects are directly identified, there is still a problem of low recognition efficiency. Therefore, in order to further improve the recognition efficiency, as an implementation method, after obtaining the moving objects, the moving objects can be further screened based on the intra-frame features of the moving objects to obtain the moving objects to be identified, thereby further reducing the moving objects that are actually used for subsequent identification.
[0062] Among them, intra-frame features refer to features obtained from a video frame.
[0063] S140 , identifying a target moving object from the moving objects to be identified based on an identification rule corresponding to the number of moving objects to be identified.
[0064] It is understood that after the moving objects are screened in step S130 to obtain the moving objects to be identified, the number of moving objects to be identified is further reduced. At this point, the number of moving objects to be identified may be one or more. In this case, different recognition rules can be selected based on the number of moving objects to be identified to identify this small number of moving objects to be identified, thereby obtaining the target moving object.
[0065] Furthermore, considering that under normal circumstances, there is by default one game ball on the field, if there is only one to-be-identified moving object, this one to-be-identified moving object can be directly determined as the target moving object. In this case, identifying the target moving object from the to-be-identified moving objects based on the identification rule corresponding to the number of to-be-identified moving objects includes: determining the to-be-identified moving object as the target moving object when the number of to-be-identified moving objects screened from the moving objects is only one.
[0066] In addition, considering that under normal circumstances, in addition to the game ball, there may also be athletes waiting to be identified as moving objects in the stadium. Similarly, since under normal circumstances there is a game ball by default in the stadium, when the number of moving objects to be identified is greater than one, the moving object to be identified cannot be directly determined as the target moving object, and each moving object to be identified needs to be identified separately. In this case, based on the recognition rules corresponding to the number of moving objects to be identified, the target moving object is identified from the moving objects to be identified, including: when the number of moving objects to be identified screened from the moving objects is greater than one, the images corresponding to each moving object to be identified are input into the object classifier, so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from sample video frames.
[0067] In this embodiment, since the moving object to be identified is a partial image within the target video frame, resolution scaling is not performed, and the resolution meets the input data requirements of the neural network model. Therefore, the object classifier can be a trained neural network model. Using the neural network model to identify the moving object improves recognition accuracy.
[0068] The object classifier can be a supervised model, that is, trained using sample moving objects with classification labels. The sample moving objects can be obtained through the above-mentioned steps S110-S130, that is, first obtaining sample video frames, then performing motion detection on the sample video frames to obtain moving objects in the sample video frames, and then filtering the moving objects to obtain sample moving objects. After obtaining the sample moving objects, sample moving objects with classification labels can be obtained through manual labeling.
[0069] Alternatively, the neural network model can adopt open source models such as TensorFlow (a symbolic mathematical system based on data flow programming, which is widely used in the programming implementation of various machine learning algorithms) and Caffe (Convolutional Architecture for Fast Feature Embedding).
[0070] In some embodiments, an object classifier can be configured to receive a moving object to be detected and then output whether the moving object is a target moving object. For example, if a moving object depicting a human head is input into the object classifier, and the target moving object is a soccer ball, the object classifier will output a negative result, indicating that the moving object is not the target moving object. In this manner, the target moving object can be identified from among multiple moving objects to be identified.
[0071] In other embodiments, an object classifier can be used to receive a moving object to be detected and then output the object type of the moving object to be detected, thereby determining a target moving object from the output object type. For example, if a moving object image of a "human head" is input into the object classifier, the object classifier will output that the moving object is a "human head." If a moving object image of a "soccer ball" is input into the object classifier, the object classifier will output that the moving object is a "soccer ball." In this way, the type of each moving object to be detected can be identified, thereby identifying a selected target moving object from among multiple moving objects to be detected.
[0072] Since the output result of the object classifier can be a target moving object or a moving object of other classifications, it is possible to obtain not only the target moving object but also the classification of other moving objects, such as the classification of athletes. Therefore, when online viewers want to pay attention to the position of an athlete, they can also frame the athlete's position and display it in the video image.
[0073] The present application provides a method for identifying a target moving object. After acquiring a target video frame, the method first performs motion detection on the target video frame to obtain a moving object in the target video frame. Then, based on the intra-frame features of the moving object, the method screens the moving objects to be identified from the moving objects. Finally, based on the recognition rules corresponding to the number of moving objects to be identified, the method identifies the target moving object from the moving objects to be identified. Since only the moving objects to be identified that have been screened out from the moving objects are identified, this is equivalent to a preliminary screening of the moving objects, which can eliminate interference from other moving objects to a certain extent, thereby improving the accuracy of target moving object identification. At the same time, since only the moving objects to be identified that have been screened out from the moving objects are identified, the number of identifications performed using preset recognition rules is reduced, thereby improving the efficiency of target moving object identification.
[0074] See also Figure 5 , Figure 5 FIG2 is a flow chart of a method for identifying a moving target object according to another embodiment of the present application. The method may include the following steps:
[0075] S210: Acquire a target video frame.
[0076] S220: Perform motion detection on the target video frame to obtain a moving object in the target video frame.
[0077] In some implementations, in order to fully select the moving object and facilitate determination of the position and actual size of the moving object, a rectangular frame may be used to select the moving object obtained by motion detection. Figure 4 As shown, various moving objects are selected by rectangular boxes.
[0078] S230 , based on the intra-frame features of the moving objects, filter out the moving objects to be identified.
[0079] There are multiple types of intra-frame features of moving objects.
[0080] In some embodiments, the intra-frame feature may be the actual size of the current position of the moving object in the target video frame. Figure 6 As shown, based on the intra-frame features of the moving objects, the moving objects to be identified are screened from the moving objects, which may specifically include the following steps:
[0081] S231A, obtaining predicted sizes of the target moving object at different positions in the target video frame, and the actual size of the moving object at the current position in the target video frame.
[0082] It is understandable that in the global image of the stadium captured by an ultra-wide-angle camera or a panoramic camera, each object shows a pattern of being larger when closer and smaller when farther away. That is to say, in actual scenes, the display size of the moving object in the final captured video image is different depending on the distance from the acquisition device. Based on this discovery, when the position of the acquisition device such as the ultra-wide-angle camera or the panoramic camera is determined, the size of the target moving object corresponding to different positions in the target video frame can be calibrated in advance. The pre-calibrated size is the predicted size of the target moving object at different positions in the target video frame.
[0083] It is understood that the moving object is located at a certain position in the target video frame, namely, the current position of the moving object. Similarly, the moving object also has an actual size in the target video frame. Therefore, the actual size of the moving object at the current position in the target video frame can be directly obtained from the target video frame.
[0084] After the moving object is selected using a rectangular frame, the actual size of the moving object can be obtained by multiplying the resolution length and width of the rectangular frame.
[0085] S232A: When the actual size of the moving object at the current position matches the predicted size of the current position, determine that the moving object is a moving object to be identified.
[0086] In some embodiments, the predicted size may be a range. In this case, the actual size of the moving object at the current position matches the predicted size at the current position, which means that the actual size is within the range of the predicted size.
[0087] Therefore, when the actual size of the moving object at its current position matches the predicted size of its current position, it can be determined that the moving object meets the requirements in terms of size, and moving objects whose sizes do not match the target moving object can be preliminarily screened out, such as the flags waving by the referee, etc., so that the moving objects that meet the requirements can be determined as the moving objects to be identified.
[0088] It can be seen that this embodiment can reduce the number of moving objects that are subsequently identified using recognition rules, eliminate interference from interfering moving objects, and improve the recognition accuracy and efficiency of target moving objects.
[0089] In other embodiments, the intra-frame feature may be the actual aspect ratio of the moving object in the target video frame. Figure 7 As shown, based on the intra-frame features of the moving objects, the moving objects to be identified are screened from the moving objects, which may specifically include the following steps:
[0090] S231B, obtaining the predicted aspect ratio of the target moving object in the target video frame, and the actual aspect ratio of the moving object in the target video frame.
[0091] It can be understood that the predicted aspect ratio of the target moving object in the target video frame refers to the aspect ratio that the target moving object should be displayed in the target video frame. For example, for ball objects such as footballs and basketballs, they should be displayed in the target video frame with an aspect ratio close to that of a square, that is, the aspect ratio is close to one to one. For athletes, they should be displayed with an aspect ratio of a long strip.
[0092] The actual aspect ratio of the moving object in the target video frame can be directly obtained from the target video frame.
[0093] After the moving object is selected by using a rectangular frame, the actual aspect ratio of the moving object can be obtained by the ratio of the resolution length to the width of the rectangular frame.
[0094] S232B: When the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object, determine that the moving object is a moving object to be identified.
[0095] In some embodiments, the predicted aspect ratio may be a range. In this case, the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object means that the actual aspect ratio is within the range of the predicted aspect ratio.
[0096] Therefore, when the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object, it can be determined that the moving object meets the requirements in terms of aspect ratio. Similarly, moving objects whose aspect ratio does not match the target moving object can be preliminarily screened out, thereby determining the moving objects that meet the requirements as the moving objects to be identified, further reducing the number of moving objects that are subsequently identified using recognition rules, and improving the recognition accuracy and efficiency of the target moving objects.
[0097] For example, when an athlete's lower limbs remain stationary and only their arms are moving, the moving object is the athlete's arm, and the actual length-to-width ratio of the arm is close to that of a rectangle. However, when the target moving object is a soccer ball, the predicted length-to-width ratio is close to that of a square. Therefore, the two do not match, and moving objects such as "arms" can be filtered out.
[0098] In other embodiments, some moving objects have symmetry, such as a game ball or an athlete's complete body, while some moving objects do not have symmetry. For example, for an athlete whose lower limbs are stationary and only one arm is moving, the identified single "arm" moving object is not symmetrical. Based on the symmetry of the moving object, in some embodiments, the intra-frame feature can be the color distribution parameter of the moving object. In this case, Figure 8 As shown, based on the intra-frame features of the moving objects, the moving objects to be identified are screened from the moving objects, which may specifically include the following steps:
[0099] S231C, obtaining color distribution parameters of the moving object.
[0100] The color distribution parameter of the moving object refers to the color distribution of the moving object, and optionally, may be the RGB color distribution of the moving object in the target video frame.
[0101] After the moving object is selected by using a rectangular frame, the color distribution parameters of the moving object may be color distribution parameters of the portion selected by the rectangular frame.
[0102] S232C: Perform symmetry detection on the moving object based on the color distribution parameter to obtain a symmetry detection result of the moving object.
[0103] It can be understood that for a moving object with symmetry, its color distribution has certain regularities, for example, a symmetrical color distribution along the axis of symmetry. Therefore, the symmetry detection of the moving object can be performed based on the color distribution parameters to obtain the symmetry detection result of the moving object.
[0104] In some embodiments, after the moving object is selected using a rectangular frame, the rectangular frame containing the moving object can be divided, for example, into four grids, nine grids, or sixteen grids. Taking the nine grids as an example, the four corners, i.e., the upper left, upper right, lower left, and lower right grids, are each subjected to RGB color comparison to calculate similarity, or the upper left grid and the upper right grid are selected to perform RGB color comparison to calculate similarity, or the lower left grid and the lower right grid are selected to perform RGB color comparison to calculate similarity.
[0105] It should be noted that the similarity calculation method is not limited in the embodiments of the present application. For example, the RGB color comparison of the entirety of the left three grids and the entirety of the right three grids in the nine-grid can be performed to calculate a single similarity. For another example, the RGB color comparison of the upper left, upper right, lower left, and lower right grids of the nine-grid can be performed in sequence to calculate multiple similarities.
[0106] After calculating the similarity, the calculated similarity (which can be one or more) can be compared with the similarity threshold. If the calculated similarity is greater than the similarity threshold, it can be determined that the symmetry detection result of the moving object is symmetrical. If the calculated similarity is less than or equal to the similarity threshold, it can be determined that the symmetry detection result of the moving object is not symmetrical.
[0107] S233C: When the symmetry detection result matches the predicted symmetry of the target moving object, determine that the moving object is a moving object to be identified.
[0108] It is understood that the predicted symmetry of the target moving object can be symmetric or non-symmetric. For example, when a game ball is selected as the target moving object, the predicted symmetry of the target moving object is symmetric. However, some moving objects do not have symmetry. For example, for an athlete whose lower limbs remain stationary and only one arm moves, the identified single "arm" moving object is non-symmetric. When a non-symmetric moving object is selected as the target moving object, the predicted symmetry of the target moving object is non-symmetric.
[0109] Therefore, when the predicted symmetry of the target motion object is symmetrical, the motion object with the symmetry detection result being symmetrical can be determined as the object to be identified, and when the predicted symmetry of the target motion object is not symmetrical, the motion object with the symmetry detection result being not symmetrical can be determined as the object to be identified.
[0110] In this embodiment, based on the color distribution parameters, a symmetry detection is performed on the moving object to obtain a symmetry detection result of the moving object, and then based on the symmetry detection result that the moving object is symmetrical, it is determined that the moving object is a moving object to be identified. It can be determined that the moving object meets the requirements in terms of the color distribution parameters. Similarly, moving objects whose color distribution parameters do not meet the requirements of the target moving object can be preliminarily screened out, so that the moving objects that meet the requirements can be determined as the moving objects to be identified, further reducing the number of moving objects that are subsequently identified using recognition rules, and improving the recognition accuracy and efficiency of the target moving objects.
[0111] It should be noted that when filtering the moving objects to be identified from the moving objects based on the intra-frame features of the moving objects, each intra-frame feature can be used separately, for example, the actual size of the moving object at its current position in the target video frame can be used separately, the actual aspect ratio of the moving object in the target video frame can be used separately, or the color distribution parameters of the moving object can be used separately. Any two or more of the intra-frame features can also be used in combination at the same time. For example, the actual size of the moving object at its current position in the target video frame and the color distribution parameters of the moving object can be used in combination, or the actual size of the moving object at its current position in the target video frame, the actual aspect ratio of the moving object in the target video frame and the color distribution parameters of the moving object can be used in combination at the same time. Moreover, when any two or more of the intra-frame features are used in combination at the same time, the various intra-frame features can be used regardless of their order.
[0112] In this embodiment, there is no specific limitation on the intra-frame features of the moving objects used to screen the moving objects to be identified.
[0113] S241 : When the number of the moving objects to be identified obtained by screening the moving objects is one, determine the moving object to be identified as a target moving object.
[0114] S242. When the number of moving objects to be identified screened from the moving objects is greater than one, the images corresponding to the moving objects to be identified are input into the object classifier, so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from sample video frames.
[0115] The target moving object recognition method of this embodiment proposes a variety of specific intra-frame features to screen moving objects in the process of screening moving objects to be identified from moving objects based on the intra-frame features of the moving objects. It has a wide range of applications and reduces the difficulty of screening moving objects.
[0116] See also Figure 9 , Figure 9 FIG2 is a flow chart of a target moving object recognition method proposed in another embodiment of the present application. The method is applied when the number of moving objects to be recognized in the target video frame is greater than one. The method may include the following steps:
[0117] S310: Acquire a target video frame.
[0118] S320: Perform motion detection on the target video frame to obtain a moving object in the target video frame.
[0119] S330 , based on the intra-frame features of the moving objects, filter out the moving objects to be identified.
[0120] In this embodiment, the number of moving objects to be identified that are screened from the moving objects based on the intra-frame features of the moving objects is greater than one.
[0121] S340 , obtaining actual distances between each to-be-identified moving object and a credible object in an adjacent target video frame corresponding to the same video frame, where the credible object is the to-be-identified moving object corresponding to one to-be-identified moving object screened from the moving objects.
[0122] It is understandable that after screening the moving objects in a certain target video frame, it is still possible to obtain multiple moving objects to be detected. However, in some cases, the number of target moving objects is limited. For example, under normal circumstances, there is only one match ball in the stadium by default. Therefore, if all the moving objects to be detected are identified in a random order based on the corresponding recognition rules, the target moving object may be identified only when the last moving object to be detected is identified, which reduces the recognition efficiency of the target moving object. Therefore, in order to further improve the recognition efficiency of the target moving object, the actual distance between each moving object to be identified and the credible object in the adjacent target video frame corresponding to the same video frame can be obtained first.
[0123] In combination with the above content, it can be seen that in some cases, for certain target video frames, based on preset screening rules, the number of moving objects to be identified screened from moving objects may be one. At this time, since there is a target moving object in the default target video frame, the moving object to be detected can be determined as the target moving object, and this target moving object is understood as a trusted object.
[0124] Furthermore, since the time interval between two adjacent target video frames is short, in actual scenes, the target moving object cannot move a long distance within a short time interval. In the two adjacent target video frames obtained by acquisition, the distance difference between the target moving objects is even smaller. Therefore, after obtaining the credible objects in the adjacent target video frames, for each moving object to be identified in the target video frame currently being identified, the actual distance between each moving object to be identified and the credible object in the adjacent video frame corresponding to the same video frame can be obtained first.
[0125] In actual scenarios, the acquisition angle of the video image acquisition device usually does not change. Therefore, the background position distribution in the acquired video frame does not change. In this case, as an implementation method, different target video frames can be mapped into the same target video frame, that is, the same coordinate system can be established in different target video frames, that is, the coordinate origin is the same in the two video frames, so as to obtain the coordinates of the trusted object and the coordinates of each moving object to be identified. Since the coordinate system is the same, the coordinates in different target video frames can be directly calculated for distance, and the actual distance between each moving object to be identified and the trusted object in the same video frame can be obtained.
[0126] S350, according to the actual distance between each moving object to be identified and the corresponding credible object in the adjacent video frame on the same video frame from small to large, the image corresponding to each moving object to be identified is input into the object classifier in sequence, so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from the sample video frames.
[0127] Since the distance difference between the target moving objects in the two adjacent target video frames obtained is small, after obtaining the actual distance between each moving object to be identified and the credible object in the adjacent target video frame corresponding to the same video frame, the moving objects to be identified and the credible object in the adjacent video frame can be sorted in ascending order according to the actual distance between the moving objects to be identified and the credible object in the adjacent video frame corresponding to the same video frame. It can be understood that the smaller the actual distance, the greater the probability that the moving object to be identified is the target moving object. Therefore, the images corresponding to each moving object to be identified can be input into the object classifier for identification in descending order according to the actual distance until the target moving object is identified from the moving objects to be identified by the object classifier.
[0128] By adopting the method of this embodiment, recognition can be started from the most likely target moving object, thereby improving the recognition efficiency of the target moving object.
[0129] It should be noted that this application provides some specific examples of possible implementations. Under the premise that they do not conflict with each other, the various embodiment examples can be arbitrarily combined to form a new target moving object recognition method. It should be understood that the new target moving object recognition method formed by combining any of the examples should fall within the scope of protection of this application.
[0130] See also Figure 10 , Figure 10 A block diagram of a target moving object recognition device 400 proposed in an embodiment of the present application is shown. The device 400 may include: a target video frame acquisition module 410, a motion detection module 420, a screening module 430 and a recognition module 440.
[0131] A target video frame acquisition module 410 is configured to acquire a target video frame;
[0132] A motion detection module 420 is configured to perform motion detection on a target video frame to obtain a moving object in the target video frame;
[0133] A screening module 430 is configured to screen the moving objects to be identified from the moving objects based on the intra-frame features of the moving objects;
[0134] The identification module 440 is configured to identify a target moving object from the moving objects to be identified based on an identification rule corresponding to the number of moving objects to be identified.
[0135] As an implementation manner, the target video frame acquisition module 410 is further configured to acquire a target video image and a video frame extraction frame rate; and extract a target video frame from the target video image based on the video frame extraction frame rate.
[0136] As an implementation manner, the motion detection module 420 is further configured to obtain a picture of a valid motion region in a target video frame; perform motion detection on the picture of the valid motion region, and obtain a moving object in the picture of the valid motion region.
[0137] As an embodiment, the screening module 430 is also used to obtain the predicted sizes of the target moving object at different positions in the target video frame, as well as the actual size of the moving object at its current position in the target video frame; when the actual size of the moving object at its current position matches the predicted size at its current position, the moving object is determined to be the moving object to be identified.
[0138] As an embodiment, the screening module 430 is also used to obtain the predicted aspect ratio of the target moving object in the target video frame, as well as the actual aspect ratio of the moving object in the target video frame; when the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object, the moving object is determined to be the moving object to be identified.
[0139] As an embodiment, the screening module 430 is also used to obtain the color distribution parameters of the moving object; based on the color distribution parameters, the moving object is subjected to symmetry detection to obtain the symmetry detection result of the moving object; when the symmetry detection result matches the predicted symmetry of the target moving object, the moving object is determined to be the moving object to be identified.
[0140] As an embodiment, the recognition module 440 is also used to determine the moving object to be identified as the target moving object when the number of moving objects to be identified screened from the moving objects is one; or, when the number of moving objects to be identified screened from the moving objects is greater than one, input the images corresponding to each moving object to be identified into the object classifier, so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from sample video frames.
[0141] As an embodiment, the number of moving objects to be identified is greater than one. In this case, the identification module 440 is also used to obtain the actual distance between each moving object to be identified and the credible object in the adjacent video frame corresponding to the same video frame, and the credible object is the moving object to be identified when the number of moving objects to be identified obtained by screening from the moving objects is one; according to the order of the actual distance between each moving object to be identified and the credible object in the adjacent video frame corresponding to the same video frame from small to large, the image corresponding to each moving object to be identified is input into the object classifier in sequence, so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from the sample video frames.
[0142] The target motion object recognition device provided by the present application only recognizes the motion objects to be recognized that have been screened out from the motion objects, which is equivalent to a preliminary screening of the motion objects. It can eliminate the interference of other motion objects to a certain extent, thereby improving the accuracy of target motion object recognition. At the same time, since only the motion objects to be recognized that have been screened out from the motion objects are recognized, the number of recognitions using preset recognition rules is reduced, thereby improving the recognition efficiency of target motion objects.
[0143] It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment and will not be repeated here.
[0144] The following will be combined Figure 11 An electronic device provided by this application is described.
[0145] See also Figure 11 Based on the above-mentioned target moving object recognition method, the embodiments of the present application also provide another electronic device 200 including a processor 104 capable of executing the above-mentioned target moving object recognition method. The electronic device 200 can be a device such as a smartphone, a tablet computer, a computer, or a portable computer. The electronic device 200 also includes a memory 104, a network module 106, and a screen 108. The memory 104 stores a program capable of executing the content of the above-mentioned embodiments, and the processor 102 can execute the program stored in the memory 104.
[0146] The processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 utilizes various interfaces and circuits to connect various components within the electronic device 200. It executes instructions, programs, code sets, or instruction sets stored in the memory 104, and accesses data stored in the memory 104 to perform various functions and process data within the electronic device 200. Optionally, the processor 102 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 102 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 102 and may be implemented separately via a communication chip.
[0147] The memory 104 may include a random access memory (RAM) or a read-only memory (ROM). The memory 104 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal 100 during use (such as a phone book, audio and video data, chat history data), etc.
[0148] The network module 106 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 106 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, etc. The network module 106 can communicate with various networks such as the Internet, an intranet, a wireless network, or communicate with other devices via a wireless network. The above-mentioned wireless network may include a cellular telephone network, a wireless local area network, or a metropolitan area network. For example, the network module 106 can exchange information with a base station.
[0149] The screen 108 can display interface content and can also be used to respond to touch gestures.
[0150] It should be noted that, in order to achieve more functions, the electronic device 200 can also protect more devices, for example, it can also protect a structured light sensor for collecting facial information or a camera for collecting irises, etc.
[0151] Please refer to Figure 12 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 1100 stores program code, which can be called by a processor to execute the method described in the above method embodiment.
[0152] Computer-readable storage medium 1100 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Alternatively, computer-readable storage medium 1100 may include a non-transitory computer-readable storage medium. Computer-readable storage medium 1100 may have storage space for program code 1110 for executing any of the method steps described above. These program codes may be read from or written to one or more computer program products. Program code 1110 may be compressed, for example, in a suitable form.
[0153] Based on the above-described moving object recognition method, according to one aspect of an embodiment of the present application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.
[0154] In summary, the embodiments of the present application provide a method, device, electronic device, storage medium, and computer program product or computer program for identifying a target moving object. Since only the moving objects to be identified that are screened out from the moving objects are identified, it is equivalent to a preliminary screening of the moving objects, which can eliminate the interference of other moving objects to a certain extent, thereby improving the accuracy of target moving object identification. At the same time, since only the moving objects to be identified that are screened out from the moving objects are identified, the number of identifications using preset identification rules is reduced, thereby improving the recognition efficiency of the target moving object.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying a moving target object, characterized in that: include: Get the target video frame; Performing motion detection on the target video frame to obtain a moving object in the target video frame; Based on the intra-frame features of the moving objects, the moving objects to be identified are screened from the moving objects; comprising: Obtaining a predicted aspect ratio of a target moving object in the target video frame and an actual aspect ratio of the moving object in the target video frame; and determining that the moving object is a moving object to be identified when the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object; or Acquiring color distribution parameters of the moving object; performing a symmetry detection on the moving object based on the color distribution parameters to obtain a symmetry detection result of the moving object; and determining that the moving object is a moving object to be identified when the symmetry detection result matches the predicted symmetry of the target moving object; Based on the recognition rules corresponding to the number of the moving objects to be recognized, a target moving object is recognized from the moving objects to be recognized.
2. The method according to claim 1, characterized in that The step of identifying a target moving object from the moving objects to be identified based on the identification rule corresponding to the number of the moving objects to be identified includes: When the number of the to-be-identified moving objects obtained by screening the moving objects is one, determining the to-be-identified moving object as a target moving object; or When the number of moving objects to be identified screened from the moving objects is greater than one, the images corresponding to each moving object to be identified are input into an object classifier so as to identify the target moving object from the moving objects to be identified through the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from sample video frames.
3. The method according to claim 1, characterized in that The number of the moving objects to be identified is greater than one, and identifying the target moving object from the moving objects to be identified based on the identification rules corresponding to the number of the moving objects to be identified includes: Obtaining an actual distance between each of the to-be-identified moving objects and a credible object in an adjacent video frame corresponding to the same video frame, wherein the credible object is the to-be-identified moving object corresponding to the to-be-identified moving object screened from the moving objects when the number of the to-be-identified moving objects is one; The images corresponding to the moving objects to be identified are input into the object classifier in order of the actual distances between the moving objects to be identified and the credible objects in the adjacent video frames on the same video frame from small to large, so as to identify the target moving object from the moving objects to be identified by the object classifier, wherein the object classifier is trained by sample moving objects with classification labels determined from sample video frames.
4. The method according to claim 1, wherein The acquiring of the target video frame comprises: Obtain target video image and video frame extraction frame rate; A target video frame is extracted from the target video image based on the video frame extraction frame rate.
5. The method according to claim 1, wherein The performing motion detection on the target video frame to obtain a moving object in the target video frame includes: Acquire a frame of a valid motion area in the target video frame; Motion detection is performed on the picture in the effective motion area to obtain the moving object in the picture in the effective motion area.
6. A target moving object recognition device, characterized in that: The device comprises: A target video frame acquisition module is used to acquire a target video frame; A motion detection module, configured to perform motion detection on the target video frame to obtain a moving object in the target video frame; A screening module, configured to screen the moving objects to be identified from the moving objects based on the intra-frame features of the moving objects, comprising: Obtaining a predicted aspect ratio of a target moving object in the target video frame and an actual aspect ratio of the moving object in the target video frame; and determining that the moving object is a moving object to be identified when the actual aspect ratio of the moving object matches the predicted aspect ratio of the target moving object; or Acquiring color distribution parameters of the moving object; performing a symmetry detection on the moving object based on the color distribution parameters to obtain a symmetry detection result of the moving object; and determining that the moving object is a moving object to be identified when the symmetry detection result matches the predicted symmetry of the target moving object; The recognition module is configured to recognize a target moving object from the moving objects to be recognized based on a recognition rule corresponding to the number of the moving objects to be recognized.
7. An electronic device, characterized in that: The method comprises a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 8.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, wherein when the program code is executed by a processor, the method according to any one of claims 1 to 8 is executed.
Citation Information
Patent Citations
Method and system for detecting balls
CN102148919A
Moving target detection and motion recognition method, device, terminal and storage medium
CN107786848A
Object motion direction identification method based on target detection
CN110070560A
Small target detection method and device, computer equipment and storage medium
CN111476064A