Object recognition device, object recognition method, and computer readable storage medium
By extracting the features and quality of each frame in the video clip in the object recognition device, dividing sub-segments that meet certain thresholds, and obtaining recognition results based on the features of these sub-segments, the problem of low recognition accuracy when the object is blocked is solved, and a more efficient recognition effect is achieved.
Patent Information
- Application Number
- CN202311446696.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, when the object is blocked, the recognition accuracy of video clips is low.
By extracting the features and quality of each frame in the video clip with respect to the object to be identified, sub-segments whose quality meets a certain threshold value, and obtaining recognition results based on the features of these sub-segments.
Improve the accuracy and speed of object recognition and reduce the demand for computing resources.
Smart Images

Figure CN119942394A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of object recognition, and in particular to an object recognition device, an object recognition method, and a computer-readable storage medium. Background Art
[0002] As we enter the information age, the application scope of technologies for identifying objects (eg, people, animals, objects, etc.) is becoming increasingly wide. Summary of the invention
[0003] A brief overview of the disclosure is given below in order to provide a basic understanding of certain aspects of the disclosure. However, it should be understood that this overview is not an exhaustive overview of the disclosure. It is not intended to identify the key or important parts of the disclosure, nor is it intended to limit the scope of the disclosure. Its purpose is simply to give certain concepts of the disclosure in a simplified form as a prelude to a more detailed description given later.
[0004] An object of the present disclosure is to provide an improved object recognition device, an object recognition method, and a computer-readable recording medium.
[0005] According to one aspect of the present disclosure, there is provided an object recognition apparatus, comprising: an object recognition apparatus, comprising: a feature extraction unit, configured to extract features of each of a plurality of frames included in a video clip about an object to be recognized; a quality estimation unit, configured to estimate the quality of each of the plurality of frames about the object to be recognized; a division unit, configured to divide the video clip into a plurality of sub-segments based on the quality of the plurality of frames; and an object recognition unit, configured to obtain a recognition result of the video clip based on features of frames of which the quality of each of the plurality of sub-segments is greater than or equal to a first predetermined threshold.
[0006] According to another aspect of the present disclosure, there is provided an object recognition method, comprising: extracting features about an object to be recognized from each of a plurality of frames included in a video clip; estimating the quality of each of the plurality of frames about the object to be recognized; dividing the video clip into a plurality of sub-segments based on the quality of the plurality of frames; and obtaining a recognition result of the video clip based on features of frames whose quality of each of the plurality of sub-segments is greater than or equal to a first predetermined threshold.
[0007] According to other aspects of the present disclosure, a computer program code and a computer program product for implementing the method according to the present disclosure are also provided, as well as a computer-readable storage medium having the computer program code for implementing the method according to the present disclosure recorded thereon.
[0008] Other aspects of the embodiments of the present disclosure are given in the following description, wherein the detailed description is used to fully disclose the preferred embodiments of the embodiments of the present disclosure without imposing limitations thereon. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The present disclosure may be better understood by referring to the detailed description given below in conjunction with the accompanying drawings, wherein the same or similar reference numerals are used throughout the drawings to represent the same or similar components. The accompanying drawings, together with the following detailed description, are included in and form a part of the present specification to further illustrate the preferred embodiments of the present disclosure and to explain the principles and advantages of the present disclosure. Among them:
[0010] Figure 1 is a block diagram showing a functional configuration example of an object recognition device according to an embodiment of the present disclosure;
[0011] Figure 2A and Figure 2B An example of a quality curve is shown;
[0012] FIG. 3A to FIG. 3C is a schematic diagram showing an example of sub-segment division;
[0013] Figure 4 is a schematic diagram illustrating an example of object recognition;
[0014] Figure 5 is a schematic diagram showing an example of a training process of a model that can be applied to an object recognition device;
[0015] Figure 6 A comparative example of the technology according to the present disclosure and the prior art is shown;
[0016] Figure 7 is a flowchart showing an example of the flow of an object recognition method according to an embodiment of the present disclosure; and
[0017] Figure 8 is a block diagram showing an example structure of a personal computer that can be employed in the embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the actual implementation are described in the specification. However, it should be understood that many implementation-specific decisions must be made in the process of developing any such actual implementation in order to achieve the developer's specific goals, such as meeting those constraints related to the system and business, and these constraints may vary from implementation to implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is only a routine task for those skilled in the art who benefit from the content of this disclosure.
[0019] It is also necessary to explain here that in order to avoid obscuring the present disclosure due to unnecessary details, only the device structure and / or processing steps closely related to the scheme according to the present disclosure are shown in the accompanying drawings, while other details that are not closely related to the present disclosure are omitted.
[0020] The embodiments according to the present disclosure are described in detail below with reference to the accompanying drawings.
[0021] First, refer to Figures 1 to 6 An implementation example of the object recognition device 100 according to an embodiment of the present disclosure is described. Figure 1 is a block diagram showing a functional configuration example of the object recognition device 100 according to an embodiment of the present disclosure. Figure 2A and Figure 2B An example of a quality curve is shown. FIG. 3A to FIG. 3C is a schematic diagram showing an example of sub-segment division. Figure 4 is a schematic diagram illustrating an example of object recognition. Figure 5 is a schematic diagram showing an example of a training process of a model that can be applied to an object recognition device. Figure 6 A comparative example of the technology according to the present disclosure and the prior art is shown.
[0022] like Figure 1 As shown, the object recognition device 100 according to an embodiment of the present disclosure may include a feature extraction unit 102 , a quality estimation unit 104 , a segmentation unit 106 and an object recognition unit 108 .
[0023] The feature extraction unit 102 may be configured to extract features of each frame of a plurality of frames included in a video segment (hereinafter, sometimes referred to as a "video segment to be identified" for ease of distinction) with respect to an object to be identified. For example, the feature extraction unit 102 may determine an object to be identified from a frame, and extract features of the object to be identified as features of the frame with respect to the object to be identified. For example, the feature extraction unit 102 may determine a target object (e.g., a pedestrian, a face, a vehicle, etc.) in a frame as an object to be identified, but is not limited thereto.
[0024] For example, the object to be identified may be a person, an object, an animal, a plant, etc. For example, in an example where the object to be identified is a person, the feature may be a person re-identification (ReID) feature, but is not limited thereto.
[0025] The quality estimation unit 104 may be configured to estimate the quality of each of the above-mentioned multiple frames with respect to the object to be identified (hereinafter referred to as "quality"). As an example, the quality of the frame may be estimated based on the clarity of the object to be identified in the frame. As another example, the quality of the frame may be estimated based on the occlusion rate of the object to be identified in the frame, so that the quality may be obtained more objectively. For example, the ratio of the area of the region where the object to be identified is occluded to the area of the region where the object to be identified is calculated as the occlusion rate. In addition, for example, the ratio of the area of the bounding box corresponding to the occluded region of the object to be identified to the area of the bounding box corresponding to the object to be identified may be calculated as the occlusion rate.
[0026] The dividing unit 106 may be configured to divide the video segment into a plurality of sub-segments based on the qualities of the plurality of frames.
[0027] The object recognition unit 108 may be configured to obtain a recognition result of the video segment based on the features of the frames whose quality of each of the multiple sub-segments divided by the division unit 106 is greater than or equal to a first predetermined threshold. For example, the first predetermined threshold may be set based on experience or obtained through a limited number of experiments.
[0028] Object recognition based on video clips can be applied to various fields, such as the field of surveillance. One object recognition technology based on video clips is to recognize each frame of the video clip, but this technology has low recognition accuracy when the object is blocked by people, obstacles, etc. As described above, the information processing device 100 according to an embodiment of the present disclosure can estimate the quality of the frame, and obtain the recognition result of the video clip based on the features of the frame whose quality is greater than or equal to the first predetermined threshold, thereby improving the recognition accuracy. In the following, for convenience, "frames whose quality is greater than or equal to the first predetermined threshold" are sometimes referred to as target frames.
[0029] In some examples, the object recognition unit 108 can obtain the features of each sub-segment based on the features of the frames whose quality is greater than or equal to the first predetermined threshold included in the sub-segment; obtain the features of the video segment based on the features of multiple sub-segments; and recognize the object to be recognized based on the features of the video segment to obtain the recognition result of the video segment. This method of recognizing the object to be recognized based on the features of the video segment can improve the recognition speed and reduce the computing resource requirements.
[0030] For example, the mean of the features of the plurality of sub-segments may be calculated as the feature of the video clip. In addition, for example, weights may be set for the sub-segments based on the quality of the target frame of the sub-segments, and the mean of the features of the plurality of sub-segments set with the weights (i.e., the weighted mean) may be calculated as the feature of the video clip.
[0031] In some examples, the object recognition unit 108 may identify the object to be recognized for each sub-segment based on features of frames included in the sub-segment whose quality is greater than or equal to a first predetermined threshold to obtain a recognition result of the sub-segment; and determine a recognition result of the video clip based on the recognition results of multiple sub-segments.
[0032] In a video clip, the object to be identified may change, for example, from the first object to be identified to the second object to be identified. In addition, for the same object to be identified, the posture in different frames may be different, for example, the front of the object to be identified is presented in some frames, while the back of the object to be identified is presented in other frames. In addition, for the same object to be identified, the lighting and corresponding background in different frames may be different. The above factors may lead to a decrease in recognition accuracy. As described above, the object recognition unit 108 can identify the object to be identified for each sub-segment, thereby further improving the recognition accuracy of each sub-segment, and further improving the accuracy of the recognition result of the video clip obtained thereby.
[0033] As an example, the object recognition unit 108 can determine the recognition result of the video clip based on the recognition results of multiple sub-segments by majority voting. For example, assuming that the video clip is divided into 5 sub-segments V1, V2, V3, V4, and V5, where the recognition results of 4 sub-segments V1, V2, V3, and V5 are object A, and the recognition result of one sub-segment V4 is object B, then object A can be determined as the recognition result of the video clip. As another example, the set of recognition results of multiple sub-segments can be determined as the recognition result of the video clip. For example, assuming that the video clip is divided into 5 sub-segments V1, V2, V3, V4, and V5, where the recognition results of 4 sub-segments V1, V2, V3, and V5 are object A, and the recognition result of one sub-segment V4 is object B, then [object A, object A, object A, object B, object A] can be determined as the recognition result of the video clip.
[0034] In some examples, the object recognition unit 108 may obtain the features of each sub-segment based on the features of multiple frames (i.e., multiple target frames) in the frames included in the sub-segment whose quality is greater than or equal to the first predetermined threshold, and recognize the object to be recognized based on the features of the sub-segment to obtain the recognition result of the sub-segment. By recognizing the object to be recognized according to each sub-segment, the object recognition speed can be further improved.
[0035] As an example, the object recognition unit 108 may obtain, for each sub-segment, an average of features of multiple target frames included in the sub-segment as the feature of the sub-segment.
[0036] As another example, the object recognition unit 108 may set a weight for the feature of each target frame in the multiple target frames included in the sub-segment based on the quality of the target frame, and obtain the mean of the multiple target frames with the weights (i.e., the weighted mean) as the feature of the sub-segment. In this way, the accuracy of the recognition result of the sub-segment can be further improved, thereby further improving the accuracy of the recognition result of the video clip.
[0037] For example, for a target frame, the higher the quality of the target frame, the greater the weight can be. For example, the quality of the target frame can be directly set as the weight of the target frame. In some examples, the quality of the target frame can be normalized, and the normalized quality can be set as the weight of the target frame. In addition, in some examples, the weight of the target frame can be set according to the range in which the quality of the target frame is located. For example, when the quality of the target frame is within a first range, a first weight corresponding to the first interval can be set for the target frame; when the quality of the target frame is within a second interval that does not overlap with the first interval and is greater than the first interval, a second weight corresponding to the second interval and greater than the first weight can be set for the target frame, and so on.
[0038] For example, the quality curve may be a curve drawn with the frame number as the horizontal axis and the quality as the vertical axis, such as Figure 2A In addition, for example, the quality curve may be a curve obtained by smoothing the above curve (for example, Gaussian smoothing), such as Figure 2B As shown. By smoothing, the adverse effect of quality estimation errors of some frames on the division can be reduced.
[0039] In some examples, the dividing unit 106 may determine a frame corresponding to a maximum value of a quality curve depicting the quality of multiple frames included in the video segment as a dividing point, and divide the video segment into multiple sub-segments based on the dividing point.
[0040] In some examples, such as Figure 3A As shown, the division unit 106 can determine the frames m1, m2, m3, m4, m5, m6, m7, m8, m9, m10, and m11 (m1 to m11 are positive integers) corresponding to the minimum values of the quality curve depicting the quality of the multiple frames included in the video clip as division points, and divide the video clip into multiple sub-segments based on the division points. Compared with the division method based on the maximum value, the similarity between the objects to be identified in the multiple frames included in the same sub-segment in the multiple sub-segments obtained by the division method based on the minimum value is higher, thereby further improving the accuracy of the recognition result of the sub-segment, and further improving the accuracy of the recognition result of the video clip.
[0041] The inventor of the present application has found through experiments that the similarity between the objects to be identified in the frames near the minimum value above the predetermined value K1 (greater than or equal to the predetermined value K1, K1>0) is relatively large. Therefore, the division points corresponding to the minimum values above the predetermined value (for example, a second predetermined threshold value that can be set based on experience or a limited number of experiments) can be removed, thereby reducing the number of sub-segments and increasing the number of frames included in the corresponding sub-segments, thereby further improving the accuracy of the recognition results of the sub-segments. For example, for Figure 3A In the example shown, the division points m1, m4, m7, and m8 corresponding to the minimum values greater than or equal to the second predetermined threshold can be removed to obtain Figure 3B For example, the second predetermined threshold value may be the same as or different from the first predetermined threshold value.
[0042] In addition, the inventor of the present application has found through experiments that the object to be identified is less likely to change between sub-segments that are close to each other. Therefore, sub-segments that are close to each other can be merged. For example, the division point corresponding to the smaller frame number among two adjacent division points whose distance between each other is less than or equal to the third predetermined threshold can be removed. For example, for Figure 3B In the example shown, the distance between adjacent partition points m2 and m3 is less than the third predetermined threshold, so the partition point m2 can be removed, thereby obtaining Figure 3C For example, a third predetermined threshold value may be set based on experience or a limited number of experiments. In some examples, only reference Figure 3B and Figure 3C One of the partition point removal operations described.
[0043] The inventor of the present application has found through experiments that the recognition accuracy of the sub-segment whose target frame number is greater than or equal to the predetermined value N1 (greater than or equal to the predetermined value N1, N1 is a positive integer) is greater than the recognition accuracy of the sub-segment whose target frame number is less than the predetermined value N1. Therefore, the recognition accuracy of the video segment can be further improved by not using the sub-segment whose target frame number is less than the predetermined value N1. For example, for Figure 3C In the example shown, the sub-segments S1, S2, S3, S4, S5, and S6 are obtained based on the division points m3, m5, m6, m9, m10, and m11. Since sub-segments S3 and S5 do not contain target frames, sub-segments S3 and S5 may not be used. In addition, since the number of target frames in sub-segment S6 is small, this sub-segment may not be used. For example Figure 3C An example of a target frame using sub-segments S1, S2, S4 is shown with shading in FIG.
[0044] In some examples, such as Figure 4As shown, the object recognition unit 108 can compare the features of the sub-segment (e.g., the mean value described above) with the features in the preset database, and recognize the object to be recognized based on the comparison result. For example, when the similarity between the features of the sub-segment and the features corresponding to object A in the preset database is greater than or equal to a predetermined value, "object A" can be determined as the recognition result of the sub-segment.
[0045] For example, the features of the sub-segment may be stored in a preset database based on the recognition result of the sub-segment, so that the data in the preset database may be expanded. For example, the features of the sub-segment may be stored in the preset database in association with the recognition result of the sub-segment.
[0046] For example, Figure 4 As shown, the features of each sub-segment of a plurality of video segments of known objects involved can be obtained in a manner similar to the above-mentioned manner of obtaining the features of each sub-segment of the video segment to be identified, and stored in a preset database for object identification. Figure 4 In the lower right part of , frames identified by the same curly brackets belong to the same subsegment. Figure 4 In the example, frames marked with “×” are frames whose quality is less than a predetermined value, and frames marked with “×” are not considered when obtaining the features of the sub-segment.
[0047] In some examples, the object recognition device can be implemented by machine learning. For example, a model can be jointly trained using a training image set labeled with a true value of the target object (e.g., an identifier of the target object) and a true value of the quality to obtain a pre-trained model. The feature extraction unit 102 can use the feature extraction layer of the pre-trained model (e.g., see below). Figure 5 The first feature extraction layer and the second feature extraction layer described in the foregoing description are used to extract features of each frame about the object to be recognized, and the quality estimation unit 104 can estimate the quality of each frame by using the quality estimation layer of the pre-trained model.
[0048] In some examples, the model can be trained using multi-task learning. Figure 5 An example of the training process of the model shown, which mainly includes three parts: a multi-task shared backbone network, a feature extraction module and a quality estimation module.
[0049] The backbone network may include a first feature extraction layer (not shown) for extracting common features of the following two: features of the target object in the training image (e.g., ReID features) and quality. The backbone network may be based on any deep learning network model, such as ResNet, Transformer, etc.
[0050] The feature extraction module can be used to extract the features of the target object in the training image based on the common features output by the backbone network. Figure 5 As shown, the feature extraction module may include a second feature extraction layer (e.g., a multi-layer fully connected layer or a convolutional layer) to further extract features from the common features output by the backbone network, and then use the final features as the features of the target object. In some examples, the feature extraction module may directly use the common features output by the backbone network as the features of the target object.
[0051] For example, the feature extraction module can be trained using only the recognition / classification loss function (see, for example, reference 1: Liao, Wentong, et al. "Triplet-based deep similarity learning for person re-identification." Proceedings of the IEEE International Conference on Computer Vision Workshops. 2017). In addition, for example, Figure 5 As shown, the recognition / classification loss function can also be used together with other auxiliary loss functions, such as the triplet loss function (see, for example, reference 1).
[0052] The quality estimation module may include a quality estimation layer for estimating the quality of the target object. In some examples, the quality estimation module uses a regression model so that the estimated quality is continuous. A mapping function (such as an exponential function or a sigmoid function) maps the result of the regression model estimation to a non-negative region to ensure that the estimated result has physical meaning. For example, the quality regression loss used when training the quality estimation module may be a mean square error loss.
[0053] although Figure 5 Not shown, the quality estimation module may include a network branch to further extract features from the common features output by the backbone network, and the quality estimation layer may estimate the quality based on the further extracted features.
[0054] In some examples, the video clip may be a video shot in real time. In another example, the video clip may be extracted in real time from a video shot in real time using a trajectory extraction method, which may be referred to as a "trajectory". According to the technology disclosed herein, the features of each frame may be automatically extracted and its quality estimated during trajectory tracking, and a quality curve of the motion trajectory may be obtained in real time using the estimated quality, and then a sub-segment with consistent appearance features and relatively high reliability may be automatically generated based on the quality curve. In addition, the features of the sub-segment may be automatically obtained for feature matching and recognition (e.g., identity authentication). The technology disclosed herein may be used to implement object tracking.
[0055] Figure 6 An example of comparison between the technology disclosed in the present invention and the prior art is shown, wherein the prior art is a technology for identifying each frame of a video clip or every few frames. Figure 6 It can be seen that for the first data set collected from the community, the second data set collected from the shopping mall, and the third data set collected from the supermarket, the recognition accuracy of the technology according to the present disclosure is improved by 6.26%, 15.87%, and 29.08% respectively compared with the other technologies. In addition, the technology according to the present disclosure can also perform high-precision recognition in crowded places (such as supermarkets).
[0056] The object recognition device according to the embodiment of the present disclosure has been described above. Corresponding to the embodiment of the object recognition device described above, the present disclosure further provides the following embodiment of the object recognition method.
[0057] Figure 7 FIG. 7 is a flow chart showing an example of the process of the object recognition method 700 according to an embodiment of the present disclosure. Figure 7 As shown, the object recognition method according to an embodiment of the present disclosure may start at a start step S701 and end at an end step S712, and may include a feature extraction step S702, a quality estimation step S704, a segmentation step S706 and an object recognition step S708.
[0058] In the feature extraction step S702, features of each of the multiple frames included in the video clip about the object to be identified can be extracted. For example, the feature extraction step S702 can be implemented by the feature extraction unit 102 described above, so the specific details can be found in the above description, and only a brief description will be given below.
[0059] In the quality estimation step S704, the quality of each of the plurality of frames with respect to the object to be identified may be estimated. For example, the quality estimation step S704 may be implemented by the quality estimation unit 104 described above, so the specific details may refer to the above description, and only a brief description will be given below.
[0060] In the dividing step S706, the video clip may be divided into multiple sub-segments based on the quality of the multiple frames. For example, the dividing step S706 may be implemented by the dividing unit 106 described above, so the specific details can be found in the above description, and only a brief description will be given below.
[0061] In the object recognition step S708, the recognition result of the video segment can be obtained based on the features of the frames whose quality of each of the multiple sub-segments obtained by the division step S706 is greater than or equal to the first predetermined threshold. For example, the object recognition step S708 can be implemented by the object recognition unit 108 described above, so the specific details can be found in the above description, and only a brief description will be given below.
[0062] Similar to the information processing device 100 according to an embodiment of the present disclosure, the information processing method according to an embodiment of the present disclosure can estimate the quality of the frame, and obtain the recognition result of the video clip based on the characteristics of the frame whose quality is greater than or equal to the first predetermined threshold, thereby improving the recognition accuracy.
[0063] In some examples, in the object recognition step S708, for each sub-segment, based on the features of the frames whose quality is greater than or equal to the first predetermined threshold included in the sub-segment, the object to be recognized can be recognized to obtain the recognition result of the sub-segment; and the recognition result of the video segment can be determined based on the recognition results of the multiple sub-segments. In this way, the recognition accuracy of each sub-segment can be further improved, and the accuracy of the recognition result of the video segment obtained thereby can be further improved.
[0064] As an example, in the object recognition step S708, the recognition result of the video segment may be determined based on the recognition results of the multiple sub-segments by majority voting.
[0065] In some examples, in the object recognition step S708, for each sub-segment, the features of the sub-segment can be obtained based on the features of multiple frames (i.e., multiple target frames) in the frames included in the sub-segment whose quality is greater than or equal to the first predetermined threshold, and the object to be recognized is recognized based on the features of the sub-segment to obtain the recognition result of the sub-segment. By recognizing the object to be recognized according to each sub-segment, the object recognition speed can be further improved.
[0066] As an example, for each sub-segment, an average value of features of multiple target frames included in the sub-segment may be obtained as the feature of the sub-segment.
[0067] As another example, for each sub-segment, a weight can be set for the feature of each target frame in the multiple target frames included in the sub-segment based on the quality of the target frame, and the average of the multiple target frames with weights is obtained as the feature of the sub-segment. In this way, the accuracy of the recognition result of the sub-segment can be further improved, and then the accuracy of the recognition result of the video clip can be further improved.
[0068] In some examples, in the division step S706, the frame corresponding to the minimum value of the quality curve describing the quality of the multiple frames included in the video clip can be determined as a division point, and the video clip is divided into multiple sub-segments based on the division point. Compared with the division method based on the maximum value, the similarity between the objects to be identified in the multiple frames included in the same sub-segment in the multiple sub-segments obtained by the division method based on the minimum value is higher, thereby further improving the accuracy of the recognition result of the sub-segment, and further improving the accuracy of the recognition result of the video clip.
[0069] In some examples, the division points corresponding to the minimum values above a predetermined value (for example, a second predetermined threshold value that can be set based on experience or a limited number of experiments) can be removed, so that the number of sub-segments can be reduced and the number of frames included in the corresponding sub-segments can be increased, thereby further improving the accuracy of the recognition results of the sub-segments.
[0070] In some examples, sub-segments that are close to each other may be merged. For example, the division point corresponding to the smaller frame number among two adjacent division points whose distance between each other is less than or equal to the third predetermined threshold may be removed.
[0071] In some examples, the object recognition method 700 can be implemented by machine learning. For example, a model can be jointly trained using a training image set labeled with the true value of the target object and the true value of the quality to obtain a pre-trained model. In the feature extraction step S702, the feature extraction layer of the pre-trained model can be used to extract the features of each frame about the object to be recognized, and in the quality estimation step S704, the quality estimation layer of the pre-trained model can be used to estimate the quality of each frame.
[0072] In some examples, in the object recognition step S708, the features of the sub-segment (eg, the mean value described above) may be compared with features in a preset database, and the object to be recognized may be recognized based on the comparison result.
[0073] For example, the technology disclosed herein can be applied to security, smart shopping malls, smart communities, smart factories, scientific research (such as primate gorilla motion trajectory recognition), smart agriculture (such as cattle motion trajectory recognition in farms), and other fields, but is not limited to these.
[0074] It should be pointed out that although the functional configuration and operation of the object recognition device and the object recognition method according to the embodiments of the present disclosure are described above, this is only an example and not a limitation, and those skilled in the art may modify the above embodiments according to the principles of the present disclosure, for example, the functional modules and operations in each embodiment may be added, deleted or combined, and such modifications shall fall within the scope of the present disclosure.
[0075] In addition, it should be pointed out that the method embodiment here corresponds to the above-mentioned device embodiment. Therefore, for the contents not described in detail in the method embodiment, please refer to the description of the corresponding parts in the device embodiment, and the description will not be repeated here.
[0076] In addition, the present disclosure also provides a storage medium and a program product. It should be understood that the machine-executable instructions in the storage medium and the program product according to the embodiments of the present disclosure can also be configured to perform the above-mentioned object recognition method, so the contents not described in detail here can refer to the description of the previous corresponding parts, and will not be repeated here.
[0077] Accordingly, the storage medium for carrying the program product including the machine executable instructions is also included in the disclosure of the present invention, including but not limited to a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, and the like.
[0078] In addition, it should be noted that the above series of processes and devices can also be implemented by software and / or firmware. In the case of being implemented by software and / or firmware, from a storage medium or a network to a computer with a dedicated hardware structure, such as Figure 8 The general-purpose personal computer 1000 shown has programs constituting the software installed therein, and when the various programs are installed therein, the computer can execute various functions and the like.
[0079] exist Figure 8 In the embodiment, a central processing unit (CPU) 1001 executes various processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 to a random access memory (RAM) 1003. In the RAM 1003, data required when the CPU 1001 executes various processes and the like is also stored as needed.
[0080] The CPU 1001, the ROM 1002, and the RAM 1003 are connected to one another via a bus 1004. To the bus 1004, an input / output interface 1005 is also connected.
[0081] The following components are connected to the input / output interface 1005: an input device 1006 including a keyboard, a mouse, etc.; an output device 1007 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage device 1008 including a hard disk, etc.; and a communication device 1009 including a network interface card such as a LAN card, a modem, etc. The communication device 1009 performs communication processing via a network such as the Internet.
[0082] A drive 1010 is also connected to the input / output interface 1005 as needed. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1010 as needed so that a computer program read therefrom is installed into the storage device 1008 as needed.
[0083] In the case where the above-described series of processing is realized by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1011 .
[0084] It should be understood by those skilled in the art that such storage media is not limited to Figure 8 The removable medium 1011 shown has a program stored therein and is distributed separately from the device to provide the program to the user. Examples of the removable medium 1011 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including minidiscs (MD) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be a ROM 1002, a hard disk included in the storage device 1008, or the like, in which the program is stored and distributed to the user together with the device containing them.
[0085] The preferred embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is certainly not limited to the above examples. Those skilled in the art may obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.
[0086] For example, a plurality of functions included in one unit in the above embodiments may be implemented by separate devices. Alternatively, a plurality of functions implemented by a plurality of units in the above embodiments may be implemented by separate devices, respectively. In addition, one of the above functions may be implemented by a plurality of units. Needless to say, such a configuration is included in the technical scope of the present disclosure.
[0087] In this specification, the steps described in the flowchart include not only the processing performed in time series in the order described, but also the processing performed in parallel or individually rather than necessarily in time series. In addition, even in the steps processed in time series, it goes without saying that the order can be appropriately changed.
[0088] Additionally, the technology according to the present disclosure may also be configured as follows.
[0089] Solution 1. An object recognition device, comprising:
[0090] A feature extraction unit, configured to extract features of each frame of a plurality of frames included in the video clip about an object to be identified;
[0091] a quality estimation unit configured to estimate the quality of each frame of the plurality of frames with respect to the object to be identified;
[0092] a dividing unit configured to divide the video segment into a plurality of sub-segments based on the qualities of the plurality of frames; and
[0093] The object recognition unit is configured to obtain a recognition result of the video segment based on features of frames of each of the multiple sub-segments whose quality is greater than or equal to a first predetermined threshold.
[0094] Solution 2. An object recognition device according to Solution 1, wherein obtaining the recognition result of the video clip comprises: for each of the multiple sub-segments, based on the features of the frames included in the sub-segment and whose quality is greater than or equal to the first predetermined threshold, identifying the object to be recognized to obtain the recognition result of the sub-segment; and determining the recognition result of the video clip based on the recognition results of the multiple sub-segments.
[0095] Solution 3. The object recognition device according to Solution 2, wherein the division unit determines a frame corresponding to a minimum value of a quality curve depicting the quality of the multiple frames as a division point, and divides the video clip into multiple sub-segments based on the division point.
[0096] Solution 4. An object recognition device according to Solution 3, wherein the division unit removes the following division points from the division points: division points corresponding to smaller frame numbers among two adjacent division points whose spacing is less than or equal to a third predetermined threshold; and / or division points corresponding to a minimum value greater than or equal to a second predetermined threshold.
[0097] Solution 5. The object recognition device according to Solution 1, wherein the feature extraction unit extracts the feature using a feature extraction layer of a pre-trained model, and the quality estimation unit estimates the quality using a quality estimation layer of the pre-trained model,
[0098] The pre-trained model is obtained by jointly training the model using a training image set labeled with the true value of the target object and the true value of the quality.
[0099] Solution 6. The object recognition device according to Solution 2, wherein the object recognition unit is configured to:
[0100] For each of the multiple subsegments: obtaining an average of features of multiple frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment as the feature of the subsegment, or setting a weight for the feature of each frame among the multiple frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment, and obtaining an average of the features of the multiple frames with the weights set as the feature of the subsegment; and
[0101] The object to be identified is identified based on the features of the sub-segment to obtain an identification result of the sub-segment.
[0102] Solution 7. According to the object recognition device of Solution 6, identifying the object to be recognized based on the feature of the sub-segment comprises: comparing the feature of the sub-segment with the feature in a preset database, and identifying the object to be recognized based on the comparison result,
[0103] The object recognition device stores the features of the sub-segment in the preset database based on the recognition result of the sub-segment.
[0104] Solution 8. The object recognition device according to Solution 2, wherein the object recognition unit determines the recognition result of the video segment based on the recognition results of the multiple sub-segments by majority voting.
[0105] Solution 9. An object recognition device according to any one of Solutions 1 to 8, wherein the video segment is extracted in real time from a video shot in real time using a trajectory extraction method.
[0106] Solution 10. An object recognition device according to any one of Solutions 1 to 8, wherein the quality of each frame is represented by the occlusion rate of the object to be recognized in the frame.
[0107] Solution 11. A method for object recognition, comprising:
[0108] Extracting features of each frame of a plurality of frames included in the video clip about an object to be identified;
[0109] estimating the quality of each frame of the plurality of frames with respect to the object to be identified;
[0110] dividing the video segment into a plurality of sub-segments based on the quality of the plurality of frames; and
[0111] The recognition result of the video segment is obtained based on the features of the frames whose quality of each sub-segment in the plurality of sub-segments is greater than or equal to a first predetermined threshold.
[0112] Scheme 12. An object recognition method according to Scheme 11, wherein obtaining the recognition result of the video clip comprises: for each of the multiple sub-segments, based on the features of the frames included in the sub-segment and whose quality is greater than or equal to the first predetermined threshold, identifying the object to be recognized to obtain the recognition result of the sub-segment; and determining the recognition result of the video clip based on the recognition results of the multiple sub-segments.
[0113] Scheme 13. An object recognition method according to Scheme 12, wherein dividing the video segment into multiple sub-segments based on the quality of the multiple frames includes: determining the frame corresponding to the minimum value of the quality curve depicting the quality of the multiple frames as a division point, and dividing the video segment into multiple sub-segments based on the division point.
[0114] Scheme 14. An object recognition method according to Scheme 13, wherein the following division points are removed from the division points: division points corresponding to smaller frame numbers among two adjacent division points whose spacing is less than or equal to a third predetermined threshold; and / or division points corresponding to a minimum value greater than or equal to a second predetermined threshold.
[0115] Solution 15. The object recognition method according to Solution 14, wherein the features are extracted using a feature extraction layer of a pre-trained model, and the quality is estimated using a quality estimation layer of the pre-trained model.
[0116] The pre-trained model is obtained by jointly training the model using a training image set labeled with the true value of the target object and the true value of the quality.
[0117] Solution 16. The object recognition method according to Solution 12, wherein obtaining the recognition result of the sub-segment comprises:
[0118] Obtaining an average of features of a plurality of frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment as the feature of the subsegment, or setting a weight for the feature of each frame among a plurality of frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment, and obtaining an average of the features of the plurality of frames with the weights set as the feature of the subsegment; and
[0119] The object to be identified is identified based on the features of the sub-segment to obtain an identification result of the sub-segment.
[0120] Solution 17. According to the object recognition method of Solution 16, identifying the object to be recognized based on the feature of the sub-segment comprises: comparing the feature of the sub-segment with the feature in a preset database, and identifying the object to be recognized based on the comparison result,
[0121] The object recognition method stores the features of the sub-segments in the preset database based on the recognition results of the sub-segments.
[0122] Solution 18. The object recognition method according to Solution 12, wherein the recognition result of the video segment is determined based on the recognition results of the multiple sub-segments by majority voting.
[0123] Solution 19. An object recognition method according to any one of Solutions 11 to 18, wherein the video clip is extracted in real time from a video shot in real time using a trajectory extraction method.
[0124] Embodiment 20. A computer-readable storage medium storing instructions, wherein when the instructions are executed by a computer, the computer performs the object recognition method according to any one of embodiments 11 to 19.
Claims
1. An object recognition device, comprising: A feature extraction unit, configured to extract features of each frame of a plurality of frames included in the video clip about an object to be identified; a quality estimation unit configured to estimate the quality of each frame of the plurality of frames with respect to the object to be identified; a dividing unit configured to divide the video segment into a plurality of sub-segments based on the quality of the plurality of frames; as well as The object recognition unit is configured to obtain a recognition result of the video segment based on features of frames of each of the multiple sub-segments whose quality is greater than or equal to a first predetermined threshold.
2. The object recognition device according to claim 1, wherein: Obtaining the recognition result of the video clip includes: for each of the multiple sub-segments, based on the characteristics of the frames included in the sub-segment and whose quality is greater than or equal to the first predetermined threshold, identifying the object to be identified to obtain the recognition result of the sub-segment; and determining the recognition result of the video clip based on the recognition results of the multiple sub-segments.
3. The object recognition device according to claim 2, wherein: The dividing unit determines a frame corresponding to a minimum value of a quality curve describing the quality of the plurality of frames as a dividing point, and divides the video segment into a plurality of sub-segments based on the dividing point.
4. The object recognition device according to claim 3, wherein: The division unit removes the following division points from the division points: division points corresponding to smaller frame numbers among two adjacent division points whose spacing is less than or equal to a third predetermined threshold; and / or division points corresponding to a minimum value greater than or equal to a second predetermined threshold.
5. The object recognition device according to claim 1, wherein: The feature extraction unit extracts the feature using a feature extraction layer of a pre-trained model, and the quality estimation unit estimates the quality using a quality estimation layer of the pre-trained model, The pre-trained model is obtained by jointly training the model using a training image set labeled with the true value of the target object and the true value of the quality.
6. The object recognition device according to claim 2, wherein: The object recognition unit is configured to: For each of the multiple subsegments: obtaining an average of features of multiple frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment as the feature of the subsegment, or setting a weight for the feature of each frame among the multiple frames whose quality is greater than or equal to the first predetermined threshold among the frames included in the subsegment, and obtaining an average of the features of the multiple frames with the weights set as the feature of the subsegment; as well as The object to be identified is identified based on the features of the sub-segment to obtain an identification result of the sub-segment.
7. The object recognition device according to claim 6, identifying the object to be recognized based on the feature of the sub-segment comprises: The feature of the sub-segment is compared with the features in the preset database, and the object to be identified is identified based on the comparison result, The object recognition device stores the features of the sub-segment in the preset database based on the recognition result of the sub-segment.
8. The object recognition device according to any one of claims 1 to 7, wherein: The video clips are extracted in real time from the video shot in real time by using a trajectory extraction method.
9. A method for object recognition, comprising: Extracting features of each frame of a plurality of frames included in the video clip about an object to be identified; estimating the quality of each frame of the plurality of frames with respect to the object to be identified; dividing the video segment into a plurality of sub-segments based on the quality of the plurality of frames; as well as The recognition result of the video segment is obtained based on the features of the frames whose quality of each sub-segment in the plurality of sub-segments is greater than or equal to a first predetermined threshold. 10 . A computer-readable storage medium storing instructions, wherein when the instructions are executed by a computer, the computer is caused to perform the object recognition method according to claim 9 .