Behavior recognition method and device, electronic equipment, vehicle and storage medium
By combining the objects in pedestrian behavior recognition and combining multiple scale description labels to generate behavior labels, the problem of low pedestrian behavior recognition accuracy in the prior art is solved, and higher behavior recognition accuracy and cost reduction are achieved.
Patent Information
- Application Number
- CN202311697992.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, pedestrian behavior recognition relies on manual design features and machine learning algorithms, with low detection accuracy and sensitivity to lighting and occlusion, resulting in low accuracy in behavior classification.
By tracking and identifying the objects in the video to be identified, the object's enclosing box and numbered label are obtained, the object's description labels under multiple scales are combined, the behavior labels are generated, and the behavior labels of the same numbered label are associated to obtain the behavior recognition results.
It avoids multiple iterative optimizations of sample data acquisition, labeling, and model training, reduces the cost of image recognition and improves the accuracy of behavior recognition.
Smart Images

Figure CN120148098A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a behavior recognition method, apparatus, electronic device, vehicle, and storage medium. Background Art
[0002] Pedestrian detection is an important research direction in the field of computer vision. In the field of autonomous vehicle driving, pedestrian behavior detection can identify and predict pedestrian actions, help autonomous driving planning and control make decisions, and ensure the driving safety in scenarios involving vulnerable road users (VRU). In the commercial field, pedestrian detection can help merchants understand customers' behavior habits, such as monitoring customer flow in shopping malls, monitoring customers' shopping behavior in supermarkets, etc., so as to optimize business strategies. In short, video pedestrian detection has a wide range of applications in many fields, which can help people better understand and control the surrounding environment, and improve the efficiency and safety of life and work. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a behavior recognition method, apparatus, electronic device, vehicle, and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a behavior recognition method, including:
[0005] Performing tracking and recognition on an object in a video to be recognized to obtain a bounding box of the object and a numbered label for each object;
[0006] Combining description labels of the object in a target video frame of the video to be recognized at multiple scales to obtain a behavior label corresponding to the object in the target video frame, where the description label is added to the object covered by the expanded bounding box after expanding the bounding box according to scales of different ratios;
[0007] Associating the behavior labels of the object with the same numbered label in each target video frame to obtain a behavior recognition result of the object corresponding to each numbered label.
[0008] According to a second aspect of an embodiment of the present disclosure, there is provided a behavior recognition apparatus, including:
[0009] A recognition module, configured to perform tracking and recognition on an object in a video to be recognized to obtain a bounding box of the object and a numbered label for each object;
[0010] A combination module, configured to combine the description tags of the object in the target video frame of the video to be recognized at multiple scales to obtain a behavior tag corresponding to the object in the target video frame, where the description tag is added to the object covered by the expanded bounding box after expanding the bounding box according to scales of different ratios;
[0011] An association module, configured to associate the behavior tags of the object in each target video frame with the same numbered tag to obtain a behavior recognition result of the object corresponding to each numbered tag.
[0012] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0013] A processor;
[0014] A memory for storing executable instructions executable by the processor;
[0015] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspect.
[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0017] According to a fifth aspect of the embodiments of the present disclosure, there is provided a chip, including a processor and an interface; the processor is used to read instructions to execute the method according to any one of the first aspect.
[0018] According to a sixth aspect of the embodiments of the present disclosure, there is provided a vehicle, including:
[0019] A processor;
[0020] A memory for storing executable instructions executable by the processor;
[0021] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspect.
[0022] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0023] Track and identify the objects in the video to be recognized, and obtain the bounding boxes of the objects and the numbered labels of each object; combine the description labels of the objects in the target video frames of the video to be recognized at multiple scales to obtain the behavior labels corresponding to the objects in the target video frames, where the description labels are added to the objects covered by the expanded bounding boxes after expanding the bounding boxes according to different scales; associate the behavior labels of the objects with the same numbered label in each target video frame to obtain the behavior recognition result of each object corresponding to the numbered label. In this way, after expanding according to different scales, description labels are added to the objects covered by the expanded bounding boxes, and then the description labels are combined to obtain behavior labels, which can avoid multiple rounds of iterative optimization of sample data collection and annotation and model training, thereby reducing the cost of image recognition. At the same time, more description labels of the object can be obtained after expanding according to different scales, and the behavior labels of the objects with the same numbered label in each target video frame are associated to improve the accuracy of behavior recognition.
[0024] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0026] Figure 1 It is a flowchart of a behavior recognition method shown according to an exemplary embodiment.
[0027] Figure 2 It is an implementation shown according to an exemplary embodiment Figure 1 of step S12 in the flowchart.
[0028] Figure 3 It is another implementation shown according to an exemplary embodiment Figure 1 of step S12 in the flowchart.
[0029] Figure 4 It is another implementation shown according to an exemplary embodiment Figure 1 of step S12 in the flowchart.
[0030] Figure 5 It is an implementation shown according to an exemplary embodiment Figure 1 of step S11 in the flowchart.
[0031] Figure 6 It is a block diagram of a behavior recognition device shown according to an exemplary embodiment.
[0032] Figure 7 It is a block diagram of a device for behavior recognition shown according to an exemplary embodiment.
[0033] Figure 8 It is a block diagram of another device for behavior recognition shown according to an exemplary embodiment.
[0034] Figure 9 It is a schematic diagram of a functional block diagram of a vehicle shown according to an exemplary embodiment. Detailed implementation manners
[0035] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0036] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0037] Before introducing a behavior recognition method, device, electronic device, and storage medium provided by the present disclosure, the technical defects existing in the related scenarios are described. Behavior recognition mainly relies on manually designed features, such as Haar features, and machine learning algorithms, such as SVM classifiers. However, there are problems with low detection accuracy in feature extraction and behavior classification, and behavior recognition is sensitive to factors such as light and occlusion, which also results in low accuracy of behavior classification.
[0038] Behavior recognition based on deep learning can automatically learn features from data and has better robustness and accuracy. For example, video pedestrian behavior detection based on a convolutional neural network (CNN), such as the convolutional neural network can be Faster R-CNN, YOLO, etc. However, the convolutional neural network model is highly dependent on data, and the models trained based on open-source datasets often have a large deviation from the behaviors in specific business scenarios. In addition, the cost of collecting and annotating sample data in special business occasions is relatively large, and the model training requires multiple rounds of iterative optimization, which is time-consuming and laborious, and also increases the cost of behavior recognition.
[0039] In view of this, the present disclosure provides a behavior recognition method, aiming to avoid the costs of sample data collection and annotation and multiple rounds of iterative optimization of model training, while improving the accuracy of behavior recognition. Furthermore, it improves the accuracy of, for example, mobile tracking and monitoring.
[0040] Figure 1 is a flowchart of a behavior recognition method shown according to an exemplary embodiment, as Figure 1 shown, the behavior recognition method includes the following steps.
[0041] In step S11, the objects in the video to be recognized are tracked and recognized to obtain the bounding boxes of the objects and the numbered labels of each object.
[0042] In an implementation manner of the present disclosure, multiple objects in the video to be recognized can be tracked by means of, for example, BoT-SORT to detect and track all objects in the scene, while adding and retaining a unique numbered label for each object, and the spatio-temporal trajectories of multiple objects in the video stream can be detected and estimated for multiple objects. Furthermore, the motion trajectories of each object can be tracked in the sequence of the entire video to be recognized.
[0043] In the embodiments of the present disclosure, before tracking and recognizing the objects in the video to be recognized, the obtained video to be recognized can also be sorted and spliced. The video images collected by sensing modules such as cameras are collected and saved separately according to a preset duration. The video can be spliced according to the collection timestamps to obtain a continuous video to be recognized.
[0044] In the embodiments of the present disclosure, tracking and recognizing the objects in the video to be recognized can also obtain the attribute information of the target objects. For example, the attribute information can be the object category, and the object category can include people, immovable objects, backpacks, animals, handbags, bicycles, etc.
[0045] In the embodiments of the present disclosure, tracking and recognizing the objects in the video to be recognized can combine the advantages of motion and surface information, add camera motion compensation and a more accurate Kalman filter state vector, so as to improve the tracking performance of objects between video frames in the video.
[0046] In step S12, the description labels of the objects in the target video frames of the video to be recognized at multiple scales are combined to obtain the behavior labels of the objects corresponding to the target video frames.
[0047] Wherein, the description labels are added to the objects covered by the expanded bounding boxes after the bounding boxes are expanded according to different scales.
[0048] In embodiments of the present disclosure, scales with different ratios can expand the bounding box in the length or width direction of the bounding box, or simultaneously in the length and width directions of the bounding box, so that the expanded bounding box can include not only the target object covered by the unexpanded bounding box but also the objects around the target object compared with the unexpanded bounding box. Thus, compared with the unexpanded bounding box, description labels of the objects around the target object can be obtained. After combining the description labels, behavior labels including more behavior features can be obtained, thereby improving the accuracy of behavior recognition.
[0049] For example, for any object in any target video frame, description labels can be added to the object covered by the unexpanded bounding box. It should be noted that the unexpanded bounding box can be understood as expanding the bounding box at a ratio of 1:1. Furthermore, the bounding box is expanded at a scale of 1:2 in the length direction of the bounding box and at a scale of 1:2 in the width direction of the bounding box. Then, description labels are added to the objects covered by the bounding box expanded at a scale of 1:2 in the length direction and description labels are added to the objects covered by the bounding box expanded at a scale of 1:2 in the width direction. Thus, description labels of the objects covered by the bounding box expanded at scales with different ratios are obtained, and the description labels are respectively classified and combined by parts of speech, for example, classified and combined as subject nouns, verbs, and objects respectively, to obtain the behavior labels of the object in the target video frame. By way of example, the description labels in the video to be recognized mainly include descriptions of people and objects. So the description labels related to people are extracted as people and male, the description labels related to verbs are sitting and squatting, and the corresponding description labels of the objects are the nouns warehouse, cabinet, stool, and cabinet. After combination, it can be obtained that the male is sitting on the chair beside the warehouse cabinet.
[0050] In some embodiments of the present disclosure, since most are for human behavior recognition, a rectangular bounding box is usually added with the height direction of a person as the length direction of the bounding box and the left - right direction of a person as the width direction of the bounding box.
[0051] In one implementation, since usually in the process of communication or interaction, the objects communicating or interacting with a person are usually in the left - right direction of the person. Therefore, scales with different ratios can be the scales for expanding the bounding box in the left - right direction of a person (the width direction of the bounding box). For example, for any object in any target video frame, description labels can be added to the object covered by the unexpanded bounding box, and the bounding box is expanded at scales of 1:2 and 1:1.5 in the width direction of the bounding box. In this way, the behavior of a person can be recognized more accurately.
[0052] In step S13, associate the behavior labels of the object in each of the target video frames with the same numbered label to obtain the behavior recognition result of the object corresponding to each numbered label.
[0053] In the embodiments of the present disclosure, there may be objects with the same numbered label in different target video frames. Therefore, there may be multiple behavior labels for the object with the same numbered label in different target video frames. Therefore, by associating the behavior labels of the same object in different target video frames with the numbered label, one or more behavior labels of the same object in the video to be recognized can be obtained. If the behavior labels of the same object in different target video frames are the same, there will be only one behavior in the finally obtained behavior recognition result. If the behavior labels of the same object in different target video frames are different, there will be multiple behaviors in the finally obtained behavior recognition result.
[0054] The above technical solution performs tracking and recognition on the objects in the video to be recognized to obtain the bounding boxes of the objects and the numbered label of each object; combines the description labels of the objects in the target video frames of the video to be recognized at multiple scales to obtain the behavior labels corresponding to the objects in the target video frames, where the description labels are added to the objects covered by the expanded bounding boxes after expanding the bounding boxes according to different scales; associates the behavior labels of the objects in each of the target video frames with the same numbered label to obtain the behavior recognition result of the object corresponding to each numbered label. In this way, after expanding according to different scales, description labels are added to the objects covered by the expanded bounding boxes, and then the description labels are combined to obtain the behavior labels, which can avoid sample data collection and annotation, and multiple rounds of iterative optimization of model training, thereby reducing the cost of image recognition. At the same time, more description labels of the object can be obtained after expanding according to different scales, and the behavior labels of the objects in each target video frame with the same numbered label are associated, improving the accuracy of behavior recognition.
[0055] Optionally, as shown in Figure 2 In step S12, the combining the description labels of the objects in the target video frames of the video to be recognized at multiple scales to obtain the behavior labels corresponding to the objects in the target video frames includes:
[0056] In step S121, extract the first video frame from the video to be recognized according to a preset frame extraction interval length.
[0057] In the embodiments of the present disclosure, one frame is extracted as the first video frame every preset frame extraction interval length. The preset frame extraction interval length can be determined according to the total length of the video to be recognized. For example, if there are 30 frames in one second of the video to be recognized, one frame can be extracted every 5 frames as the first video frame.
[0058] In the embodiments of the present disclosure, the preset frame extraction interval length is positively correlated with the total length of the video to be recognized. The longer the total length of the video to be recognized, the longer the preset frame extraction interval length. The frame extraction interval length range can also be set according to the total length of the video to be recognized. For example, there is a one-to-one correspondence between each preset frame extraction interval length and a total length range. For example, the total length range corresponding to the preset frame extraction interval length of 5 seconds is 30 to 100 seconds, and the total length range corresponding to the preset frame extraction interval length of 10 seconds is 100 to 200 seconds.
[0059] In step S122, after expanding the bounding boxes in the first video frame according to different scales, description labels are added to the objects covered by the expanded bounding boxes in the first video frame respectively.
[0060] In the embodiments of the present disclosure, the implementation manner in step S12 can be realized to expand the bounding box of each object in the first video frame.
[0061] In the embodiments of the present disclosure, it can be automated tag addition based on tags. Description labels are added to the objects covered by each bounding box. It can be understood that after expanding the bounding boxes in the first video frame according to different scales, each object in each frame will be covered by multiple bounding boxes.
[0062] Furthermore, the description labels added to each bounding box are screened. For example, screening is performed based on confidence. The description labels with a confidence higher than or equal to the preset confidence threshold are retained, and the description labels with a confidence lower than the preset confidence threshold are excluded. And based on the principle of voting, a tag result is obtained from the high-confidence description labels. The tag result can include multiple description labels corresponding to each bounding box.
[0063] In step S123, the description labels of the object in the first video frame at multiple scales are combined to obtain the behavior label corresponding to the object in the first video frame.
[0064] In the embodiments of the present disclosure, each description tag in the tag result can be combined. For example, they can be combined in the subject-predicate-object manner to obtain multiple candidate behavior tags of the object in the first video frame. Then, the multiple candidate behavior tags are screened for rationality and confidence to obtain a behavior tag for each object in the first video frame. In this way, after expanding according to different scales, description tags are added to the objects covered by the expanded bounding box, and then the description tags are combined to obtain behavior tags.
[0065] In this way, it is possible to avoid multiple rounds of iterative optimization of sample data collection, annotation, and model training, thereby not only reducing the cost of image recognition, but also quickly forming a closed loop for behavior recognition, reducing the cost of storing samples, etc., and at the same time avoiding the problem of low accuracy caused by the recognition of a single bounding box.
[0066] In step S124, it is determined whether the behavior tags of adjacent first video frames are the same, where the adjacent first video frames are two video frames on both sides of the same frame extraction interval length.
[0067] For example, for 30 frames of images in any second, when one frame is extracted every 5 frames to obtain the first video frame, the 1st, 6th, 11th, 16th, 21st, and 26th frames are the extracted first video frames. Then, the 1st and 6th frames are video frames on both sides of the same frame extraction interval length, so the 1st and 6th frames are adjacent first video frames. Similarly, the 6th and 11th frames are video frames on both sides of the same frame extraction interval length, so the 6th and 11th frames are adjacent first video frames. The 11th and 16th frames are video frames on both sides of the same frame extraction interval length, so the 11th and 16th frames are adjacent first video frames. The 16th and 21st frames are video frames on both sides of the same frame extraction interval length, so the 16th and 21st frames are adjacent first video frames. The 21st and 26th frames are video frames on both sides of the same frame extraction interval length, so the 21st and 26th frames are also adjacent first video frames.
[0068] In step S125, when the behavior tags of adjacent first video frames are the same, the first video frame is used as the target video frame to obtain the behavior tag of the object in the target video frame.
[0069] In the embodiments of the present disclosure, if the behavior tags of adjacent first video frames are the same, it can be considered that in the video to be recognized, among the video frames between the adjacent first video frames, the behavior tag of the object is the same as that of the first video frame, that is, the behavior tag of the first video frame represents the behavior tag of the object during this period.
[0070] In the above technical solution, when the behavior labels of adjacent first video frames are the same, the first video frame is used as the target video frame to obtain the behavior label of the object in the target video frame. In this way, by extracting frames, the accuracy of behavior recognition can be ensured while reducing the amount of calculation, thereby quickly obtaining the behavior recognition result.
[0071] Alternatively, see Figure 3 As shown, the method also includes:
[0072] In step S126, when the behavior labels of adjacent first video frames are different, the video frames between the adjacent first video frames in the video to be identified are determined as second video frames.
[0073] In the disclosed embodiment, if the behavior labels of adjacent first video frames are not the same, it means that there are at least two different behavior labels between the adjacent first video frames and the video frames extracted between the adjacent first video frames. However, due to the presence of extracted video frames in the middle, it is not possible to determine when the behavior represented by the behavior label corresponding to the previous video frame in the adjacent first video frames ends, and it is not possible to determine when the behavior represented by the behavior label corresponding to the next video frame in the adjacent first video frames begins. Therefore, it is necessary to perform behavior recognition on the video frames between the adjacent first video frames in the video to be recognized, so as to determine when the behavior represented by the behavior label corresponding to the previous video frame in the adjacent first video frames ends, and to determine when the behavior represented by the behavior label corresponding to the next video frame in the adjacent first video frames begins.
[0074] In step S127, after the enclosing box in the second video frame is expanded according to different scales, description labels are added to the objects covered by the expanded enclosing box in the second video frame.
[0075] In the disclosed embodiment, the implementation in step S12 may be implemented to expand the bounding box of each object in the second video frame.
[0076] Based on the same principle as step S122, descriptive tags can be added to objects covered by the bounding box of each object in the second video frame based on automatic tag addition, and then the descriptive tags added to each bounding box are screened. For example, based on confidence, description tags with confidence higher than or equal to a preset confidence threshold are retained, and description tags with confidence lower than the preset confidence threshold are removed. Based on the principle of voting, tag results are obtained from high-confidence description tags. The tag results may include multiple description tags corresponding to each bounding box in the second video frame.
[0077] In step S128, the description tags of the object in the second video frame at multiple scales are combined to obtain a behavior tag corresponding to the object in the second video frame.
[0078] In the embodiments of the present disclosure, the description tags in the tag results can also be combined. For example, they can be combined in the subject-verb-object manner to obtain multiple candidate behavior tags of the object in the second video frame. Then, the multiple candidate behavior tags are further screened for rationality and confidence to obtain a behavior tag for each object in the second video frame. In this way, after expanding according to scales of different ratios, description tags are added to the objects covered by the expanded bounding boxes, and then the description tags are combined to obtain behavior tags.
[0079] In step S129, the first video frame and the second video frame are used as the target video frames to obtain the behavior tags of the objects in the target video frames.
[0080] In the embodiments of the present disclosure, when the behavior tags of adjacent first video frames are different, the behavior tags of the video frames between the adjacent first video frames in the video to be recognized can be used to sequentially determine which one or which several video frames have the same or different behavior tags of the objects as the behavior tag of the previous first video frame among the adjacent first video frames. Similarly, it can be determined which one or which several video frames have the same or different behavior tags of the objects as the behavior tag of the subsequent first video frame among the adjacent first video frames. In this way, it can be avoided that it is impossible to determine when the behavior represented by different behavior tags ends and starts. Furthermore, the accuracy of behavior recognition is improved.
[0081] In the embodiments of the present disclosure, as shown in Figure 4 the flowchart shown, the behavior tags of the objects in the target video frames in the video to be recognized can be determined. First, the video to be recognized is input, and the objects in the video to be recognized are tracked and recognized through, for example, the BoT-SORT model to obtain the bounding box and numbered ID tag of each object. Further, each numbered ID tag is polled to perform bounding box positioning on the objects in each video frame.
[0082] Further, for any bounding box, it is expanded according to multiple scales of different ratios, and behavior features are obtained for the local images covered by the bounding boxes corresponding to different scales. Then, description tags for the behavior features are automatically added based on tags. As shown in the figure, behavior features are obtained for the local image covered by the original bounding box, the local image covered by the bounding box expanded at a scale of 1:2 in the length direction, and the local image covered by the bounding box expanded at a scale of 1:2 in the width direction.
[0083] Further, according to a preset confidence threshold, description tags with low confidence can be screened out based on the confidence of the description tags, and description tags with high confidence can be retained. Voting is performed on the description tags with high confidence to obtain the tag results corresponding to the description tags. Then, the description tags in the tag results are combined to obtain the unique behavior tag of the object in a video frame. Furthermore, according to the acquisition timestamp of the video frame, the start time and end time of the behavior corresponding to the behavior tag can be determined based on multiple behavior tags of the same object in different video frames.
[0084] Optionally, the method further includes:
[0085] According to the attributes of the object represented by the numbered tag, multiple scales with different ratios are determined from a preset scale.
[0086] In the embodiments of the present disclosure, the attributes of the object are used to represent the type of the object. For example, the object is a person or what type of object. Each attribute can correspond to a different preset scale. For example, for a person, scales with ratios of 1:2 and 1:1.5 in the width direction can be determined from the preset scale. Another example is that for a handbag, scales with ratios of 1:2 in both the length and width directions and 1:3 in the length direction can be determined from the preset scale.
[0087] According to the determined multiple scales with different ratios, the bounding box is expanded respectively.
[0088] In the embodiments of the present disclosure, the bounding box is expanded by multiple scales with different ratios respectively, so that for different attributes of the object, different bounding boxes can be obtained by expansion, and thus the behavior recognition in different scenarios can be adapted.
[0089] Optionally, the step of expanding the bounding box according to the determined multiple scales with different ratios respectively includes:
[0090] Determine the expansion direction of the bounding box under each of the scales.
[0091] In the embodiments of the present disclosure, the expansion direction of the bounding box under each scale can be determined according to the orientation or facing direction of the object. For example, for a person facing left, the direction the person is facing can be determined as the expansion direction, or the direction the person is facing away from can also be determined as the expansion direction. Another example is that for a bicycle, since the person pushing the bicycle is usually parallel to the bicycle and higher than the bicycle, the height direction of the bicycle can be determined as the expansion direction.
[0092] According to the determined multiple scales with different ratios and the corresponding expansion directions, the bounding box is expanded respectively.
[0093] In this way, for different objects and the same object in different scenarios, different extension directions can be determined, so that richer behavior labels can be extracted through the bounding box, thereby improving the accuracy of behavior recognition.
[0094] Optionally, the method further includes:
[0095] Determine the time of the target video frame in the video to be recognized as the time of the behavior label corresponding to the target video frame.
[0096] It can be understood that the time of the target video frame in the video to be recognized is the acquisition time of the video frame, and then this acquisition time can be determined as the time of the behavior label corresponding to the target video frame. Thus, it can be determined when the object performed this behavior.
[0097] In step S13, associating the behavior labels of the object in each of the target video frames with the same numbered label to obtain the behavior recognition result of the object corresponding to each numbered label includes:
[0098] Associate the behavior labels of the object in each of the target video frames with the same numbered label and the time of the behavior label to obtain the behavior recognition result of the object corresponding to each numbered label.
[0099] In the embodiments of the present disclosure, based on the behavior labels of the same numbered label in each target video frame and the time corresponding to the behavior label, it can be determined when the same behavior and different behaviors start and end. For example, if the behavior labels between adjacent target video frames are different, it indicates that a behavior has been converted to another behavior between these adjacent target video frames. If the behavior labels between adjacent target video frames are the same, it indicates that the behavior of the object has not changed. Then, based on the behavior labels of other adjacent target video frames, it can be determined when this behavior starts and ends.
[0100] In this way, the time of the behavior represented by the behavior label can be determined according to the time of the video frame, and then when a behavior ends and starts can be determined according to the times of different behavior labels. Thereby improving the accuracy of behavior recognition.
[0101] Optionally, as shown in Figure 5 In step S11, the object in the video to be recognized is tracked and recognized to obtain the bounding box of the object and the numbered label of each object, including:
[0102] In step S111, the objects in each video frame of the video to be recognized are tracked and recognized to determine the coordinate positions of each object in each video frame.
[0103] Among them, the coordinate position can represent the position of the object in the video frame, which is conducive to tracking and identifying the object in the video frame according to the change of the coordinate position.
[0104] In step S112, a bounding box is added to the object according to the coordinate position, and the bounding boxes of the objects in each video frame are obtained.
[0105] Among them, one bounding box only contains one object.
[0106] In the embodiments of the present disclosure, a rectangular bounding box can be added to each object in the video frame centered on the coordinate position, and adjacent objects will be added with a bounding box respectively, so that one bounding box only contains one object. In this way, one bounding box only containing one object can distinguish multiple objects in the same video frame, so that descriptive labels can be added and descriptive label combinations can be made for each object separately, thereby improving the accuracy of behavior recognition.
[0107] In step S113, a numbered label is marked for the object after adding the bounding box, and the numbered label of each object.
[0108] Among them, the same numbered label is marked for the same object in the video to be recognized.
[0109] In the embodiments of the present disclosure, the numbered label can be a pure digital label, an alphabetical label, or a label obtained by combining letters, characters, or numbers. Usually, there are different categories of objects, so a set of numbering methods can be adopted for objects of the same category. For example, for people, pure digital labels can be used, starting from number 1. For objects such as bicycles and backpacks, they can be numbered with English letters plus numbers, where the English letters of different bicycles in the same video frame or different bicycles in different video frames are the same, but the numbers are different.
[0110] By marking the same numbered label for the same object, the behavior labels with the same numbered label can be associated, so that one or more behaviors of the same object can be determined.
[0111] The embodiments of the present disclosure also provide a behavior recognition device. Refer to Figure 6 As shown, the behavior recognition device includes: an identification module 610, a combination module 620, and an association module 630.
[0112] Among them, the identification module 610 is configured to track and identify the objects in the video to be recognized, and obtain the bounding boxes of the objects and the numbered label of each object;
[0113] The combination module 620 is configured to combine the description tags of the object in the target video frame of the video to be recognized at multiple scales to obtain a behavior tag corresponding to the object in the target video frame, where the description tag is added to the object covered by the expanded bounding box after expanding the bounding box according to scales of different ratios;
[0114] The association module 630 is configured to associate the behavior tags of the object with the same numbered tag in each target video frame to obtain a behavior recognition result corresponding to each numbered tag.
[0115] Optionally, the combination module 620 is configured to:
[0116] Extract a first video frame from the video to be recognized according to a preset frame extraction interval length;
[0117] After expanding the bounding box in the first video frame according to scales of different ratios, add description tags to the objects covered by the expanded bounding boxes in the first video frame respectively;
[0118] Combine the description tags of the object in the first video frame at multiple scales to obtain a behavior tag corresponding to the object in the first video frame;
[0119] Determine whether the behavior tags of adjacent first video frames are the same, where the adjacent first video frames are two video frames on both sides of the same frame extraction interval length;
[0120] In the case where the behavior tags of adjacent first video frames are the same, use the first video frame as the target video frame to obtain the behavior tag of the object in the target video frame.
[0121] Optionally, the combination module 620 is configured to:
[0122] In the case where the behavior tags of adjacent first video frames are different, determine the video frames between the adjacent first video frames in the video to be recognized as second video frames;
[0123] After expanding the bounding box in the second video frame according to scales of different ratios, add description tags to the objects covered by the expanded bounding boxes in the second video frame respectively;
[0124] Combine the description tags of the object in the second video frame at multiple scales to obtain a behavior tag corresponding to the object in the second video frame;
[0125] Use the first video frame and the second video frame as the target video frame to obtain the behavior label of the object in the target video frame.
[0126] Optionally, the combination module 620 is configured to:
[0127] Determine multiple scales with different ratios from a preset scale according to the attributes of the object characterized by the number label;
[0128] Expand the bounding box respectively according to the determined multiple scales with different ratios.
[0129] Optionally, the combination module is configured to:
[0130] Determine the expansion direction of the bounding box at each of the scales;
[0131] Expand the bounding box respectively according to the determined multiple scales with different ratios and the corresponding expansion directions.
[0132] Optionally, the association module 630 is configured to:
[0133] Determine the time of the target video frame in the video to be recognized as the time of the behavior label corresponding to the target video frame;
[0134] Associate the behavior labels of the objects in each target video frame with the same number label and the time of the behavior labels to obtain the behavior recognition result of the object corresponding to each number label.
[0135] Optionally, the recognition module 610 is configured to:
[0136] Track and recognize the objects in each video frame of the video to be recognized, and determine the coordinate positions of each object in each video frame;
[0137] Add a bounding box to the object according to the coordinate position to obtain the bounding box of each object in each video frame, where one bounding box only contains one object;
[0138] Label the object with a number label after adding the bounding box, and the number label of each object, where the same object in the video to be recognized is labeled with the same number label.
[0139] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0140] The embodiments of the present disclosure further provide an electronic device, including:
[0141] A processor; a memory for storing processor-executable instructions;
[0142] wherein the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the foregoing embodiments.
[0143] An embodiment of the present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method according to any one of the foregoing embodiments are implemented.
[0144] An embodiment of the present disclosure also provides a chip, including a processor and an interface; the processor is configured to read instructions to execute the method according to any one of the foregoing embodiments.
[0145] Figure 7 is a block diagram of a device 800 for behavior recognition shown according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0146] Referring to Figure 7 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output interface 812, a sensor component 814, and a communication component 816.
[0147] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0148] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0149] The power supply component 806 provides power to various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0150] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0151] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0152] The input / output interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0153] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0154] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0155] In an exemplary embodiment, the device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described behavior recognition method.
[0156] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided. The above instructions can be executed by the processor 820 of the device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0157] In addition to being an independent electronic device, the above-mentioned device can also be a part of an independent electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip. The integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip), etc. The above-mentioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the above-mentioned behavior recognition method. The executable instructions can be stored in the integrated circuit or chip, or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, a memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-mentioned behavior recognition method is implemented. Or, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-mentioned behavior recognition method.
[0158] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned behavior recognition method when executed by the programmable device.
[0159] Figure 8 is a block diagram of a device 1900 for behavior recognition shown according to an exemplary embodiment. For example, the device 1900 can be provided as a server. Referring to Figure 8 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-mentioned behavior recognition method.
[0160] The apparatus 1900 may further include a power supply component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958. The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0161] According to an embodiment of the present disclosure, a vehicle is further provided, including:
[0162] a processor; a memory for storing processor-executable instructions;
[0163] wherein the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the foregoing embodiments.
[0164] It can be noted that the vehicle in the embodiment of the present disclosure can execute the method according to any one of the foregoing embodiments, so as to quickly identify the collected video during the driving of the vehicle, provide a relatively accurate image basis for, for example, autonomous driving, automatic parking, automatic obstacle avoidance, etc., and improve the driving safety of the vehicle.
[0165] Figure 9 FIG. is a block diagram of a vehicle 600 shown according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, or may be a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0166] Referring to Figure 9 , the vehicle 600 may include various subsystems. For example, an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. Among them, the vehicle 600 may further include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wired or wireless means.
[0167] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, a navigation system, etc.
[0168] The perception system 620 may include several types of sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device.
[0169] The decision-making and control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0170] The drive system 640 may include components that provide powered movement for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a powertrain, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0171] Some or all functions of the vehicle 600 are controlled by the computing platform 650. The computing platform 650 may include at least one third processor 651 and a third memory 652, and the third processor 651 may execute instructions 653 stored in the third memory 652.
[0172] The third processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0173] The third memory 652 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0174] In addition to the instructions 653, the third memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the third memory 652 can be used by the computing platform 650.
[0175] In an embodiment of the present disclosure, the third processor 651 may execute an instruction 653 to complete all or part of the steps of the above-mentioned behavior recognition method.
[0176] After considering the specification and practicing the present disclosure, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0177] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A behavior recognition method, characterized in that, it includes: Tracking and recognizing the objects in the video to be recognized, obtaining the bounding boxes of the objects and the numbered labels of each object; Combining the description labels of the objects in the target video frames of the video to be recognized at multiple scales, obtaining the behavior labels of the objects corresponding to the target video frames, wherein the description labels are added to the objects covered by the expanded bounding boxes after expanding the bounding boxes according to scales of different ratios; Associating the behavior labels of the objects with the same numbered label in each of the target video frames, obtaining the behavior recognition results of the objects corresponding to each numbered label.
2. The method according to claim 1, characterized in that, The step of combining the description labels of the objects in the target video frames of the video to be recognized at multiple scales, obtaining the behavior labels of the objects corresponding to the target video frames, includes: Extracting a first video frame from the video to be recognized according to a preset frame extraction interval length; After expanding the bounding boxes in the first video frame according to scales of different ratios, adding description labels to the objects covered by the expanded bounding boxes in the first video frame respectively; Combining the description labels of the objects in the first video frame at multiple scales, obtaining the behavior labels of the objects corresponding to the first video frame; Determining whether the behavior labels of adjacent first video frames are the same, where the adjacent first video frames are two video frames on both sides of the same frame extraction interval length; In the case where the behavior labels of adjacent first video frames are the same, taking the first video frame as the target video frame, obtaining the behavior labels of the objects in the target video frame.
3. The method according to claim 2, characterized in that, The method further includes: In the case where the behavior labels of adjacent first video frames are different, determining the video frames between the adjacent first video frames in the video to be recognized as second video frames; After expanding the bounding boxes in the second video frames according to scales of different ratios, adding description labels to the objects covered by the expanded bounding boxes in the second video frames respectively; Combining the description labels of the objects in the second video frames at multiple scales, obtaining the behavior labels of the objects corresponding to the second video frames; Taking the first video frame and the second video frame as the target video frames, obtaining the behavior labels of the objects in the target video frames.
4. The method according to claim 1, characterized in that, The method further includes: Determining multiple scales of different ratios from a preset scale according to the attributes of the objects characterized by the numbered labels; Expanding the bounding boxes respectively according to the determined multiple scales of different ratios.
5. The method according to claim 4, characterized in that, The step of expanding the bounding boxes respectively according to the determined multiple scales of different ratios includes: Determining the expansion directions of the bounding boxes at each scale; Expand the bounding box according to the determined scales with multiple different ratios and the corresponding expansion directions respectively.
6. The method according to claim 1, wherein, the method further comprises: Determine the time of the target video frame in the video to be recognized as the time of the behavior label corresponding to the target video frame; The associating the behavior labels of the objects with the same numbered label in each of the target video frames to obtain the behavior recognition result of the object corresponding to each numbered label includes: Associating the behavior labels of the objects with the same numbered label in each of the target video frames and the time of the behavior labels to obtain the behavior recognition result of the object corresponding to each numbered label.
7. The method according to any one of claims 1-6, wherein, The performing tracking and recognition on the objects in the video to be recognized to obtain the bounding boxes of the objects and the numbered label of each object includes: Performing tracking and recognition on the objects in each video frame of the video to be recognized to determine the coordinate positions of each object in each video frame; Adding a bounding box to the object according to the coordinate position to obtain the bounding boxes of the objects in each video frame, wherein one bounding box only contains one object; Labeling the numbered label to the object after adding the bounding box, and the numbered label of each object, wherein the same object in the video to be recognized is labeled with the same numbered label.
8. A behavior recognition device, wherein, comprises: An identification module configured to perform tracking and recognition on the objects in the video to be recognized to obtain the bounding boxes of the objects and the numbered label of each object; A combination module configured to combine the description labels of the objects in the target video frames in the video to be recognized at multiple scales to obtain the behavior labels corresponding to the objects in the target video frames, wherein the description labels are added to the objects covered by the expanded bounding box after expanding the bounding box according to scales with different ratios; An association module configured to associate the behavior labels of the objects with the same numbered label in each of the target video frames to obtain the behavior recognition result of the object corresponding to each numbered label.
9. An electronic device, wherein, comprises: A processor; A memory for storing executable instructions executable by the processor; wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-7.
10. A computer-readable storage medium, on which computer program instructions are stored, wherein, the program instructions, when executed by a processor, implement the steps of the method according to any one of claims 1-7.
11. A chip, wherein, comprises a processor and an interface; the processor is used to read instructions to execute the method according to any one of claims 1-7.
12. A vehicle, wherein, comprises: A processor; A memory for storing executable instructions executable by the processor; Wherein, the processor is configured to execute executable instructions stored in the memory to implement the method according to any one of claims 1-7.