Pedestrian Recognition Method, Device, Intelligent Robot and Storage Medium

The method enhances smart robot rowan recognition by extracting image features, performing semantic recognition, and instance segmentation to accurately describe human behavior activities, addressing the limitation of current methods in understanding posture and actions.

CN114255476BActive Publication Date: 2025-07-15中原动力智能机器人有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111488654.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-07-15
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

The prior art cannot deeply identify pedestrian behavioral activities, and can only calibrate pedestrian positions but cannot express a deep understanding of pedestrian behavioral activities.

Method used

By extracting the image features of the target image, semantic recognition and instance segmentation are performed, the target image is segmented with a preset instance segmentation network, and superimposed by a pedestrian class and a pedestrian instance mask, so as to achieve a detailed description of the pedestrian behavior posture.

Benefits of technology

It realizes an accurate description of pedestrian behavior activities, can distinguish the corresponding behavior activities of different pedestrians, and improves the accuracy and efficiency of pedestrian identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255476B_ABST
    Figure CN114255476B_ABST
Patent Text Reader

Abstract

The present application discloses a pedestrian recognition method, device, intelligent robot and storage medium. By extracting the image features of a target image, semantic recognition is performed on the target image according to the image features to obtain the pedestrian category of the target image, so as to analyze the number of pedestrians in the target image; and a preset instance segmentation network is used to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image, so as to perform fine segmentation on pedestrians in an instance segmentation manner, including limb segmentation, so that the behavior postures of pedestrians can be described according to the pedestrian mask, so as to better recognize the behavior activities of pedestrians; finally, the pedestrian category and the pedestrian instance mask are superimposed to obtain the pedestrian instance segmentation result, so as to distinguish the pedestrian instance masks corresponding to different pedestrians, so as to accurately describe the behavior activities of each pedestrian.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a pedestrian recognition method, device, intelligent robot, and storage medium. Background Art

[0002] With the rise of artificial intelligence, intelligent robots are increasingly widely used in public places such as shopping malls, airports, and stations. Identifying buildings, green belts, pedestrians, vehicles, etc. has become an essential function for them. Among them, to further develop robot applications, while the robot identifies pedestrians, it is also necessary to understand the behavior activities of pedestrians.

[0003] Currently, robots mainly identify pedestrians based on object detection. By understanding the foreground and background of pictures, it is necessary to separate the target of interest from the background, determine the category and position of this target, and then use a rectangular box to represent the pedestrians in the image. To a certain extent, it can understand pedestrian behavior, such as identifying that a pedestrian is moving and the moving direction. However, it can only calibrate the position of pedestrians and cannot represent a deep understanding of pedestrian behavior activities, such as the activity posture of pedestrians. It can be seen that the current pedestrian recognition method has the problem of being unable to deeply identify pedestrian behavior activities. Summary of the Invention

[0004] This application provides a pedestrian recognition method, device, intelligent robot, and storage medium to solve the technical problem that the current pedestrian recognition method cannot deeply identify pedestrian behavior activities.

[0005] To solve the above technical problem, in a first aspect, an embodiment of this application provides a pedestrian recognition method, including:

[0006] Extracting the image features of the target image collected by the intelligent robot;

[0007] Performing semantic recognition on the target image according to the image features to obtain the pedestrian category of the target image;

[0008] Using a preset instance segmentation network, performing instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image;

[0009] Overlaying the pedestrian category and the pedestrian instance mask to obtain the pedestrian instance segmentation result.

[0010] In this embodiment, the image features of the target image are extracted, and based on the image features, semantic recognition is performed on the target image to obtain the human category of the target image, so as to analyze the number of pedestrians in the target image; and a preset instance segmentation network is used to perform instance segmentation on the target image based on the image features to obtain the pedestrian instance mask of the target image, so as to perform fine segmentation on pedestrians in an instance segmentation manner, including limb segmentation, so that the behavior postures of pedestrians can be described according to the pedestrian mask, so as to better recognize the behavior activities of pedestrians; finally, the human category and the pedestrian instance mask are superimposed to obtain the pedestrian instance segmentation result, so as to distinguish the pedestrian instance masks corresponding to different pedestrians, so as to accurately describe the behavior activities of each pedestrian.

[0011] In one embodiment, the extraction of the image features of the target image collected by the intelligent robot includes:

[0012] Obtain the target image collected by the intelligent robot;

[0013] Divide the target image into multiple image grids;

[0014] Determine the image grid corresponding to the position of the pedestrian centroid in the target image;

[0015] Extract features from the image grid to obtain the image features.

[0016] In this embodiment, the image grid is determined by the pedestrian centroid, and features are extracted from the image grid to reduce the computational complexity in the feature extraction process and improve the pedestrian recognition efficiency.

[0017] In one embodiment, the use of a preset instance segmentation network to perform instance segmentation on the target image based on the image features to obtain the pedestrian instance mask of the target image includes:

[0018] Use a preset instance segmentation network to perform instance segmentation on the target image based on the image features to obtain the pedestrian instance mask data of the target image;

[0019] Perform edge detection on the target image based on the image features to obtain the pedestrian edge data of the target image;

[0020] Fuse the pedestrian instance mask data and the pedestrian edge data to obtain the pedestrian instance mask.

[0021] In this embodiment, by fusing the pedestrian instance mask data and the pedestrian edge data, the segmented pedestrian boundary is made more refined, so that behavior activities such as the limb movements of pedestrians can be extracted.

[0022] In one embodiment, the step of using a preset instance segmentation network to perform instance segmentation on the target image according to the image features to obtain pedestrian instance mask data of the target image includes:

[0023] Using the Maskkernelbranch network layer in the Mask Brach network to determine a convolution kernel for instance segmentation by the Maskfeaturebranch network layer in the Mask Brach network according to the image features;

[0024] Using the Maskfeaturebranch network layer to perform instance segmentation on the target image according to the convolution kernel to obtain pedestrian instance mask data of the target image.

[0025] In this embodiment, instance segmentation of multiple image grids is achieved by determining the convolution kernel.

[0026] In one embodiment, the step of performing edge detection on the target image according to the image features to obtain pedestrian edge data of the target image includes:

[0027] Using the Canny edge algorithm to perform edge detection on the target image according to the image features to obtain pedestrian edge data of the target image.

[0028] In one embodiment, the target image includes multiple image grids. The step of superimposing the pedestrian category and the pedestrian instance mask to obtain a pedestrian instance segmentation result includes:

[0029] According to the multiple image grids in the target image, superimposing the pedestrian category and the pedestrian instance mask at the grid positions to obtain the instance segmentation result.

[0030] In one embodiment, before using a preset instance segmentation network to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image, it further includes:

[0031] Collecting a road pedestrian video through a camera on an intelligent robot;

[0032] Based on the cosin similarity algorithm, screening the data of each video frame of the road pedestrian video to obtain multiple video frames with a similarity greater than a preset value, and combining the multiple video frames into a data set;

[0033] Using the data set to perform iterative training on a preset deep learning network until the preset deep learning network reaches a preset convergence condition, stopping the iteration, and obtaining the preset instance segmentation network.

[0034] In this embodiment, data is collected and filtered by an intelligent robot to improve the accuracy of the training data set and enrich the data set samples, thereby solving the problem of few samples for pedestrian instance segmentation. At the same time, data is collected by the robot and the network is trained to reduce the image differences between the training stage and the actual application stage, so as to effectively avoid the problem that the network recognition accuracy decreases due to large differences in data collected by different devices.

[0035] In a second aspect, an embodiment of the present application provides a pedestrian recognition device, including:

[0036] An extraction module, configured to extract image features of a target image collected by the intelligent robot;

[0037] A recognition module, configured to perform semantic recognition on the target image according to the image features to obtain the pedestrian category of the target image;

[0038] A segmentation module, configured to perform instance segmentation on the target image according to the image features by using a preset instance segmentation network to obtain a pedestrian instance mask of the target image;

[0039] An overlay module, configured to overlay the pedestrian category and the pedestrian instance mask to obtain a pedestrian instance segmentation result.

[0040] In a third aspect, an embodiment of the present application provides an intelligent robot, including a processor and a memory, where the memory is used to store a computer program, and when the computer program is executed by the processor, the pedestrian recognition method described in the first aspect is implemented.

[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the pedestrian recognition method described in the first aspect is implemented.

[0042] It should be noted that for the beneficial effects of the second to fourth aspects above, please refer to the relevant descriptions of the first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flowchart of the pedestrian recognition method provided by an embodiment of the present application;

[0044] Figure 2 It is a schematic structural diagram of the pedestrian recognition device provided by an embodiment of the present application;

[0045] Figure 3 It is a schematic structural diagram of the intelligent robot provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0047] As recorded in the related art, currently, robots mainly identify pedestrians based on object detection. By understanding the foreground and background of images, it is necessary to separate the target of interest from the background, determine the category and location of this target, and then use a rectangular box to represent the pedestrians in the image, which can understand pedestrian behavior to a certain extent, such as identifying that a pedestrian is moving and the moving direction. However, it can only calibrate the position of pedestrians and cannot represent a deep understanding of pedestrian behavior activities, such as the activity postures of pedestrians. It can be seen that the current pedestrian recognition method has the problem of being unable to deeply recognize pedestrian behavior activities.

[0048] Therefore, the embodiments of the present application provide a pedestrian recognition method, device, intelligent robot, and storage medium. By extracting the image features of a target image, semantic recognition is performed on the target image according to the image features to obtain the pedestrian category of the target image, so as to analyze the number of pedestrians in the target image; and a preset instance segmentation network is used to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image, so as to perform fine segmentation on pedestrians in an instance segmentation manner, including limb segmentation, so as to be able to describe the behavior postures of pedestrians according to the pedestrian mask, so as to better recognize pedestrian behavior activities; finally, the pedestrian category and the pedestrian instance mask are superimposed to obtain the pedestrian instance segmentation result, so as to distinguish the pedestrian instance masks corresponding to different pedestrians, so as to accurately describe the behavior activities of each pedestrian.

[0049] Please refer to Figure 1 , Figure 1Schematic flow chart of a behavior recognition method provided by an embodiment of the present application. The behavior recognition method of the embodiment of the present application can be applied to an intelligent robot, which includes but is not limited to inspection robots, reception robots, etc. The intelligent robot is provided with a data acquisition module, a communication module, a data storage module, and a central control module. Among them, the data acquisition module can be a camera, which is used to receive instructions from the central control module, complete the setting of parameters such as the height and inclination angle of the camera device; obtain image data, and transmit the data to the data storage module; and feedback the current working status information to the central control module. The communication module serves as the communication interface between the robot system and the outside world, and can complete two-way communication with the central control module. It can not only receive external control instructions, but also feedback the current working status information of the robot to the outside world. The data storage module can be a memory, which is used to receive instructions from the central control module, feedback the current working status information to the central control module, and store a large amount of data from the data acquisition module. The central control module can be a processor, which serves as the decision-making center of the robot system. It can not only obtain the working status information of the other modules, but also send instructions to the other modules to coordinate the work of each module. As Figure 1 shown, the pedestrian recognition method includes steps S101 to S104, which are described in detail as follows:

[0050] Step S101, extract the image features of the target image collected by the intelligent robot.

[0051] In this step, the target image is the image collected by the camera of the intelligent robot. Optionally, the overall features of the target image can be extracted, or the target image can be divided into multiple image grids for feature extraction respectively. Optionally, the feature pyramid FCN network is used to extract the features of the target image.

[0052] Step S102, perform semantic recognition on the target image according to the image features to obtain the pedestrian category of the target image.

[0053] In this step, there may be multiple pedestrians in the target image, so each pedestrian in the image is recognized, that is, one pedestrian corresponds to one category. Optionally, semantic recognition is performed based on the Category Brach network. After the features are extracted by the FCN network, the classification result is obtained through convolution of the convolution layer of the Category Brach network, downsampling of the pooling layer, activation function, and output of the fully connected layer.

[0054] Step S103, use a preset instance segmentation network to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image.

[0055] In this step, the instance segmentation network can be the Mask Brach network. The Mask Brach is decomposed into a Mask kernel branch and a Mask feature branch. The Mask kernel branch is used to predict the convolution kernel, and the Mask feature branch is used for feature expression, that is, instance segmentation. Optionally, when the target image is for overall feature extraction and the overall feature is obtained, the Mask feature branch can be directly used for feature expression. When the target image is multiple segmented image grids, first use the Mask kernel branch to predict the convolution kernel, and then use the Mask feature branch for feature expression.

[0056] Step S104: Superimpose the pedestrian category and the pedestrian instance mask to obtain the pedestrian instance segmentation result.

[0057] In this step, since there may be multiple pedestrians in the target image and the behavior of each pedestrian needs to be recognized, the pedestrian category is superimposed with the pedestrian instance mask to obtain an independent mask for each pedestrian.

[0058] In one embodiment, based on Figure 1 the embodiment shown, step S101 specifically includes:

[0059] Obtain the target image collected by the intelligent robot;

[0060] Divide the target image into multiple image grids;

[0061] Determine the image grid corresponding to the position of the pedestrian centroid in the target image;

[0062] Extract features from the image grid to obtain the image features.

[0063] In this embodiment, the pedestrian centroid is the center point of the pedestrian in the target image. Exemplarily, first scale the size of the target image to meet the requirements of network input, and then divide the target image into S×S image grids. If the center point of the pedestrian falls into a certain image grid, then extract features from this image grid, and perform semantic recognition and instance segmentation on this image grid subsequently. In this embodiment, the image grid is determined by the pedestrian centroid, and features are extracted from the image grid to reduce the computational complexity in the feature extraction process and improve the pedestrian recognition efficiency.

[0064] In one embodiment, based on Figure 1 the embodiment shown, step S103 specifically includes:

[0065] Using a preset instance segmentation network, according to the image features, perform instance segmentation on the target image to obtain pedestrian instance mask data of the target image;

[0066] According to the image features, perform edge detection on the target image to obtain pedestrian edge data of the target image;

[0067] Fuse the pedestrian instance mask data and the pedestrian edge data to obtain the pedestrian instance mask.

[0068] In this embodiment, by fusing the pedestrian instance mask data and the pedestrian edge data, the segmented pedestrian boundary is made more refined, so that behavioral activities such as the limb movements of pedestrians can be extracted.

[0069] Optionally, the using a preset instance segmentation network, according to the image features, perform instance segmentation on the target image to obtain pedestrian instance mask data of the target image includes:

[0070] Using the Maskkernelbranch network layer in the Mask Brach network, according to the image features, determine the convolution kernel for performing instance segmentation on the Maskfeaturebranch network layer in the Mask Brach network;

[0071] Using the Maskfeaturebranch network layer, according to the convolution kernel, perform instance segmentation on the target image to obtain pedestrian instance mask data of the target image.

[0072] In this optional manner, Mask Brach is decomposed into two branches: Mask kernel branch and Maskfeature branch, which respectively predict the convolution kernel and the feature expression, and the outputs of the two branches are finally combined into the output of the entire Maskbranch.

[0073] Among them, the Mask kernel branch is used to learn and predict the convolution kernel, that is, the weight of the classifier. For example, if the input is a feature of H×W×E, where E is the number of channels of the input feature, then the output is a convolution kernel of S×S×D, where S is the number of divided grids and D is the number of channels of the convolution kernel. The corresponding relationship is as follows: for a 1×1×E convolution kernel, then D = E; for a 3×3×E convolution kernel, then D = 9E, and so on. It can be understood that this branch does not require an activation function. In this embodiment, by determining the convolution kernel, instance segmentation of multiple image grids is achieved.

[0074] Optionally, the according to the image features, perform edge detection on the target image to obtain pedestrian edge data of the target image includes:

[0075] Using the Canny edge algorithm, based on the image features, perform edge detection on the target image to obtain the pedestrian edge data of the target image.

[0076] In this alternative, the target image is grayscaled by the Canny edge algorithm, Gaussian filtering is performed on the grayscaled target image, the gradient magnitude and direction are calculated using the finite difference of the first-order partial derivative, non-maximum suppression is performed on the gradient magnitude, and finally the double-threshold algorithm is used to detect and connect the edges to obtain the edge data.

[0077] In one embodiment, on the basis of Figure 1 the embodiment shown, step S104 specifically includes:

[0078] According to multiple image grids in the target image, perform grid position superposition on the pedestrian category and the pedestrian instance mask to obtain the instance segmentation result.

[0079] In this embodiment, since the target image is divided into S×S image grids, and each image grid obtains the pedestrian category and the pedestrian instance mask based on semantic recognition and instance segmentation, the pedestrian instance masks in the image grids corresponding to the same pedestrian category are superposed in position to obtain the instance segmentation result.

[0080] In one embodiment, on the basis of Figure 1 the embodiment shown, before step S103, it further includes:

[0081] Collect the road pedestrian video through the camera on the intelligent robot;

[0082] Based on the cosin similarity algorithm, perform data screening on each video frame of the road pedestrian video to obtain multiple video frames with a similarity greater than a preset value, and combine the multiple video frames into a data set;

[0083] Use the data set to perform iterative training on the preset deep learning network until the preset deep learning network reaches the preset convergence condition, stop the iteration, and obtain the preset instance segmentation network.

[0084] In this embodiment, exemplarily, an instruction to collect data is sent to the intelligent robot. After the communication module of the intelligent robot receives the instruction, it conveys the instruction to the central control module. The central control module obtains the states of each relevant module, initializes each relevant module, and starts real-time monitoring of the states of each relevant module, such as the data storage amount, data transmission rate, and so on. The data acquisition module starts to collect road pedestrian videos and transmits the road pedestrian videos to the data storage module in real time. The data storage module starts to store the road pedestrian videos from the data acquisition module. After the data acquisition work is completed, based on the cosine similarity algorithm, the road pedestrian videos are screened to obtain multiple video frames with a similarity greater than the preset value, and the multiple video frames are combined into a data set, and the data set in the data storage module is labeled.

[0085] Initialize the classifier network FCN, the instance segmentation network Mask Brach, and the semantic recognition network Category Brach. Divide the image frames in the data set into multiple image grids and input them into the FCN, and forward them to Mask Brach, Category Brach, and the Canny edge algorithm. The image frames are feature-extracted by the FCN, and the feature images are passed into CategoryBrach to obtain the category information at the corresponding positions. The feature images are passed into Mask Brach to obtain the instance mask information at the corresponding positions. The feature images are passed into the Canny edge algorithm to obtain the edge information of the pedestrians to be segmented. The instance mask information and the edge information are fused to obtain a fine pedestrian instance mask. The pedestrian instance mask and the category information are superimposed in position to obtain the instance segmentation result. Use the instance segmentation result and the original label to solve the loss function, and backpropagate and iterate continuously until the model converges.

[0086] Optionally, the model is compressed by pruning to remove unimportant layers and parameters to make the model as lightweight as possible. Then use TensorRT to accelerate the inference of the model, and then deploy it on the edge device.

[0087] In this embodiment, the intelligent robot is used to collect data and screen data, which improves the accuracy of the training data set and enriches the data set samples, and solves the problem of few samples for pedestrian instance segmentation. At the same time, based on the data collected by the robot and the training network, the image difference between the training stage and the actual application stage is reduced, so as to effectively avoid the problem that the network recognition accuracy decreases due to the large difference in the data collected by different devices.

[0088] In order to execute the corresponding pedestrian recognition method in the above method embodiment to achieve the corresponding functions and technical effects. See Figure 2 , Figure 2The structural block diagram of a pedestrian recognition device provided by an embodiment of the present application is shown. For ease of description, only the parts related to this embodiment are shown. The pedestrian recognition device provided by the embodiment of the present application includes:

[0089] An extraction module 201, configured to extract image features of a target image collected by the intelligent robot;

[0090] An identification module 202, configured to perform semantic recognition on the target image according to the image features to obtain the pedestrian category of the target image;

[0091] A segmentation module 203, configured to perform instance segmentation on the target image according to the image features by using a preset instance segmentation network to obtain a pedestrian instance mask of the target image;

[0092] An overlay module 204, configured to overlay the pedestrian category and the pedestrian instance mask to obtain a pedestrian instance segmentation result.

[0093] In one embodiment, the extraction module 201 includes:

[0094] An acquisition unit, configured to acquire a target image collected by the intelligent robot;

[0095] A division unit, configured to divide the target image into a plurality of image grids;

[0096] A determination unit, configured to determine the image grid corresponding to the position of the pedestrian centroid in the target image;

[0097] An extraction unit, configured to perform feature extraction on the image grid to obtain the image features.

[0098] In one embodiment, the segmentation module 203 includes:

[0099] A segmentation unit, configured to perform instance segmentation on the target image according to the image features by using a preset instance segmentation network to obtain pedestrian instance mask data of the target image;

[0100] A detection unit, configured to perform edge detection on the target image according to the image features to obtain pedestrian edge data of the target image;

[0101] A fusion unit, configured to fuse the pedestrian instance mask data and the pedestrian edge data to obtain the pedestrian instance mask.

[0102] In one embodiment, the segmentation unit includes:

[0103] A determination subunit, configured to use the Maskkernelbranch network layer in the Mask Brach network to determine, according to the image features, a convolution kernel for instance segmentation in the Maskfeaturebranch network layer of the Mask Brach network;

[0104] A segmentation subunit, configured to use the Maskfeaturebranch network layer to perform instance segmentation on the target image according to the convolution kernel, so as to obtain pedestrian instance mask data of the target image.

[0105] In one embodiment, the detection unit is specifically configured to:

[0106] Use the Canny edge algorithm to perform edge detection on the target image according to the image features, so as to obtain pedestrian edge data of the target image.

[0107] In one embodiment, the superimposing module 204 includes:

[0108] A superimposing unit, configured to perform grid position superimposition on the pedestrian category and the pedestrian instance mask according to multiple image grids in the target image, so as to obtain the instance segmentation result.

[0109] In one embodiment, the pedestrian recognition device further includes:

[0110] An acquisition module, configured to acquire a road pedestrian video through a camera on the intelligent robot;

[0111] A screening module, configured to perform data screening on each video frame of the road pedestrian video based on the cosin similarity algorithm, obtain multiple video frames with a similarity greater than a preset value, and combine the multiple video frames into a data set;

[0112] A training module, configured to use the data set to perform iterative training on a preset deep learning network until the preset deep learning network reaches a preset convergence condition, stop iteration, and obtain the preset instance segmentation network.

[0113] The above-mentioned pedestrian recognition device can implement the pedestrian recognition method in the above method embodiment. The optional items in the above method embodiment are also applicable to this embodiment, which will not be elaborated here. The remaining content of the embodiment of the present application can refer to the content of the above method embodiment, and will not be repeated in this embodiment.

[0114] Figure 3 It is a schematic structural diagram of an intelligent robot provided in an embodiment of the present application. As Figure 3 shown, the intelligent robot 3 in this embodiment includes: at least one processor 30 ( Figure 3Only one) processor, a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30 are shown. When the processor 30 executes the computer program 32, the steps in any of the above method embodiments are implemented.

[0115] The intelligent robot 3 can be an inspection robot, a reception robot, etc. The intelligent robot may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 This is only an example of the intelligent robot 3 and does not constitute a limitation on the intelligent robot 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0116] The so-called processor 30 may be a central processing unit (CPU). The processor 30 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0117] In some embodiments, the memory 31 may be an internal storage unit of the intelligent robot 3, such as the hard disk or memory of the intelligent robot 3. In other embodiments, the memory 31 may also be an external storage device of the intelligent robot 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the intelligent robot 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the intelligent robot 3. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 31 may also be used to temporarily store data that has been output or will be output.

[0118] In addition, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0119] An embodiment of the present application provides a computer program product. When the computer program product runs on an intelligent robot, the intelligent robot is caused to implement the steps in each of the above method embodiments when executed.

[0120] In several embodiments provided by the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0121] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a terminal device to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0122] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only specific embodiments of the present application and is not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A pedestrian recognition method, characterized in that, Applied to an intelligent robot, the method includes: Extracting the image features of the target image collected by the intelligent robot, including: obtaining the target image collected by the intelligent robot; dividing the target image into multiple image grids; determining the image grid corresponding to the position of the pedestrian centroid in the target image; extracting features from the image grid to obtain the image features; According to the image features, performing semantic recognition on the target image to recognize each pedestrian in the image and obtain the pedestrian category of the target image, with one pedestrian corresponding to one category; Using a preset instance segmentation network, according to the image features, performing instance segmentation on the target image to obtain the pedestrian instance mask of the target image, including: Using the preset instance segmentation network, according to the image features, performing instance segmentation on the target image to obtain the pedestrian instance mask data of the target image; according to the image features, performing edge detection on the target image to obtain the pedestrian edge data of the target image; fusing the pedestrian instance mask data and the pedestrian edge data to obtain the pedestrian instance mask; The step of using the preset instance segmentation network to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask data of the target image includes: using the Mask kernelbranch network layer in the Mask Brach network to determine the convolution kernel for instance segmentation by the Mask featurebranch network layer in the Mask Brach network according to the image features; using the Mask featurebranch network layer to perform instance segmentation on the target image according to the convolution kernel to obtain the pedestrian instance mask data of the target image; The step of performing edge detection on the target image according to the image features to obtain the pedestrian edge data of the target image includes: using the Canny edge algorithm to perform edge detection on the target image according to the image features to obtain the pedestrian edge data of the target image; Superimposing the pedestrian category and the pedestrian instance mask to obtain the pedestrian instance segmentation result, including: According to the multiple image grids in the target image, performing grid position superimposition on the pedestrian category and the pedestrian instance mask to obtain the instance segmentation result.

2. The pedestrian recognition method according to claim 1, wherein, Before using the preset instance segmentation network to perform instance segmentation on the target image according to the image features to obtain the pedestrian instance mask of the target image, it further includes: Collecting the road pedestrian video through the camera on the intelligent robot; Based on the cosin similarity algorithm, performing data screening on each video frame of the road pedestrian video to obtain multiple video frames with similarity greater than a preset value, and combining the multiple video frames into a data set; Using the data set to perform iterative training on the preset deep learning network until the preset deep learning network reaches the preset convergence condition, stopping the iteration, and obtaining the preset instance segmentation network.

3. A pedestrian recognition device, characterized in that, Including: An extraction module, configured to extract image features of a target image collected by an intelligent robot, including: obtaining the target image collected by the intelligent robot; dividing the target image into a plurality of image grids; determining the image grid corresponding to the position of the pedestrian centroid in the target image; performing feature extraction on the image grid to obtain the image features; An identification module, configured to perform semantic identification on the target image according to the image features, identify each pedestrian in the image, and obtain the pedestrian category of the target image, where one pedestrian corresponds to one category; A segmentation module, configured to perform instance segmentation on the target image according to the image features by using a preset instance segmentation network, to obtain a pedestrian instance mask of the target image, including: performing instance segmentation on the target image according to the image features by using a preset instance segmentation network, to obtain pedestrian instance mask data of the target image; performing edge detection on the target image according to the image features, to obtain pedestrian edge data of the target image; fusing the pedestrian instance mask data and the pedestrian edge data to obtain the pedestrian instance mask; The performing instance segmentation on the target image according to the image features by using a preset instance segmentation network to obtain the pedestrian instance mask data of the target image includes: determining, according to the image features, a convolution kernel for performing instance segmentation on a Mask feature branch network layer in a Mask Brach network by using a Mask kernel branch network layer in the Mask Brach network; performing instance segmentation on the target image according to the convolution kernel by using the Mask feature branch network layer to obtain the pedestrian instance mask data of the target image; The performing edge detection on the target image according to the image features to obtain the pedestrian edge data of the target image includes: performing edge detection on the target image according to the image features by using a Canny edge algorithm to obtain the pedestrian edge data of the target image; An overlay module, configured to overlay the pedestrian category and the pedestrian instance mask to obtain a pedestrian instance segmentation result, including: performing grid position overlay on the pedestrian category and the pedestrian instance mask according to a plurality of image grids in the target image to obtain the instance segmentation result.

4. An intelligent robot, characterized in that, It includes a processor and a memory, where the memory is configured to store a computer program, and when the computer program is executed by the processor, the pedestrian recognition method according to any one of claims 1 to 2 is implemented.

5. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, the pedestrian recognition method according to any one of claims 1 to 2 is implemented.

Citation Information

Patent Citations

  • Portrait extraction method and device, electronic equipment and storage medium

    CN112802037A

  • Target instance segmentation method based on traffic monitoring video

    CN112989942A