A multimodal unsupervised pedestrian pixel-level semantic annotation method and system

Through the multimodal unsupervised pedestrian pixel-level semantic labeling method and combined with a variety of image acquisition equipment, automated pedestrian detection is realized, solving the problem of time-consuming and labor-consuming manual labeling in the prior art, and improving detection efficiency and cost-effectiveness.

CN112766061BActive Publication Date: 2025-05-16ROPEOK TECHNOLOGY GROUP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011615688.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-05-16
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

The prior art requires a lot of manpower, material resources and financial resources to mark the position of pedestrians in the picture, which makes the process of marking the training data for pedestrian detection cumbersome and time-consuming.

Method used

A multimodal unsupervised pedestrian pixel-level semantic labeling method is proposed. By combining Tof image acquisition equipment, infrared image acquisition equipment and RGB image acquisition equipment, three-dimensional reconstruction and image processing are automatically performed to obtain human area collections, and pedestrian detection without manual labeling is realized.

Benefits of technology

This method can automatically extract human pixel points in the scene, reduce the need for manual labeling, and improve the efficiency and cost-effectiveness of pedestrian detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112766061B_ABST
    Figure CN112766061B_ABST
Patent Text Reader

Abstract

The present invention provides a multimodal unsupervised pixel-level semantic annotation method and system for pedestrians, including three-dimensional reconstruction of an unmanned monitoring scene to obtain the initial point cloud information of the monitoring scene; using a Tof image acquisition device to obtain the first point cloud information in the monitoring scene, aligning it with the initial point cloud information and performing a set difference operation to obtain the second point cloud information, and projecting the second point cloud information on a horizontal plane to obtain a set of personnel point cloud information; dilating and corroding the thresholded binary image of the scene information obtained by the infrared image acquisition device to obtain a connected area information set; projecting the personnel point cloud information set and the connected area information set into the image plane space of the RGB image acquisition device respectively using the positional relationship between the calibrated cameras to perform a set intersection operation, and in response to the common pixel exceeding the first threshold, obtaining the corresponding human body area set. The method and system fully integrate the advantages of cameras of different modalities and can effectively extract human body pixel points in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a multimodal unsupervised pedestrian pixel-level semantic labeling method and system. Background Art

[0002] Pedestrian detection is a classic problem in computer vision, and its related technologies can be applied in fields such as video surveillance and autonomous driving. The current common method is to first take a large number of samples containing pedestrians, and then manually mark the positions of pedestrians in the pictures as training data; finally, use supervised learning methods (such as support vector machines, deep learning) to train a classifier to distinguish between pedestrians and non-pedestrian areas. With the development of random deep learning technology, the number of training samples required is getting larger and larger. Labeling a large number of samples is a time-consuming and labor-intensive task.

[0003] According to the format of input data, pedestrian detection technology can be divided into methods based on two-dimensional images (including color and grayscale); methods based on three-dimensional point clouds; and methods based on infrared imaging. From a technical point of view, it can be divided into: overall method, part method, and local block method. Most of the above methods use supervised classification technology in machine learning. Supervised classification technology requires marking the position of pedestrians in the picture, so it requires a lot of manpower, material resources, and financial resources. Summary of the invention

[0004] In order to solve the technical problem in the prior art that a lot of manpower, material resources and financial resources are needed to mark the positions of pedestrians in pictures, the present invention proposes a multimodal unsupervised pedestrian pixel-level semantic labeling method and system, which avoids the trouble of manual labeling of pedestrian samples.

[0005] According to one aspect of the present invention, a multimodal unsupervised pedestrian pixel-level semantic annotation method is proposed, comprising:

[0006] S1: Perform 3D reconstruction of the unmanned monitoring scene to obtain the initial point cloud information of the monitoring scene;

[0007] S2: using the Tof image acquisition device to obtain the first point cloud information in the monitoring scene, aligning it with the initial point cloud information and performing a set difference operation to obtain the second point cloud information, and projecting the second point cloud information on the horizontal plane to obtain a set of personnel point cloud information;

[0008] S3: dilate and erode the binary image obtained by the infrared image acquisition device after thresholding the scene information to obtain a connected area information set; and

[0009] S4: Projecting the personnel point cloud information set and the connected area information set respectively into the image plane space of the RGB image acquisition device using the positional relationship between the calibrated cameras to perform an intersection operation of the sets, and obtaining the corresponding human body area set in response to the common pixels exceeding the first threshold.

[0010] In some specific embodiments, step S1 specifically includes:

[0011] Select an origin at random in an unmanned surveillance scene and establish a three-dimensional coordinate system;

[0012] m*n points are set at intervals in the x-axis and z-axis directions as image acquisition positions of the RGB image acquisition device, and shooting angles are selected at intervals of k degrees for the pitch angle, yaw angle, and roll angle, respectively, to acquire M=m*n*(180 / k)*(180 / k)*(180 / k) images;

[0013] The 3D reconstruction algorithm of Structure from Motion is used to reconstruct the 3D scene of the M images and obtain the initial point cloud information. The STM algorithm can be used to recover the 3D structure from the 2D motion field of the projection of a moving object or scene.

[0014] In some specific embodiments, the first point cloud information and the initial point cloud information are registered using an iterative closest point algorithm. This step can be used to register images acquired by different acquisition devices.

[0015] In some specific embodiments, the second point cloud information is projected on the XY plane of the three-dimensional coordinate system, and a number of circular areas are obtained based on the Hough transform, and the point cloud information corresponding to the same circular area is included in the personnel point cloud information set.

[0016] In some specific embodiments, after the binary image is expanded and eroded in step S3, the process further includes removing the area whose pixel area is smaller than the second threshold. With this step, the image can be processed to obtain the connected area.

[0017] In some specific embodiments, the first threshold is within the range of 20*40-80*160, and the second threshold is within the range of 1000-8196.

[0018] In some specific embodiments, a Tof image acquisition device, an infrared image acquisition device, and an RGB image acquisition device are installed in the monitoring scene, and the positional relationship and posture information of the Tof image acquisition device, the infrared image acquisition device, and the RGB image acquisition device are calculated using the initial point cloud information. The positional relationship and posture information of the three modal image acquisition devices can facilitate the subsequent conversion of feature point clouds.

[0019] In some specific embodiments, the specific method of obtaining the position relationship and posture information includes:

[0020] The Tof image acquisition device is used to obtain the depth image of the monitoring scene, and the iterative closest point algorithm is used to obtain the degree of freedom position of the Tof image acquisition device in the monitoring scene in combination with the initial point cloud information;

[0021] Infrared image acquisition devices and RGB image acquisition devices are used to obtain color images of the monitoring scene. SIFT descriptors and Bag of words feature description algorithms are used to obtain the position and posture information of the infrared image acquisition devices and RGB image acquisition devices based on the BundleAdjustment bundle adjustment algorithm according to the acquired images and initial point cloud information.

[0022] In some specific embodiments, based on the position relationship, posture information and internal parameters of the image acquisition device, the first transformation matrix of the Tof image acquisition device and the infrared image acquisition device, the second transformation matrix of the Tof image acquisition device and the RGB image acquisition device, and the third transformation matrix of the infrared image acquisition device and the RGB image acquisition device are obtained.

[0023] In some specific embodiments, step S4 specifically includes:

[0024] Using the second transformation matrix, the personnel point cloud information is projected onto the image plane space of the RGB image acquisition device according to the camera imaging principle to obtain a first projection area set;

[0025] Using a third transformation matrix, the connected region information is projected onto an image plane space of an RGB image acquisition device to obtain a second projection region set;

[0026] An intersection operation is performed on pixels of the first projection area and set and pixels of the second projection area and set.

[0027] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, any of the above methods is implemented.

[0028] According to a third aspect of the present invention, a multimodal unsupervised pedestrian pixel-level semantic annotation system is proposed, the system comprising:

[0029] Initial point cloud information acquisition unit: configured to perform three-dimensional reconstruction of an unmanned monitoring scene and acquire initial point cloud information of the monitoring scene;

[0030] Personnel point cloud information set acquisition unit: configured to use the Tof image acquisition device to acquire the first point cloud information in the monitoring scene, perform a set difference operation after registering it with the initial point cloud information, obtain the second point cloud information, and project the second point cloud information on the horizontal plane to obtain the personnel point cloud information set;

[0031] A connected region information set acquisition unit: configured to dilate and erode the binary image obtained by the infrared image acquisition device after the scene information is thresholded, so as to obtain a connected region information set; and

[0032] A human body region set acquisition unit is configured to project the personnel point cloud information set and the connected area information set into the image plane space of the RGB image acquisition device, respectively, using the positional relationship between the calibrated cameras to perform set intersection operations, and acquire the corresponding human body region set in response to the common pixels exceeding a first threshold.

[0033] The present invention proposes a multimodal unsupervised pedestrian pixel-level semantic annotation method and system, which integrates the advantages of different modal cameras of Tof image acquisition devices, infrared image acquisition devices and RGB image acquisition devices, and can effectively extract human pixel points in the scene. In pedestrian detection tasks, pixel-level annotation information can be automatically provided for use by machine learning algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and together with the description are used to explain the principles of the present invention. Other embodiments and many expected advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. Other features, objects and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments made with reference to the following drawings:

[0035] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0036] Figure 2 This is a flow chart of a multimodal unsupervised pedestrian pixel-level semantic annotation method according to an embodiment of the present application;

[0037] Figure 3 It is a framework diagram of a multimodal unsupervised pedestrian pixel-level semantic annotation system according to an embodiment of the present application;

[0038] Figure 4 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0039] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0040] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0041] Figure 1 An exemplary system architecture 100 is shown to which the multimodal unsupervised pedestrian pixel-level semantic labeling method according to an embodiment of the present application can be applied.

[0042] like Figure 1 As shown, system architecture 100 may include data server 101, network 102 and main server 103. Network 102 is used to provide a medium for a communication link between data server 101 and main server 103. Network 102 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0043] The main server 103 may be a server that provides various services, such as a data processing server that processes the information uploaded by the data server 101. The data processing server may detect pedestrians and associate and store the detection results in a database.

[0044] It should be noted that the multimodal unsupervised pedestrian pixel-level semantic labeling method provided in the embodiment of the present application is generally executed by the main server 103, and accordingly, the device for semantic analysis of small data sets is generally set in the main server 103.

[0045] It should be noted that the data server and the main server can be hardware or software. When it is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When it is software, it can be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it can be implemented as a single software or software module.

[0046] It should be understood that Figure 1 The numbers of data servers, networks and main servers in the embodiment are only for illustration purposes. Any number of terminal devices, networks and servers may be provided as required.

[0047] According to a multimodal unsupervised pedestrian pixel-level semantic annotation method according to an embodiment of the present application, Figure 2 FIG. 4 is a flowchart of a multimodal unsupervised pedestrian pixel-level semantic annotation method according to an embodiment of the present application. Figure 2 As shown, the method includes:

[0048] S201: Perform 3D reconstruction of the unmanned surveillance scene and obtain the initial point cloud information of the surveillance scene. In the unmanned surveillance scene, the 3D reconstruction algorithm of Structure from Motion is used to reconstruct the 3D surveillance scene of M images and obtain the initial point cloud information. The goal of Structure from Motion (SfM) is to automatically recover the camera motion and scene structure using two or more scenes. It is a self-calibration technology that can automatically complete camera tracking and motion matching.

[0049] S202: Use the Tof image acquisition device to obtain the first point cloud information in the monitoring scene, align it with the initial point cloud information, and perform a set difference operation to obtain the second point cloud information, and project the second point cloud information on the horizontal plane to obtain a set of personnel point cloud information. Use the iterative closest point algorithm to align the first point cloud information with the initial point cloud information. In this step, pedestrians are allowed to enter the monitoring scene, and the second point cloud information is projected on the XY plane of the three-dimensional coordinate system of step S201. Based on the Hough transform, several circular areas are obtained, and the point cloud information corresponding to the same circular area is included in the personnel point cloud information set.

[0050] S203: dilate and erode the binary image obtained by the infrared image acquisition device after the scene information is thresholded to obtain a connected region information set. After dilating and eroding the binary image, further remove the area where the pixel area is less than the second threshold to obtain the connected region information set, wherein the second threshold is within the range of 1000-8196.

[0051] S204: Projecting the personnel point cloud information set and the connected area information set respectively into the image plane space of the RGB image acquisition device using the positional relationship between the calibrated cameras to perform set intersection operations, and obtaining the corresponding human body area set in response to the common pixels exceeding a first threshold.

[0052] In a specific embodiment, a Tof image acquisition device, an infrared image acquisition device, and an RGB image acquisition device are respectively installed in the monitoring scene, and the position relationship and posture information of the Tof image acquisition device, the infrared image acquisition device, and the RGB image acquisition device are respectively calculated using the initial point cloud information:

[0053] The Tof image acquisition device is used to obtain the depth image of the monitoring scene, and the iterative closest point algorithm is used to obtain the degree of freedom position of the Tof image acquisition device in the monitoring scene in combination with the initial point cloud information;

[0054] Infrared image acquisition devices and RGB image acquisition devices are used to obtain color images of the monitoring scene. SIFT descriptors and Bag of words feature description algorithms are used to obtain the position and posture information of the infrared image acquisition devices and RGB image acquisition devices based on the BundleAdjustment bundle adjustment algorithm according to the acquired images and initial point cloud information.

[0055] According to the position relationship, posture information and internal parameters of the image acquisition device, the first transformation matrix of the Tof image acquisition device and the infrared image acquisition device, the second transformation matrix of the Tof image acquisition device and the RGB image acquisition device, and the third transformation matrix of the infrared image acquisition device and the RGB image acquisition device are obtained. The second transformation matrix is ​​used to project the personnel point cloud information to the image plane space of the RGB image acquisition device according to the camera imaging principle to obtain the first projection area set; the third transformation matrix is ​​used to project the connected area information to the image plane space of the RGB image acquisition device to obtain the second projection area set; the pixels of the first projection area and set and the pixels of the second projection area and set are intersected, and through common judgment, the area where the common pixels exceed the first threshold is obtained as the human body area set, and the first threshold is taken in the range of 20*40-80*160.

[0056] According to a specific embodiment of the present invention, a multimodal unsupervised pixel-level semantic annotation method for pedestrians can be implemented by the following steps to detect pedestrians and automatically annotate them. In this method, three acquisition devices are required: a time-of-flight camera (camera A), an infrared thermal imaging camera (camera B), and an RGB color camera (camera C). In the following description, camera A, camera B, and camera C are used to replace them respectively.

[0057] Step 1: Select a monitoring scene, take a point P on the ground, and establish a three-dimensional coordinate system XYZ, where the X-axis is in the horizontal plane and points to a certain direction, the Z-axis is perpendicular to the ground and points to the center of the earth, and the Y-axis is perpendicular to the X-axis in the horizontal plane, and its direction is determined by the right-hand rule;

[0058] Step 2: In the X-axis direction, select a point every 100 cm, and select m points in total as the horizontal shooting position of camera C, denoted as Q1, Q2, ..., Qm; in the Z-axis direction, select a point every 50 cm, and select P1, P2, ..., Pn as the vertical shooting height of the camera; at m*n positions, select a shooting angle every k degrees for the pitch angle, yaw angle, and roll angle. Camera C collects a total of M=m*n*(180 / k)*(180 / k)*(180 / k) images at different positions and angles.

[0059] Step 3: For the M images obtained in step 2, use the Structure-from-Motion technology to reconstruct the scene in three dimensions, so as to obtain the point cloud information of the scene, which is recorded as Scene_Point_Cloud_BG.

[0060] Step 4: Install cameras A, B, and C in the scene respectively, and calculate the relative positional relationship between cameras ABC using the scene point cloud information from step 3. In this step, it must be ensured that there are no moving objects in the scene.

[0061] 4a) After installing camera A, obtain the depth image of the scene, recorded as Depth_Image; use Depth_Image and Scene_Point_Cloud_BG as input, and use the iterative closest point algorithm (ICP) to solve the 6-DOF position (three rotation angles and three translation coordinate information) of camera A in the scene.

[0062] 4b) After cameras B and C are installed, the color image information of the scene is obtained, which are recorded as Color_Image_B and Color_Image_C respectively. The position and posture information of cameras B and C are solved based on the BundleAdjustment algorithm according to the M collected images and the point cloud information of the scene using the SIFT descriptor and Bag-of-Word feature description method.

[0063] 4c) According to the camera pose information obtained in step 4a and step 4b, and the internal parameters of cameras ABC, the transformation matrix Tab between camera A and camera B, the transformation matrix Tac between camera A and camera C, and the transformation matrix Tbc between camera B and camera C are obtained.

[0064] Step 5: After completing steps 1-4 above, open the scene and allow pedestrians to enter. Use camera A to obtain the three-dimensional point cloud information Scene_Point_Cloud_New in the scene, and use the ICP algorithm again to align Scene_Point_Cloud_New and Scene_Point_Cloud_BG; after alignment, perform a set difference operation on the two point cloud sets to obtain a new point cloud Scene_Point_Cloud_FG. Project Scene_Point_Cloud_FG on the XY plane, and obtain several circular areas C1, C2, ..., Cp based on the Hough transform. The point cloud information corresponding to the same circular area Ci is recorded as Person_i.

[0065] Step 6: Using the scene information captured by camera B, a binary image Camera_B_Image_Binary is obtained through thresholding. Dilation and erosion operations are performed on Camera_B_Image_Binary, and areas with pixel areas smaller than a threshold thr (thr is set according to the actual scene and ranges from 1000 to 8196) are removed. The obtained connected area information is recorded as R1, R2, …, Rq.

[0066] Step 7: Based on the transformation matrix Tac between cameras AC obtained in step 4, the point cloud information Person_i (i=1, 2, ..., p) of step 5 is projected into the image plane space of camera C according to the camera imaging principle, and the corresponding region is recorded as Region_From_A_i (i=1, 2, ..., p). Based on the transformation matrix Tbc between cameras BC obtained in step 4, the region Rj (j=1, 2, ..., q) obtained in step 6 is projected into the image plane space of camera C, and the corresponding region is recorded as Region_From_B_j (j=1, 2, ..., q).

[0067] Step 8: Perform set intersection operation on the two region sets obtained in step 7: {Region_From_A_1, ..., Region_From_A_p} and {Region_From_B_1, ..., Region_From_B_q}. When the number of common pixels between Region_From_A_i and Region_From_B_j exceeds the threshold thr_region (set to 20x40 to 80x160), the corresponding human body region Region_From_C_k is obtained, where k≥1 and k≤min(p,q).

[0068] Continue to refer Figure 3 , Figure 3 The framework diagram of a multimodal unsupervised pedestrian pixel-level semantic annotation system according to an embodiment of the present application is shown. The system specifically includes an initial point cloud information acquisition unit 301, a personnel point cloud information set acquisition unit 302, a connected region information set acquisition unit 303, and a human region set acquisition unit 304.

[0069] In a specific embodiment, the initial point cloud information acquisition unit 301 is configured to perform three-dimensional reconstruction of an unmanned monitoring scene and obtain the initial point cloud information of the monitoring scene; the personnel point cloud information set acquisition unit 302 is configured to use a Tof image acquisition device to acquire the first point cloud information in the monitoring scene, align it with the initial point cloud information, and perform a set difference operation to obtain the second point cloud information, and project the second point cloud information on a horizontal plane to obtain a personnel point cloud information set; the connected area information set acquisition unit 303 is configured to dilate and erode the thresholded binary image of the scene information acquired by the infrared image acquisition device to obtain a connected area information set; the human body area set acquisition unit 304 is configured to project the personnel point cloud information set and the connected area information set into the image plane space of the RGB image acquisition device, respectively, to perform a set intersection operation, and in response to the common pixels exceeding the first threshold, obtain the corresponding human body area set.

[0070] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system 400 suitable for implementing an electronic device of an embodiment of the present application. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0071] like Figure 4 As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0072] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.

[0073] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.

[0074] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0075] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0076] The modules involved in the embodiments of the present application may be implemented by software or by hardware.

[0077] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device: includes three-dimensional reconstruction of the unmanned monitoring scene, and obtains the initial point cloud information of the monitoring scene; uses the Tof image acquisition device to obtain the first point cloud information in the monitoring scene, aligns it with the initial point cloud information, and performs a set difference operation to obtain the second point cloud information, and projects the second point cloud information on the horizontal plane to obtain a set of personnel point cloud information; dilates and erodes the thresholded binary image of the scene information obtained by the infrared image acquisition device to obtain a connected area information set; projects the personnel point cloud information set and the connected area information set into the image plane space of the RGB image acquisition device respectively to perform a set intersection operation, and obtains the corresponding human body area set in response to the common pixel exceeding the first threshold.

[0078] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.

Claims

1. A multimodal unsupervised pedestrian pixel-level semantic annotation method, characterized in that: include: S1: Perform three-dimensional reconstruction of an unmanned monitoring scene to obtain initial point cloud information of the monitoring scene; S2: using a Tof image acquisition device to obtain first point cloud information in the monitoring scene, aligning it with the initial point cloud information and performing a set difference operation to obtain second point cloud information, and projecting the second point cloud information on a horizontal plane to obtain a personnel point cloud information set; S3: dilate and erode the binary image obtained by the infrared image acquisition device after thresholding the scene information to obtain a connected region information set; and S4: projecting the person point cloud information set and the connected area information set respectively into the image plane space of the RGB image acquisition device using the positional relationship between the calibrated cameras to perform set intersection operations, and acquiring a corresponding human body area set in response to a common pixel exceeding a first threshold; According to the position relationship, posture information and internal parameters of the image acquisition device, a first transformation matrix between the Tof image acquisition device and the infrared image acquisition device, a second transformation matrix between the Tof image acquisition device and the RGB image acquisition device, and a third transformation matrix between the infrared image acquisition device and the RGB image acquisition device are obtained; S4 specifically includes: Using the second transformation matrix, the personnel point cloud information is projected onto the image plane space of the RGB image acquisition device according to the camera imaging principle to obtain a first projection area set; Projecting the connected region information onto the image plane space of the RGB image acquisition device using the third transformation matrix to obtain a second projection region set; An intersection operation is performed on pixels of the first projection area and set and pixels of the second projection area and set.

2. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 1, characterized in that: The S1 specifically includes: Select an origin at random in the unmanned monitoring scene to establish a three-dimensional coordinate system; m*n points are set at intervals in the x-axis and z-axis directions as image acquisition positions of the RGB image acquisition device, and shooting angles are selected at intervals of k degrees for the pitch angle, yaw angle, and roll angle, respectively, to acquire M=m*n*(180 / k)*(180 / k)*(180 / k) images; The three-dimensional reconstruction of the monitoring scene is performed on the M images using the three-dimensional reconstruction algorithm of Structure from motion, and the initial point cloud information is obtained.

3. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 1, characterized in that: The first point cloud information and the initial point cloud information are registered using an iterative closest point algorithm.

4. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 2, characterized in that: The second point cloud information is projected on the XY plane of the three-dimensional coordinate system, and a number of circular areas are obtained based on the Hough transform, and the point cloud information corresponding to the same circular area is included in the personnel point cloud information set.

5. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 1, characterized in that: After the binary image is expanded and eroded in S3, the process further includes removing areas where the pixel area is smaller than the second threshold.

6. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 5, characterized in that: The first threshold is within the range of 20*40-80*160, and the second threshold is within the range of 1000-8196.

7. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 1, characterized in that: A Tof image acquisition device, an infrared image acquisition device and an RGB image acquisition device are respectively installed in the monitoring scene, and the position relationship and posture information of the Tof image acquisition device, the infrared image acquisition device and the RGB image acquisition device are respectively calculated using the initial point cloud information.

8. The multimodal unsupervised pedestrian pixel-level semantic labeling method according to claim 7, characterized in that: The specific methods for obtaining the position relationship and posture information include: Using a Tof image acquisition device to acquire a depth image of the monitoring scene, and using an iterative closest point algorithm in combination with initial point cloud information to acquire the degree of freedom position of the Tof image acquisition device in the monitoring scene; The infrared image acquisition device and the RGB image acquisition device are used to acquire a color image of the monitoring scene, and the position and posture information of the infrared image acquisition device and the RGB image acquisition device are acquired based on the Bundle Adjustment algorithm according to the acquired image and the initial point cloud information using a SIFT descriptor and a Bag of words feature description algorithm.

9. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 8 is implemented.

10. A multimodal unsupervised pedestrian pixel-level semantic annotation system, characterized in that: The system comprises: An initial point cloud information acquisition unit: configured to perform three-dimensional reconstruction on an unmanned monitoring scene and acquire initial point cloud information of the monitoring scene; A personnel point cloud information set acquisition unit: configured to acquire the first point cloud information in the monitoring scene by using a Tof image acquisition device, perform a set difference operation after registering the first point cloud information with the initial point cloud information to obtain the second point cloud information, and project the second point cloud information on a horizontal plane to obtain a personnel point cloud information set; A connected region information set acquisition unit: configured to dilate and erode the binary image obtained by the infrared image acquisition device after the scene information is thresholded, so as to obtain a connected region information set; and A human body region set acquisition unit: configured to project the personnel point cloud information set and the connected region information set into the image plane space of the RGB image acquisition device respectively, using the positional relationship between the calibrated cameras, to perform a set intersection operation, and acquire a corresponding human body region set in response to a common pixel exceeding a first threshold; According to the position relationship, posture information and internal parameters of the image acquisition device, a first transformation matrix between the Tof image acquisition device and the infrared image acquisition device, a second transformation matrix between the Tof image acquisition device and the RGB image acquisition device, and a third transformation matrix between the infrared image acquisition device and the RGB image acquisition device are obtained; the human body region set acquisition unit is specifically configured to: Using the second transformation matrix, the personnel point cloud information is projected onto the image plane space of the RGB image acquisition device according to the camera imaging principle to obtain a first projection area set; Projecting the connected region information onto the image plane space of the RGB image acquisition device using the third transformation matrix to obtain a second projection region set; An intersection operation is performed on pixels of the first projection area and set and pixels of the second projection area and set.

Citation Information

Patent Citations

  • Plane detection method, computing device and circuit system

    CN110458805A

  • Map construction method and device

    CN111882611A