A target gesture recognition method, device, equipment and medium
By using multiple cameras and image processing technology, combined with the position and angle information of the cameras, the three-dimensional attitude trajectory of targets such as ammunition is determined, which solves the problem that the two-dimensional attitude trajectory in the existing technology cannot reflect the true three-dimensional attitude, and realizes accurate three-dimensional attitude analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA ORDNANCE SCI INST
- Filing Date
- 2022-10-19
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies cannot accurately determine the attitude trajectory of targets such as munitions in three-dimensional space, resulting in insufficient accuracy in analysis.
By using multiple cameras at different locations to capture multiple images in a short period of time, and combining the camera coordinates, azimuth angle, and pitch angle, the central axis and spatial plane of the target are determined in the world coordinate system. The central axis of the target is extracted using the moment method and image processing technology. Finally, the three-dimensional attitude trajectory is determined by combining the spatial plane information at multiple time points.
It enables precise measurement of the attitude trajectory of targets such as munitions in three-dimensional space, improving the accuracy of attitude analysis.
Smart Images

Figure CN115457135B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a target pose recognition method, apparatus, device, and medium. Background Technology
[0002] Currently, by using high-speed cameras to photograph test munitions, multiple frames of images of the munition can be obtained in a very short time. By combining the position of the munition in the multiple frames, the attitude trajectory of the munition in the imaging plane can be obtained.
[0003] However, this method only yields a two-dimensional attitude trajectory within the imaging plane, which cannot fully reflect the actual attitude trajectory of the ammunition in real three-dimensional space, thus affecting the accuracy of users' attitude trajectory analysis of ammunition, etc. Therefore, how to accurately determine the attitude trajectory of ammunition, etc. in three-dimensional space is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides a target attitude recognition method, device, equipment, and medium for accurately determining the attitude trajectory of targets such as ammunition in three-dimensional space.
[0005] Firstly, this application provides a target pose recognition method, the method comprising:
[0006] During the first time period, M*N images are captured by M cameras, wherein the M cameras are located at different positions, and the N images captured by each of the M cameras correspond to N moments within the first time period, where M and N are integers greater than or equal to 2.
[0007] For each of the M cameras, determine the central axis of the target in each of the N images captured by the camera;
[0008] For each of the M cameras, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images, N spatial planes including the central axis of the target and the camera are determined in the world coordinate system.
[0009] For each of the N time points, based on the M spatial planes corresponding to the M images captured by the M cameras at that time point, the attitude information of the central axis of the target at that time point in the world coordinate system is determined;
[0010] Based on the attitude information of the target's central axis in the world coordinate system at each of the N time points, the attitude trajectory of the target in the world coordinate system during the first time period is determined.
[0011] Further, determining the central axis of the target in each of the N images captured by each of the M cameras includes:
[0012] Acquire at least one image captured by the camera that does not include the target;
[0013] Based on the at least one image that does not include the target, determine the background of the N images captured by the camera;
[0014] Based on the background of the N images captured by the camera, perform differential processing on the N images captured by the camera to obtain N images with the background removed;
[0015] Determine the central axis of the target in each of the N background-removed images.
[0016] Further, determining the central axis of the target in each of the N background-removed images includes:
[0017] The moment method is used to determine the target region containing the target in each of the N background-removed images;
[0018] Based on the target region in each of the N background-removed images, determine the central axis of the target in each of the N background-removed images.
[0019] Further, the step of using the moment method to determine the target region containing the target in each of the N background-removed images includes:
[0020] Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region;
[0021] Step 2: Based on the cumulative and distributed projected gray values of the first region along the central axis and perpendicular to the central axis of the first region, respectively, and the selected proportions of the cumulative projected gray values along the central axis and perpendicular to the central axis of the first region, adjust the size of the first region to obtain the second region; within the second region, use the moment method to determine the third region containing the target, and determine the central axis of the third region;
[0022] Step 3: If the angle between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the return to Step 2.
[0023] Furthermore, the method also includes:
[0024] The N background-removed images are preprocessed, and the preprocessing includes one or more of the following: filtering, binarization, and line smoothing.
[0025] Secondly, this application provides a target pose recognition device, the device comprising:
[0026] The acquisition module is used to capture M*N images through M cameras within a first time period, wherein the M cameras are located at different positions, and the N images captured by each of the M cameras correspond to N moments within the first time period, where M and N are integers greater than or equal to 2.
[0027] The processing module is used to determine the central axis of the target in each of the N images captured by each of the M cameras.
[0028] The processing module is further configured to, for each of the M cameras, determine N spatial planes in the world coordinate system that include the central axis of the target and the camera, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images.
[0029] The processing module is further configured to, for each of the N time moments, determine the attitude information of the central axis of the target in the world coordinate system based on the M spatial planes corresponding to the M images captured by the M cameras at that time moment;
[0030] The processing module is further configured to determine the attitude trajectory of the target in the world coordinate system within the first time period based on the attitude information of the target's central axis in the world coordinate system at each of the N time periods.
[0031] Furthermore, the processing module, for each of the M cameras, determines the central axis of the target in each of the N images captured by the camera, specifically for:
[0032] Acquire at least one image captured by the camera that does not include the target;
[0033] Based on the at least one image that does not include the target, determine the background of the N images captured by the camera;
[0034] Based on the background of the N images captured by the camera, perform differential processing on the N images captured by the camera to obtain N images with the background removed;
[0035] Determine the central axis of the target in each of the N background-removed images.
[0036] Furthermore, the processing module is also used for:
[0037] The N background-removed images are preprocessed, and the preprocessing includes one or more of the following: filtering, binarization, and line smoothing.
[0038] Furthermore, when the processing module determines the central axis of the target in each of the N background-removed images, it specifically performs the following:
[0039] The moment method is used to determine the target region containing the target in each of the N background-removed images;
[0040] Based on the target region in each of the N background-removed images, determine the central axis of the target in each of the N background-removed images.
[0041] Furthermore, when the processing module uses the moment method to determine the target region containing the target in each of the N background-removed images, it specifically performs the following:
[0042] Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region;
[0043] Step 2: Based on the cumulative and distributed projected gray values of the first region along the central axis and perpendicular to the central axis of the first region, respectively, and the selected proportions of the cumulative projected gray values along the central axis and perpendicular to the central axis of the first region, adjust the size of the first region to obtain the second region; within the second region, use the moment method to determine the third region containing the target, and determine the central axis of the third region;
[0044] Step 3: If the angle between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the return to Step 2.
[0045] Thirdly, this application provides an electronic device, which includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of any of the target pose recognition methods described above.
[0046] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the target pose recognition methods described above.
[0047] In this application, the target's attitude is captured by M cameras at different positions within a first time period. The N images captured by each camera at N time points are processed and analyzed to determine the target's central axis in the images. Then, combining the position, azimuth, and elevation angles of each camera, N spatial planes including the target's central axis and the cameras are determined in the world coordinate system. Finally, based on the M spatial planes corresponding to the M images captured by the M cameras at each time point, the target's attitude information in the world coordinate system at that time is determined, thereby accurately determining the target's trajectory in three-dimensional space. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 A flowchart illustrating a target pose recognition method provided in this application embodiment;
[0050] Figure 2 A schematic diagram illustrating how the central axis of a target is determined using a spatial plane corresponding to two cameras, as provided in this application embodiment;
[0051] Figure 3 A flowchart of an algorithm for determining the centerline of a target in an image is provided in this application embodiment;
[0052] Figure 4aThis is a schematic diagram illustrating how the length and width of a second region are determined based on the projected grayscale values along the X-axis, as provided in an embodiment of this application.
[0053] Figure 4b This is a schematic diagram illustrating how the length and width of a second region are determined based on the projected grayscale values along the Y-axis, as provided in an embodiment of this application.
[0054] Figure 5 This application provides an embodiment of a schematic diagram illustrating the acquisition of a target region containing a target using the method described in this application.
[0055] Figure 6 This is a schematic diagram of the structure of a target pose recognition device provided in an embodiment of this application;
[0056] Figure 7 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation
[0057] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0058] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0059] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0060] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0061] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0063] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
[0064] Currently, by using high-speed cameras to photograph targets such as test munitions, multiple frames of images of the munitions and other targets can be obtained in a very short time. By combining the positions of the munitions and other targets in the multiple frames, the attitude trajectory of the munitions and other targets in the imaging plane can be obtained.
[0065] However, this method only yields a two-dimensional attitude trajectory within the imaging plane, which cannot fully and accurately reflect the attitude trajectory of targets such as munitions in real three-dimensional space, thus affecting the accuracy of users' attitude trajectory analysis of such targets. Therefore, how to accurately determine the attitude trajectory of targets such as munitions in three-dimensional space is a technical problem that urgently needs to be solved.
[0066] In order to accurately determine the attitude trajectory of a target in three-dimensional space, embodiments of this application provide a target attitude recognition method, apparatus, device, and storage medium.
[0067] Example 1:
[0068] Figure 1 A flowchart of a target pose recognition method provided in this application embodiment. This method is mainly applied to electronic devices, such as computers and servers. The method includes:
[0069] S101: During the first time period, M*N images are captured by M cameras, wherein the M cameras are located at different positions, and the N images captured by each of the M cameras correspond to N moments within the first time period, where M and N are integers greater than or equal to 2.
[0070] It should be understood that in the embodiments of this application, the target for which attitude estimation needs to be obtained can be an object such as ammunition or an aircraft, and this application does not limit it.
[0071] During the first time period, electronic devices use M cameras to record video. After recording, each camera generates N images from the video, with each image corresponding to a moment within the first time period. For example, the first camera starts recording at 8:00:00, capturing a 1-minute video. This 1-minute video is then used to generate one image per second (i.e., the first image is generated at 8:00:01, the second image at 8:00:02, and so on). The other cameras also start recording at 8:00:00, capturing a 1-minute video, and generating one image per second from this 1-minute video.
[0072] The first time period can include the period when the target enters the camera's field of view and the period before the target enters the camera's field of view (i.e., the camera starts shooting before the target enters the camera's field of view); or it can be just the period when the target enters the camera's field of view.
[0073] Before using the camera to photograph the target, the camera can be placed in a designated position and fixed. Adjust the camera's shooting direction so that it can capture more of the target's trajectory. The camera's coordinates in the world coordinate system, as well as the camera's (specifically, the camera lens or the camera's optical center) azimuth and elevation angles in the world coordinate system, can be recorded.
[0074] Optionally, in order to make the attitude trajectory of the target such as ammunition obtained by this method more accurate in three-dimensional space, a high-speed camera can be used to photograph the target in this embodiment of the application.
[0075] S102: For each of the M cameras, determine the central axis of the target in each of the N images captured by the camera;
[0076] Electronic devices can use methods such as contour recognition and Hough transform to determine the central axis of the target in each of the N images captured by each of the M cameras.
[0077] In some implementations, to more accurately determine the central axis of the target in each image, the electronic device determines the central axis of the target in each of the N images captured by the camera, including:
[0078] The moment method is used to determine the region containing the target in each of the N images captured by the camera;
[0079] Based on the region containing the target in each of the N images captured by the camera, determine the central axis of the target in each of the N images captured by the camera.
[0080] The method of moments for determining the region containing the target in each image can be performed using the following steps:
[0081] Step 1: The locations of the parallel bounding boxes of the detected targets in the image are designated as regions of interest (ROIs).
[0082] Step 2: Preprocess the ROI region (grayscale conversion, binarization);
[0083] Step 3: Calculate the moments of the object using the moment method; where the moments of the object are obtained based on the pixels of the object in the image.
[0084] Step 4: Use moments to calculate the center point, deflection angle, and major and minor axes of the object.
[0085] S103: For each of the M cameras, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images, determine N spatial planes in the world coordinate system that include the central axis of the target and the camera.
[0086] After acquiring the central axis of the target in each image for each camera, the electronic device can determine the position of the target's central axis in the world coordinate system (e.g., the equation of the target's central axis in each image) based on the recorded coordinates of the camera in the world coordinate system and the azimuth and elevation angles of each of the N images taken by the camera. Based on the camera's coordinates in the world coordinate system (i.e., the optical center coordinates of the camera, i.e., a point) and the position of the target's central axis in the world coordinate system (i.e., a straight line), N spatial planes corresponding to the N images can be determined, including the target's central axis in the spatial coordinate system and the camera's spatial plane.
[0087] S104: For each of the N time moments, based on the M spatial planes corresponding to the M images captured by the M cameras at that time moment, determine the attitude information of the central axis of the target in the world coordinate system at that time moment;
[0088] Taking the example of determining the attitude information of the target's central axis in the world coordinate system at a certain moment by using the spatial plane corresponding to two cameras. Figure 2 This is a schematic diagram illustrating the determination of the target's central axis using the spatial plane corresponding to the two cameras. (Example) Figure 2As shown, XYZ is the world coordinate system, and O is the origin of the world coordinate system. O1 is the coordinate point of the first camera (e.g., the optical center of the first camera) in the world coordinate system. The plane at O1 is the image plane of the first camera, and l on the image plane of the first camera is the image of target L on this image plane. Similarly, O2 is the coordinate point of the second camera in the world coordinate system. The plane at O2 is the image plane of the second camera, and l′ on the image plane of the second camera is the image of target L on this image plane. After obtaining the spatial plane containing O1 and L and the spatial plane containing O2 and L, the intersection of the two spatial planes is the central axis of target L in the world coordinate system. From this, the attitude information of the target in the world coordinate system at this moment can be known. Among them, the attitude information of the target in the world coordinate system at this moment includes the angle of attack, angle of attack, and other information of the target in the world coordinate system at this moment, and may also include the position information of the target in the world coordinate system.
[0089] When M images are captured simultaneously by M cameras, and each of these M images corresponds to one of M spatial planes, a central axis of the target in the world coordinate system is determined using each pair of these M spatial planes. This yields a total of... The central axis of the target in the world coordinate system is then obtained. The attitude information of each target in the world coordinate system will The final attitude information of a target in the world coordinate system is obtained by averaging the attitude information of several targets in the world coordinate system and combining them according to certain weights.
[0090] S105: Based on the attitude information of the target's central axis in the world coordinate system at each of the N time points, determine the target's attitude trajectory in the world coordinate system during the first time period.
[0091] In this application, the target's attitude is captured by M cameras at different positions within a first time period. The N images captured by each camera at N time points are processed and analyzed to determine the target's central axis in the images. Then, combining the position, azimuth, and pitch angles of each camera, N spatial planes, including the target's central axis and the camera's position, are determined in the world coordinate system. Finally, based on the M spatial planes corresponding to the M images captured by the M cameras at each time point, the attitude information of the target's central axis in the world coordinate system at that time point is determined. Furthermore, the attitude trajectory of the high-speed target in the world coordinate system is obtained based on the attitude information of the target's central axis in the world coordinate system at each time point. This application enhances the ability to recognize the attitude of high-speed targets by capturing images of the high-speed target's attitude using multiple high-speed cameras and combining image fusion processing and analysis, thereby achieving accurate measurement of the high-speed target's attitude.
[0092] Furthermore, to enhance the accuracy of high-speed target pose recognition, based on the above embodiments, in this embodiment, determining the central axis of the target in each of the N images captured by each of the M cameras includes:
[0093] Acquire at least one image captured by the camera that does not include the target;
[0094] Based on the at least one image that does not include the target, determine the background of the N images captured by the camera;
[0095] Based on the background of the N images captured by the camera, perform differential processing on the N images captured by the camera to obtain N images with the background removed;
[0096] Determine the central axis of the target in each of the N background-removed images.
[0097] For each camera, at least one additional image excluding the target can be obtained outside the first time period, or at least one image excluding the target can be selected from the N images captured by the camera (i.e., images captured a short period before the second time period). The average pixel value of this at least one image excluding the target is then calculated and used as the background for the N images captured by the camera. The first few images excluding the target from the N images captured by the camera can be selected manually, or the pixel values of two images taken at adjacent times can be subtracted. If the pixel value change exceeds a certain threshold, the image at the next time period is considered to contain the target; that is, the first few images captured at that time period are images excluding the target.
[0098] For each camera, N images captured by that camera are subtracted from the background of N other images captured by the same camera (i.e., the pixel values of the two images are subtracted), resulting in N background-removed images. Similarly, inter-frame differencing can be used to determine whether a target has entered the camera's field of view. The pixel values of each pair of adjacent background-removed images are subtracted; if the pixel value change exceeds a certain threshold, the differenced image is considered to contain the target. Thus, for each camera, there are P differenced images containing the target, where P is a positive integer less than N.
[0099] Furthermore, in order to reduce noise in the image and optimize the shape of the target in the image, thereby better extracting the central axis of the target, the method further includes:
[0100] The N background-removed images are preprocessed, and the preprocessing includes one or more of the following: filtering, binarization, and line smoothing.
[0101] For each camera, Gaussian smoothing filtering is applied to P images containing the target, after N background removal images, to eliminate isolated noise points. Binarization is then performed using the Otsu's maximum inter-class variance (OTC) thresholding method, followed by morphological opening and closing operations to smooth the target edges. Finally, Hough transform is used to fit the target's edge lines in the image, removing smaller edges from the target's edge lines, and constructing a regular rectangular target shape using the remaining edge lines.
[0102] Because the target is ammunition, the dynamic damage range of the ammunition is large, the high-speed camera is far from the target, the target size in the image is small, and the image quality is poor. Furthermore, the target's local area also contains noise and other interfering objects. Directly using the moment method to obtain the central axis of the region containing the target in the image is ineffective. To further improve the accuracy of the moment method in extracting the central axis of the target region, in this embodiment, the step of using the moment method to determine the target region containing the target in each of the N background-removed images includes:
[0103] Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region;
[0104] Step 2: Based on the cumulative and distributed projected gray values of the first region along the central axis and perpendicular to the central axis of the first region, respectively, and the selected proportions of the cumulative projected gray values along the central axis and perpendicular to the central axis of the first region, adjust the size of the first region to obtain the second region; within the second region, use the moment method to determine the third region containing the target, and determine the central axis of the third region;
[0105] Step 3: If the angle between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the return to Step 2.
[0106] Because the moment method calculates all grayscale values in an image, it will not yield accurate results if the area containing the target is near noise or other interfering objects. In actual captured images, the camera is far from the target, and the target size is very small. The target can be considered as a long, narrow rectangle in the image, with its grayscale values symmetrically and uniformly distributed along the central axis and Gaussian distributed along the perpendicular direction of the central axis. Based on this, this application provides an algorithm to more accurately determine the region containing the target in an image. Figure 3 The flowchart of the algorithm for determining the centerline of a target in an image is as follows: Figure 3 As shown:
[0107] S301: Determine the first region and its central axis;
[0108] For each of the N background-removed images, the moment method is first used to obtain an initial first region containing the target and the central axis of the first region, which is a rectangular box containing the target.
[0109] S302: Based on the projection distribution of the gray values of the first region onto the X and Y axes, obtain the length, width, and centroid coordinates of the second region;
[0110] Establish X-axis and Y-axis along the central axis and perpendicular line of the central axis of the first region, respectively, and project the gray values in the X-axis and Y-axis directions.
[0111] The length of the second region is determined by selecting a coordinate range where the sum of the projected grayscale values in the X-direction is greater than a selection ratio along the central axis of the first region. The center of this coordinate range is then used as the x-coordinate of the centroid of the second region. Since the grayscale values of the target are symmetrically and uniformly distributed along the central axis in each of the N background-removed images, the selection ratio of the sum of the projected grayscale values along the central axis of the first region can be 2 / 3 of the maximum sum of projected grayscale values. This selection ratio can be set according to the actual situation.
[0112] The width of the second region is defined as the range where the sum of the projected grayscale values in the Y-direction is greater than a selection ratio along the perpendicular direction of the central axis of the first region. The center of this range is then used as the ordinate of the centroid of the second region. Since the grayscale values of the target in each of the N background-removed images exhibit a Gaussian distribution along the perpendicular direction of the central axis, the selection ratio of the sum of the projected grayscale values along the perpendicular direction of the central axis of the first region can be 1 / 3 of the maximum sum of projected grayscale values. This selection ratio can be set according to actual conditions.
[0113] S303: Determine the third region containing the target within the second region and its central axis;
[0114] The second region is determined in the image based on its length and width, as well as the x and y coordinates of its centroid. The moment method is then used to obtain the third region containing the target within the second region, and the central axis of the third region.
[0115] S304: Determine whether the angle difference between the axis of the third region and the axis of the first region is less than a set threshold.
[0116] Calculate the angle difference between the central axis of the third region and the central axis of the first region. If the angle difference is less than or equal to a set threshold, select the third region as the target region containing the target. If the angle difference is greater than the set threshold, use the third region as the first region containing the target and return to step S302 until the angle difference between the central axis of the third region and the central axis of the first region is less than or equal to the set threshold. Alternatively, after repeating the steps a set number of times, end the algorithm and use the final obtained third region as the target region containing the target.
[0117] S305: Outputs the centerline of the target region containing the target.
[0118] Figure 4a This is a schematic diagram illustrating the determination of the length and width of the second region based on the projected grayscale values along the X-axis. Figure 4b This is a schematic diagram illustrating the determination of the length and width of the second region based on the projected grayscale values along the Y-axis. (See diagram below.) Figure 4a The vertical axis represents the sum of the projected gray values of the first region projected onto the X-axis, and the horizontal axis represents the X-axis. Since the gray values of the target in each of the N background-removed images are symmetrically and uniformly distributed along the central axis (i.e., the X-axis direction), the selection ratio of the projected gray value sum along the central axis direction of the first region can be 2 / 3 of the maximum projected gray value sum. That is, the portion in the figure where the projected gray value sum is greater than 5500 (i.e., the portion of length H in the figure) is selected as the coordinate range where the projected gray value sum projected onto the X-axis is similar to the plateau function distribution. Figure 4b The vertical axis represents the sum of the projected gray values of the first region projected onto the Y-axis, and the horizontal axis represents the Y-axis. Since the gray values of the target in each of the N background-removed images are Gaussian distributed along the vertical direction of the central axis (i.e., the Y-axis direction), the selection ratio of the projected gray value sum along the vertical direction of the central axis of the first region can be 1 / 3 of the maximum projected gray value sum. That is, the portion in the figure where the projected gray value sum is greater than 3300 (i.e., the portion of length W in the figure) is selected as the coordinate range where the projected gray value sum projected onto the Y-axis is similar to the Gaussian function distribution.
[0119] Figure 5 This is a schematic diagram illustrating the acquisition of a target region containing a target using the method described in the embodiments of this application. For example... Figure 5As shown, the initial region (i.e., the first region) may contain noise other than the target, which will affect the moment method's calculation of the region containing the target. After using the algorithm of this application embodiment to determine the target's central axis in the image, the accuracy is continuously optimized by updating the first region to obtain the third region, thus obtaining the target region containing the target and its central axis, which can improve the accuracy of the moment method in extracting the central axis of the target region.
[0120] Based on the above embodiments, this application also provides a target pose recognition device. Figure 6 This is a schematic diagram of the structure of a target pose recognition device provided in an embodiment of this application, as shown below. Figure 6 As shown, the device includes:
[0121] The acquisition module 601 is used to capture M*N images through M cameras within a first time period, wherein the M cameras are located at different positions, and the N images captured by each of the M cameras correspond to N times within the first time period, and M and N are integers greater than or equal to 2.
[0122] Processing module 602 is used to determine the central axis of the target in each of the N images captured by each of the M cameras.
[0123] The processing module 602 is further configured to, for each of the M cameras, determine N spatial planes in the world coordinate system that include the central axis of the target and the camera, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images.
[0124] The processing module 602 is further configured to, for each of the N times, determine the attitude information of the central axis of the target in the world coordinate system at that time based on the M spatial planes corresponding to the M images captured by the M cameras at that time;
[0125] The processing module 602 is further configured to determine the attitude trajectory of the target in the world coordinate system during the first time period based on the attitude information of the target's central axis in the world coordinate system at each of the N time periods.
[0126] The processing module 602, for each of the M cameras, determines the central axis of the target in each of the N images captured by the camera, specifically for:
[0127] Acquire at least one image captured by the camera that does not include the target;
[0128] Based on the at least one image that does not include the target, determine the background of the N images captured by the camera;
[0129] Based on the background of the N images captured by the camera, perform differential processing on the N images captured by the camera to obtain N images with the background removed;
[0130] Determine the central axis of the target in each of the N background-removed images.
[0131] The processing module 602 is further configured to:
[0132] The N background-removed images are preprocessed, and the preprocessing includes one or more of the following: filtering, binarization, and line smoothing.
[0133] When the processing module 602 determines the central axis of the target in each of the N background-removed images, it is specifically used for:
[0134] The moment method is used to determine the target region containing the target in each of the N background-removed images;
[0135] Based on the target region in each of the N background-removed images, determine the central axis of the target in each of the N background-removed images.
[0136] When the processing module 602 uses the moment method to determine the target region containing the target in each of the N background-removed images, it is specifically used for:
[0137] Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region;
[0138] Step 2: Based on the cumulative and distributed projected gray values of the first region along the central axis and perpendicular to the central axis of the first region, respectively, and the selected proportions of the cumulative projected gray values along the central axis and perpendicular to the central axis of the first region, adjust the size of the first region to obtain the second region; within the second region, use the moment method to determine the third region containing the target, and determine the central axis of the third region;
[0139] Step 3: If the angle between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the return to Step 2.
[0140] The device can be deployed in a terminal computing device connected to a high-speed camera.
[0141] Based on the above embodiments, this application also provides an electronic device. Figure 7 This is a schematic diagram of an electronic device structure provided in this application. Figure 7 As shown, it includes: processor 701, communication interface 702, memory 703 and communication bus 704, wherein processor 701, communication interface 702 and memory 703 communicate with each other through communication bus 704.
[0142] The memory 703 stores a computer program, which, when executed by the processor 701, causes the processor 701 to complete the steps of any of the target pose recognition methods described above.
[0143] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0144] The communication interface 702 is used for communication between the above-mentioned electronic device and other devices.
[0145] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0146] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0147] Based on the above embodiments, the present invention provides a computer-readable storage medium storing a computer program executable by an electronic device, wherein computer-executable instructions are used to cause a computer to execute the process performed by any of the aforementioned target pose recognition methods.
[0148] The aforementioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor in an electronic device, including but not limited to magnetic storage such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), optical storage such as CDs, DVDs, BDs, HVDs, etc., and semiconductor storage such as ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs), etc.
[0149] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0151] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0152] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0153] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A target pose recognition method, characterized in that, The method includes: During the first time period, M was filmed using M cameras. N images, wherein the positions of the M cameras are different, and the N images taken by each of the M cameras correspond to N moments within the first time period, wherein M and N are integers greater than or equal to 2; the M cameras take pictures at the same time, and the position of each camera is fixed; wherein the M cameras are all high-speed cameras. For each of the M cameras, determine the central axis of the target in each of the N images captured by the camera; wherein the target includes ammunition; For each of the M cameras, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images, N spatial planes including the central axis of the target and the camera are determined in the world coordinate system. For each of the N time points, based on the M spatial planes corresponding to the M images captured by the M cameras at that time point, the attitude information of the central axis of the target at that time point in the world coordinate system is determined; Based on the attitude information of the target's central axis in the world coordinate system at each of the N time points, determine the target's attitude trajectory in the world coordinate system during the first time period; Determining the central axis of the target in each of the N images captured by each of the M cameras includes: Acquire at least one image captured by the camera that does not include the target; determine the background of the N images captured by the camera based on the at least one image that does not include the target; perform differential processing on the N images captured by the camera based on the background of the N images captured by the camera to obtain N images with the background removed; The N background-removed images are preprocessed, including: Gaussian smoothing filtering of the images containing the target and after background removal to eliminate isolated noise points; binarization based on the maximum inter-class variance threshold segmentation method, and morphological opening and closing operations to smooth the edges of the target; fitting the edge lines of the target in the image using Hough transform, removing some edges on the target edge lines, and constructing a regular target rectangle with the remaining edge lines. The moment method is used to determine the target region containing the target in each of the N preprocessed background-removed images; based on the target region in each of the N background-removed images, the central axis of the target in each of the N background-removed images is determined; The ammunition is configured as a long, narrow rectangle in the image, and its grayscale values are symmetrically and uniformly distributed along the central axis and Gaussian distributed along the perpendicular line to the central axis. The method of moments is used to determine the target region containing the target in each of the N background-removed images, including: Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region; Step 2: Select a first coordinate range where the sum of the projected grayscale values along the central axis of the first region is greater than a selected proportion of the sum of the projected grayscale values along the central axis of the first region, and use the center of the first coordinate range as the abscissa of the centroid of the second region; select a second coordinate range where the sum of the projected grayscale values along the perpendicular line of the central axis of the first region is greater than a selected proportion of the sum of the projected grayscale values along the perpendicular line of the central axis of the first region, and use the center of the second coordinate range as the ordinate of the centroid of the second region; adjust the size of the first region according to the length and width of the second region, and the abscissa and ordinate of the centroid of the second region to obtain the second region; within the second region, use the moment method to determine a third region containing the target, and determine the central axis of the third region; Step 3: If the angle difference between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle difference between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the return to Step 2.
2. A target posture recognition device, characterized in that, The device includes: The acquisition module is used to capture M images from M cameras within a first time period. N images, wherein the positions of the M cameras are different, and the N images taken by each of the M cameras correspond to N moments within the first time period, wherein M and N are integers greater than or equal to 2; the M cameras take pictures at the same time, and the position of each camera is fixed; wherein the M cameras are all high-speed cameras. The processing module is configured to, for each of the M cameras, determine the central axis of the target in each of the N images captured by the camera; and, for each of the M cameras, determine N spatial planes in the world coordinate system that include the central axis of the target and the camera, based on the central axis of the target in each of the N images captured by the camera, the coordinates of the camera in the world coordinate system, and the azimuth and elevation angles of the camera when capturing each of the N images; wherein the target includes ammunition; The processing module is further configured to, for each of the N time moments, determine the attitude information of the target's central axis in the world coordinate system based on the M spatial planes corresponding to the M images captured by the M cameras at that time moment; and determine the attitude trajectory of the target in the world coordinate system within the first time period based on the attitude information of the target's central axis in the world coordinate system at each of the N time moments. The processing module is specifically used for: Acquire at least one image captured by the camera that does not include the target; determine the background of the N images captured by the camera based on the at least one image that does not include the target; perform differential processing on the N images captured by the camera based on the background of the N images captured by the camera to obtain N images with the background removed; The N background-removed images are preprocessed, including: Gaussian smoothing filtering of the images containing the target and after background removal to eliminate isolated noise points; binarization based on the maximum inter-class variance threshold segmentation method, and morphological opening and closing operations to smooth the edges of the target; fitting the edge lines of the target in the image using Hough transform, removing some edges on the target edge lines, and constructing a regular target rectangle with the remaining edge lines. Determine the central axis of the target in each of the N preprocessed background-removed images; The processing module is specifically used for: Step 1: For each of the N background-removed images, use the moment method to determine the first region containing the target, and determine the central axis of the first region; Step 2: Select a first coordinate range where the sum of the projected grayscale values along the central axis of the first region is greater than a selected proportion of the sum of the projected grayscale values along the central axis of the first region, and use the center of the first coordinate range as the abscissa of the centroid of the second region; select a second coordinate range where the sum of the projected grayscale values along the perpendicular line of the central axis of the first region is greater than a selected proportion of the sum of the projected grayscale values along the perpendicular line of the central axis of the first region, and use the center of the second coordinate range as the ordinate of the centroid of the second region; adjust the size of the first region according to the length and width of the second region, and the abscissa and ordinate of the centroid of the second region to obtain the second region; within the second region, use the moment method to determine a third region containing the target, and determine the central axis of the third region; Step 3: If the angle difference between the central axis of the third region and the central axis of the first region is less than a set threshold, the third region is designated as the target region containing the target; if the angle difference between the central axis of the third region and the central axis of the first region is greater than the set threshold, the third region and its central axis are designated as the central axis of the first region and the first region, and the process returns to Step 2. The ammunition is configured as a long, narrow rectangle in the image, and the grayscale values of the ammunition are symmetrically and uniformly distributed along the central axis and Gaussian distributed along the perpendicular line of the central axis.
3. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is used to implement the steps of the target pose recognition method as described in claim 1 when executing a computer program stored in the memory.
4. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the target pose recognition method as described in claim 1.