Video processing device, video processing method, video processing program, and computer program product

By integrating 2D image data with 3D scanned data, the device accurately estimates 3D posture information, addressing inaccuracies in conventional methods and enhancing precision and robustness in complex scenarios.

JP2025187858APending Publication Date: 2025-12-25DENSO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024096951
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Conventional video processing devices struggle to accurately estimate 3D posture information due to difficulties in recognizing the distance to feature points using 2D camera images, leading to inaccuracies in 3D posture recognition.

Method used

The device integrates two-dimensional captured image data with three-dimensional scanned data, utilizing a recognition unit to identify key points in the 2D images and matching them with a scanning point group in 3D data to generate accurate 3D posture information.

Benefits of technology

This approach enhances the accuracy of 3D posture estimation by reducing errors in distance estimation, eliminating ambiguity in keypoint detection, and improving robustness in crowded environments, while maintaining precision across varying distances and reducing computational load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187858000001_ABST
    Figure 2025187858000001_ABST
Patent Text Reader

Abstract

To provide a video processing device, video processing method, video processing program, and computer program product, which enable acquisition of accurate 3D posture information.SOLUTION: A video processing device (10) is provided, comprising: an input unit (11a) for inputting two-dimensional captured image data and three-dimensional scanning data; a recognition unit (11b) for recognizing multiple key points representing characteristic points of a posture of a person through recognition processing performed on the two-dimensional captured image data; and an output unit (11c) for outputting recognition data obtained by three-dimensionally recognizing the posture of the person through matching processing with the key points performed on a scanning point cloud of the three-dimensional scanning data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a video processing device, a video processing method, a video processing program, and a computer program product. [Background technology]

[0002] The video processing device disclosed in Patent Document 1 can accurately estimate 3D posture information by suppressing left-right inversion of skeletal coordinates in 2D posture information. The video processing device generates 2D posture information of a subject from video data of the subject, corrects the 2D posture information, and generates 3D posture information of the subject from the corrected 2D posture information. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2023 / 223508 Summary of the Invention [Problem to be solved by the invention]

[0004] However, because the device in Patent Document 1 uses multiple 2D camera images, even if it can recognize feature points of a person's posture, it is difficult to accurately recognize the distance to those feature points using 2D camera images. Therefore, there is room for improvement in the conventional technology in terms of obtaining accurate 3D posture information.

[0005] The present disclosure aims to provide an image processing device, an image processing method, an image processing program, and a computer program product that are capable of obtaining accurate 3D posture information. [Means for solving the problem]

[0006] A video processing device (10) according to a first aspect of the present disclosure includes an input unit (11a) that inputs two-dimensional captured image data and three-dimensional scanned data, a recognition unit (11b) that recognizes a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data, and an output unit (11c) that outputs recognition data that three-dimensionally recognizes the posture of the person by performing a matching process on the key points with a scanning point group of the three-dimensional scanned data.

[0007] A video processing method according to a second aspect of the present disclosure executes a process in which at least one processor (11A) inputs two-dimensional captured image data and three-dimensional scanned data, recognizes a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data, and outputs recognition data that three-dimensionally recognizes the posture of the person by performing a matching process on the key points with a scanning point group of the three-dimensional scanned data.

[0008] A video processing program (13A) according to a third aspect of the present disclosure causes at least one processor (11A) to execute processing including inputting two-dimensional captured image data and three-dimensional scanned data, recognizing a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data, and outputting recognition data that three-dimensionally recognizes the posture of the person by performing a matching process on the key points with a scanning point group of the three-dimensional scanned data.

[0009] A computer program product (13) according to a fourth aspect of the present disclosure includes a program (13A) for execution, which includes inputting two-dimensional captured image data and three-dimensional scanned data into at least one processor (11A), recognizing a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data, and outputting recognition data that three-dimensionally recognizes the posture of the person by performing a matching process on the key points with a scanning point group of the three-dimensional scanned data. [Effects of the Invention]

[0010] According to the present disclosure, an image processing device, an image processing method, an image processing program, and a computer program product are provided that are capable of obtaining accurate 3D posture information. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a video processing system according to the present disclosure. [Figure 2] FIG. 2 is a diagram showing an example of the hardware configuration of the video processing device 10. As shown in FIG. [Figure 3] FIG. 3 is a block diagram showing an example of the functional configuration of the CPU 11A of the video processing device 10. As shown in FIG. [Figure 4] FIG. 4 is a flowchart for explaining the operation of the video processing device 10. [Figure 5] FIG. 4 is a diagram showing an outline of the processing contents of the video processing device 10. As shown in FIG. [Figure 6A] FIG. 6A is a diagram for explaining the operation of the video processing device 10. In FIG. [Figure 6B] FIG. 6B is a diagram for explaining the operation of the video processing device 10. [Figure 6C] FIG. 6C is a diagram for explaining the operation of the video processing device 10. [Figure 6D] FIG. 6D is a diagram for explaining the operation of the video processing device 10. [Figure 7] FIG. 7 is a diagram showing an outline of the processing contents of the video processing device 10. As shown in FIG. [Figure 8] FIG. 8 is a diagram for explaining 2D association. [Figure 9] FIG. 9 is a diagram for explaining narrowing down of the candidate point group. [Figure 10] FIG. 10 is a diagram for explaining the selection of 3D joint points. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, the present embodiment will be described with reference to the accompanying drawings. To facilitate understanding of the description, the same components in the drawings will be denoted by the same reference numerals as much as possible, and duplicated descriptions will be omitted.

[0013] FIG. 1 is a diagram illustrating a schematic configuration of a video processing system according to the present disclosure. The video processing system 100 may include multiple video processing devices 10 and a cloud server 20. The multiple video processing devices 10 may be installed in vehicles, infrastructure, and the like installed at, for example, a new product event venue. The multiple video processing devices 10 may generate and output recognition data that recognizes the posture of a person in three dimensions, using, for example, two-dimensional (2D) image data of a person present at the event venue captured by a camera and three-dimensional (3D) scanning data detected by, for example, LiDAR (Light Detection and Ranging). The cloud server 20 may have a function for integrating the multiple recognition data and displaying the posture of the person using the integrated data. Note that the vehicle is not limited to being installed at the event venue and may be located outdoors. For example, if the video processing device 10 is installed in a vehicle traveling outdoors, the invention described below can be applied to recognize people and ensure the safety of people and vehicles by combining the video processing device 10 of the vehicle with infrastructure sensors installed outdoors to cover the vehicle's blind spots.

[0014] 2 is a diagram showing an example of the hardware configuration of the video processing device 10. The video processing device 10 includes a processing unit 11, a communication unit 12, and a storage unit 13.

[0015] The processing unit 11 is configured as a device including a general computer. The processing unit 11 has a CPU 11A, a ROM 11B, a RAM 11C, and an input / output interface (I / O) 11D. The CPU 11A, the ROM 11B, the RAM 11C, and the I / O 11D are connected to each other via a bus 11E. The bus 11E includes a control bus, an address bus, a data bus, etc.

[0016] The communication unit 12, the storage unit 13, the camera 1, and the LiDAR 2 are connected to the I / O 11D.

[0017] The communication unit 12 is an interface for communicating with external devices such as the camera 1 and the LiDAR 2. The camera 1 may be considered as an imaging device that generates and outputs two-dimensional captured image data of a person. The LiDAR 2 may be considered as a sensor that detects surrounding conditions such as vehicles and infrastructure, for example, people present at an event venue, as a point cloud. The two-dimensional captured image data is input to the video processing device 10, and the point cloud is input to the video processing device 10 as three-dimensional scanned data.

[0018] The storage unit 13 is configured as a non-volatile external storage device such as a hard disk. The storage unit 13 stores a video processing program 13A. The storage unit 13 may be interpreted as a computer program product of the present disclosure.

[0019] The CPU 11A is an example of a computer. The term "computer" here refers to a processor in a broad sense, and includes a general-purpose processor (e.g., the CPU 11A) and a dedicated processor (e.g., a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, etc.).

[0020] The video processing program 13A may be stored in a non-volatile non-transitory recording medium or distributed via a network and appropriately installed in the video processing device 10, thereby being stored in the storage unit 13. The video processing program 13A may also be appropriately updated by so-called OTA (Over The Air).

[0021] Examples of non-volatile non-transitory recording media include CD-ROMs (Compact Disc Read Only Memory), magneto-optical disks, HDDs (Hard Disk Drives), DVD-ROMs (Digital Versatile Disc Read Only Memory), flash memory, and memory cards.

[0022] Next, an example of the functional configuration of the CPU 11A of the video processing device 10 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the functional configuration of the CPU 11A of the video processing device 10. The CPU 11A functions as each functional unit shown in Fig. 3 by reading and executing a video processing program 13A stored in the storage unit 13 (see Fig. 2). The CPU 11A includes an input unit 11a, a recognition unit 11b, and an output unit 11c.

[0023] (input unit 11a) The input unit 11a may receive as input two-dimensional captured image data transmitted from the camera 1 and three-dimensional scanned data transmitted from the LiDAR 2.

[0024] (Recognition unit 11b) The recognition unit 11b may recognize a plurality of key points, which are characteristic points of the posture of a person, by performing recognition processing on the two-dimensional captured image data.

[0025] The recognition unit 11b may set a person whose difference between the out-of-area rate, which is the rate at which the detection area (bounding box) surrounding the person goes outside the extraction area set in the two-dimensional captured image data, and the specific person detection score is within the detection allowance range, as a target for recognizing multiple key points. The extraction area may be interpreted as an area for extracting a specific person as a posture recognition target, in order to reduce the processing load of the CPU 11A by excluding people who are not originally desired to be included in the posture recognition target from the recognition target.

[0026] Specifically, the recognition unit 11b does not set a person whose difference between the out-of-area ratio and the specific person detection score is outside the detection allowable range as a target for recognizing multiple key points. The recognition unit 11b sets a person whose difference between the out-of-area ratio and the specific person detection score is within the detection allowable range as a target for recognizing multiple key points.

[0027] The recognition unit 11b may dynamically change the detection permission range in multiple stages based on specific conditions.

[0028] The recognition unit 11b may narrow the detection allowable range as the specific person detection score becomes smaller, thereby enabling detection of an object with low detection confidence in an area closer to the set area.

[0029] (Output section 11c) The output unit 11c may output recognition data obtained by three-dimensionally recognizing the posture of a person by matching the scanning point group of the three-dimensional scanning data with key points of the two-dimensional captured image data.

[0030] The output unit 11c may map key points of the two-dimensional captured image data onto the three-dimensional space of the three-dimensional scanned data, and perform matching between a group of points in the three-dimensional space and the key points.

[0031] The output unit 11c may perform mapping processing of key points of the two-dimensional captured image data using specific parameters between the two-dimensional captured image data and the three-dimensional scanned data.

[0032] The output unit 11c may select points to be used in the matching process according to dynamic programming from candidate points included in the search range of key points mapped onto the three-dimensional space of the three-dimensional scanning data in the current frame.

[0033] The output unit 11c may limit a group of candidate points included in a specific search range centered on a matching point recognized in the three-dimensional scanning data of a previous frame from a group of specific candidate points included in a search range of key points in the two-dimensional captured image data, and select a matching point from the limited group of candidate points.

[0034] The output unit 11c may map key points whose positional error between the two-dimensional captured image data of the current frame and the two-dimensional captured image data of the previous frame is within a specific recognition allowance range into the three-dimensional space of the three-dimensional scanning data.

[0035] The output unit 11c may remove, by filtering, matching points whose variation frequency between the three-dimensional scanning data from the past frame to the current frame is within a specific noise range.

[0036] The output unit 11c may remove, by filtering, matching points whose variation frequency between the three-dimensional scanning data from the past frame to the current frame is within a specific noise range.

[0037] Next, the operation of the video processing device 10 will be described with reference to Figures 4 to 10. Figure 4 is a flowchart for explaining the operation of the video processing device 10.

[0038] In step S1, the CPU 11A may detect multiple people and filter out (filter) a person to be detected from the multiple people. For example, as shown in FIGS. 6A and 6B, there may be two people within a specific extraction area, with the entire bounding box of one person being within the extraction area, and the majority of the other person being within the extraction area, although a portion of the bounding box is outside the extraction area (FIG. 6A). For example, if a person with only their toes outside the extraction area is excluded from the posture recognition target, it will feel very strange. Also, there may be cases where the majority of the other person is outside the extraction area, although a portion of the other person is within the extraction area (FIG. 6B). In this way, it will also feel strange if a specific person is not excluded from the posture recognition target until they are completely outside the extraction area.

[0039] Therefore, CPU 11A may prioritize posture recognition of people within the extraction area. Specifically, CPU 11A may prioritize outputting as recognition results targets with a high person detection score and a low proportion of people outside the extraction area. For example, as shown in FIG. 6C, if the proportion of people outside the extraction area (out-of-area proportion) is 0.6 and the person detection confidence score (specific person detection score) is 0.8, and the difference between these is smaller than a specific threshold (detection allowance range), CPU 11A may set the target outside of posture recognition, i.e., not set the target for multiple keypoint recognition. On the other hand, if the difference between these is larger than a specific threshold (detection allowance range), CPU 11A may set the target as a posture recognition target.

[0040] The extraction area and threshold may not be static, but may be dynamically changed in multiple stages depending on the scene. For example, if there are few people in the entire image and the CPU 11A has ample computing resources, the CPU 11A may reduce the threshold and widen the extraction area.

[0041] In narrowing down (filtering) posture recognition targets, as shown in FIG. 6D, a person who is closest to the center point of the extraction area may be preferentially set as a posture recognition target. Specifically, for each person, the CPU 11A may calculate the distance from the center point of the extraction area to the center point of the bounding box, compare the distances for each person, and set the upper limit number of people as posture recognition targets in order of shortest distance. The CPU 11A may continue to set a person who has been set as a recognition target because of their short distance from the point of interest (the center point of the extraction area) as a priority recognition target until the distance from that person increases by a certain threshold or more. This makes it possible to prevent a change in the recognition output target, which could lead to an unnatural feeling, even when the distance changes significantly as a person walks around.

[0042] In step S2, the CPU 11A executes an algorithm for estimating posture keypoints within the filtered detection area, thereby estimating, for example, 18 keypoints. That is, the CPU 11A recognizes a plurality of keypoints, which are characteristic points of the posture of the person, through recognition processing of the two-dimensional captured image data.

[0043] In step S3, the CPU 11A executes 2D association. Because 2D posture recognition results are output without a person ID, 2D association links the new recognition result with the person ID of a previous recognition result. For example, when multiple people exist, specifically, when another person is detected near a specific person (e.g., previous frame: k-1) stored as shown in FIG. 8, it is necessary to set specific criteria for determining whether the person in the current frame (current frame: k) is associated with the specific person. For example, the distance between each point (each joint position) of the 2D captured image data (k) and the stored 2D captured image data (k-1) may be calculated, and the sum of the squared errors of these distances may be used as the cost of association. The 2D captured image data (k) with the lowest cost may be adopted as the most reliable data.

[0044] In step S4, the CPU 11A may limit a group of candidate points included in a specific search range centered on a matching point recognized in the 3D scan data of the previous frame from among a group of specific candidate points included in the search range of the key point, and extract and select matching points (3D candidate points) from the limited group of candidate points. Specifically, as shown in Fig. 9, (1) for the group of points included in the specific search range of the 2D associated 2D posture described above, (2) by further narrowing down the group of candidate points included in the candidate search range from the group of 3D points constituting the 3D posture of the previous frame, the calculation of human posture estimation can be accelerated.

[0045] In step S5, the CPU 11A uses dynamic programming to select a point from a group of candidate points for each joint (keypoint) that minimizes the "sum of squared errors between each link length and the reference value," as shown in FIG. 10. For example, the CPU 11A selects a matching point (3D joint point) to be used in the matching process from candidate points included in the keypoint search range according to dynamic programming. Measurement values ​​from a human body dimension database may be used as the reference value for the link length between joints. Note that due to limitations of dynamic programming, the existence of a closed loop in the graph structure makes the problem unsolvable (leading to long calculation times), so shoulder width and waist width are not taken into account. Joint angles do not need to be considered. While link lengths can be calculated from information on two joints, determining the angles requires information on three or more joint points, making the problem unsuccessful. Note also that optimization does not work well for people whose link lengths deviate significantly from the average, such as children.

[0046] After the processing in step S5, a filtering process (noise removal) described later may be performed. This will result in smooth 3D posture data. For example, if there is fluctuation in the solution obtained by the dynamic programming method described above, high-frequency noise will be included within a certain distance range.

[0047] To address this issue, as shown in Figure 7, the data from which 3D joint points have been selected can be passed through an LPF (low-pass filter) to eliminate the high-frequency noise that we want to remove. For example, by assuming the speed of human movement (waving, walking, etc.) and setting the cutoff frequency to around 2 Hz, much of the movement will not be removed even after filtering. Furthermore, even when people cross paths, they can be correctly identified, and calculations (speed and orientation estimation) can be performed using the 3D pose recognition results from previous frames.

[0048] (Use of calibration parameters) The CPU 11A may map key points using calibration parameters (specific parameters) between the 2D captured image data and the 3D scanned data. The calibration parameters may be interpreted as parameters for converting the coordinate systems of different sensors, such as camera 1 and LiDAR 2, or parameters for linking the coordinate systems of camera 1 and LiDAR 2. The CPU 11A may acquire internal camera parameters and external camera-LiDAR parameters as calibration parameters and use these parameters to project the projected pose estimation results (e.g., 18 points + center of gravity) onto the 3D point cloud. When converting the pose obtained in 2D into 3D, the calibration parameters and a specific formula can be substituted to determine where the 3D point cloud is located on the 2D plane. For example, a typical parameter calculation using a checkerboard (using open-source software) may be used. Additionally, multiple images can be taken by changing the distance between the camera and the board, the position of the board within the camera, and the board's posture, and the parameters can be automatically acquired within the software (Reference: M. Velas, "Calibration of RGB Camera With Velodyne LiDAR").

[0049] As described above, the image processing device 10 of the present disclosure can recognize multiple key points, which are characteristic points of a person's posture, by performing recognition processing on two-dimensional captured image data, and output recognition data that recognizes the person's posture in three dimensions by performing matching processing on the key points with the scanning point group of three-dimensional scanning data.

[0050] This configuration can provide the following effects.

[0051] (1) In conventional technology, it is necessary to match key points (skeleton points) from multiple different poses, and matching errors significantly affect 3D pose estimation errors. Therefore, when using images, the accuracy of distance estimation decreases significantly at long distances (for example, inversely proportional to the square of the distance). In contrast, according to the present disclosure, by using LiDAR point clouds, estimation accuracy does not decrease significantly depending on the distance.

[0052] (2) In conventional technologies, keypoints of the same type point at slightly different positions across multiple images. This means that there is ambiguity in keypoint detection. In contrast, this disclosure matches 2D image keypoints with LiDAR point clouds, meaning that keypoints are not matched to each other, eliminating the ambiguity in keypoint detection.

[0053] (3) In conventional technology, reducing depth errors requires a long baseline length (distance between cameras). Therefore, the greater the distance, the greater the change in appearance across multiple images, making matching more difficult. In contrast, according to the present disclosure, the distance between sensors is shortened and the target person is sensed from as close to the same direction as possible, so the appearance of the target does not change significantly.

[0054] (4) In conventional technology, matching becomes more difficult in crowded situations where multiple people are present. In contrast, according to the present disclosure, even when multiple people are present, the LiDAR and the camera view the same image, so matching between the sensors does not become significantly more difficult.

[0055] (5) According to the present disclosure, by setting a recognition priority for people, calculations are limited to necessary posture recognition processing. This prevents the posture recognition processing speed from becoming proportional to the number of target people when narrowing down posture recognition targets in 2D posture recognition, making it impossible to calculate the postures of all people within the time limit. It also prevents the sensor from sensing unnecessary distant areas and performing posture recognition on unnecessary targets.

[0056] (6) According to the present disclosure, by deriving a matching candidate point cloud and solving an optimization problem from it, the estimation accuracy in 3D pose generation can be improved, and robust 3D pose recognition results with less unnaturalness can be achieved. In other words, it is possible to prevent unnatural 3D pose recognition results from being output due to misalignment between the posture keypoints on the image and the corresponding point cloud positions.

[0057] (7) According to the present disclosure, by passing the 3D posture recognition results through a low-pass filter to remove high-frequency noise components, it is possible to suppress noise in the 3D posture recognition results due to fluctuations in posture keypoint detection between frames and sensing noise in the point cloud.

[0058] Although the present embodiment has been described above, the present disclosure is not limited to the above-described embodiments, and various modifications and applications are possible within the scope of the gist of the present disclosure.

[0059] Furthermore, the configuration of the video processing device 10 described in the above embodiment (see FIG. 2) is an example, and it goes without saying that unnecessary parts may be deleted or new parts may be added within the scope of the present disclosure.

[0060] Furthermore, the processing flow of the video processing program 13A described in the above embodiment is also an example, and it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged within the scope of the present disclosure.

[0061] The controller and methods described herein may be implemented by a special-purpose computer having a processor programmed to perform one or more functions embodied in a computer program. Alternatively, the apparatus and methods described herein may be implemented by a special-purpose computer having a processor configured with dedicated hardware logic circuitry. Alternatively, the apparatus and methods described herein may be implemented by one or more special-purpose computers configured by a combination of a processor executing a computer program and one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by a computer on a computer-readable non-transitory storage medium.

[0062] The present invention can also be applied to a program and a program product (storage unit 13).

[0063] The following notes are provided regarding the technology of the present disclosure.

[0064] (Appendix 1) an input unit (11a) for inputting two-dimensional photographed image data and three-dimensional scanned data; a recognition unit (11b) that recognizes a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data; an output unit (11c) that outputs recognition data obtained by three-dimensionally recognizing the posture of the person by matching the key points with the scanning point cloud of the three-dimensional scanning data; A video processing device (10) comprising: (Appendix 2) The image processing device described in Appendix 1, wherein the recognition unit sets a person as a target for recognizing multiple key points if the difference between the out-of-area rate, which is the rate at which the detection area surrounding the person leaves the extraction area set in the two-dimensional captured image data, and the specific person detection score is within the detection allowance range. (Appendix 3) The video processing device described in Appendix 2, wherein the recognition unit does not set a person whose difference between the out-of-area ratio and the specific person detection score is outside the detection allowance range as a target for recognizing multiple key points. (Appendix 4) 3. The image processing device according to claim 2, wherein the recognition unit dynamically changes the detection permission range in multiple stages based on specific conditions. (Appendix 5) The image processing device according to claim 4, wherein the recognition unit narrows the detection allowable range as the specific person detection score becomes smaller. (Appendix 6) The image processing device according to claim 1, wherein the output unit performs a mapping process of the key points onto a three-dimensional space of the three-dimensional scanning data, and performs the matching process between a point cloud in the three-dimensional space and the key points. (Appendix 7) The image processing device according to claim 1, wherein the output unit performs mapping processing on the key points using specific parameters between the two-dimensional captured image data and the three-dimensional scanned data. (Appendix 8) The image processing device described in Appendix 1, wherein the output unit selects points to be used in the matching process from candidate points included in a search range of the key points mapped onto the three-dimensional space of the three-dimensional scanning data in the current frame according to a dynamic plan. (Appendix 9) The image processing device described in Appendix 8, wherein the output unit limits a group of candidate points included in a specific search range centered on a matching point recognized in the 3D scanning data of a past frame from a specific group of candidate points included in a search range of the key point, and selects a matching point from the limited group of candidate points. (Appendix 10) The image processing device according to claim 1, wherein the output unit maps the key points whose positional error between the two-dimensional captured image data of the current frame and the two-dimensional captured image data of the past frame is within a specific recognition allowance range into the three-dimensional space of the three-dimensional scanning data. (Appendix 11) 2. The image processing device according to claim 1, wherein the output unit removes, by filtering, matching points whose variation frequency between the three-dimensional scanning data from the past frame to the current frame is within a specific noise range. (Appendix 12) At least one processor (11A) Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A video processing method that performs processing including: (Appendix 13) At least one processor (11A) has Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A video processing program (13A) for executing a process including the above. (Appendix 14) At least one processor (11A) has Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A computer program product (13) including a program (13A) for executing a process including the steps of: [Explanation of symbols]

[0065] 1 camera 2. LiDAR 10. Video Processing Device 11 Processing section 11E Bus 11a Input section 11b Recognition part 11c Output section 12 Communications Department 13 Storage section 13A Image Processing Program 20 Cloud Server 100 Video Processing System

Claims

1. an input unit (11a) for inputting two-dimensional photographed image data and three-dimensional scanned data; a recognition unit (11b) that recognizes a plurality of key points that are characteristic points of a person's posture by performing a recognition process on the two-dimensional captured image data; an output unit (11c) that outputs recognition data obtained by three-dimensionally recognizing the posture of the person by matching the key points with the scanning point group of the three-dimensional scanning data; A video processing device (10) comprising:

2. The image processing device of claim 1, wherein the recognition unit sets a person as a target for recognizing multiple key points if the difference between the out-of-area rate, which is the rate at which the detection area surrounding the person leaves the extraction area set in the two-dimensional captured image data, and the specific person detection score is within a detection allowance range.

3. The image processing device according to claim 2 , wherein the recognition unit does not set a person for which a difference between the out-of-area ratio and the specific person detection score is outside the detection allowable range as a target for recognizing a plurality of the key points.

4. The image processing device according to claim 2 , wherein the recognition unit dynamically changes the detection permission range in multiple stages based on specific conditions.

5. The image processing device according to claim 4 , wherein the recognition unit narrows the detection allowable range as the specific person detection score becomes smaller.

6. The image processing device according to claim 1 , wherein the output unit performs a mapping process of the key points onto a three-dimensional space of the three-dimensional scanning data, and performs the matching process between the point group in the three-dimensional space and the key points.

7. The image processing device according to claim 1 , wherein the output unit performs mapping processing of the key points using specific parameters between the two-dimensional captured image data and the three-dimensional scanned data.

8. 2. The image processing device according to claim 1, wherein the output unit selects points to be used in the matching process from candidate points included in a search range of the key points mapped onto the three-dimensional space of the three-dimensional scanning data in the current frame according to a dynamic program.

9. The image processing device according to claim 8, wherein the output unit limits a group of candidate points included in a specific search range centered on a matching point recognized in the 3D scanning data of a past frame from a group of specific candidate points included in a search range of the key point, and selects a matching point from the limited group of candidate points.

10. 2. The image processing device according to claim 1, wherein the output unit maps the key points whose positional error between the two-dimensional captured image data of the current frame and the two-dimensional captured image data of the past frame is within a specific recognition allowance range into the three-dimensional space of the three-dimensional scanning data.

11. The image processing device according to claim 1 , wherein the output unit removes, by filtering, matching points whose fluctuation frequency between the three-dimensional scanning data from the past frame to the current frame is within a specific noise range.

12. At least one processor (11A) Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A video processing method that performs processing including:

13. At least one processor (11A) Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A video processing program (13A) for executing a process including the above.

14. At least one processor (11A) Input two-dimensional photographed image data and three-dimensional scanned data, Recognizing a plurality of key points, which are characteristic points of the posture of the person, by performing a recognition process on the two-dimensional captured image data; outputting recognition data that recognizes the posture of the person in three dimensions by performing a matching process with the key points on the scanning point cloud of the three-dimensional scanning data; A computer program product (13) including a program (13A) for executing a process including the steps of:

Citation Information

Patent Citations

  • Video processing device, video processing method, and program

    WO2023223508A1