Operation action recognition method, device and computer storage medium

CN115774497BActive Publication Date: 2026-09-25SHENZHEN PENGXING INTELLIGENT RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211515588.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-09-25
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

[0003]但是,由于深度传感器的精度有限,因此仅通过深度数据的识别过程鲁棒性比较低,识别得到的触控端的操作动作准确度较低,并影响后续执行相应的操作指令,从而影响用户体验

Benefits of technology

[0015]本申请实施例的有益效果至少包括:通过图像数据协同深度数据去获得触控端于投影区域的目标空间坐标,使该目标空间坐标更加准确,从而提高识别触控端操作动作的鲁棒性。并且,通过利用深度数据构件投影平面的平面方程,可以平滑投影平面的凹凸因素,在利用平面方程结合目标空间坐标去识别触控端的操作动作时,可以提高该识别过程的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115774497B_ABST
    Figure CN115774497B_ABST
Patent Text Reader

Abstract

The application discloses an operation action recognition method, device and computer storage medium. The operation action recognition method comprises the following steps: acquiring image data and depth data corresponding to a projection area, wherein the projection area comprises a projection plane obtained through projection; determining target space coordinates of a touch end in the projection area according to the image data and the depth data corresponding to the projection area; calculating a plane equation corresponding to the projection plane based on the depth data; and recognizing an operation action of the touch end on the projection plane according to the target space coordinates and the plane equation. The application can improve the robustness of the operation action recognition process and the accuracy of the recognized operation action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, specifically to a method, apparatus, and computer storage medium for recognizing operational actions. Background Technology

[0002] In related technologies, projection touch control is generally implemented through depth sensing. For example, a projection touch control system may include a control module, a projection unit, and a depth sensing unit. The projection unit projects video or control images onto the projection area. The depth sensing unit detects touch data from the touch terminal on the projection area and transmits the touch data to the control module. The control module receives the touch data and identifies the touch terminal's operation based on the touch data.

[0003] However, due to the limited accuracy of depth sensors, the recognition process based solely on depth data is relatively unreliable, resulting in low accuracy of the recognized touch actions and affecting the subsequent execution of corresponding operation commands, thus impacting the user experience. Summary of the Invention

[0004] In view of this, this application provides an operation action recognition method, apparatus and computer storage medium to improve the robustness of the operation action recognition process and the accuracy of the recognized operation actions.

[0005] In a first aspect, this application provides an operation action recognition method, the method comprising: acquiring image data and depth data corresponding to a projection area, the projection area including a projection plane obtained by projection; determining the target spatial coordinates of a touch terminal in the projection area based on the image data and the depth data corresponding to the projection area; calculating the plane equation corresponding to the projection plane based on the depth data; and recognizing the operation action of the touch terminal on the projection plane based on the target spatial coordinates and the plane equation.

[0006] In one embodiment, obtaining the spatial coordinates of the touch device in the projection area based on the image data, the depth data, and a preset coordinate detection algorithm includes: determining the first spatial coordinates of the touch device in the projection area based on the image data; determining the second spatial coordinates of the touch device in the projection area based on the depth data; and determining the target spatial coordinates based on the first spatial coordinates and the second spatial coordinates.

[0007] In one embodiment, determining the first spatial coordinates of the touch device in the projection area based on the image data includes: performing keypoint detection on the touch device based on the image data to obtain multiple two-dimensional initial coordinates corresponding to the touch device, and the confidence level corresponding to each of the multiple two-dimensional initial coordinates; determining a target matrix based on the multiple two-dimensional initial coordinates and the confidence level corresponding to each of the multiple two-dimensional initial coordinates; inputting the target matrix into a completion model, performing keypoint completion on the two-dimensional initial coordinates based on the target matrix to obtain two-dimensional completed coordinates; and converting the two-dimensional completed coordinates into the first spatial coordinates of the touch device in the projection area.

[0008] In one embodiment, the step of performing key point detection on the touch terminal based on the image data to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates includes: converting the image data into brightness-normalized color space data; inputting the color space data into a two-dimensional key point detection model to obtain the multiple two-dimensional initial coordinates and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates.

[0009] In one embodiment, determining the second spatial coordinates of the touch terminal in the projection area based on the depth data includes: obtaining the pixel coordinates of the touch terminal; obtaining the second spatial coordinates based on the pixel coordinates, a preset pose relationship, and the depth data, wherein the preset pose relationship includes the relative pose relationship between the camera and the depth sensor.

[0010] In one embodiment, the step of calculating the plane equation corresponding to the projection plane based on the depth data includes: obtaining a preset number of depth data points corresponding to the projection plane; establishing a solution equation corresponding to the projection plane; obtaining the equation coefficients corresponding to the solution equation using the preset number of depth data points corresponding to the projection plane; and obtaining the plane equation corresponding to the projection plane based on the equation coefficients and the solution equation.

[0011] In one embodiment, the method further includes: generating an operation instruction based on the operation action and the projection plane; and displaying the operation result obtained in response to the operation instruction through the projection plane.

[0012] Secondly, this application provides an operation action recognition device, the device comprising: a data acquisition module for acquiring image data and depth data corresponding to a projection area, the projection area including a projection plane obtained by projection; a target coordinate acquisition module for determining the target spatial coordinates of a touch terminal in the projection area based on the image data and the depth data corresponding to the projection area; a plane equation acquisition module for calculating the plane equation corresponding to the projection plane based on the depth data; and an operation action acquisition module for recognizing the operation action of the touch terminal on the projection plane based on the target spatial coordinates and the plane equation.

[0013] Thirdly, this application provides a robot, which includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the operation action recognition method.

[0014] Fourthly, this application provides a computer storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the operation action recognition method.

[0015] The beneficial effects of the embodiments of this application include at least the following: obtaining the target spatial coordinates of the touch device in the projection area by combining image data with depth data makes the target spatial coordinates more accurate, thereby improving the robustness of recognizing touch device operation actions. Furthermore, by using depth data to construct the plane equation of the projection plane, the concavity and convexity of the projection plane can be smoothed, and the accuracy of the recognition process can be improved when using the plane equation in combination with the target spatial coordinates to recognize touch device operation actions. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the hardware structure of a robot provided in this application.

[0017] Figure 2 This is a schematic diagram of the mechanical structure of a robot provided in this application.

[0018] Figure 3 This is a schematic diagram of the hardware structure of another robot provided in an embodiment of this application.

[0019] Figure 4 This is a flowchart illustrating an operation action recognition method provided in an embodiment of this application.

[0020] Figure 5 This is a flowchart illustrating a method for determining target spatial coordinates provided in an embodiment of this application.

[0021] Figure 6This is a flowchart illustrating a first spatial coordinate acquisition method provided in an embodiment of this application.

[0022] Figure 7 This is a flowchart illustrating a planar equation calculation method provided in an embodiment of this application.

[0023] Figure 8 This is a flowchart illustrating another operation action recognition method provided in the embodiments of this application.

[0024] Figure 9 This is a flowchart illustrating a preferred operation action recognition method provided in this application.

[0025] Figure 10 This is a schematic diagram of the structure of an operation action recognition device provided in an embodiment of this application.

[0026] Explanation of main component symbols

[0027] 100-Robot; 101-Mechanical Unit; 1011-Drive Board; 1012-Electric Motor; 1013-Mechanical Structure; 1014-Main Body; 1015-Legs; 1016-Foot; 1017-Head Structure; 1018-Tail Structure; 1019-Carrying Structure; 1020-Saddle Structure; 1021-Camera Structure; 102-Communication Unit; 103-Sensing Unit; 104-Interface Unit; 105-Storage Unit; 106-Display Unit; 1061-Display Panel; 107-Input Unit; 1071-Touch Panel; 1072-Other Input Devices; 1073-Touch Detection Device; 1074-Touch Controller; 110-Control Module; 111-Power Supply. Detailed Implementation

[0028] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.

[0029] It should also be noted that the methods disclosed in the embodiments of this application or the methods shown in the flowcharts include one or more steps for implementing the method. Without departing from the scope of the claims, the execution order of multiple steps can be interchanged, and some steps can also be deleted.

[0030] Please see Figure 1 , Figure 1 This is a schematic diagram of the hardware structure of a robot 100 according to one embodiment of this application. Figure 1 In the illustrated embodiment, robot 100 includes a mechanical unit 101, a communication unit 102, a sensing unit 103, an interface unit 104, a storage unit 105, a control module 110, and a power supply 111. The various components of robot 100 can be connected in any way, including wired or wireless connections. Those skilled in the art will understand that... Figure 1 The specific structure of the robot 100 shown does not constitute a limitation on the robot 100. The robot 100 may include more or fewer parts than shown. Some parts are not essential components of the robot 100 and may be omitted or combined as needed without changing the nature of the invention.

[0031] The following is combined Figure 1 A detailed introduction to each component of Robot 100:

[0032] Mechanical unit 101 is the hardware of robot 100. For example... Figure 1 As shown, the mechanical unit 101 may include a drive board 1011, a motor 1012, and a mechanical structure 1013, such as... Figure 2 As shown, the mechanical structure 1013 may include a main body 1014, extendable legs 1015, and feet 1016. In other embodiments, the mechanical structure 1013 may also include an extendable robotic arm (not shown), a rotatable head structure 1017, a rocking tail structure 1018, a cargo-carrying structure 1019, a saddle structure 1020, a camera structure 1021, etc. It should be noted that the various component modules of the mechanical unit 101 can be one or multiple, depending on the specific situation. For example, there may be four legs 1015, and each leg 1015 may be equipped with three motors 1012, resulting in a total of twelve motors 1012.

[0033] The communication unit 102 can be used for receiving and sending signals, and can also communicate with networks and other devices. For example, it can receive instructions from a remote control or other robot 100 to move in a specific direction at a specific speed according to a specific gait, and then transmit these instructions to the control module 110 for processing. The communication unit 102 includes modules such as WiFi, 4G, 5G, Bluetooth, and infrared modules.

[0034] The sensing unit 103 is used to acquire information data about the environment surrounding the robot 100 and to monitor parameter data of various components inside the robot 100, and then sends this data to the control module 110. The sensing unit 103 includes various sensors, such as sensors for acquiring information about the surrounding environment: lidar (for remote object detection, distance determination, and / or velocity determination), radar (for short-range object detection, distance determination, and / or velocity determination), cameras, infrared cameras, and Global Navigation Satellite System (GNSS). Sensors for monitoring various components inside the robot 100 include: an inertial measurement unit (IMU) (for measuring velocity, acceleration, and angular velocity values), foot sensors (for monitoring the position of the foot's contact point, foot posture, magnitude and direction of the contact force), and temperature sensors (for detecting component temperature). Other sensors that can be configured on the robot 100, such as load sensors, touch sensors, motor angle sensors, and torque sensors, are not detailed here.

[0035] The interface unit 104 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within the robot 100, or it can be used to output to external devices (e.g., data, power, etc.). The interface unit 104 may include a power port, a data port (such as a USB port), a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, etc.

[0036] Storage unit 105 is used to store software programs and various data. Storage unit 105 may mainly include a program storage area and a data storage area. The program storage area may store operating system programs, motion control programs, application programs (such as text editors), etc.; the data storage area may store data generated by the robot 100 during use (such as various sensor data acquired by the sensing unit 103, log file data, etc.). Furthermore, storage unit 105 may include high-speed random access memory, and may also include non-volatile memory, such as disk storage, flash memory, or other volatile solid-state memory.

[0037] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0038] Input unit 107 can be used to receive input numerical or character information. Specifically, input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect user touch operations (such as operations performed by the user using their palm, fingers, or suitable accessories on or near touch panel 1071) and drive corresponding connected devices according to a pre-set program. Touch panel 1071 may include two parts: touch detection device 1073 and touch controller 1074. Touch detection device 1073 detects the user's touch position and the signal generated by the touch operation, and transmits the signal to touch controller 1074; touch controller 1074 receives touch information from touch detection device 1073, converts it into touch point coordinates, sends it to control module 110, and can receive and execute commands from control module 110. In addition to touch panel 1071, input unit 107 may also include other input devices 1072. Specifically, other input devices 1072 may include, but are not limited to, one or more of the following: remote control handles, etc., without any specific limitation here.

[0039] Furthermore, the touch panel 1071 can cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the control module 110 to determine the type of touch event. Subsequently, the control module 110 provides corresponding visual output on the display panel 1061 according to the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components that implement input and output functions respectively. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to implement input and output functions. The specific implementation is not limited here.

[0040] The control module 110 is the control center of the robot 100. It connects all the components of the robot 100 through various interfaces and lines. It controls the robot 100 as a whole by running or executing the software program stored in the storage unit 105 and calling the data stored in the storage unit 105.

[0041] Power supply 111 supplies power to various components. Power supply 111 may include a battery and a power control board. The power control board controls battery charging, discharging, and power consumption management. Figure 1 In the illustrated embodiment, power supply 111 is electrically connected to control module 110. In other embodiments, power supply 111 may also be electrically connected to sensing unit 103 (such as camera, radar, speaker, etc.) and motor 1012 respectively. It should be noted that each component may be connected to a different power supply 111, or be powered by the same power supply 111.

[0042] Based on the above embodiments, specifically, in some embodiments, a terminal device can be used to communicate with the robot 100. When the terminal device communicates with the robot 100, it can send instruction information to the robot 100. The robot 100 can receive the instruction information through the communication unit 102 and, upon receiving the instruction information, can transmit it to the control module 110, so that the control module 110 can process the instruction information to obtain the target speed value. The terminal device includes, but is not limited to, mobile phones, tablets, servers, personal computers, wearable smart devices, and other electrical appliances with image capture capabilities.

[0043] The instruction information can be determined based on preset conditions. In one embodiment, the robot 100 may include a sensing unit 103, which can generate instruction information based on the current environment of the robot 100. The control module 110 can determine whether the current speed value of the robot 100 meets the corresponding preset conditions based on the instruction information. If it does, the robot 100 will maintain its current speed value and current gait; if it does not, the control module 110 will determine a target speed value and a corresponding target gait based on the corresponding preset conditions, thereby controlling the robot 100 to move at the target speed value and the corresponding target gait. Environmental sensors may include temperature sensors, air pressure sensors, vision sensors, and sound sensors. Instruction information may include temperature information, air pressure information, image information, and sound information. The communication method between the environmental sensors and the control module 110 can be wired or wireless. Wireless communication methods include, but are not limited to: wireless networks, mobile communication networks (3G, 4G, 5G, etc.), Bluetooth, and infrared.

[0044] Please see Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of another robot provided in an embodiment of this application. The robot 300 includes a control module 310, a projection unit 320, and a projection touch detection unit 330. The projection unit 320 projects video or control images onto a projection area. The projection touch detection unit 330 detects touch data from a touch terminal on the projection area and transmits the touch data to the control module 310. The control module 310 receives the touch data, identifies the operation actions of the touch terminal based on the touch data, obtains corresponding operation instructions, and executes them.

[0045] The aforementioned touch terminal includes the user's hand, a stylus, or a laser pointer projection point, etc., and is not limited here. The projection touch detection unit 330 includes a depth sensor. The control module 310 includes a processor.

[0046] In existing related technologies, after receiving depth data from the depth sensor, the control module's touch recognition process involves: acquiring the corresponding depth data of the touch device, using the depth data to identify whether the touch device has performed a touch operation on the projection plane, and recognizing the touch operation. That is, the touch recognition process relies solely on depth data to identify whether the touch device has performed a touch operation on the projection plane and to recognize the operation. However, due to the limited accuracy of depth sensors, the recognition process based solely on depth data has low robustness, resulting in low accuracy of the identified touch device operation and affecting the subsequent execution of corresponding operation commands, thus impacting the user experience.

[0047] This application provides a method for recognizing operation actions, which uses image data and depth data to identify whether a touch terminal has performed a touch operation on the projection plane in the projection area, and to identify the operation action, thereby improving the robustness of the recognition process, increasing the accuracy of operation action recognition, and improving the user experience. This operation action recognition method can be applied to a robot equipped with the aforementioned control module, projection unit, and projection touch detection unit. The projection touch detection unit includes a camera and a depth sensor; the camera is used to acquire image data, and the depth sensor is used to acquire depth data.

[0048] Please see Figure 4 , Figure 4 A flowchart illustrating an operation action recognition method provided for embodiments of this application includes the following steps:

[0049] Step S41: Obtain the image data and depth data corresponding to the projection area, which includes the projection plane obtained by projection.

[0050] In this embodiment, the projection area refers to the planar area projected by the projection unit. It is understood that when this application is applied to a robot, the robot can automatically find a suitable planar area as the projection area and then drive the projection unit to project onto the projection plane. Alternatively, the robot can travel to and face the target planar area according to user instructions or operation commands, and then drive the projection unit to project. For example, the robot is equipped with a depth sensor. Data from the depth sensor can detect relatively flat planar areas in the surrounding environment, allowing the robot to automatically travel to that area or provide location prompts for that area, awaiting further user instructions. The planar area includes structures with a flat surface, such as a screen or wall, and is not limited here.

[0051] While the projection unit projects the projection area onto the projection plane, the robot can simultaneously drive the aforementioned projection touch detection unit to acquire image data and depth data within the projection area. This projection touch detection unit can include image sensors such as cameras, as well as depth sensors. The detection range of this projection touch detection unit can be larger than the projection area to cover the entire projection area for data acquisition, ensuring the quality of the acquired image and depth data.

[0052] The projection plane includes images, videos, and user-defined interfaces. Users can control the content displayed on the projection plane and the robot by clicking and swiping; the specific implementation will be explained later. If the robot has a touchscreen, the content displayed on that screen can be synchronized to the projection plane in real time. While projecting onto the projection plane, the robot can also perform human-computer interactions related to the projection plane, such as playing sounds associated with it; this is not limited to these specific actions.

[0053] Step S42: Determine the target spatial coordinates of the touch terminal in the projection area based on the image data and depth data corresponding to the projection area.

[0054] In this embodiment, the aforementioned touch terminal is used to operate an object on the projection plane, or a user's hand and fingers. The control module can receive the image data and depth data in real time and perform real-time processing and calculation on the image data and depth data to determine the target spatial coordinates of the touch terminal and the projection area. For example, when the control module drives the camera and depth sensor to acquire initial image data and initial depth data, it can establish a first coordinate system for the projection area based on the initial image data and a second coordinate system for the projection area based on the initial depth data. After receiving subsequent image data and depth data, the control module can locate the first coordinate of the touch terminal based on the first coordinate system and the image data, and locate the second coordinate of the touch terminal based on the second coordinate system and the depth data. Finally, the target spatial coordinates are determined based on the first coordinate and the second coordinate. The aforementioned first coordinate system includes a camera coordinate system and a pixel coordinate system, etc., and the corresponding first coordinate can be the camera coordinate and pixel coordinate of the touch terminal, etc., which is not limited here.

[0055] Step S43: Calculate the plane equation corresponding to the projection plane based on the depth data.

[0056] In everyday robot use cases, the projection area is typically a flat structure in an indoor environment, such as a wall. These flat structures are only approximations of flat surfaces, and their surfaces still have some unevenness. Alternatively, the robot's everyday use scenarios may involve near-flat structures serving as the projection area. In these scenarios, the unevenness of the projection area causes the projection plane itself to be uneven as well. This affects the process of detecting whether the touch device has touched the projection plane based on the target's spatial coordinates, thus impacting the accuracy of touch recognition.

[0057] In this embodiment of the application, to solve the above-mentioned problems, before using depth data to detect the coordinates of the touch terminal, a plane equation of the projection plane can be constructed based on the depth data. By constructing the plane equation to smooth the uneven projection plane, and by combining the plane equation with the target spatial coordinates of the touch terminal, the accuracy of touch recognition can be improved.

[0058] For example, when the control module first drives the projection unit, camera, and depth sensor, it can project an initial image onto the projection area through the projection unit, then acquire depth data from this initial image through the depth sensor, and finally calculate the plane equation using this depth data. It is understandable that if the robot changes the projection area, it can repeat the above process to obtain the corresponding plane equation, or it can update the plane equation using real-time depth data; this is not a limitation.

[0059] Step S44: Based on the target spatial coordinates and the plane equation, identify the operation actions of the touch terminal on the projection plane.

[0060] In this embodiment, the aforementioned operations include clicking and swiping. After obtaining the plane equation of the projection plane, the control module can use subsequently acquired image data and depth data to obtain the target spatial coordinates of the touch device. Furthermore, after obtaining the latest target spatial coordinates of the touch device, these coordinates can be input into the plane equation to detect whether the target spatial coordinates lie on the projection plane. Moreover, by detecting consecutive target spatial coordinates on the projection plane, the corresponding operation can be identified.

[0061] In this embodiment, as can be seen from the steps of obtaining the target spatial coordinates and the plane equation described above, obtaining the target spatial coordinates of the touch device in the projection area by combining image data with depth data makes the target spatial coordinates more accurate, thereby improving the robustness of recognizing touch device operation actions. Furthermore, by using depth data to construct the plane equation of the projection plane, the unevenness of the projection plane can be smoothed out, and the accuracy of the recognition process can be improved when using the plane equation in conjunction with the target spatial coordinates to recognize touch device operation actions.

[0062] For example, such as Figure 5The diagram shown is a flowchart illustrating a method for determining target spatial coordinates according to an embodiment of this application. This method is one embodiment of step S42 described above and specifically includes the following steps:

[0063] Step S51: Determine the first spatial coordinates of the touch terminal in the projection area based on the image data.

[0064] In this embodiment, a first spatial coordinate system for the projection area can first be established, and the first spatial coordinates of the touch terminal on the first spatial coordinate system can be obtained through image data. The process of establishing the first spatial coordinate system for the projection area includes defining the origin, X-axis, Y-axis, and Z-axis of the first spatial coordinate system on the projection area. For example, the control module can project an initial image onto the projection area by driving the projection unit. This initial image includes a target point. After acquiring image data through the camera, the target point can be used as the origin of the first spatial coordinate system, and then mutually perpendicular X-axis, Y-axis, and Z-axis can be defined and extended outwards from the origin.

[0065] Alternatively, the first spatial coordinate system can also be the camera coordinate system. That is, after acquiring the image data described above, the pixel coordinates of the touchscreen can be determined first using the image data, and then the first spatial coordinates can be obtained based on the pixel coordinates and the camera's internal parameters. These internal parameters are the camera's intrinsic parameters, which are parameters known at the time of manufacture or obtained through Zhang Zhengyou's calibration method. If the camera's extrinsic parameters are known, the first spatial coordinates of the camera coordinate system can be converted to spatial coordinates of the world coordinate system for subsequent calculations; this is not limited here.

[0066] Step S52: Determine the second spatial coordinates of the touch terminal in the projection area based on the depth data.

[0067] In this embodiment, a second spatial coordinate system can also be established for the projection area, and the second spatial coordinates of the touch device on the second spatial coordinate system can be obtained through depth data. The process of establishing the second spatial coordinate system for the projection area also includes defining the origin, X-axis, Y-axis, and Z-axis of the second spatial coordinate system on the projection area. For example, after determining the projection area, the control module drives the depth sensing unit to acquire the initial depth data of the projection area, and obtains the center point of the projection area based on the initial depth data. This center point is used as the origin of the second spatial coordinate system, and then mutually perpendicular X-axis, Y-axis, and Z-axis are defined and extended outwards from the origin. The origin, X-axis, Y-axis, and Z-axis of the first spatial coordinate system and the second spatial coordinate system may coincide, partially coincide, or not coincide; this is not limited here.

[0068] In establishing the first and second spatial coordinate systems, the relative pose between the camera and the depth sensor can also be considered. Based on the relative pose, the origin, X-axis, Y-axis and Z-axis of the first and second spatial coordinate systems can be made to coincide, thereby reducing the amount of calculation required to determine the target spatial coordinates based on the first and second spatial coordinates.

[0069] Step S53: Determine the target spatial coordinates based on the first spatial coordinates and the second spatial coordinates.

[0070] In this embodiment, after obtaining the first and second spatial coordinates of the touch terminal, an algorithm can be used to reconstruct and determine the target spatial coordinates. For example, assuming the origins, X-axis, Y-axis, and Z-axis of the established first and second spatial coordinate systems coincide, the algorithm can include an averaging algorithm. That is, the target spatial coordinates are obtained by averaging the first and second spatial coordinates.

[0071] Since the first spatial coordinates are taken from image data and the second spatial coordinates are taken from depth data, the reconstructed target spatial coordinates are more robust than those obtained solely from depth data, thereby improving the accuracy of touch screen operation recognition.

[0072] For example, such as Figure 6 The diagram shown is a flowchart illustrating a first spatial coordinate acquisition method provided in this application. This method is one embodiment of step S51 described above and specifically includes the following steps:

[0073] Step S61: Perform key point detection on the touch terminal based on the image data to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal, and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates.

[0074] In this embodiment, as a preferred method for obtaining the first spatial coordinates, after acquiring the image data of the projection area, key point detection can first be performed on the touch terminal image in the image data. These key points can be the key feature points of the touch terminal; for example, when the touch terminal is a user's finger, the color and texture of the finger will be represented as corresponding key feature points in the image. By detecting and locating the key feature points, multiple two-dimensional initial coordinates and corresponding confidence levels corresponding to the touch terminal can be obtained. The confidence level represents the degree of confidence that the key feature point corresponding to the two-dimensional initial coordinates is the touch terminal. The two-dimensional initial coordinates can be pixel coordinates based on a pixel coordinate system.

[0075] The keypoint detection process described above can utilize a keypoint filter, which filters out pixels other than keypoints in the image data to obtain the keypoint output and then locate the two-dimensional coordinates of the keypoint. For example, when the touch device is a user's finger, the keypoint filter is constructed based on features such as the finger's color and texture; this is not a limitation here.

[0076] Step S62: Determine the target matrix based on multiple two-dimensional initial coordinates and the confidence levels corresponding to each of the multiple two-dimensional initial coordinates.

[0077] In this embodiment, the data structure of the target matrix is ​​the target data structure of the input data required by the completion model. Specifically, the arrangement and position of the multiple two-dimensional initial coordinates and their corresponding confidence scores in the target matrix are all based on the target data structure of the input data required by the completion model. Therefore, after obtaining the multiple two-dimensional initial coordinates and their corresponding confidence scores, the control module can generate the target matrix according to this target data structure.

[0078] Step S63: Input the target matrix into the completion model, and complete the key points of the two-dimensional initial coordinates according to the target matrix to obtain the two-dimensional completed coordinates.

[0079] In this embodiment, the aforementioned completion model, namely the 3D keypoint completion model, is used to complete the initial 2D coordinates of the touch terminal into 2D completed coordinates. These 2D completed coordinates are actually 3D coordinates, but their coordinate system is not transformed. For example, if the initial 2D coordinates are in a pixel coordinate system, the completed 2D coordinates are also in a pixel coordinate system. This completion model can be a deep learning model, meaning it can be obtained by inputting a large number of 2D sample coordinates, the corresponding confidence scores, and the corresponding 3D sample coordinates into the deep learning model for a deep learning process.

[0080] Step S64: Convert the two-dimensional completed coordinates into the first spatial coordinates of the touch terminal in the projection area.

[0081] In this embodiment, since the two-dimensional completed coordinate transformation and the two-dimensional initial coordinates are in the same coordinate system, it is also necessary to transform the two-dimensional completed coordinates to the first spatial coordinate system to obtain the first spatial coordinates. For example, when the two-dimensional completed coordinates are in the pixel coordinate system, the two-dimensional completed coordinates can be transformed into the first spatial coordinates using the camera's intrinsic and extrinsic parameters.

[0082] For example, an embodiment of this application provides a key point detection method, which is one embodiment of step S61 above, specifically including the following steps:

[0083] Step S611: Convert the image data into luminance-normalized color space data.

[0084] Projection units typically use high-power spotlights to ensure the clarity of the image on the projection plane. As a result, the acquired image data may have local or global high brightness, which can affect the accuracy of key point detection.

[0085] In this embodiment of the application, to solve the above-mentioned problems, the image data is converted into luminance-normalized color space data, that is, the luminance of locally high or globally high parts of the image is normalized. For example, when the image data is RGB data, it can first be converted into YCbCr data, and then all Y components in the YCbCr data are normalized, and the Y component is the luminance component of the YCbCr data.

[0086] Step S612: Input the color space data into the two-dimensional key point detection model to obtain multiple two-dimensional initial coordinates and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates.

[0087] In this embodiment of the application, the two-dimensional key point detection model can be a deep learning model. That is, a large number of image data samples and the corresponding key point data in the image data samples can be input into the deep learning model to perform a deep learning process and finally obtain the two-dimensional key point detection model.

[0088] For example, an embodiment of this application provides a second spatial coordinate acquisition method, which is one embodiment of step S52 above, and specifically includes the following steps:

[0089] Step S521: Obtain the pixel coordinates of the touch screen.

[0090] In this embodiment, to obtain the depth data corresponding to the touch device, it is first necessary to locate the current position of the touch device. This can be achieved by obtaining the pixel coordinates of the touch device. Specifically, for example, image data can be input into a preset touch device detection model to obtain the pixel coordinates of the touch device. Alternatively, to further improve the accuracy of the pixel coordinates, the image data can be converted into luminance-normalized color space data, and then the color space data can be input into the preset touch device detection model to obtain the pixel coordinates of the touch device.

[0091] The touch detection model can also be a deep learning model. This means that a large number of image data samples and the corresponding touch pixel coordinates in the image data samples can be input into the deep learning model to perform a deep learning process and finally obtain the touch detection model.

[0092] Step S522: Obtain the second spatial coordinates based on the pixel coordinates, the preset pose relationship, and the depth data. The preset pose relationship includes the relative pose relationship between the camera and the depth sensor.

[0093] In this embodiment of the application, the depth data corresponding to the touch terminal can be extracted from the depth data by using the pixel coordinates and the relative pose relationship, and then the second spatial coordinates can be obtained based on the depth data corresponding to the touch terminal.

[0094] For example, such as Figure 7 The diagram shown is a flowchart illustrating a planar equation calculation method provided in this application. This method is one embodiment of step S53 described above and specifically includes the following steps:

[0095] Step S71: Obtain a preset number of depth data points corresponding to the projection plane.

[0096] In this embodiment, the control module acquires depth data points covering the entire projection plane using a depth sensor, and then selects a preset number of depth data points from all the depth data points to reduce the computational workload of establishing the plane equation. To ensure the quality of the selected depth data points, the selection process can ensure that the depth data points are evenly distributed on the projection plane. For example, the distance between adjacent depth data points can be fixed, or the number of other depth data points between adjacent depth data points can be consistent, etc., without limitation.

[0097] Step S72: Establish the solution equations corresponding to the projection plane.

[0098] In this embodiment of the application, the above-mentioned solution equation can be a plane equation for which the coefficients have not been solved, such as Ax+By+Cz+D=0, where A, B, C and D are coefficients for which the coefficients have not been solved, while x, y and z can be the corresponding coordinate values ​​obtained from the depth data points.

[0099] Step S73: Obtain the equation coefficients corresponding to the solution equation by using a preset number of depth data points corresponding to the projection plane.

[0100] Step S74: Obtain the plane equation corresponding to the projection plane based on the equation coefficients and the solved equation.

[0101] In this embodiment of the application, taking the above-mentioned solution equation Ax+By+Cz+D=0 as an example, the depth data points can first be converted into three-dimensional spatial coordinates. The process is the same as or similar to the process of obtaining the second spatial coordinates, and is not limited here. The coordinate values ​​of the multiple three-dimensional spatial coordinates are input into the solution equation, and finally the values ​​of coefficients A, B, C and D are obtained by solving, thus obtaining the plane equation corresponding to the projection plane.

[0102] Please see Figure 8 , Figure 8 This is a flowchart illustrating another operation action recognition method provided in an embodiment of this application. The operation action recognition method includes steps S101 to S104, wherein steps S101 to S104 are... Figure 4 Steps S41 to S44 are the same or similar; please refer to [link / reference needed]. Figure 4 The descriptions of steps S41 to S44 are omitted here. It is understandable that... Figure 10 The operation action recognition method shown is the same as Figure 4 Compared to the operation action recognition method shown, its difference lies in that, Figure 10 The operation action recognition method shown also includes steps S85 and S86.

[0103] Step S85: Generate operation instructions based on the operation action and the projection plane.

[0104] Step S86: Display the operation result obtained in response to the operation command through the projection plane.

[0105] In this embodiment, after recognizing an operation action, the control module can match the operation action and the projection plane with a pre-set operation instruction library to obtain the operation instruction corresponding to the operation action from the operation instruction library and execute the operation instruction. Simultaneously, the control module will also project the result of executing the operation instruction on the projection plane onto the projection plane via the projection unit. For example, when the operation action is to swipe the screen, the control module will execute the swiping instruction and project the swiping animation onto the projection plane.

[0106] Please see Figure 9 , Figure 9 This is a flowchart illustrating a preferred operation action recognition method provided for an embodiment of this application. The operation action recognition method specifically includes the following steps:

[0107] Step S901: Obtain depth data.

[0108] Step S902: Obtain RGB image data.

[0109] The detailed description of steps S901 and S902 can be found in step S41 of the aforementioned embodiment, and will not be repeated here.

[0110] Step S903: Convert RGB image data to YCbCr data.

[0111] Step S904: Normalize the brightness of the YCbCr data.

[0112] The detailed description of steps S903 and S904 can be found in step S611 of the aforementioned embodiment, and will not be repeated here.

[0113] Step S905: Establish plane equations based on depth data.

[0114] The detailed description of step S905 can be found in steps S71 to S74 of the aforementioned embodiments, and will not be repeated here.

[0115] Step S906: Input the brightness-normalized YCbCr data into the two-dimensional key point detection model to obtain the confidence level of the two-dimensional initial coordinates.

[0116] Step S907: Combine the initial two-dimensional coordinates with the depth data to obtain the second spatial coordinates.

[0117] In this embodiment of the application, the two-dimensional key point detection model is used as the touch terminal detection model in step S521 above, and the two-dimensional initial coordinates are used as the pixel coordinates in step S521 above.

[0118] Step S908: Input the confidence scores of the initial two-dimensional coordinates into the three-dimensional keypoint completion model to obtain the two-dimensional completed coordinates.

[0119] Step S909: Use the PnP algorithm to convert the two-dimensional completed coordinates into first spatial coordinates.

[0120] The detailed description of steps S908 and S909 can be found in steps S61 to S64 of the aforementioned embodiments, and will not be repeated here.

[0121] Step S910: Obtain the first spatial coordinates.

[0122] Step S911: Obtain the target spatial coordinates based on the first spatial coordinates and the second spatial coordinates.

[0123] The detailed description of step S911 can be found in steps S51 to S53 of the aforementioned embodiments, and will not be repeated here.

[0124] Step S912: Identify the operation action based on the plane equation and the target space coordinates.

[0125] The detailed description of step S912 can be found in step S44 of the aforementioned embodiment, and will not be repeated here.

[0126] In this embodiment, since the first spatial coordinates are derived from image data and the second spatial coordinates are derived from depth data, the reconstructed target spatial coordinates are more robust than those obtained solely from depth data, thereby improving the accuracy of touch-screen operation recognition. Converting image data into brightness-normalized color space data—that is, normalizing locally high or globally high brightness areas of the image—improves the accuracy of keypoint detection. Smoothing the uneven projection plane by constructing a plane equation, and combining this plane equation with the target spatial coordinates of the touch screen, further enhances touch recognition accuracy.

[0127] Please see Figure 10 , Figure 10 This is a schematic diagram of an operation action recognition device provided in an embodiment of this application. The operation action recognition device 1000 includes: a data acquisition module 1010, used to acquire image data and depth data corresponding to a projection area, the projection area including a projection plane; a target coordinate acquisition module 1020, used to determine the target spatial coordinates of the touch terminal in the projection area based on the image data and depth data corresponding to the projection area; a plane equation acquisition module 1030, used to calculate the plane equation corresponding to the projection plane based on the depth data; and an operation action acquisition module 1040, used to identify the operation action of the touch terminal on the projection plane based on the target spatial coordinates and the plane equation.

[0128] The target coordinate acquisition module 1020 includes: a first spatial coordinate acquisition unit, used to determine the first spatial coordinates of the touch terminal in the projection area based on image data; a second spatial coordinate acquisition unit, used to determine the second spatial coordinates of the touch terminal in the projection area based on depth data; and a target spatial coordinate acquisition unit, used to determine the target spatial coordinates based on the first spatial coordinates and the second spatial coordinates.

[0129] The first spatial coordinate acquisition unit includes: a key point detection subunit, used to perform key point detection on the touch terminal based on image data to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal, and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates; a target matrix determination subunit, used to determine a target matrix based on the multiple two-dimensional initial coordinates and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates; a two-dimensional completion coordinate acquisition subunit, used to input the target matrix into the completion model, and perform key point completion on the two-dimensional initial coordinates based on the target matrix to obtain two-dimensional completion coordinates; and a transformation subunit, used to convert the two-dimensional completion coordinates into first spatial coordinates.

[0130] The keypoint detection subunit includes: an image data conversion subunit, used to convert image data into brightness-normalized color space data; and a two-dimensional keypoint detection subunit, used to input color space data into a two-dimensional keypoint detection model to obtain multiple two-dimensional initial coordinates and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates.

[0131] The second spatial coordinate acquisition unit includes: a pixel coordinate acquisition subunit for acquiring the pixel coordinates of the touch terminal; and a second spatial coordinate acquisition subunit for acquiring the second spatial coordinates based on the pixel coordinates, a preset pose relationship, and depth data, wherein the preset pose relationship includes the relative pose relationship between the camera and the depth sensor.

[0132] The plane equation acquisition module 1030 includes: a depth data point acquisition subunit, used to acquire a preset number of depth data points corresponding to the projection plane; an equation establishment subunit, used to establish the equation corresponding to the projection plane; an equation coefficient solving subunit, used to acquire the equation coefficients corresponding to the equation using the preset number of depth data points corresponding to the projection plane; and a plane equation acquisition subunit, used to acquire the plane equation corresponding to the projection plane based on the equation coefficients and the equation.

[0133] The operation action recognition device 1000 further includes an operation command execution module, which is used to obtain the operation command corresponding to the operation action, execute the operation command, and project the result of executing the operation command onto a projection plane.

[0134] In the embodiments of this application, more detailed functional descriptions of the above-mentioned modules, units and sub-units can be found in the corresponding content of the foregoing section, and will not be repeated here.

[0135] This application also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to perform the aforementioned operation action recognition method. If the various modules of the aforementioned operation action recognition device are implemented as software functional units and sold or used as independent products, they can be stored in the storage medium.

[0136] This application also provides a robot, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the operation action recognition method described above.

[0137] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.

[0139] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.

Claims

1. A method for recognizing operational actions, characterized in that, The method includes: Acquire image data and depth data corresponding to the projection area, wherein the projection area includes the projection plane obtained by projection; Based on the image data and depth data corresponding to the projection area, the target spatial coordinates of the touch terminal in the projection area are determined; The plane equation corresponding to the projection plane is calculated based on the depth data. Based on the target spatial coordinates and the plane equation, identify the operation actions of the touch terminal on the projection plane; Determining the target spatial coordinates of the touch terminal in the projection area based on the image data and depth data corresponding to the projection area includes: The first spatial coordinates of the touch terminal in the first spatial coordinate system of the projection area are determined based on the image data. The second spatial coordinates of the touch terminal in the second spatial coordinate system of the projection area are determined based on the depth data. The target spatial coordinates are determined based on the first spatial coordinates and the second spatial coordinates; Determining the first spatial coordinates of the touch terminal in the first spatial coordinate system of the projection area based on the image data includes: Based on the image data, key point detection is performed on the touch terminal to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal, and the confidence level corresponding to each of the multiple two-dimensional initial coordinates; The target matrix is ​​determined based on the plurality of two-dimensional initial coordinates and the confidence scores corresponding to each of the plurality of two-dimensional initial coordinates; The target matrix is ​​input into the 3D keypoint completion model. Based on the target matrix, the 2D initial coordinates are completed with keypoints to obtain 2D completed coordinates, which are 3D coordinates. Based on the camera's internal and external parameters, the two-dimensional completed coordinates are converted into the first spatial coordinates of the touch device in the projection area.

2. The operation action recognition method as described in claim 1, characterized in that, The step of performing key point detection on the touch terminal based on the image data to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal, and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates, includes: The image data is converted into luminance-normalized color space data; The color space data is input into the two-dimensional key point detection model to obtain the multiple two-dimensional initial coordinates and the confidence scores corresponding to each of the multiple two-dimensional initial coordinates.

3. The operation action recognition method as described in claim 1, characterized in that, Determining the second spatial coordinates of the touch terminal in the projection area based on the depth data includes: Obtain the pixel coordinates of the touch terminal; The second spatial coordinates are obtained based on the pixel coordinates, the preset pose relationship, and the depth data. The preset pose relationship includes the relative pose relationship between the camera and the depth sensor.

4. The operation action recognition method as described in claim 1, characterized in that, The calculation of the plane equation corresponding to the projection plane based on the depth data includes: Obtain a preset number of depth data points corresponding to the projection plane; Establish the solution equations corresponding to the projection plane; The equation coefficients corresponding to the solution equation are obtained by using a preset number of depth data points corresponding to the projection plane; Based on the equation coefficients and the solution equation, the plane equation corresponding to the projection plane is obtained.

5. The operation action recognition method as described in claim 1, characterized in that, The method further includes: Based on the operation action and the projection plane, an operation command is generated; The projection plane displays the operation results obtained in response to the operation command.

6. An operation action recognition device, characterized in that, The device includes: The data acquisition module is used to acquire image data and depth data corresponding to the projection area, wherein the projection area includes the projection plane obtained by projection; The target coordinate acquisition module is used to determine the target spatial coordinates of the touch terminal in the projection area based on the image data and depth data corresponding to the projection area; A plane equation acquisition module is used to calculate the plane equation corresponding to the projection plane based on the depth data; An operation action acquisition module is used to identify the operation actions of the touch terminal on the projection plane based on the target spatial coordinates and the plane equation; Determining the target spatial coordinates of the touch terminal in the projection area based on the image data and depth data corresponding to the projection area includes: The first spatial coordinates of the touch terminal in the first spatial coordinate system of the projection area are determined based on the image data. The second spatial coordinates of the touch terminal in the second spatial coordinate system of the projection area are determined based on the depth data. The target spatial coordinates are determined based on the first spatial coordinates and the second spatial coordinates; Determining the first spatial coordinates of the touch terminal in the first spatial coordinate system of the projection area based on the image data includes: Based on the image data, key point detection is performed on the touch terminal to obtain multiple two-dimensional initial coordinates corresponding to the touch terminal, and the confidence level corresponding to each of the multiple two-dimensional initial coordinates; The target matrix is ​​determined based on the plurality of two-dimensional initial coordinates and the confidence scores corresponding to each of the plurality of two-dimensional initial coordinates; The target matrix is ​​input into the 3D keypoint completion model. Based on the target matrix, the 2D initial coordinates are completed with keypoints to obtain 2D completed coordinates, which are 3D coordinates. Based on the camera's internal and external parameters, the two-dimensional completed coordinates are converted into the first spatial coordinates of the touch device in the projection area.

7. A robot, characterized in that, The robot includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the operation recognition method as described in any one of claims 1 to 5.

8. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the operation action recognition method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Depth estimation method and apparatus of monocular video, terminal, and storage medium

    CN108765481A

  • Projection control method and device, projection interaction system and storage medium

    CN109816723A