Dynamic object identification method and device
By using a depth camera in VR equipment to acquire multi-frame depth images and identify dynamic objects in preset areas, the problem of the existing technology being unable to recognize dynamic objects is solved, and the accurate identification and position determination of dynamic objects are achieved, which improves user security.
Patent Information
- Application Number
- CN202311833969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
Existing VR devices have shortcomings in identifying dynamic objects and cannot effectively identify dynamic objects in the preset area, resulting in users who may encounter dynamic objects when using VR devices, which poses safety risks.
By using a depth camera to acquire multiple frames of depth images in VR devices, determine the spatial points corresponding to the pixel points in each frame of depth image, and determine whether there is a dynamic object in the grid based on whether the grid is connected to the light center of the depth camera and the spatial point line.
Accurate identification and position determination of dynamic objects in the preset area is realized, effectively avoiding collision between users and dynamic objects, and improving user safety.
Smart Images

Figure CN120220040A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of VR, and in particular, to a method and device for identifying dynamic objects. Background Art
[0002] VR (Virtual Reality) devices are mainly used in scenarios such as gaming and making friends. When a user uses a VR device, they can see the content of the virtual world through a display screen, but cannot observe the real world. When the user moves during the use of the VR device, it is easy to touch objects such as walls and tables.
[0003] The current method is that the user first calibrates a safe area, and when the user walks out of the safe area, the game and other interfaces will be forced to exit. Although this method can prevent the user from touching these static objects, if a dynamic object suddenly appears in the safe area, the current technology cannot identify the dynamic object, so it will also pose a certain safety hazard to the user. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method and device for identifying dynamic objects to identify dynamic objects in a preset area.
[0005] In a first aspect, embodiments of the present disclosure provide a method for identifying dynamic objects, which is applied to a VR device. The VR device includes a depth camera. The method for identifying dynamic objects includes: determining a preset area, where the preset area includes a plurality of grids; obtaining multiple frames of depth images taken in the preset area; for each frame of depth image, determining a first spatial point corresponding to a pixel point in the depth image in the preset area; determining that the grid passed through by the line connecting the camera optical center of the depth camera to the first spatial point when the depth image is taken is an intermediate grid; determining the grid where the first spatial point is located as an end grid; for each grid in the plurality of grids, determining whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images.
[0006] In a second aspect, embodiments of the present disclosure provide a device for identifying dynamic objects, which is applied to a VR device. The VR device includes a depth camera. The device for identifying dynamic objects includes:
[0007] A first determination unit, configured to determine a preset area, where the preset area includes a plurality of grids;
[0008] An obtaining unit, configured to obtain multiple frames of depth images taken in the preset area;
[0009] A second determination unit, configured to determine, for each frame of depth image, a first spatial point corresponding to a pixel point in the depth image in the preset area;
[0010] A third determination unit, configured to determine that the grid passed through by the line connecting the optical center of the depth camera to the first spatial point when capturing a depth image is the middle grid; and determine that the grid where the first spatial point is located is the end grid.
[0011] A fourth determination unit, configured to determine, for each grid among multiple grids, whether the grid is a dynamic grid with a dynamic object according to whether the grid is a middle grid and / or an end grid based on multiple depth images.
[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0013] The memory stores computer-executable instructions;
[0014] At least one processor executes the computer-executable instructions stored in the memory, so that at least one processor executes the dynamic object recognition method provided in the first aspect above.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the dynamic object recognition method provided in the first aspect above is implemented.
[0016] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer-executable instructions. When a processor executes the computer-executable instructions, the dynamic object recognition method provided in the first aspect above is implemented.
[0017] For the dynamic object recognition method and device provided in this embodiment, the present disclosure determines a preset area, which includes multiple grids; acquires multiple depth images captured in the preset area; for each depth image, determines a first spatial point corresponding to a pixel point in the depth image in the preset area; determines that the grid passed through by the line connecting the optical center of the depth camera to the first spatial point when capturing the depth image is the middle grid; determines that the grid where the first spatial point is located is the end grid; and determines, for each grid among multiple grids, whether the grid is a dynamic grid with a dynamic object according to whether the grid is a middle grid and / or an end grid based on multiple depth images, so as to accurately identify the dynamic object and the position of the dynamic object in the preset area. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 The application scenario diagram of a dynamic object recognition method provided by an embodiment of the present disclosure;
[0020] Figure 2 The flowchart of steps of a dynamic object recognition method provided by an embodiment of the present disclosure;
[0021] Figure 3 The schematic diagram of determining an intermediate grid provided by an embodiment of the present disclosure;
[0022] Figure 4 The flowchart of steps of another dynamic object recognition method provided by an embodiment of the present disclosure;
[0023] Figure 5 The structural block diagram of a dynamic object recognition device provided by an embodiment of the present disclosure;
[0024] Figure 6 The schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0026] In the related art, it is the user who first calibrates a safe area, and when the user walks out of the safe area, the game and other interfaces will be forced to exit. Although this method can prevent the user from touching these static objects, if a dynamic object suddenly appears in the safe area, the current technology cannot recognize this dynamic object, so it will also bring certain safety hazards to the user.
[0027] Based on the above problems, the present disclosure provides a dynamic object recognition method. By rasterizing a preset area, then collecting a depth image in the preset area, the grid passed through by the connection line between the first spatial point corresponding to the pixel point in the depth image in the preset area and the camera optical center of the depth camera is the intermediate grid, and the end grid where the first spatial point is located. According to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images, it is possible to efficiently and accurately determine whether there is a dynamic object in the grid and the dynamic grid where the dynamic object exists.
[0028] An application scenario of the present disclosure is referred to Figure 1, including a preset area which is divided into multiple grids. Among them, the user wears a VR device and can move in the preset area. During the movement, the VR device can collect depth images, and then can determine which grid contains a dynamic object to prevent the user from hitting the dynamic object.
[0029] Reference Figure 2 , which is a schematic flowchart of the dynamic object recognition method provided by the embodiment of the present disclosure. As Figure 2 shown, this dynamic object recognition method is applied to a VR device. The VR device includes a depth camera. The dynamic object recognition method specifically includes the following steps:
[0030] S201. Determine the preset area.
[0031] In the present disclosure, the preset area is preset by the user and includes multiple grids. Specifically, when the user uses the VR device, a safety area will be drawn, and this safety area is the preset area.
[0032] Furthermore, multiple grids are determined according to the area of the preset area. Among them, each grid represents a sub-area of the preset area.
[0033] Specifically, after obtaining the preset area, determine the area of the preset area, and initialize a section of memory for storing the representation of the preset area and the grids. The preset area can be represented by coordinates in a three-dimensional space coordinate system, and the grids are also represented by coordinates in the three-dimensional space coordinate system.
[0034] For example, the area of the preset area is 30 square meters, and the area × the first value × the second value of grids are initialized in the memory. Among them, the first value represents how many grids there are per square meter on the X1Y1 plane as Figure 1 shown, and the second value represents how many layers of grids there are in the Z1 axis direction as Figure 1 shown.
[0035] Among them, the first value and the second value are preset. If the first value is 100 and the second value is 5, then 15,000 grids need to be initialized for 30 square meters, and the size (length × width × height) of each grid can be 10 cm × 10 cm × 40 cm.
[0036] Furthermore, in the present disclosure, since the VR device is applied in a preset scene and the area of the preset scene has an upper limit, the preset area drawn by the user also has an upper limit. Therefore, the initial number of grids in the memory also has an upper limit value.
[0037] S202. Obtain multiple frames of depth images taken in the preset area.
[0038] Among them, the depth image is captured by the depth camera of the VR device. In the embodiments of the present disclosure, a depth camera is provided on the VR device, and the depth camera can capture a frame of depth image every first preset time. The pixel size of the depth image is not limited.
[0039] Further, it can be to capture a frame of depth image every preset time, and obtain multiple frames of depth images captured within a preset duration. Exemplarily, a frame of depth image can be captured every 1 s. If the preset duration is 10 s, then 10 frames of depth images are obtained.
[0040] In addition, it can also be that a frame of depth image is captured every time the user takes a step, and multiple frames of depth images captured within a preset number of steps of the user are obtained. The present disclosure does not limit how to obtain multiple frames of depth images.
[0041] S203. For each frame of depth image, determine the first spatial point corresponding to the pixel point in the depth image in the preset area
[0042] Specifically, for each frame of depth image, first determine the image coordinates of the pixel point in the depth image, and then convert the image coordinates according to the conversion relationship from the depth camera coordinate system to the three-dimensional space coordinate system to obtain the first spatial point of the pixel point in the preset area (the first spatial point in the three-dimensional space coordinate system is used to represent the first spatial point). Refer to Figure 1 , the three-dimensional space coordinate system is such as X1Y1Z1.
[0043] In the disclosure, the corresponding first spatial point can be determined for each pixel point of each frame of depth image, and the corresponding first spatial point can be determined for each pixel point in some pixel points of each frame of depth image.
[0044] S204. Determine that the grid passed through by the line connecting the camera optical center of the depth camera to the first spatial point when the depth image is captured is the middle grid.
[0045] In the embodiments of the present disclosure, taking the grid having only one layer as an example, refer to Figure 3 , for example, the preset area includes grids a1 to a20. At time T1 (corresponding to a depth image), the camera optical center of the depth camera is at point P1 in the preset area, and a pixel point of the depth image corresponds to point Z1 in the preset area (the position where the first spatial point is located), then the middle grids are grids a10 and a15. Another pixel point corresponds to point Z2 in the preset area (the position where the first spatial point is located), then the middle grid is grid a6.
[0046] At time T2 (corresponding to another depth image), the optical center of the depth camera is at point P2 in the preset area, and a pixel point of the depth image corresponds to point Z3 (the position where the first spatial point is located) in the preset area. Then the intermediate grids are grid a10, grid a11, grid a7, and grid a8. Another pixel point corresponds to point Z4 (the position where the first spatial point is located) in the preset area, and there is no intermediate grid.
[0047] S205. Determine the grid where the first spatial point is located as the end grid.
[0048] Refer to Figure 3 In the depth image, there will be multiple pixel points, so there will be multiple first spatial points. Furthermore, for each intermediate grid corresponding to the line connecting the camera optical center to the first spatial point, the grid where each first spatial point is located is the end grid. For example, at time T1, a first spatial point corresponding to a pixel point on the depth image is point Z1, then the end grid is grid a20. Another pixel point on the depth image corresponds to a first spatial point Z2, then the end grid is a7. At time T2, a first spatial point corresponding to a pixel point on the depth image is point Z3, then the end grid is grid a4. Another pixel point on the depth image corresponds to a first spatial point Z4, then the end grid is a11.
[0049] S206. For each grid among the multiple grids, determine whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images.
[0050] Determining whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images includes: if the grid is both an intermediate grid and an end grid based on multiple frames of depth images, then determine that the grid is a dynamic grid with a dynamic object.
[0051] It can be understood that for any frame of depth image, the grids not occupied by objects are mostly intermediate grids, and the grids occupied by objects are end grids. Then, after multiple shootings to obtain multiple frames of depth images, if a grid is always an end grid based on different depth images, it can be determined that the grid is occupied by a static object. If a grid is always an intermediate grid based on different depth images, it can be determined that the grid is not occupied by an object. If a grid is both an intermediate grid and an end grid based on different depth images, it can be determined that there may have been a dynamic object in the grid.
[0052] Such as Figure 3 Grids a7 and a11. Among them, grid a7 is an end grid based on the depth image at time T1 and an intermediate grid based on the depth image at time T2. Then it can be determined that there has been a dynamic object in grid a7, and it can be determined that grid a7 is a dynamic grid. Similarly, grid a11 can be determined as a dynamic grid.
[0053] With the above method, the present disclosure can determine whether each grid in the preset area division is a dynamic grid.
[0054] Further, it can be set that if a grid is both an intermediate grid and an end grid based on multiple frames of depth images, it is determined that the grid is a dynamic grid with a dynamic object, including: if the number of times the grid is an intermediate grid and the number of times it is an end grid in multiple frames of depth images satisfy a preset proportional relationship, it is determined that the grid is a dynamic grid with a dynamic object.
[0055] Specifically, the preset proportional relationship is preset and satisfies being less than or equal to 4 and greater than or equal to 0.5. Among them, if the ratio of the number of times as an intermediate grid to the number of times as an end grid is x, and 0 satisfies 0.5 ≤ x ≤ 4, then it is determined that the grid is a dynamic grid.
[0056] In another alternative embodiment, the grid supports a first variable and a second variable, and the initial values of the first variable and the second variable are the same. For each grid among multiple grids, according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images, determining whether the grid is a dynamic grid with a dynamic object includes: obtaining the first number of times the grid is an intermediate grid based on multiple frames of depth images, and adding a first preset value to the first variable and subtracting the first preset value from the second variable based on the first number; obtaining the second number of times the grid is an end grid based on multiple frames of depth images, and adding a first preset value to the first variable and adding a first preset value to the second variable based on the second number; when the current first variable and second variable of the grid satisfy a preset relationship, it is determined that the grid is a dynamic grid with a dynamic object.
[0057] Specifically, for a grid not occupied by an object, since the first preset value is added to the first variable and the first preset value is subtracted from the second variable multiple times, the first difference between the first variable and the second variable becomes larger and larger. The grid occupied by a static object is mostly an end grid (the grid where the first spatial point is located). Since the first preset value is added to the first variable and the first preset value is also added to the second variable for this end grid, after multiple calculations, the second difference between the first variable and the second variable of this end grid is still very small or there is no difference. Since the dynamic object is moving, the corresponding grid is sometimes an intermediate grid and sometimes a grid occupied by the dynamic object. Then, the first variable of this grid is always added with the first preset value, and the second variable is sometimes added with the first preset value and sometimes subtracted with the first preset value. Then, the third difference between the first variable and the second variable will be between the first difference and the second difference. Furthermore, within a period of time (such as 10 s), by shooting multiple frames of depth images, it is possible to determine the dynamic grids with dynamic objects.
[0058] In an alternative embodiment, for an intermediate grid determined for a depth image, a first preset value may be added once to the first variable of the intermediate grid, and the first preset value may be subtracted once from the second variable. For example, for the depth image X1 captured at time T1, the intermediate grids corresponding to the pixel point D1 of the depth image are determined to be grid a10 and grid a15, the intermediate grids corresponding to the pixel point D2 are determined to be a6 and a7, and the intermediate grids corresponding to the pixel point D3 are determined to be a6 and a10. Among them, grid a6 and grid a10 are determined as intermediate grids multiple times, but the first preset value is added only once to the first variable of grid a6 and grid a10, and the first preset value is subtracted only once from the second variable.
[0059] In another alternative embodiment, for an intermediate grid determined for a depth image, if the intermediate grid is determined to be an intermediate grid n times, the corresponding first variable is added n times with the first preset value, and the second variable is subtracted n times with the first preset value. For example, for the depth image X1 captured at time T1, the intermediate grids corresponding to the pixel point D1 of the depth image are determined to be grid a10 and grid a15, the intermediate grids corresponding to the pixel point D2 are determined to be a6 and a7, and the intermediate grids corresponding to the pixel point D3 are determined to be a6 and a10. Among them, grid a6 and grid a10 are determined as intermediate grids 2 times, but the first preset value is added 2 times to the first variable of grid a6 and grid a10, and the first preset value is subtracted 2 times from the second variable.
[0060] Among them, the first preset value is a positive number, and the initial first variable and the initial second variable of the grid are the same.
[0061] It can be understood that by adopting this calculation method, most of the grids not occupied by objects are intermediate grids. Then, after capturing multiple depth images for calculation, for the grids not occupied by objects, since the first preset value is added multiple times to the first variable and subtracted multiple times from the second variable, the first difference between the first variable and the second variable becomes larger and larger. Most of the grids occupied by static objects are end grids (the grids where the first spatial points are located). For this end grid, since the first preset value is added to the first variable and the first preset value is also added to the second variable, after multiple calculations, the second difference between the first variable and the second variable of this end grid is still very small or there is no difference. Since the dynamic object is moving, the corresponding grid is sometimes an intermediate grid and sometimes a grid occupied by the dynamic object. Then, the first variable of this grid is always added with the first preset value, and the second variable is sometimes added with the first preset value and sometimes subtracted with the first preset value. Then, the third difference between the first variable and the second variable will be between the first difference and the second difference. Furthermore, within a certain period of time (such as 10 s), by capturing multiple frames of depth images, the dynamic grids where dynamic objects exist can be determined.
[0062] Specifically, when the difference between the first variable and the second variable of a grid is within a preset range, it is determined that there is a dynamic object in this grid.
[0063] Specifically, the initial values of both the first variable and the second variable are 0, and the first preset value is 1.
[0064] In the present disclosure, during the process of a user using a VR device, the depth camera scans a preset area at intervals to obtain a depth image. For each or some of the pixel points in the depth camera, they can be converted into a point in a three-dimensional space coordinate system. Taking this point as the end point, and the optical center of the depth camera is determined as the starting point in the three-dimensional space coordinate system. If the line connecting the starting point to the end point successively passes through many grids, these grids are used as intermediate grids, and the first variable of each passed grid is incremented by 1, and the second variable is decremented by 1. For the grid where the end point is located, the first variable is incremented by 1 and the second variable is incremented by 1.
[0065] During initialization, the first variable and the second variable of all grids are set to 0. Then, within a second preset time (such as 10 s), the first variable and the second variable are iteratively updated according to multiple captured depth images. After determining the dynamic object based on the first variable and the second variable, the first variable and the second variable of all grids are initialized to 0 again. It can be understood that in the present disclosure, an initialization operation is performed on all grids every second preset time.
[0066] Furthermore, for each of the multiple grids, if the difference between the current first variable and the current second variable of the grid is less than a first threshold, it is determined that the grid is the first grid occupied by a static object.
[0067] In the embodiments of the present disclosure, first, traverse the multiple grids, and determine that the grids with the first variable greater than a preset threshold (such as 3) are grids to be determined. Among the grids to be determined, the first grid, the second grid, and the dynamic grids are determined. Further, the first threshold is preset, such as 1 or 2. If the difference between the current first variable and the current second variable of the grid is less than the first threshold, it can be understood that the first variable and the second variable are very close, then it can be determined that this grid is occupied by a static object.
[0068] In an alternative embodiment, for example, it can be set that when "the second variable > (0.7 × the first variable)", the grid is occupied by a static object.
[0069] Wherein, if the current second variable of the grid is negative, and the difference between the absolute value of the second variable and the current first variable is less than a second threshold, it is determined that the grid is the second grid without an object occupying it.
[0070] Specifically, the second threshold is preset, such as 1 or 2. If the difference between the absolute value of the second variable and the current first variable is less than the second threshold, it can be understood that the absolute value of the second variable and the first variable are very close, then it can be determined that this is not occupied by a static object. In an alternative embodiment, for example, it can be set that when "the second variable < (-0.7 × the first variable)", the grid has no object occupying it.
[0071] Among them, if the grid is neither the first grid nor the second grid, the grid is determined to be a dynamic grid.
[0072] It can be understood that if the grid to be determined is neither the first grid nor the second grid, it is a dynamic grid, and the dynamic grid is occupied by a dynamic object.
[0073] In the present disclosure, a depth image can be taken every first preset time (e.g., 1 s), and the first variable and the second variable are initialized once every second preset time (10 s). Then, 10 frames of depth images can be taken within the second preset time. Among them, within the second preset time, before the first frame of depth image is taken, the first variable and the second variable of each grid are both 0. After the first frame of depth image is taken, the first variable and the second variable are updated for the first frame of depth image. Among them, it is impossible to determine the dynamic object based on the first variable and the second variable after the first frame of depth image is taken. After the second frame of depth image is taken, the first variable and the second variable after the first frame of depth image are updated for the second frame of depth image, and then the dynamic object can be determined based on the first variable and the second variable. After the third frame of depth image is taken, the first variable and the second variable after the second frame of depth image are updated for the third frame of depth image, and then the dynamic object can be determined based on the first variable and the second variable, that is, it can be determined whether there is a dynamic object for each frame of depth image.
[0074] Further, if the dynamic grid can be continuously determined for each frame of depth image, the time for initializing the first variable and the second variable can be extended, such as extending 10 s to initialize the first variable and the second variable. In the present disclosure, the initialization interval duration of the first variable and the second variable is adjustable and is not limited herein.
[0075] In the embodiment of the present disclosure, by rasterizing the preset area, and then collecting a depth image in the preset area, the grid through which the connection line between the corresponding first spatial point in the preset area and the camera optical center of the depth camera passes is the intermediate grid, and the end grid where the first spatial point is located. According to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images, it is possible to efficiently and accurately determine whether there is a dynamic object in the grid and the dynamic grid where the dynamic object exists.
[0076] Reference Figure 4 , in Figure 2 After a dynamic object recognition method provided, the following steps are further included:
[0077] S401. When the number of dynamic grids is greater than the preset number, determine the current distance between the dynamic grid and the user.
[0078] Among them, the preset quantity is set in advance, such as 10. If the number of dynamic grids is less than the preset quantity, it can be determined that the volume of the dynamic object understood is very small and will not pose an obstacle to the user, so it can be ignored. If the number of dynamic grids is greater than the preset quantity, the current distance between the currently determined dynamic grids and the user is determined.
[0079] S402. If the current distance is less than the threshold distance, a prompt message is output.
[0080] Among them, it is determined whether the dynamic grid is the grid where the first spatial point is located. If so, the coordinates of the center or centroid of the dynamic grid and the current distance of the user are calculated. If the current distance is less than the threshold distance (such as 1m), a prompt message is output for alarm to prevent the user from hitting the dynamic object.
[0081] Furthermore, the prompt message may include: the azimuth information and distance information of the dynamic object relative to the user.
[0082] S403. Determine the target dynamic grid.
[0083] The target dynamic grid is the dynamic grid that serves as the end grid in the last frame of target depth image among multiple frames of depth images.
[0084] Specifically, the target dynamic grid can be determined among all the determined dynamic grids.
[0085] S404. Obtain the grayscale image corresponding to the last frame of target depth image captured by the grayscale camera
[0086] In the present disclosure, a grayscale camera is installed on the VR device. The grayscale camera is used to collect images simultaneously with the depth camera, and the grayscale camera is used to collect the grayscale image with the same viewing angle as the depth camera.
[0087] S405. Determine the dynamic object recognition area according to the image coordinates of the target dynamic grid in the grayscale image.
[0088] In the embodiments of the present disclosure, the dynamic object recognition area is the area surrounded by connecting all the target dynamic grids, and the dynamic object recognition area includes all the target dynamic grids.
[0089] Among them, the image coordinates are determined by the following method: obtain the spatial coordinates of the dynamic grid in the preset area; according to the preset conversion formula, convert the spatial coordinates into the image coordinates in the grayscale image.
[0090] Furthermore, the preset conversion formula is expressed as:
[0091]
[0092] Among them, z represents the Z-axis coordinate of the camera coordinate system of the grayscale camera of the VR device, Tcw represents the external parameter matrix of the grayscale camera, K represents the internal parameter matrix of the grayscale camera, and p w represents the spatial coordinates of the center point of the target dynamic grid in the preset area, and p uv represents the image coordinates.
[0093] Specifically, project the centers or centroids of all target dynamic grids onto the pixel coordinates in the grayscale image (image coordinates), then determine the minimum bounding box of multiple image coordinates in the grayscale image as the dynamic object recognition area to enclose the projection points of the target dynamic grids in the grayscale image, then extract the regional image within the bounding box in the grayscale image, and identify the regional image through a pre-trained deep learning network to determine the category of the dynamic object included in the regional image.
[0094] Furthermore, the category of the dynamic object can be output to the prompt message together to prompt the user.
[0095] In the present disclosure, by rasterizing the preset area and determining the dynamic grid according to the grid as the intermediate grid and / or the end grid, and further determining the distance between the dynamic grid and the user to remind the user of the dynamic object. Furthermore, by identifying the category of the dynamic object in the grayscale image, the user can obtain in real time the dynamic object existing in the preset area, as well as the distance and azimuth information between the dynamic object and the user, which can effectively avoid the user colliding with the dynamic object and protect the safety of the user using the VR device.
[0096] Corresponding to the dynamic object recognition method in the above embodiment, Figure 5 is the structural block diagram of the dynamic object recognition device 50 provided by the embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiment of the present disclosure are shown. As Figure 5 shown, the dynamic object recognition device 50 is applied to a VR device, and the VR device includes a depth camera, specifically including:
[0097] The first determination unit 501 is used to determine a preset area, and the preset area includes a plurality of grids;
[0098] The acquisition unit 502 is used to acquire multiple frames of depth images taken in the preset area;
[0099] The second determination unit 503 is used to, for each frame of depth image, determine the first spatial point corresponding to the pixel point in the depth image in the preset area;
[0100] The third determination unit 504 is used to determine that the grid passed through by the line connecting the camera optical center of the depth camera to the first spatial point when taking the depth image is the intermediate grid; determine the grid where the first spatial point is located as the end grid;
[0101] A fourth determination unit 505, configured to determine, for each grid among a plurality of grids, whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images.
[0102] In some embodiments, the fourth determination unit 505 is specifically configured to determine that the grid is a dynamic grid with a dynamic object if the grid is both an intermediate grid and an end grid based on multiple frames of depth images.
[0103] In some embodiments, when the fourth determination unit 505 determines that the grid is a dynamic grid with a dynamic object if the grid is both an intermediate grid and an end grid based on multiple frames of depth images, it is specifically configured to determine that the grid is a dynamic grid with a dynamic object if the number of times the grid serves as an intermediate grid and the number of times it serves as an end grid in multiple frames of depth images satisfy a preset proportional relationship.
[0104] In some embodiments, it further includes: an output unit (not shown), configured to determine a current distance between the dynamic grid and the user when the number of dynamic grids is greater than a preset number; and output a prompt message if the current distance is less than a threshold distance.
[0105] In some embodiments, it further includes: an identification unit (not shown), configured to determine a target dynamic grid, where the target dynamic grid is a dynamic grid that serves as an end grid in the last frame of target depth image in multiple frames of depth images; obtain a grayscale image captured by a grayscale camera corresponding to the last frame of target depth image; and determine a dynamic object recognition area according to the image coordinates of the target dynamic grid in the grayscale image.
[0106] In some embodiments, the identification unit determines the image coordinates in the following manner:
[0107] Obtain the spatial coordinates of the dynamic grid in a preset area;
[0108] Convert the spatial coordinates according to a preset conversion formula to obtain the image coordinates in the grayscale image.
[0109] In some embodiments, the preset conversion formula is expressed as:
[0110]
[0111] where z represents the Z-axis coordinate of the camera coordinate system of the grayscale camera of the VR device, T cw represents the external parameter matrix of the grayscale camera, K represents the internal parameter matrix of the grayscale camera, p w represents the spatial coordinates of the center point of the target dynamic grid in the preset area, and p uv represents the image coordinates.
[0112] The dynamic object recognition device provided in this embodiment can be used to implement the technical solutions of the embodiments of the above-mentioned dynamic object recognition method. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0113] Reference Figure 6 , which shows a schematic structural diagram of an electronic device 60 suitable for implementing the embodiments of the present disclosure. The electronic device 60 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0114] As Figure 6 shown, the electronic device 60 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 61, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 62 or the program loaded from the storage device 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 are also stored. The processing device 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. The input / output (I / O) interface 65 is also connected to the bus 64.
[0115] Generally, the following devices can be connected to the I / O interface 65: an input device 66 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 67 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 68 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 69. The communication device 69 can allow the electronic device 60 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 shows the electronic device 60 having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0116] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 69, or installed from a storage device 68, or installed from a ROM 92. When the computer program is executed by a processing device 61, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0117] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0118] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.
[0119] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the methods shown in the above embodiments.
[0120] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a unit, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0122] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0123] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0124] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM) or a flash memory, an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a method for dynamic object recognition, which is applied to a VR device. The VR device includes a depth camera. The method for dynamic object recognition includes: determining a preset area, where the preset area includes a plurality of grids; acquiring multiple frames of depth images captured in the preset area; for each frame of depth image, determining a first spatial point corresponding to a pixel point in the depth image in the preset area; determining that the grid passed through by the line connecting the camera optical center of the depth camera to the first spatial point when the depth image is captured is an intermediate grid; determining the grid where the first spatial point is located as an end grid; for each grid among the plurality of grids, determining whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on multiple frames of depth images.
[0126] According to one or more embodiments of the present disclosure, determining whether a grid is a dynamic grid with a dynamic object based on whether the grid based on multi-frame depth images is an intermediate grid and / or an end grid includes: if the grid is both an intermediate grid and an end grid based on multi-frame depth images, determining that the grid is a dynamic grid with a dynamic object.
[0127] According to one or more embodiments of the present disclosure, if the grid is both an intermediate grid and an end grid based on multi-frame depth images, determining that the grid is a dynamic grid with a dynamic object includes: if the number of times the grid serves as an intermediate grid and the number of times it serves as an end grid in the multi-frame depth images satisfy a preset proportional relationship, determining that the grid is a dynamic grid with a dynamic object.
[0128] According to one or more embodiments of the present disclosure, it further includes: when the number of dynamic grids is greater than a preset number, determining the current distance between the dynamic grid and the user; if the current distance is less than a threshold distance, outputting a prompt message.
[0129] According to one or more embodiments of the present disclosure, the VR device further includes a grayscale camera, and the method further includes: determining a target dynamic grid, where the target dynamic grid is a dynamic grid that serves as an end grid in the last frame of target depth image in the multi-frame depth images; obtaining a grayscale image captured by the grayscale camera corresponding to the last frame of target depth image; determining a dynamic object recognition region according to the image coordinates of the target dynamic grid in the grayscale image.
[0130] According to one or more embodiments of the present disclosure, the image coordinates are determined by the following method:
[0131] Obtaining the spatial coordinates of the dynamic grid in a preset region;
[0132] Converting the spatial coordinates according to a preset conversion formula to obtain the image coordinates in the grayscale image.
[0133] According to one or more embodiments of the present disclosure, the preset conversion formula is expressed as:
[0134]
[0135] where z represents the Z-axis coordinate of the camera coordinate system of the grayscale camera, T cw represents the external parameter matrix of the grayscale camera, K represents the internal parameter matrix of the grayscale camera, p w represents the spatial coordinates of the center point of the target dynamic grid in the preset region, and p uv represents the image coordinates.
[0136] In a second aspect, according to one or more embodiments of the present disclosure, a dynamic object recognition device is provided.
[0137] Applied to a VR device, the VR device includes a depth camera, and the dynamic object recognition device includes:
[0138] A first determination unit, configured to determine a preset area, where the preset area includes a plurality of grids;
[0139] An acquisition unit, configured to acquire multiple frames of depth images captured in the preset area;
[0140] A second determination unit, configured to, for each frame of depth image, determine a first spatial point corresponding to a pixel point in the depth image in the preset area;
[0141] A third determination unit, configured to determine that the grid through which the connection line from the camera optical center of the depth camera to the first spatial point passes when the depth image is captured is the middle grid; and determine the grid where the first spatial point is located as the end grid;
[0142] A fourth determination unit, configured to, for each grid among the multiple grids, determine whether the grid is a dynamic grid with a dynamic object according to whether the grid is a middle grid and / or an end grid based on the multiple frames of depth images.
[0143] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;
[0144] The memory stores computer-executable instructions;
[0145] At least one processor executes the computer-executable instructions stored in the memory, so that at least one processor executes the dynamic object recognition method provided in the first aspect above.
[0146] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the dynamic object recognition method provided in the first aspect above is implemented.
[0147] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product, the computer program product includes computer-executable instructions, and when a processor executes the computer-executable instructions, the dynamic object recognition method provided in the first aspect above is implemented.
[0148] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0149] In addition, although the operations are depicted in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0150] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for identifying dynamic objects, applied to a VR device, the VR device including a depth camera, the method for identifying dynamic objects comprising: Determine a preset area, the preset area including a plurality of grids; Obtain multiple depth images captured in the preset area; For each depth image, determine a first spatial point corresponding to a pixel point in the depth image in the preset area; Determine that the grid through which the line connecting the camera optical center of the depth camera to the first spatial point passes when the depth image is captured is the intermediate grid; Determine the grid where the first spatial point is located as the end grid; For each grid among the plurality of grids, determine whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on the multiple depth images.
2. The method for identifying dynamic objects according to claim 1, the determining whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on the multiple depth images comprises: If the grid is both an intermediate grid and an end grid based on the multiple depth images, determine that the grid is a dynamic grid with a dynamic object.
3. The method for identifying dynamic objects according to claim 2, the if the grid is both an intermediate grid and an end grid based on the multiple depth images, determine that the grid is a dynamic grid with a dynamic object, comprises: If the number of times the grid is an intermediate grid and the number of times the grid is an end grid in the multiple depth images satisfy a preset proportional relationship, determine that the grid is a dynamic grid with a dynamic object.
4. The method for identifying dynamic objects according to any one of claims 1 to 3, further comprises: When the number of the dynamic grids is greater than a preset number, determine the current distance between the dynamic grids and the user; If the current distance is less than a threshold distance, output a prompt message.
5. The method for identifying dynamic objects according to any one of claims 1 to 3, the VR device further includes the grayscale camera, the method further comprises: Determine a target dynamic grid, the target dynamic grid being a dynamic grid that is an end grid in the last frame of target depth image in the multiple depth images; Obtain a grayscale image captured by the grayscale camera corresponding to the last frame of target depth image; Determine a dynamic object recognition area according to the image coordinates of the target dynamic grid in the grayscale image.
6. The method for identifying dynamic objects according to claim 5, the image coordinates are determined by the following method: Obtain the spatial coordinates of the dynamic grid in the preset area; According to a preset conversion formula, convert the spatial coordinates to obtain the image coordinates in the grayscale image.
7. The preset conversion formula according to claim 6 is expressed as: Among them, where \(z\) represents the \(Z\)-axis coordinate of the camera coordinate system of the grayscale camera, and \(T\) cw represents the external parameter matrix of the grayscale camera, \(K\) represents the internal parameter matrix of the grayscale camera, and \(p\) w represents the spatial coordinate of the center point of the target dynamic grid in the preset area, and \(p\) uv represents the image coordinate.
8. A dynamic object recognition device, applied to a VR device, the VR device including a depth camera, the dynamic object recognition device comprising: A first determination unit, configured to determine a preset area, the preset area including a plurality of grids; An acquisition unit, configured to acquire multiple depth images captured in the preset area; A second determination unit, configured to determine, for each frame of depth image, a first spatial point corresponding to a pixel point in the depth image in the preset area; A third determination unit, configured to determine that a grid through which a connection line from the camera optical center of the depth camera to the first spatial point passes when the depth image is captured is an intermediate grid; and determine that a grid where the first spatial point is located is an end grid; A fourth determination unit, configured to, for each grid in the plurality of grids, determine whether the grid is a dynamic grid with a dynamic object according to whether the grid is an intermediate grid and / or an end grid based on the multi-frame depth images.
9. An electronic device, comprising: At least one processor and a memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the dynamic object recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the dynamic object recognition method according to any one of claims 1 to 7 is implemented.