Terminal device

By using depth maps to compare pixel values within a threshold, the method reduces computational costs in collision detection, addressing the inefficiencies of conventional polygon-based methods.

JP7704135B2Active Publication Date: 2025-07-08TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022212649
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-07-08
Estimated Expiration
2042-12-28

Smart Images

  • Figure 0007704135000001
    Figure 0007704135000001
  • Figure 0007704135000002
    Figure 0007704135000002
  • Figure 0007704135000003
    Figure 0007704135000003
Patent Text Reader

Abstract

To provide a terminal capable of reducing the computational cost in collision determination.SOLUTION: A terminal 10A includes a control unit 17. A first depth map 20 includes multiple first pixels. The first pixel has a pixel value according to the distance to a subject in a 3D space. A second depth map 22 has multiple second pixels. The second pixel has a pixel value according to the distance to a predetermined line set in the 3-dimensional space. A control unit 17 is configured so as to, when the difference between the pixel value of the first pixels and the pixel value of the second pixels at the same coordinate is within a threshold, detect the coordinates as a collision point between the subject and the predetermined line.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a terminal device.

Background Art

[0002] Conventionally, collision detection for determining whether or not two objects in a three-dimensional virtual space collide is known (see Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional collision detection is performed by detecting an intersection point, i.e., a collision point, between a polygon model of one of two objects and a straight line extended from the other object. The other object is, for example, a wall or a polygon model of another object. The number of polygonal faces constituting the polygon model is enormous. Therefore, there has been a problem that the calculation cost becomes high in conventional collision detection.

[0005] In view of such a point, an object of the present disclosure is to reduce the calculation cost in collision detection.

Means for Solving the Problems

[0006] A terminal device according to an embodiment of the present disclosure includes a control unit. The first depth map includes a plurality of first pixels. The first pixel has a pixel value corresponding to the distance to a subject in a three-dimensional space. The second depth map includes a plurality of second pixels. The second pixel has a pixel value corresponding to the distance to a predetermined straight line set in the three-dimensional space. When the difference between the pixel value of the first pixel and the pixel value of the second pixel at the same coordinate is within a threshold value, the control unit detects the coordinate as a collision point between the subject and the predetermined straight line.

Advantages of the Invention

[0007] According to an embodiment of the present disclosure, the calculation cost in collision determination can be reduced.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings.

[0010] A system 1 as shown in FIG. 1 provides a virtual event using a virtual space. The system 1 includes a terminal device 10A and a terminal device 10B. The terminal device 10A and the terminal device 10B can communicate via a network 2. The network 2 may be any network including a mobile communication network and the Internet. Hereinafter, when the terminal devices 10A and 10B are not particularly distinguished, they are also simply referred to as "terminal device 10".

[0011] The terminal device 10A is used by the user 3A. The user 3A participates as a participant in the virtual event using the terminal device 10A. The terminal device 10B is used by the user 3B. The user 3B participates as a participant in the virtual event using the terminal device 10B. The terminal device 10 is, for example, a terminal device such as a desktop PC (Personal Computer), a tablet PC, a notebook PC, or a smartphone.

[0012] The terminal device 10 includes a communication unit 11, an input unit 12, an output unit 13, a camera 14, a distance measuring sensor 15, a storage unit 16, and a control unit 17.

[0013] The communication unit 11 is configured to include at least one communication module connectable to the network 2. The communication module is, for example, a communication module compatible with a standard such as a wired LAN or a wireless LAN, or a communication module compatible with a mobile communication standard such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation).

[0014] The input unit 12 is capable of receiving an input from the user. The input unit 12 is configured to include at least one input interface capable of receiving an input from the user. The input interface is, for example, a physical key, a capacitive key, a pointing device, a touch screen provided integrally with the display, or a microphone.

[0015] The output unit 13 is capable of outputting data. The output unit 13 is configured to include at least one output interface capable of outputting data. The output interface is, for example, a display or a speaker. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display.

[0016] The camera 14 can image a subject and generate an imaged image. The camera 14 is, for example, a visible light camera. The camera 14 continuously images the subject at a frame rate of, for example, 15 to 30 [fps]. The camera 14 is disposed at a position where it can image a user facing the display of the output unit 13 as a subject.

[0017] The distance measurement sensor 15 can measure the distance to the subject. The distance measurement sensor 15 is disposed at a position where it can measure the distance to the user as a subject. For example, the distance measurement sensor 15 is disposed on the frame of the display of the output unit 13. The distance measurement sensor 15 generates a depth map. The depth map includes a plurality of pixels. Each pixel of the depth map has a pixel value corresponding to the distance from the distance measurement sensor 15 to the user. The distance measurement sensor 15 includes, for example, a ToF (Time Of Flight) camera, LiDAR (Light Detection And Ranging), or a stereo camera. The distance measurement sensor is also referred to as a "depth camera".

[0018] The storage unit 16 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The storage unit 16 may function as a main storage device, an auxiliary storage device, or a cache memory. The storage unit 16 stores data used for the operation of the terminal device 10 and data obtained by the operation of the terminal device 10. The storage unit 16 of the terminal device 10A stores the camera parameters of the distance measurement sensor 15 of the terminal device 10B, which will be described later.

[0019] The control unit 17 includes at least one processor, at least one dedicated circuit, or a combination of these. The processor is, for example, a general-purpose processor such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing. The control unit 17 executes processes related to the operation of the terminal device 10A while controlling each part of the terminal device 10A.

[0020] Figure 2 is a flowchart showing an operation example of the terminal device 10A shown in FIG. 1. When the virtual event starts, the control unit 17 starts the process of step S1.

[0021] In the process of step S1, the control unit 17 receives, from the terminal device 10B via the network 2, the captured image of the user 3B and the data of the first depth map 20. The control unit 17 acquires the captured image of the user 3B and the data of the first depth map 20 by receiving them. The captured image of the user 3B is generated by the camera 14 of the terminal device 10B. The first depth map 20 is generated by the distance measuring sensor 15 of the terminal device 10B. The first depth map 20 includes a plurality of first pixels. The first pixel has a pixel value corresponding to the distance from a predetermined point in the three-dimensional space to the subject. In the present embodiment, in the three-dimensional space, it is the space where the user 3B is located. In the present embodiment, the predetermined point is the position of the distance measuring sensor 15 of the terminal device 10B. That is, in the present embodiment, the first pixel has a pixel value corresponding to the distance from the distance measuring sensor 15 of the terminal device 10B as the predetermined point to the user 3B as the subject. FIG. 3 shows the first depth map 20. In FIG. 3, the difference in pixel values is shown by hatching.

[0022] In the process of step S2, the control unit 17 sets a predetermined straight line 21 as shown in FIG. 3 in the three-dimensional space. In the present embodiment, as described above, the three-dimensional space is the space where the user 3B is located. The control unit 17 may set the predetermined straight line 21 based on the line of sight of the user 3A. The control unit 17 may detect the line of sight of the user 3A by analyzing the captured image generated by the camera 14 of the terminal device 10A. When the control unit 17 repeatedly executes the processes from step S1 to step S7, the control unit 17 may detect the line of sight of the user 3A with respect to the two-dimensional model displayed on the display of the output unit 13 in the process of step S7 described later. In this case, the control unit 17 may specify the line of sight of the user 3A in the three-dimensional space where the user 3B is located, and set the specified line of sight of the user 3A as the predetermined straight line 21. With such a configuration, among the two-dimensional models of the user 3B displayed on the display of the output unit 13 in the process of step S7 described later, the portion that the user 3A is looking at can be detected as a collision point in the process of step S5 described later.

[0023] In the process of step S3, the control unit 17 generates a second depth map 22 as shown in FIG. 3 based on a predetermined straight line 21 set in the three-dimensional space. The second depth map 22 includes a plurality of second pixels. The second pixel has a pixel value corresponding to the distance from a predetermined point in the three-dimensional space to the predetermined straight line 21. In the present embodiment, as described above, the three-dimensional space is the space where the user 3B is located. The predetermined point is the position of the distance measuring sensor 15 of the terminal device 10B. That is, the second image has a pixel value corresponding to the distance from the distance measuring sensor 15 of the terminal device 10B as a predetermined point to the predetermined straight line 21. The control unit 17 may acquire the camera data of the distance measuring sensor 15 of the terminal device 10B from the storage unit 16 and generate the second depth map 22 based on the camera parameters of the distance measuring sensor 15 of the terminal device 10B. The camera parameters of the distance measuring sensor 15 of the terminal device 10B may include at least one of the internal parameters and external parameters of the distance measuring sensor 15 of the terminal device 10B. In FIG. 3, the difference in pixel values is shown by hatching. Here, the same coordinate system is used for the first depth map 20 and the second depth map 22. Further, since the first depth map 20 and the second depth map 22 correspond to the same three-dimensional space, in the first depth map 20 and the second depth map 22, the same coordinates correspond to the same position in the three-dimensional space.

[0024] In the process of step S4, the control unit 17 compares the pixel value of the first pixel and the pixel value of the second pixel with the same coordinates. The control unit 17 determines whether the difference between the pixel value of the first pixel and the pixel value of the second pixel with the same coordinates is within a threshold value. The threshold value may be set based on the variation in the pixel values corresponding to the distance from a predetermined point in the three-dimensional space for the same object. That is, when the difference between the pixel value of the first pixel and the pixel value of the second pixel is within the threshold value, the distance from the predetermined point in the three-dimensional space of the object corresponding to the first pixel and the distance from the predetermined point in the three-dimensional space of the object corresponding to the second pixel may be regarded as the same distance.

[0025] When the control unit 17 determines that the difference between the pixel value of the first pixel at the same coordinate and the pixel value of the second image exceeds the threshold (step S4: NO), it shifts the coordinate and re-executes the process of step S4. In the process of step S4 to be re-executed, the control unit 17 determines whether the difference between the pixel value of the first pixel at the shifted coordinate and the pixel value of the second image at the same coordinate is within the threshold. When repeating the process of step S4, if the control unit 17 does not determine that the difference between the pixel value of the first pixel at the same coordinate and the pixel value of the second image at the same coordinate is within the threshold for all coordinates of the first depth map 20 and the second depth map 22, it may proceed to the process of step S6.

[0026] When the control unit 17 determines that the difference between the pixel value of the first pixel at the same coordinate and the pixel value of the second image is within the threshold (step S4: YES), it detects that coordinate as the collision point between the subject and the predetermined straight line 21 (step S5). For example, in FIG. 3, for the sake of convenience of explanation, a depth map 23 in which the first depth map 20 and the second depth map are superimposed is shown. The pixel value of the first pixel and the pixel value of the second pixel at coordinate 24 are within the threshold. The control unit 17 detects coordinate 24 as the collision point between the user 3B as the subject and the predetermined straight line 21.

[0027] In the process of step S6, the control unit 17 generates a three-dimensional model of the user 3B. As an example, the control unit 17 generates a polygon model using the first depth map 20. Further, the control unit 17 generates a three-dimensional model of the user 3B by performing texture mapping using the data of the captured image of the user 3B on the polygon model.

[0028] In the process of step S7, the control unit 17 generates a 2D model by two-dimensionally converting the 3D model generated in the process of step S6. The control unit 17 causes the generated 2D model to be displayed on the display of the output unit 13. In the 2D model to be displayed on the output unit 13, the control unit 17 may reduce the resolution of portions other than the portion corresponding to the collision point detected in the process of step S5 to be lower than the set value. With such a configuration, when the control unit 17 detects, as a collision point, the portion that user 3A is viewing among the 2D models of user 3B displayed on the display of the output unit 13 in the process of step S5, the resolution of portions other than the portion that the user is viewing can be reduced. In the 2D model to be displayed on the output unit 13, the control unit 17 may increase the resolution of the portion corresponding to the collision point detected in the process of step S5 to be higher than the set value. With such a configuration, when the control unit 17 detects, as a collision point, the portion that user 3A is viewing among the 2D models of user 3B displayed on the display of the output unit 13 in the process of step S5, the resolution of the portion that the user is viewing can be increased. The set value may be a default value.

[0029] After the process of step S7, the control unit 17 returns to the process of step S1. The control unit 17 may repeatedly execute the processes from step S1 to step S7 until the virtual event ends.

[0030] As described above, in the terminal device 10A according to this embodiment, when the difference between the pixel value of the first pixel of the first depth map 20 and the pixel value of the second pixel of the second depth map 22 at the same coordinate is within the threshold value, the control unit 17 detects that coordinate as the collision point between the subject and the predetermined straight line 21. By detecting the collision point using the first depth map 20 before generating the polygon model in this way, the computational cost can be reduced compared to the case of detecting the intersection point between the polygon surface and the straight line that constitutes the polygon model.

[0031] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each component or each step, etc. can be rearranged so as not to be logically contradictory, and a plurality of components or steps, etc. can be combined into one or divided.

Description of Reference Numerals

[0032] 1: System, 2: Network, 3A, 3B: Users, 10A, 10B: Terminal Devices, 11: Communication Unit, 12: Input Unit, 13: Output Unit, 14: Camera, 15: Distance Measuring Sensor, 16: Storage Unit, 17: Control Unit, 20: First Depth Map, 21: Predetermined Straight Line, 22: Second Depth Map, 23: Depth Map, 24: Coordinates

Claims

【Claim 1】 A terminal device comprising a communication unit and a control unit, wherein the control unit: receives a first depth map from another terminal device via the communication unit, the first depth map including a plurality of first pixels, the first pixels having pixel values corresponding to distances to a subject in a three-dimensional space, the subject being a user of the other terminal device; sets a predetermined straight line based on the line of sight of the user of the terminal device; generates a second depth map based on the predetermined straight line, the second depth map including a plurality of second pixels, the second pixels having pixel values corresponding to distances to the predetermined straight line set in the three-dimensional space; and detects, as a collision point between the subject and the predetermined straight line, a coordinate when a difference between the pixel value of the first pixel and the pixel value of the second pixel at the same coordinate is within a threshold value.

Citation Information

Patent Citations

  • Device and method for deciding collision and medium where collision deciding method is recorded

    JP1999328445A

  • Collision detecting method and collision detecting apparatus

    JP2005327125A

  • Interacting with 3D Virtual Objects Using Attitude and Multiple DOF Controllers

    JP2022159417A

  • Holographic palm raycasting for targeting virtual objects

    US20210383594A1