Information processing method, information processing device, and non-volatile storage medium

The method addresses positional inaccuracies in AR by adjusting effect positions and detecting triggers, ensuring seamless integration of AR objects with real-world elements.

JP7836773B2Active Publication Date: 2026-03-27SONY SEMICON SOLUTIONS CORP
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing AR technologies struggle to maintain an accurate positional relationship between real and augmented objects, leading to viewer discomfort due to misapplication of effects.

Method used

A computer-based information processing method that extracts segments from camera footage, adjusts effect application positions based on camera orientation changes, and detects triggers using acceleration data to ensure precise alignment and initiation of effects.

Benefits of technology

Ensures accurate and smooth integration of AR effects by maintaining alignment despite camera movements, reducing jarring images and enhancing viewer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836773000001
    Figure 0007836773000001
  • Figure 0007836773000002
    Figure 0007836773000002
  • Figure 0007836773000003
    Figure 0007836773000003
Patent Text Reader

Abstract

This information processing method comprises segment extraction processing and adjustment processing. In the segment extraction processing, a segment corresponding to a label associated with an effect is extracted from video from a camera (20). In the adjustment processing, the position to which the effect is applied is adjusted in accordance with a change in the orientation of the camera (20) detected using information including acceleration data, so that the position to which the effect is applied does not shift from the extracted segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing method, an information processing apparatus, and a non-volatile storage medium.

Background Art

[0002] Techniques for applying effects to photos and videos using AR (Augmented Reality) technology are known. For example, an AR object generated by CG (Computer Graphics) is superimposed on a video of a real object captured by a camera. Thereby, a video as if in a different space is generated.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to generate a realistic video, an accurate positional relationship between a real object and an AR object is required. If the application position of the effect is shifted, it may give a sense of discomfort to the viewer.

[0005] Therefore, the present disclosure proposes an information processing method, an information processing apparatus, and a non-volatile storage medium capable of appropriately applying an effect.

Means for Solving the Problems

[0006] This disclosure provides a computer-based information processing method that extracts segments corresponding to labels associated with effects from camera footage, and adjusts the application position of the effects in accordance with changes in the camera's orientation detected using information including acceleration data, so that the application position of the effects does not deviate from the extracted segments. This disclosure also provides an information processing device that performs the information processing of the information processing method, and a non-volatile storage medium that stores a program that enables the computer to perform the information processing. [Brief explanation of the drawing]

[0007] [Figure 1] This is a diagram showing the schematic configuration of an information processing device. [Figure 2] This is a functional block diagram of an information processing device. [Figure 3] This figure shows an example of the effect. [Figure 4] This figure shows another example of effect processing. [Figure 5] This figure shows another example of effect processing. [Figure 6] This figure shows an example of a trigger detection method. [Figure 7] This figure shows an example of a trigger detection method. [Figure 8] This figure shows an example of a gesture detected by the motion detection unit. [Figure 9] This figure shows an example of information processing performed by an information processing device. [Figure 10] This figure shows an example of information processing performed by an information processing device. [Figure 11] This figure shows an example of the hardware configuration of an information processing device. [Modes for carrying out the invention]

[0008] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0009] The explanation will proceed in the following order. [1. Configuration of the Information Processing Device] [2. Information Processing Methods] [3. Hardware Configuration Example] [4. Effects]

[0010] [1. Configuration of the Information Processing Device] Figure 1 shows a schematic configuration of the information processing device 1.

[0011] The information processing device 1 is an electronic device that processes various types of information, such as photographs and videos. For example, the information processing device 1 has a display 50 on the front side and a camera 20 and a ToF (Time of Flight) sensor 30 on the rear side.

[0012] Camera 20 is, for example, a compound eye camera capable of switching between ultra-wide-angle, wide-angle, and telephoto. ToF sensor 30 is, for example, a distance image sensor that detects distance information (depth information) for each pixel. The distance measurement method can be either dToF (direct ToF) or iToF (indirect ToF). As a method for detecting distance information, only the data from ToF sensor 30 may be used, or the data from both ToF sensor 30 and camera 20 may be used for detection, or the distance information may be calculated from the data from camera 30 using AI (Artificial Intelligence) technology.

[0013] As the display 50, known displays such as LCD (Liquid Crystal Display) and OLED (Organic Light Emitting Diode) are used. The display 50 has, for example, a touch-operable screen SCR.

[0014] FIG. 1 shows a smartphone as an example of the information processing apparatus 1, but the information processing apparatus 1 is not limited to a smartphone. The information processing apparatus 1 may be a tablet terminal, a notebook computer, a desktop computer, a digital camera, or the like.

[0015] FIG. 2 is a functional block diagram of the information processing apparatus 1.

[0016] The information processing apparatus 1 includes, for example, a processing unit 10, a camera 20, a ToF sensor 30, an IMU (Inertial Measurement Unit) 40, a display 50, an effect information storage unit 60, a gesture model storage unit 70, and a program storage unit 80.

[0017] The processing unit 10 applies an effect to the video of the camera 20 based on the measurement data of the ToF sensor 30 and the IMU 40. The processing unit 10 includes, for example, a position information detection unit 11, an attitude detection unit 12, an effect processing unit 13, and a trigger detection unit 14.

[0018] The position information detection unit 11 acquires depth data measured by the ToF sensor 30. The depth data includes depth information for each pixel. The position information detection unit 11 detects distance information of a real object existing in the real space based on the depth data.

[0019] The attitude detection unit 12 acquires video data of the camera 20 and IMU data measured by the IMU 40. The video data includes data of photos and videos. The IMU data includes information regarding three-dimensional angular velocity and acceleration. The attitude detection unit 12 detects the attitude of the camera 20 using the video data and the IMU data. Attitude information regarding the attitude of the camera 20 is detected using a known method such as SLAM (Simultaneous Localization and Mapping).

[0020] In this disclosure, posture information is detected using video data and IMU data, but the method of detecting posture information is not limited to this. Posture information can be detected if acceleration data from the camera 20 is available. Therefore, the posture detection unit 12 can detect the posture of the camera 20 using information that includes at least acceleration data. Accurate posture information can be detected by fusing acceleration data with other sensor information. Therefore, in this disclosure, the posture information of the camera 20 is detected using such a sensor fusion method.

[0021] The effect processing unit 13 applies effects to the video captured by the camera 20. Various information related to the effects, such as the content of the effects and the location where the effects are applied, is stored in the effect information storage unit 60 as effect information 61. The effect processing unit 13 processes the effects based on the effect information 61.

[0022] Figure 3 shows an example of the effect.

[0023] The effect processing is performed, for example, using CG-generated AR objects AROB. In the example in Figure 3, multiple spherical AR objects AROB are displayed superimposed on the real object ROB. The effect processing unit 13 detects the positional relationship between the real object ROB and the AR objects AROB based on the detected distance information of the real object ROB. Based on the positional relationship between the real object ROB and the AR objects AROB, the effect processing unit 13 performs occlusion processing on the real object ROB using the AR objects AROB. The display 50 displays the result of the occlusion processing.

[0024] Occlusion refers to the state in which an object in the foreground hides an object behind it. Occlusion processing refers to the process of detecting the front-to-back relationship between objects and, based on the detected relationship, superimposing the objects while hiding the object behind with the object in front. For example, if the real object ROB is closer to the ToF sensor 30 than the AR object AROB, the effect processing unit 13 will perform occlusion processing by superimposing the real object ROB in front of the AR object AROB so that the AR object AROB is hidden by the real object ROB.

[0025] Figure 4 shows another example of effect processing.

[0026] In the example in Figure 4, a hole leading to another dimension is displayed as the AR object AROB. Figure 4 shows how the video CM from camera 20 shifts due to camera shake. The effect processing unit 13 adjusts the position where the effect is applied to the video CM based on the orientation of camera 20 so that no shift occurs between the video CM and the AR object AROB.

[0027] Figure 5 shows another example of effect processing.

[0028] In the example in Figure 5, the effect is selectively applied to a specific segment SG of the video commercial. For example, the video commercial from camera 20 is divided into a first segment SG1 labeled "sky," a second segment SG2 labeled "buildings," and a third segment SG3 labeled "hands." The effect is selectively applied to the first segment SG1.

[0029] The discrepancy between the video CM and the AR object AROB is easily noticeable when the effect is selectively applied to a specific segment. For example, if the effect is applied to the first segment SG1, and camera shake causes the effect to spread to the second segment SG2 or the third segment SG3, an unnatural image will be generated. Therefore, the effect processing unit 13 adjusts the application position of the effect according to the discrepancy in the video CM.

[0030] For example, as shown in Figure 2, the effect processing unit 13 includes a segment extraction unit 131 and an adjustment unit 132. The segment extraction unit 131 extracts segments SG corresponding to labels associated with effects from the video CM of the camera 20. The extraction of segments SG is performed using known methods such as semantic segmentation. The adjustment unit 132 adjusts the application position of the effect in accordance with changes in the posture of the camera 20 so that the application position of the effect does not shift from the extracted segments SG. For example, if the video CM shifts to the left due to camera shake, the application position of the effect within the screen SCR is also shifted to the left.

[0031] The trigger detection unit 14 detects a trigger to start processing the effect. The trigger can be anything. For example, the trigger detection unit 14 determines that a trigger has been detected when it detects a specific object (trigger object) or when the trigger object performs a specific action. The trigger object may be a real object ROB or an AR object AROB. The trigger is detected, for example, based on depth data. Trigger information regarding the trigger object and action is included in the effect information 61.

[0032] Figures 6 and 7 show an example of a trigger detection method.

[0033] In the example shown in Figure 6, gestures such as hands and fingers are detected as triggers. The trigger detection unit 14 detects the movement of the real object ROB, such as a hand or fingers, which is the trigger object TOB, based on depth data. The trigger detection unit 14 then determines whether the movement of the trigger object TOB corresponds to the trigger gesture.

[0034] For example, as shown in Figure 2, the trigger detection unit 14 includes a depth map generation unit 141, a joint information detection unit 142, an motion detection unit 143, and a determination unit 144.

[0035] As shown in Figure 7, the depth map generation unit 141 generates a depth map DM of the trigger object TOB. The depth map DM is an image in which a distance value (depth) is assigned to each pixel. The depth map generation unit 141 generates depth maps DM of the trigger object TOB at multiple time points using time-series depth data acquired from the ToF sensor 30.

[0036] The joint information detection unit 142 extracts joint information of the trigger object TOB at multiple time points based on the depth map DM at multiple time points. The joint information includes information about the arrangement of multiple joints JT set on the trigger object TOB.

[0037] The parts that become joints (JTs) are set for each trigger object (TOB). For example, if the trigger object (TOB) is a human hand, the center of the palm, the base of the thumb, the center of the thumb, the tip of the thumb, the center of the index finger, the tip of the index finger, the center of the middle finger, the tip of the middle finger, the center of the ring finger, the tip of the ring finger, the center of the little finger, the tip of the little finger, and the two joints of the wrist are each set as joints (JTs). The joint information detection unit 142 extracts the 3D coordinate information of each joint (JT) as joint information.

[0038] Note that the parts that can be designated as joint JTs are not limited to those listed above. For example, the tips of the thumb, index finger, middle finger, ring finger, little finger, and the two joints of the wrist may each be set as joint JTs. In addition to the 14 joint JTs mentioned above, other parts such as the base of the index finger, middle finger, ring finger, and little finger may also be set as joints. By setting a number of joint JTs close to the number of joints in the hand (21), gestures can be detected with high accuracy.

[0039] The motion detection unit 143 detects the motion of the trigger object TOB based on joint information from multiple time points. For example, the motion detection unit 143 applies joint information from multiple time points to the gesture model 71. The gesture model 71 is an analysis model that learns the relationship between time-series joint information and gestures using RNN (Recurrent Neural Network) and LSTM (Long Short-Term Memory), etc. The gesture model 71 is stored in the gesture model storage unit 70. The motion detection unit 143 analyzes the motion of the trigger object TOB using the gesture model 71 and detects the gesture corresponding to the motion of the trigger object TOB based on the analysis results.

[0040] The determination unit 144 compares the gesture corresponding to the action of the trigger object TOB with the gesture defined in the effect information 61. Based on this, the determination unit 144 determines whether the action of the trigger object TOB corresponds to the trigger gesture.

[0041] Figure 8 shows an example of a gesture detected by the motion detection unit 143.

[0042] Figure 8 shows "Air Tap," "Bloom," and "Snap Finger" as examples of gestures. "Air Tap" is a gesture of raising the index finger and tilting it straight down. "Air Tap" corresponds to a mouse click and a tap on a touch panel. "Bloom" is a gesture of opening the hand with the palm facing up. "Bloom" is used when closing an application or opening the Start menu. "Snap Finger" is a gesture of snapping the fingers by rubbing the thumb and middle finger together. "Snap Finger" is used when starting the processing of an effect.

[0043] Returning to Figure 2, the program storage unit 80 stores the program 81 that the processing unit 10 will execute. The program 81 is a program that causes a computer to execute the information processing according to this disclosure. The processing unit 10 performs various processes according to the program 81 stored in the program storage unit 80. The program storage unit 80 includes, for example, any non-temporary, non-volatile storage medium such as a semiconductor storage medium and a magnetic storage medium. The program storage unit 80 is configured to include, for example, an optical disk, a magneto-optical disk, or flash memory. The program 81 is stored, for example, in a non-temporary storage medium that can be read by a computer.

[0044] [2. Information Processing Methods] Figures 9 and 10 show an example of information processing performed by the information processing device 1. Figure 10 is a diagram showing the processing flow, and Figure 9 shows the display content for each step.

[0045] In step SA1, the processing unit 10 displays the effect selection screen ES on the display 50. For example, the effect selection screen ES displays a list of effects. The user selects the desired effect from the list of effects.

[0046] In step SA2, the processing unit 10 displays an explanation of how to operate the effect on the display 50. In the example in Figure 9, for example, it is explained that the effect processing is started by a "Snap Finger" gesture, the effect scene is switched by performing the "Snap Finger" gesture again, and the process of making the effect disappear is started by performing a gesture of shaking the hand from side to side.

[0047] In step SA3, the user performs a "Snap Finger" gesture in front of camera 20. When processing unit 10 detects the "Snap Finger" gesture, it starts processing the effect in step SA4. In the example in Figure 9, pink clouds rise from behind the building and gradually take on a horse-like shape.

[0048] In step SA5, the user performs the "Snap Finger" gesture again in front of the camera 20. When the processing unit 10 detects the "Snap Finger" gesture, it switches the effect scene in step SA6. In the example in Figure 9, the sky is suddenly covered with a pink filter, and a glowing unicorn made of clouds appears in the sky.

[0049] In step SA7, the user makes a gesture of shaking their hands from side to side in front of camera 0. When processing unit 10 detects the hand-shaking gesture, in step SA8, it starts the process of removing the effect. In the example in Figure 9, the unicorn smiles gently at the user, then lowers its head and flies away, its long eyelashes fluttering. Hearts spread around the unicorn as it flies away. After the unicorn flies away, the system returns to the normal state before the effect processing began.

[0050] In step SA9, the processing unit 10 determines when the effect has finished. If an termination flag is detected, such as pressing the effect termination button, it is determined that the effect has finished. If it is determined in step SA9 that the effect has finished (step SA9:Yes), the processing unit 10 terminates the effect processing. If it is not determined in step SA9 that the effect has finished (step SA9:No), the process returns to step SA3 and the above process is repeated until an termination flag is detected.

[0051] [3. Hardware Configuration Example] Figure 11 shows an example of the hardware configuration of the information processing device 1.

[0052] The information processing device 1 includes a CPU (Central Processing Unit) 1001, a ROM (Read Only Memory) 1002, a RAM (Random Access Memory) 1003, an internal bus 1004, an interface 1005, an input device 1006, an output device 1007, a storage device 1008, a sensor device 1009, and a communication device 1010.

[0053] The CPU 1001 is configured as an example of the processing unit 10. The CPU 1001 functions as an arithmetic processing unit and a control unit, and controls the overall operation of the information processing unit 1 according to various programs. The CPU 1001 may be a microprocessor.

[0054] ROM 1002 stores programs and calculation parameters used by CPU 1001. RAM 1003 temporarily stores programs used in the execution of CPU 1001 and parameters that change as needed during its execution. CPU 1001, ROM 1002, and RAM 1003 are interconnected by an internal bus 1004, which consists of a CPU bus and other components.

[0055] Interface 1005 connects the input device 1006, output device 1007, storage device 1008, sensor device 1009, and communication device 1010 to the internal bus 1004. For example, the input device 1006 exchanges data with the CPU 1001 and other devices via interface 1005 and the internal bus 1004.

[0056] The input device 1006 includes input means for the user to input information, such as a touch panel, buttons, a microphone, and switches, and an input control circuit that generates an input signal based on the user's input and outputs it to the CPU 1001. By operating the input device 1006, the user can input various types of data to the information processing device 1 or instruct it to perform processing operations.

[0057] The output device 1007 includes a display 50 and audio output devices such as speakers and headphones. For example, the display 50 displays images captured by the camera 20 or images generated by the processing unit 10. The audio output devices convert audio data and the like into audio and output it.

[0058] The storage device 1008 includes an effect information storage unit 60, a gesture model storage unit 70, and a program storage unit 80. The storage device 1008 also includes a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deletion device for deleting data recorded on the storage medium. The storage device 1008 stores the program 81 executed by the CPU 1001 and various data.

[0059] The sensor device 1009 includes, for example, a camera 20, a ToF sensor 30, and an IMU 40. The sensor device 1009 may also include a GPS (Global Positioning System) receiving function, a clock function, an accelerometer, a gyroscope, a barometric pressure sensor, and a geomagnetic sensor.

[0060] The communication device 1010 is a communication interface composed of communication devices, etc., for connecting to the communication network NT. The communication device 1010 may be a wireless LAN-compatible communication device or an LTE (Long Term Evolution)-compatible communication device.

[0061] [4. Effects] The information processing device 1 includes a posture detection unit 12, a segment extraction unit 131, and an adjustment unit 132. The posture detection unit 12 detects the posture of the camera 20 using information including acceleration data. The segment extraction unit 131 extracts segments SG corresponding to labels associated with effects from the video CM of the camera 20. The adjustment unit 132 adjusts the application position of the effect in accordance with changes in the posture of the camera 20 so that the application position of the effect does not deviate from the extracted segments SG. In this embodiment, the information processing method of the information processing device 1 described above is executed by a computer. The non-volatile storage medium (program storage unit 80) of this embodiment stores a program 81 that causes the computer to implement the processing of the information processing device 1 described above.

[0062] With this configuration, the effect will be applied to the correct position even if the camera 20's orientation changes.

[0063] The information processing device 1 has a trigger detection unit 14. The trigger detection unit 14 detects a trigger to start processing an effect based on depth data.

[0064] With this configuration, the trigger is detected with high accuracy.

[0065] The trigger detection unit 14 detects the gesture of the trigger object TOB as a trigger.

[0066] With this configuration, effects can be initiated using gestures.

[0067] The trigger detection unit 14 includes a depth map generation unit 141, a joint information detection unit 142, a motion detection unit 143, and a determination unit 144. The depth map generation unit 141 generates depth maps DM of the trigger object TOB at multiple time points using depth data. The joint information detection unit 142 detects joint information of the trigger object TOB at multiple time points based on the depth maps DM at multiple time points. The motion detection unit 143 detects the motion of the trigger object TOB at multiple time points based on the joint information at multiple time points. The determination unit 144 determines whether the motion of the trigger object TOB corresponds to a trigger gesture.

[0068] With this configuration, gestures can be detected with high accuracy.

[0069] The information processing device 1 includes a position information detection unit 11 and an effect processing unit 13. The position information detection unit 11 detects distance information of the real object ROB based on depth data acquired by the ToF sensor 30. The effect processing unit 13 performs occlusion processing on the real object ROB and the AR object AROB, which is generated by CG for effects, based on the distance information of the real object ROB.

[0070] With this configuration, the positional relationship between the real object ROB and the AR object AROB is accurately detected based on depth data. Because occlusion processing can be performed appropriately, the viewer is provided with images that are less jarring.

[0071] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0072] [Note] Furthermore, this technology can also be configured as follows. (1) Extract segments from the camera footage that correspond to labels associated with effects, To ensure that the application position of the effect does not deviate from the extracted segment, the application position of the effect is adjusted according to the change in the camera's posture detected using information including acceleration data. A method of information processing performed by a computer. (2) A posture detection unit that detects the camera's posture using information including acceleration data, A segment extraction unit extracts segments corresponding to labels associated with effects from the video footage of the aforementioned camera, An adjustment unit adjusts the application position of the effect in accordance with changes in the camera's posture so that the application position of the effect does not shift from the extracted segment. An information processing device having (3) The system includes a trigger detection unit that detects a trigger to initiate the processing of the effect based on depth data. The information processing device described in (2) above. (4) The aforementioned depth data is acquired by a ToF sensor. The information processing device described in (3) above. (5) The trigger detection unit detects the gesture of the trigger object as the trigger. The information processing device described in (3) or (4) above. (6) The trigger detection unit, A depth map generation unit generates depth maps of the trigger object at multiple time points using the aforementioned depth data, A joint information detection unit detects joint information of the trigger object at multiple time points based on the depth maps at multiple time points, Based on the joint information of the multiple time points, an motion detection unit detects the movement of the trigger object, A determination unit that determines whether the action of the trigger object corresponds to the gesture that acts as the trigger, The information processing apparatus described in (5) above, having the following: (7) It has a position information detection unit that detects distance information of a real object based on depth data acquired by a ToF sensor, An effect processing unit performs occlusion processing on the real object and the AR object for the effect generated by CG, based on the distance information of the real object. An information processing apparatus according to any one of (2) to (6) above, having the following: (8) Extract segments from the camera footage that correspond to labels associated with effects, To ensure that the application position of the effect does not deviate from the extracted segment, the application position of the effect is adjusted according to the change in the camera's posture detected using information including acceleration data. A non-volatile storage medium that stores programs that enable a computer to perform a certain action. [Explanation of Symbols]

[0073] 1. Information Processing Device 11 Location Information Detection Unit 12 Attitude detection unit 13. Effects Processing Unit 131 Segment extraction unit 132 Adjustment section 14 Trigger detection unit 141 Depth Map Generation Unit 142 Joint Information Detection Unit 143 Motion detection unit 144 Judgment section 20 cameras 30 ToF sensors 80 Program memory unit (non-volatile storage medium) 81 Programs AROB ARobject CM video DM Depth Map ROB (Real Object) SG segment TOB Trigger Object

Claims

1. Extract segments from the camera footage that correspond to labels associated with effects, To prevent the application position of the effect from shifting from the extracted segment due to camera shake, the application position of the effect is adjusted according to the change in the camera's posture detected using information including acceleration data. Based on depth data, the gesture of the trigger object is detected as a trigger to initiate the processing of the effect. This is done by a computer. The process for detecting the aforementioned trigger is: Using the aforementioned depth data, depth maps of the trigger object at multiple time points are generated. Based on the depth maps at multiple time points, joint information of the trigger object at multiple time points is detected. Based on the joint information at multiple time points, the movement of the trigger object is detected. The system includes determining whether the action of the trigger object corresponds to the gesture that acts as the trigger. The joint information includes information regarding the arrangement of multiple joints set on the trigger object. The process for detecting the joint information involves extracting the three-dimensional coordinate information of each joint as the joint information. Information processing methods.

2. A posture detection unit that detects the camera's posture using information including acceleration data, A segment extraction unit extracts segments corresponding to labels associated with effects from the video footage of the aforementioned camera, An adjustment unit adjusts the application position of the effect in accordance with changes in the camera's posture so that the application position of the effect does not shift from the extracted segment due to camera shake. A trigger detection unit detects the gesture of a trigger object based on depth data as a trigger to start processing the effect, It has, The trigger detection unit, A depth map generation unit generates depth maps of the trigger object at multiple time points using the aforementioned depth data, A joint information detection unit detects joint information of the trigger object at multiple time points based on the depth maps at multiple time points, Based on the joint information of the multiple time points, an motion detection unit detects the movement of the trigger object, A determination unit that determines whether the action of the trigger object corresponds to the gesture that acts as the trigger, It has, The joint information includes information regarding the arrangement of multiple joints set on the trigger object. The joint information detection unit extracts the three-dimensional coordinate information of each joint as the joint information. Information processing device.

3. The aforementioned depth data is acquired by a ToF sensor. The information processing apparatus according to claim 2.

4. A position information detection unit that detects distance information of a real object based on depth data acquired by a ToF sensor, An effect processing unit performs occlusion processing on the real object and the CG-generated AR object for the effect based on the distance information of the real object. The information processing apparatus according to claim 2, having the following:

5. Extract segments from the camera footage that correspond to labels associated with effects, To prevent the application position of the effect from shifting from the extracted segment due to camera shake, the application position of the effect is adjusted according to the change in the camera's posture detected using information including acceleration data. Based on depth data, the gesture of the trigger object is detected as a trigger to initiate the processing of the effect. Store a program in the computer that makes this possible. The process for detecting the aforementioned trigger is: Using the aforementioned depth data, depth maps of the trigger object at multiple time points are generated. Based on the depth maps at multiple time points, joint information of the trigger object at multiple time points is detected. Based on the joint information at multiple time points, the movement of the trigger object is detected. The system includes determining whether the action of the trigger object corresponds to the gesture that acts as the trigger. The joint information includes information regarding the arrangement of multiple joints set on the trigger object. The process for detecting the joint information involves extracting the three-dimensional coordinate information of each joint as the joint information. Non-volatile storage medium.

Citation Information

Patent Citations

  • Method and apparatus for representing a physical scene

    JP2016534461A

  • Machine learning gesture detection

    US20130077820A1

  • Information processing apparatus and information processing method, display apparatus and display method, and information processing system

    US20160093107A1

  • Six degree of freedom tracking with scale recovery and obstacle avoidance

    US20180348854A1

  • Interactive augmented reality

    US20200050259A1