Data processing apparatus and program
The data processing device addresses unnatural composite images by using multiple cameras and depth sensors to synchronize virtual objects with user movements and sounds, enhancing the composite image experience through accurate and interactive display.
Patent Information
- Application Number
- JP2025211509
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-02-16
AI Technical Summary
Conventional techniques for displaying composite images with virtual objects superimposed on real-space images often appear unnatural.
A data processing device equipped with multiple cameras and depth sensors captures images and movement information to generate composite images with a virtual object that changes based on user movement and sound, allowing for synchronized display on a device screen.
The solution creates a more natural and engaging composite image experience by accurately reflecting user movements and sounds, enabling real-time distribution and interaction.
Smart Images

Figure 2026026260000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a data processing device and a program. [Background technology]
[0002] There is known a technique for displaying a composite image in which a virtual object is superimposed on an image of real space (see, for example, Patent Document 1). In the conventional technique, the composite image may look unnatural. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-184689 Summary of the Invention
[0004] According to one aspect of the present invention, a first acquisition unit for acquiring image information of a first subject; A data processing device is provided, comprising: a second acquisition unit for acquiring movement information relating to the movement of a second subject different from the first subject; a third acquisition unit for acquiring sound information relating to sound; and a processing unit for displaying on a display surface a first image based on the image information, a virtual object that changes based on the movement information, and a second image based on the sound information.
[0005] According to another aspect of the present invention, there is provided a program that causes a data processing device to acquire image information of a first subject, acquire movement information relating to movement of a second subject different from the first subject, acquire sound information relating to sound, and display on a display surface a first image based on the image information, a virtual object that changes based on the movement information, and a second image based on the sound information. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a diagram showing the overall configuration of a distribution system. [Figure 2]FIG. 2 is a block diagram of a user terminal device. [Figure 3] Block diagram of the server. [Figure 4] 10 is a flowchart of a distribution preparation process. [Figure 5] 10 is a flowchart of a distribution process when the user is stationary. [Figure 6] (a) is a diagram defining the orientation of an avatar object, (b) is a transition diagram when an avatar object is rotated 90 degrees to the right from a state facing away, and (c) is a transition diagram when an avatar object is rotated 270 degrees to the left from a state facing away. [Figure 7] 10A and 10B are diagrams showing changes in acceleration when a user moves from the front to the back. [Figure 8] 10 is a flowchart of a distribution process when a user moves by walking or running. [Figure 9] 10A is a diagram showing an example of adding an effect object in addition to an avatar object when a specific action is detected, and FIG. 10B is a diagram showing an example of adding an effect object. [Figure 10] FIG. 10 is a block diagram of a user terminal device having depth sensors on both the outer and inner surfaces. [Figure 11] (a) is a schematic diagram when a 360-degree camera is used, (b) is a schematic diagram when two hemispherical cameras are used, (c) is a schematic diagram when one hemispherical camera is used, and (d) is a schematic diagram when the image receiving element of the outer and inner cameras is shared and used as a single camera. DETAILED DESCRIPTION OF THE INVENTION
[0007] A distribution system according to one embodiment of the present invention will be described below with reference to FIGS. [Outline of the distribution system] 1, the distribution system 1 includes a user terminal device 10, a server 30, and a viewer terminal device 50. There may be multiple user terminal devices 10 and multiple viewer terminal devices 50, each of which can be connected to the server 30.
[0008] The user terminal device 10 is a data processing device equipped with multiple cameras. Examples of the user terminal device 10 include a smartphone and a tablet. A user is a person who operates the user terminal device 10, a photographer who uses the camera of the user terminal device 10, and a distributor who distributes images using the user terminal device 10.
[0009] The user terminal device 10 has an outer surface facing away from the user holding the user terminal device 10 and an inner surface facing the user. The user terminal device 10 has an outer camera 11 on the outer surface and an inner camera 13 on the inner surface. The outer camera 11 captures an image of a subject (first subject) such as a scene in the user's line of sight, which is the first direction, for example. Therefore, the subject captured by the outer camera 11 does not include the user.
[0010] On the other hand, as an example, the in-camera 13 captures an image of a subject in a second direction, which is the opposite direction to the first direction. That is, the in-camera 13 captures an image of a subject (second subject) in the direction of the user. Therefore, the subject captured by the in-camera 13 includes the user. The user terminal device 10 can be attached with a support member 3 having an operation unit 22. By attaching the support member 3 to the user terminal device 10, the out-camera 11 can capture an image of a subject in the first direction from a position farther away from the user. Furthermore, the in-camera 13 can capture an image of a subject in the second direction from a position farther away from the user.
[0011] At least one of the outer camera 11 and the inner camera 13 may be an external camera connected to the user terminal device 10 by wire or wirelessly. The user terminal device 10 includes a depth sensor 15. The depth sensor 15 is disposed on the inner surface where the in-camera 13 is disposed. The depth sensor 15 is disposed near the in-camera 13. As an example, the depth sensor 15, like the in-camera 13, has a detection range in the direction of the user, which is the second direction, and detects the distance between the depth sensor 15 and the user. As an example, the depth sensor 15 detects the user's head, face, torso, hands, arms, legs, etc. As an example, the depth sensor 15 detects the user's movement within the detection range, and generates and outputs motion information based on the user's movement. The depth sensor 15 generates motion information at the same time as the image 101 captured by the out-camera 11. That is, the motion information is synchronized with the image 101, for example.
[0012] The user terminal device 10 includes a sensor 17. Examples of the sensor 17 include an acceleration sensor and a gyro sensor. The user terminal device 10 detects the direction of movement and tilt of the user terminal device 10 using the sensor 17 as a position information detection unit. This allows the user terminal device 10 to calculate position information for the user terminal device 10. The user terminal device 10 also functions as a pedometer by using the acceleration, angle, and angular velocity data detected by the sensor 17. The pedometer can be realized by installing an application program in the user terminal device 10. The user terminal device 10 has a GNSS function and can identify its current location.
[0013] The user terminal device 10 includes a display unit 19 having a display surface on which an image is displayed, and a touch panel 20 provided on the display surface. The display unit 19 and the touch panel 20 are disposed on the inner surface of the user terminal device 10 where the in-camera 13 is disposed. The display surface of the display unit 19 displays an image of a subject captured by the out-camera 11 and the in-camera 13. As an example, the image of the subject captured by the out-camera 11 and a virtual object based on movement information of the subject captured by the in-camera 13 are displayed. Specifically, a composite image is displayed in which a virtual object based on movement information of the user detected by the in-camera 13 is superimposed on the image of the subject captured by the out-camera 11. The virtual object is, for example, an avatar object 102. On the display surface, the display position of the avatar object 102 corresponds to the position where the user is captured by the in-camera 13. The display position of the avatar object 102 also corresponds to the position of the user relative to the depth sensor 15. The display position of avatar object 102 is a position designated on touch panel 20 by touching a predetermined position on the display surface, and avatar object 102 is displayed at the designated position on the display surface.
[0014] The avatar object 102 is an incarnation of the user operating the user terminal device 10. Examples of the avatar object 102 include a person object imitating a famous person, a famous person, an ordinary person, etc., an animal object imitating an animal, a character object appearing in an animation, and an original object created by the user.
[0015] The avatar object 102 changes according to the user's movement information detected by the depth sensor 15. The avatar object 102 also changes according to the detection result of the sensor 17. For example, when the user detected by the depth sensor 15 moves their hand, the avatar object 102 moves in the same manner according to the hand movement information of the subject detected by the depth sensor 15. For example, when the user is walking, the avatar object 102 also moves as if walking according to the detection result of the sensor 17, and when the user is running, the avatar object 102 also moves as if running. In other words, when the user is moving from place to place by walking or running, the movement direction and movement speed are detected according to the sensor 17 and the GNSS in addition to the user's movement information detected by the depth sensor 15, and the avatar object 102 moves as if walking or running.
[0016] The user terminal device 10 is a communication device equipped with a communication function for communicating with the server 30. The user can distribute composite data 100 of a composite image in which an avatar object 102 is superimposed as a virtual object on an image 101 of the user's current location, such as a tourist spot, famous place, historical site, or landscape, captured by the outer camera 11. The composite data 100 generated by the user terminal device 10 is data for distribution that is transmitted to the server 30 and distributed to the viewer terminal device 50. When the user terminal device 10 transmits such composite data 100 to the server 30, the server 30 live-distributes the composite data 100. The viewer terminal device 50 can receive and view the composite data 100 distributed by the server 30 in real time. The server 30 also has a program guide and can distribute programs at specified time slots.
[0017] The viewer terminal device 50 is a data processing device such as a smartphone or tablet. The viewer terminal device 50 has at least a communication function capable of communicating with the server 30 and a display unit having a display surface on which an image is displayed. When the server 30 is live streaming the composite data 100, the viewer terminal device 50 displays a composite image of the composite data 100 on the display surface, allowing the viewer to view the image in real time. The viewer terminal device 50 can also access the server 30 and check the program guide to view an image of the composite data 100 at the distribution time. For example, during live streaming, the viewer terminal device 50 can send request comments to the user terminal device 10 via the server 30. This allows the user and the viewer to converse with each other. For example, the user can capture an image in accordance with the request comment from the viewer.
[0018] [User terminal device 10] As shown in FIG. 2, the user terminal device 10 includes an outer camera 11, an outer camera interface (hereinafter, "interface" will be simply referred to as "IF") 12, an inner camera 13, an inner camera IF 14, a depth sensor 15, a depth sensor IF 16, a sensor 17, and a sensor IF 18. The user terminal device 10 also includes a display unit 19, a touch panel 20, and a communication IF 21. An operation unit 22 of the support member 3 is connected to the communication IF 21. The user terminal device 10 also includes a GNSS receiver 23, an audio input device 24, an audio output device 25, and an audio IF 26. The user terminal device 10 also includes a network IF 27, a storage unit 28, and a processing unit 29.
[0019] The outer camera 11 and the inner camera 13 are equipped with imaging elements such as a CCD sensor or a CMOS sensor. The outer camera 11 is disposed on the outer surface opposite the user holding the user terminal device 10. The inner camera 13 is disposed on the inner surface facing the user holding the user terminal device 10. A display unit 19 is disposed on the inner surface, and the inner camera 13 is disposed around the display unit 19.
[0020] Outer camera 11 as a first detection unit is connected to outer camera IF 12 as a first acquisition unit, and outputs imaging data to outer camera IF 12. Outer camera IF 12 outputs imaging data of a subject imaged by outer camera 11 to processing unit 29. Inner camera 13 is connected to inner camera IF 14, and outputs imaging data to inner camera IF 14. Inner camera IF 14 outputs imaging data of a subject imaged by inner camera 13 to processing unit 29.
[0021] Out-camera 11 captures an image of a subject such as a scene in the user's line of sight, which is a first direction. Out-camera 11 captures an image of a subject that does not include the user. In-camera 13 captures an image in the user's direction, which is a second direction opposite to the user's line of sight and different from the first direction, allowing for self-portraits. The subject captured by in-camera 13 includes the user, but the subject captured by out-camera 11 does not include the user.
[0022] Outer camera 11 has a first resolution, and in-camera 13 has a second resolution. Outer camera 11 and in-camera 13 may have the same resolution, or one of the cameras may have a higher resolution than the other. As an example, outer camera 11, which often captures scenery or the like as a subject, has a higher resolution than in-camera 13, which often captures a user as a subject.
[0023] Outer camera 11 has a first angle of view, and in-camera 13 has a second angle of view. Outer camera 11 and in-camera 13 may have the same angle of view, or one of the cameras may have a wider angle than the other camera.
[0024] The outer camera 11 and the inner camera 13 are equipped with zoom lenses, and the focal length can be changed by changing the zoom lens in the optical axis direction, thereby enlarging or reducing the captured image. The outer camera 11 and the inner camera 13 may also be equipped with a digital zoom function.
[0025] The outer camera 11 is an imaging unit that captures an image of a scene in the user's line of sight as a subject. A plurality of outer cameras 11 may be provided. As an example, when two outer cameras 11 are provided, two images (an image for the right eye and an image for the left eye) for stereoscopic display with parallax can be generated. A plurality of inner cameras 13 may also be provided. In such a case, when the inner camera 13 is used, two images for stereoscopic display with parallax can be generated.
[0026] Depth sensor 15 is a detection unit and is disposed near in-camera 13. The detection range of depth sensor 15 is approximately the same as the imaging range of in-camera 13. Note that the detection range of depth sensor 15 may be wider or narrower than the imaging range of in-camera 13.
[0027] The depth sensor 15 is, for example, an infrared camera. The infrared camera includes a light-projecting unit that projects infrared rays and an infrared detection unit that detects the infrared rays. The depth sensor 15 acquires depth information, such as three-dimensional position information in real space, from the time it takes for an infrared pulse projected from the light-projecting unit to be reflected and returned. The depth sensor 15 acquires depth information, such as the distance from the depth sensor 15 to the subject. For example, the depth sensor 15 acquires depth information, such as the distance between the depth sensor 15 and each part of the user (head, upper body, arms, hands, lower body, feet, etc.). The depth sensor 15 then generates movement information according to changes in the depth information of each part. The depth sensor 15 is connected to a depth sensor IF 16 serving as a second acquisition unit and outputs the movement information to the depth sensor IF 16. The depth sensor IF 16 outputs the movement information of the subject captured by the depth sensor 15 to a processing unit 29. The depth sensor 15 can obtain depth information of various parts of the user's body without the user having to attach motion sensors to various parts of the body.
[0028] Specifically, the depth sensor 15 extracts the user's human region and divides it into a human region and a non-human region. For example, it extracts regions based on the amount of infrared light and determines regions with shapes that can be recognized as human bodies as human regions. For example, the depth sensor 15 acquires 25 skeletal positions of each user as skeletal data and further calculates depth information for each skeletal position. Examples of skeletal positions include the left and right hands, head, neck, left and right shoulders, left and right elbows, left and right knees, and left and right feet. The number of acquired skeletal positions is preferably 25, but is not limited to this. If the human region falls outside the detection range of the depth sensor 15, the number of skeletal positions will be fewer than 25. Increasing the number of skeletal positions beyond 25 allows for more accurate reproduction of movements.
[0029] Here, the depth information is, for example, the distance from the objective lens or sensor surface in front of the depth sensor 15 to the user's skeletal position. The skeletal position is determined by acquiring depth information for each part in the person area, identifying the real-world parts of the person in the person area (such as the left and right hands, head, neck, left and right shoulders, left and right elbows, left and right knees, and left and right feet) based on the depth and shape features, and setting the skeletal position at the center of each part. Furthermore, the skeletal position is determined by comparing the features determined from the person area with the features of each part registered in the feature dictionary stored in the storage unit, thereby identifying each part in the person area. The depth sensor 15 also detects the person's contours (such as the shape of the palm, the contours of clothing wrinkles, the contours of the hairstyle, etc.). This enables the depth sensor 15 to track each part and generate movement information according to changes in the depth information for each part (the skeleton and the movement of each skeleton = movement information). The motion capture process described above is performed without attaching markers to the user, making it easy to acquire movement information.
[0030] Note that depth information may be calculated by reading a projected infrared pattern and obtaining depth information from the distortion of the pattern (light coding method). Depth information may also be calculated from parallax information obtained by a twin-lens camera or multiple cameras. Depth information can also be calculated by performing image recognition and image analysis on the video captured by the in-camera 13. In this case, the in-camera 13 functions as a detection unit and the in-camera IF 14 functions as a second acquisition unit, eliminating the need for the depth sensor 15 and the depth sensor IF 16. When detecting the skeletal position from the image of the subject captured by the in-camera 13, the skeletal position can be detected by using AI that utilizes deep learning. The detection and recognition process for the user's motion (the movement of the skeleton and each skeleton = action) in the motion capture process can also be performed by subjecting the image of the subject captured by the in-camera 13 to recognition processing based on learning results such as deep learning.
[0031] The sensor 17 is, for example, an acceleration sensor or a gyro sensor, and serves as an acceleration detection unit. The sensor 17 also functions as a step detection unit in conjunction with a pedometer program. The sensor 17 generates detection data, such as acceleration data, angle data, and angular velocity data, of the user terminal device 10 according to the outputs from various sensor elements. The sensor 17 is connected to the sensor IF 18 and outputs the detection data to the sensor IF 18. The sensor IF 18 outputs the detection data detected by the sensor 17 to the processing unit 29. The user terminal device 10 is carried and used by the user. Therefore, the sensor 17 calculates detection data that serves as movement information of the user terminal device 10 as the user moves, such as walking or running, and outputs the data to the processing unit 29. Furthermore, even when the user is stationary, the acceleration data, angle data, angular velocity data, etc. of the user terminal device 10 are detected, and detection data that serves as posture information of the user terminal device 10 (information such as whether the user terminal device 10 is facing upward, downward, rightward or leftward, and whether it has moved up and down, forward and backward, or left and right) is generated and output to the processing unit 29.
[0032] The display unit 19 is disposed on the inner surface of the user terminal device 10, located on the side of the user holding the user terminal device 10. The display unit 19 is disposed on most of the inner surface. The display unit 19 is, for example, a flat display panel such as a liquid crystal display panel or an organic EL panel. The display unit 19 has a display surface. The display surface is provided with a touch panel 20 as a position input unit for inputting a position on the display surface. On the touch panel 20, an input position may be specified by the user's finger or a stylus pen. The touch panel 20 may be of a capacitance type, a resistive film type, a surface acoustic wave type, an infrared type, an electromagnetic induction type, or the like. The capacitance type may be a surface type or a projected type. The projected type allows multiple inputs using multi-touch. For example, the size of the avatar object 102 can be adjusted by opening and closing two fingers (e.g., a thumb and an index finger).
[0033] The display surface displays an image or video of a subject captured by outer camera 11 or inner camera 13. An avatar object 102 is displayed superimposed on image 101 of the subject captured by the outer camera. The position where avatar object 102 is displayed can be specified using touch panel 20. Avatar object 102 moves on the display surface in accordance with information about the user's movement detected by depth sensor 15.
[0034] The communication IF 21 is a wired or wireless communication unit that communicates with other devices, and examples thereof include USBIF and Bluetooth (registered trademark). A support member 3 can be attached to the user terminal device 10. The support member 3 is provided with an operation unit 22. The operation unit 22 is connected to the user terminal device 10 via the communication IF 21 and enables operation input. The operation unit 22 is composed of, for example, a plurality of push buttons, a slide switch, a rotary switch, a cross key, a joystick, a trackball, and the like. For example, when capturing an image of a subject using the outer camera 11 and / or the inner camera 13, operations such as start recording, stop recording, pause, start playback, stop playback, start distribution, stop distribution, and zoom can be performed from the operation unit 22.
[0035] It should be noted that other external devices can also be connected to the communication IF 21. For example, a camera that replaces the outer camera 11, the inner camera 13, or the depth sensor 15 is connected. In this case, the communication IF 21 functions as an acquisition unit such as a first acquisition unit that acquires the image 101 and a second acquisition unit that acquires movement information.
[0036] The GNSS receiver 23 detects the current position of the user terminal device 10 by receiving signals from a GNSS (Global Navigation Satellite System). One example of a GNSS is the Global Positioning System (GPS). Such a positioning system makes it possible to identify the current position of the user terminal device 10 and to ascertain its movement history. Furthermore, it is possible to calculate the speed at which the user is moving.
[0037] The audio input device 24 is an acoustic-electrical conversion element such as a microphone, and is connected to the audio IF 26 as an acoustic acquisition unit. The audio input device 24 may be a device built into the user terminal device 10, or may be an external device. If the audio input device 24 is an external device, it is connected to the audio IF 26 via a wired or wireless connection. When an image of a subject is being captured using the outer camera 11 and / or the inner camera 13, the audio input device 24 simultaneously collects sound at the location where the image is captured.
[0038] The audio output device 25 is an electroacoustic conversion device such as a speaker, earphone, or headphone, and is connected to the audio IF 26. The speaker of the audio output device 25 may be a device built into the user terminal device 10 or may be an external device. If it is an external device, it is connected to the audio IF 26 via a wired or wireless connection. The audio output device 25 emits an acoustic signal in synchronization with the video being played. It also emits an acoustic signal based on the voice uttered by the user during image capture and the surrounding environmental sounds.
[0039] The network IF 27 is a communication unit that communicates with the server 30 via a network 2 such as a wide area network or a local area network. The network IF 27 communicates with the server 30 via a mobile communication network such as a wireless LAN (Wi-Fi) using the IEEE802.11 standard, 3G, 4G / LTE, or 5G. The user terminal device 10 transmits the composite data 100 to the server 30, and the composite data 100 can be distributed via the server 30 live or at a predetermined time.
[0040] The storage unit 28 is a storage medium such as a nonvolatile memory, and examples thereof include a hard disk and flash memory. The storage unit 28 may be configured as a storage medium built into the user terminal device 10, or as a removable storage medium that is detachable from the user terminal device 10. The processing unit 29 is configured with a CPU, RAM, ROM, etc., and controls the overall operation of the user terminal device 10 according to a program.
[0041] A distribution program 28a that distributes an image of a subject captured by the outer camera 11 is installed in the storage unit 28. The storage unit 28 also stores images and videos of the subject captured by the outer camera 11 and the inner camera 13. The storage unit 28 also stores three-dimensional data of an avatar object 102 to be superimposed on the image 101 in association with the distribution program 28a. The avatar object 102 has nodes set that correspond to the skeleton positions detected by the depth sensor 15.
[0042] The distribution program 28a has a function of acquiring user movement information, detection data from the sensor 17, GNSS location information, and sound at the location where the image 101 is captured, using the depth sensor 15, simultaneously with the image 101 captured by the outer camera 11. The distribution program 28a also has a function of generating composite data 100 that overlays an avatar object 102 on the image 101 captured by the outer camera 11 and changes the avatar object 102 based on information and data such as the user's movement information. The distribution program 28a also has a function of transmitting the composite data 100 to the server 30 while displaying the composite data 100 on the display surface of the display unit 19. The distribution program 28a also has a function of viewing the composite data 100 on the viewer terminal device 50. The processing unit 29 executes the functions of the distribution program 28a in accordance with a user's operation.
[0043] The distribution program 28a installed in the user terminal device 10 uses the functional portion for distribution, and the distribution program 28a installed in the viewer terminal device 50 uses the functional portion for viewing and the functional portion for sending comments to the user terminal device 10.
[0044] Furthermore, a pedometer program 28b is installed in the storage unit 28. The pedometer program 28b has thresholds, such as a first threshold for determining whether the user is standing still or walking and a second threshold for determining whether the user is walking or running, at least with respect to acceleration. The pedometer program 28b can determine whether the user is walking or standing still using the first threshold and can determine whether the user is walking or running using the second threshold. When the pedometer program 28b is started, the processing unit 29 detects up-down and forward-backward movements based on the acceleration, angle, and angular velocity detected by the sensor 17, and counts the number of steps based on the results. Note that the pedometer may be external to the user terminal device 10. The angle and angular velocity may also have first and second thresholds.
[0045] As an example, the distribution program 28a and the pedometer program 28b are linked together. For example, the pedometer program 28b runs in the background of the distribution program 28a. By linking the distribution program 28a and the pedometer program 28b, the movement of the user and the movement of the user terminal device 10 can be detected separately. As an example, when a user holding the user terminal device 10 is moving the user terminal device 10 while standing still and moving it back and forth and left and right, the pedometer program 28b does not count the number of steps even if the sensor 17 detects acceleration, etc. In contrast, when the user holding the user terminal device 10 is walking or running, the sensor 17 detects acceleration, etc., and the pedometer program 28b counts the number of steps when the acceleration exceeds a first threshold. The processing unit 29 performs processing in accordance with the distribution program 28a to overlay an avatar object 102 that changes in accordance with the user's movement information on an image 101 captured by the outer camera 11, thereby generating composite data 100. At this time, when the processing unit 29 detects that the user is walking or running according to the pedometer program 28b, it generates the synthetic data 100 in which the avatar object 102 changes to appear as if it is walking or running.
[0046] Furthermore, a voice recognition program 28c is installed in the storage unit 28. The voice recognition program 28c performs voice recognition processing on the user's voice data collected by the audio input device 24 to generate text data. As an example, the voice recognition program 28c identifies phonemes from a voice waveform, matches the sequence of phonemes with a pre-registered dictionary, converts the phonemes into words, and generates text data of the converted sentence. As an example, the voice recognition program 28c also works in conjunction with the distribution program 28a. While the processing unit 29 performs processing in accordance with the distribution program 28a to overlay an avatar object 102 on an image 101 captured by the outer camera 11 and generate composite data 100, the processing unit 29 simultaneously performs voice recognition processing to generate text data and add an effect object to the image 101.
[0047] [Viewer terminal device 50] The viewer terminal device 50 is a data processing device having a configuration substantially similar to that of the user terminal device 10. Hereinafter, the components of the viewer terminal device 50 are denoted by the same reference numerals as those of the user terminal device 10, and detailed descriptions thereof will be omitted. The viewer terminal device 50 is a terminal for viewing composite data 100, which is delivered by the user terminal device 10 via the server 30 and which shows an image 101 captured by the user with the out-camera 11 and an avatar object 102 that changes in accordance with the user's movement information superimposed on the image 101. The composite data 100 may be generated by the viewer terminal device 50 or by the server 30, which will be described next. Therefore, the viewer terminal device 50 is required to have at least the functions of a communication device and a display device. Furthermore, the viewer terminal device 50 can also be used as the user terminal device 10, as long as it has a configuration substantially similar to that of the user terminal device 10.
[0048] [Server 30] The server 30 provides information and processing results in response to requests from the user terminal devices 10 and viewer terminal devices 50 that serve as clients.
[0049] As shown in FIG. 3, the server 30 is a device that manages the distribution system 1, and includes a distribution management database 31, a storage unit 32, a network IF 33, and a control unit . The distribution management database 31 manages a content ID that uniquely identifies the content to be distributed, a distribution time, and a viewer account that uniquely identifies the distribution destination for the viewer, in association with a distributor account that uniquely identifies the user who is the owner of the user terminal device 10, i.e., the distributor. Examples of the distributor account include a distributor ID and a distributor password. Examples of the viewer account include a viewer ID and a viewer password.
[0050] For example, it is managed that distributor account "AAA" is live streaming content ID "A123" through user terminal device 10. It is also managed that viewer terminal IDs "XXX," "YYY," and "ZZZ" are viewing content ID "A123." It is also managed that distributor account "BBB" is streaming content ID "B456" from 5:00 PM to 6:00 PM on December 25, 2018. In this case, the content ID "B456" to be streamed is uploaded from user terminal device 10 in advance and stored in storage unit 32. It is also managed that viewer accounts "XYZ," "PPP," and "QQQ" are viewing content ID "B456" through viewer terminal device 50 during the playback streaming time period. By associating viewer accounts in this way during streaming, it is possible, for example, to send comments sent from viewer terminal device 50 to user terminal device 10.
[0051] The storage unit 32 is a non-volatile memory, such as a hard disk or flash memory, and stores the distribution management program 32a and the above-mentioned distribution management database 31. The network IF 33 is an IF for communicating with the user terminal device 10 and the viewer terminal device 50 via the network 2. The control unit 34 is composed of a CPU, RAM, ROM, etc., and controls the overall operation of the user terminal device 10 in accordance with a program.
[0052] The control unit 34 is a distribution unit that controls distribution of the composite data 100 transmitted from the user terminal device 10 when it receives a distribution request from the user terminal device 10 via the network IF 33. For example, when the control unit 34 receives a live distribution request, it enables distribution of the composite data 100 sequentially transmitted from the user terminal device 10 on a predetermined channel. Furthermore, when the control unit 34 receives a distribution request for distribution at a predetermined time and time slot, it registers the channel and distribution time slot in a program guide and enables distribution at the registered predetermined time and time slot. When the control unit 34 receives a viewing request from the viewer terminal device 50, it enables viewing of the composite data 100 of the content ID corresponding to the received viewing request on the viewer terminal device 50. The viewer terminal device 50 performs streaming playback. Of course, the viewer terminal device 50 may download the composite data 100 and store it in the storage unit 28.
[0053] The operation of the distribution system will now be described. [Distribution preparation process] A user takes an image of a location where they are currently located, such as a tourist spot, famous place, historical site, or landscape, and prepares to distribute the image 101 of that location via the server 30. First, the user connects the support member 3 to the user terminal device 10 so that they can capture images in a first direction and a second direction from a position farther away from the user. Note that the support member 3 does not have to be connected.
[0054] The user operates the user terminal device 10 owned by the user to start the distribution program 28a. In response to this, the processing unit 29 of the user terminal device 10 starts the depth sensor 15. The depth sensor 15 identifies each part of the user's body in the person area by identifying the user's skeletal position. The depth sensor 15 then acquires depth information indicating the distance between the depth sensor 15 and each part of the user (head, upper body, arms, hands, lower body, feet, etc.), becomes able to track each part, and generates movement information according to changes in the depth information of each part.
[0055] 4, in step S1, the processing unit 29 acquires user movement information from the depth sensor 15 via the depth sensor IF 16. That is, the processing unit 29 detects the user with the depth sensor 15, recognizes the user, and becomes able to track the user's movement. In addition, the processing unit 29 may also activate the in-camera 13 to acquire images and videos of the user detected by the depth sensor 15.
[0056] In addition to the user's body movements, facial expressions, i.e., changes in the shape of the eyes and mouth, may also be detected. The detection results can be reflected in the avatar's facial expression. Furthermore, the user's movements and facial expressions may be detected using not only the depth sensor 15 but also the image from the in-camera 13, or may be detected using information from both sensors in combination.
[0057] In step S2, the processing unit 29 repeatedly determines whether or not the motion information has been acquired through the depth sensor IF 16, and if the motion information has been acquired, the processing proceeds to step S4. In step S3, the processing unit 29 activates the outer camera 11 and acquires an image 101 in a first direction, which is the user's line of sight, through the outer camera IF 12. This image 101 becomes a background image for the avatar object 102, and is, for example, an image that the user wants to introduce to viewers through the avatar object 102. This allows the processing unit 29 to acquire motion information from the depth sensor 15 and the image 101 from the outer camera 11 at the same time, and associate them.
[0058] The processing unit 29 may generate composite data that displays an image 101 captured by the outer camera 11 on the display unit 19 while simultaneously superimposing a user image object of the user captured by the inner camera 13 and detected by the depth sensor 15 on the image 101, and display the composite data on the display surface of the display unit 19. In this case, for example, the processing unit 29 extracts the user image object from the image captured by the inner camera 13 based on the detection result from the depth sensor 15. The user image object is then displayed on the image 101. The position at which the user image object is displayed on the display surface is an initial position predetermined by the distribution program 28a on the image 101 or the display surface of the display unit 19. The avatar object 102 is a virtual object that represents the user, and the user can predict, via the user image object, where the avatar object 102 will be displayed on the display surface.
[0059] Note that a temporary avatar object may be displayed at the initial position instead of the user image object until the avatar object 102 is officially determined. Also, the process of displaying such a user image object or temporary avatar object in the image 101 may be omitted.
[0060] The processing unit 29 displays a selection screen for an avatar object 102 on the display surface of the display unit 19. As an example, a list of multiple avatar objects 102 stored in the storage unit 28 in association with the distribution program 28a is displayed. If it is not possible to display all the avatar objects at once on the display surface, the list display is scrolled up and down or left and right so that the user can check all the avatar objects that are candidates for selection. The user touches a desired avatar object using a finger or a stylus pen. In step S4, the processing unit 29 selects the avatar object at the touched position from the multiple avatar objects displayed in the list, identifies the selected avatar object, and determines the avatar object 102 to be used. Note that further candidate avatar objects may be downloaded from the server 30.
[0061] In step S5, processing unit 29 determines whether a display position for displaying avatar object 102 has been specified within a predetermined period of time. Specifically, processing unit 29 determines whether a user has touched a predetermined position on the display surface of display unit 19 using a finger, a stylus pen, or the like within the predetermined period of time. If a touch has been made, processing unit 29 specifies the touched position as the position for displaying avatar object 102. If a touch has not been made, in step S5-1, distribution program 28a sets an initial position predetermined on image 101 or the display surface of display unit 19 as the position for displaying avatar object 102.
[0062] In step S6, processing unit 29 displays background image 101 on the display surface, and generates composite data 100 by overlaying avatar object 102 on image 101 so that avatar object 102 is displayed at a predetermined display position on the display surface, and displays the composite image on the display surface of display unit 19. Specifically, when a user image object of a user captured by in-camera 13 is displayed on image 101, processing unit 29 generates composite data 100 by overlaying avatar object 102 in place of the user image object or temporary avatar object at the position where the user image object or temporary avatar object is displayed. Then, processing unit 29 displays the composite image based on composite data 100 on the display surface of display unit 19.
[0063] In this way, the avatar object 102 is not superimposed on the user object captured by the in-camera 13 in the image 101. This prevents the avatar object 102 from being displayed outside the user object. As an example, the avatar object 102 may be displayed smaller than the user image object or the temporary avatar object.
[0064] Avatar object 102 is an object that occupies a predetermined area on the display surface on which image 101 is displayed. Therefore, as an example, when a point is designated as the display position of avatar object 102, processing unit 29 displays avatar object 102 on the display surface so that the designated position is the center of avatar object 102. Also, as another example, when the display position of avatar object 102 is designated to have an area using a finger, a stylus pen, or the like, processing unit 29 calculates the center of the area and displays avatar object 102 on the display surface of display unit 19 so that the center coincides with the center of avatar object 102.
[0065] The display portion of the avatar object 102 corresponds to the range of the user detected by the depth sensor 15, for example. That is, when the depth sensor 15 detects only the upper half of the user's body, the display range of the avatar object 102 also corresponds to the upper half of the body. On the other hand, when the depth sensor 15 detects the entire body of the user, the display range of the avatar object 102 also corresponds to the entire body. On the other hand, the display portion of the avatar object 102 does not have to correspond to the range of the user detected by the depth sensor 15. In this case, even if the depth sensor 15 detects only the upper half of the user's body, the avatar object 102 displays the entire body (from head to toe). Or, even if the depth sensor 15 detects the entire body of the user, the avatar object 102 displays the upper half of the body. Furthermore, the portions not detected by the depth sensor 15 may be stationary, or may change according to the detection results of the sensor 17. Alternatively, the movement of the portions not detected by the depth sensor 15 may be used to infer and change the movement of the portions not detected by the depth sensor 15.
[0066] However, the size of the avatar object 102 in the image 101 displayed on the display surface of the display unit 19 may be too large or too small. Also, the user may want to adjust the size of the avatar object 102 to suit their own preferences. To handle such cases, in step S7, the processing unit 29 displays a size determination button on the display surface. Then, the processing unit 29 determines whether the size determination button has been touched, and in step S8, executes a process for adjusting the size of the avatar object 102 by the user until the determination button is touched.
[0067] As an example, when using the support member 3, the size of the avatar object 102 in the image 101 displayed on the display surface can be changed by changing the distance between the user and the depth sensor 15 of the user terminal device 10. Specifically, by using the support member 3 to move the user away from the depth sensor 15 of the user terminal device 10, the avatar object 102 can be made smaller in the image 101 displayed on the display surface. Furthermore, by moving the user closer to the depth sensor 15 of the user terminal device 10, the avatar object 102 can be made larger in the image 101 displayed on the display surface. Furthermore, the size of the avatar object 102 can be changed using the zoom function of the in-camera 13. For example, when the focal length of the in-camera 13 is changed toward the wide-angle direction, the avatar object 102 is reduced, and when the focal length is changed toward the telephoto direction, the avatar object 102 is enlarged.
[0068] Furthermore, if touch panel 20 is a capacitive projection type, multi-touch is possible, and the size of avatar object 102 can be adjusted by spreading or closing two fingers. For example, when two fingers are spread apart, avatar object 102 can be enlarged according to the amount by which the two fingers are spread apart, and when two fingers are closed, avatar object 102 can be reduced according to the amount by which the two fingers are closed.
[0069] If the size determination button is touched in step S7, processing unit 29 proceeds to step S9. When using support member 3 to change the distance between the user and depth sensor 15 of user terminal device 10, the same determination process as when the size determination button is touched on operation unit 22 of support member 3 can be performed. As a result, avatar object 102 is superimposed on image 101 at the size of avatar object 102 at the time of determination process.
[0070] After adjusting the size of the avatar object 102, the orientation of the avatar object 102 is set. The user is facing the depth sensor 15 of the user terminal device 10 and faces the depth sensor 15. Therefore, the avatar object 102 also faces forward, i.e., the face is visible in the image. On the other hand, when the user is moving forward, since the avatar object 102 is the user's incarnation, it is also preferable that the avatar object 102 is walking facing the same direction as the user, i.e., the avatar object 102 is shown from behind with the back of its head and back displayed. Furthermore, when the avatar object 102 is orally introducing famous places, historical sites, etc., that appear in the background image 101 on behalf of the user, it is also preferable that the avatar object 102 faces forward, i.e., the face is visible in the image. Depending on the imaging situation, it may be preferable that the avatar object 102 faces sideways, regardless of the user's orientation.
[0071] Therefore, processing unit 29 displays a direction determination button on the display surface. Then, in step S9, processing unit 29 determines whether the direction determination button has been touched, and in step S10, executes a process of adjusting the direction of avatar object 102 by the user until the direction determination button is touched.
[0072] As an example, when a support member 3 is used, the orientation of the avatar object 102 can be changed using the operation unit 22 of the support member 3. For example, the orientation of the avatar object 102 can be easily changed by using a cross key, joystick, trackball, or the like of the operation unit 22 to operate the avatar object 102 in the desired direction. The orientation of the avatar object 102 can also be changed using the touch panel 20. As an example, when the touch panel 20 is a capacitive projection type, the orientation of the avatar object 102 can be easily changed by the user tracing their finger in the desired direction. The size and orientation of the avatar object 102 set here become the basis at the start of distribution. In other words, during distribution, the display state of the set avatar object 102 changes depending on the movement information, etc.
[0073] If the orientation decision button is touched in step S9, the processing unit 29 proceeds to step S11. Note that it is also possible to perform the same decision processing as when the orientation decision button is touched on the operation unit 22 of the support member 3.
[0074] In step S11, the processing unit 29 displays a distribution start button on the display surface. When the distribution start button is touched, distribution starts. Note that the same decision processing as when the distribution start button is touched on the operation unit 22 of the support member 3 can also be performed.
[0075] The processing unit 29 associates the content with the distributor account and content ID, and transmits the distribution date and time to the server 30 via the network IF 27 to decide whether to perform live distribution or to specify the distribution date, time, and time slot for distribution. The control unit 34 of the server 30 then registers this information in the distribution management database 31.
[0076] When the distribution start process is performed, processing unit 29 generates composite data 100 of image 101 with avatar object 102 superimposed thereon, and transmits composite data 100 to server 30 while displaying it on display unit 19. Then, control unit 34 of server 30 starts distributing composite data 100 transmitted from user terminal device 10 via network IF 33 on the assigned channel (e.g., URL).
[0077] [Distribution processing example 1] Distribution process example 1 is a case where the user is stationary and not moving forward or backward. As shown in Fig. 5, in step S21, processing unit 29 acquires user movement information from depth sensor 15. In step S22, processing unit 29 generates composite data 100 including video and audio in which avatar object 102 changes in accordance with the movement information, and transmits the composite data 100 to server 30.
[0078] The depth sensor 15 detects, for example, the movement of the eyes and lips when the user is talking, and detects the movement of the hands when the user moves them. Furthermore, when the user turns sideways, the movement of the user turning sideways is detected. The processing unit 29 changes the avatar object 102 in accordance with the user movement information from the depth sensor 15. That is, when the user is talking, the lips of the avatar object 102 change in the same way as the user, and when the user moves their hands, the hands of the avatar object 102 change in the same way as the user's hand movements. Furthermore, when the user turns sideways, the avatar object 102 turns sideways. The user's movements may be acquired not only from the depth sensor 15, but also from images from the in-camera 13.
[0079] The processing unit 29 also displays the composite data 100 being transmitted to the server 30 on the display screen of the display unit 19. Therefore, the user can capture an image for distribution while checking how the avatar object 102 is displayed on his / her behalf (see FIG. 1).
[0080] Viewer terminal device 50 can view a composite image based on composite data 100 being distributed by accessing a predetermined channel of server 30. Composite data 100 is data in which the lips of avatar object 102 move in accordance with lip movement information of the user. Therefore, to the viewer, it appears as if avatar object 102 is speaking.
[0081] As shown in FIG. 6(a), the data of avatar object 102 is three-dimensional data. Therefore, when avatar object 102 faces away from the camera, the back of the head is displayed; when facing forward, the face is displayed; when facing right, the right side of the face is displayed; and when facing left, the left side of the face is displayed. For example, processing unit 29 changes avatar object 102 based on the movement information from depth sensor 15 and the detection data from sensor 17, relative to the state before the movement. As shown in FIG. 6(b), for example, when avatar object 102 faces away from the camera and rotates 90 degrees rightward, avatar object 102 rotates rightward.
[0082] FIG. 6(c) shows a case where the avatar object 102 rotates 270 degrees to the left from a state facing away to a state facing right. In this case, the avatar object 102 rotates leftward. However, the state after rotation is the same as the state where the avatar object 102 rotates 90 degrees to the right from a state facing away to a state facing rightward. In this case, as shown in FIG. 6(a), the processing unit 29 may perform processing to rotate the avatar object 102 in a direction that reduces the rotation angle of the avatar object 102, i.e., to the right. This reduces the amount of calculation processing performed by the processing unit 29 to rotate the avatar object 102. Furthermore, the movement of the avatar object 102 is reduced, allowing the avatar object 102 to look more naturally.
[0083] [Distribution processing example 2] Distribution process example 2 is performed when the user is moving. Sensor 17 is an acceleration sensor or gyro sensor, and detection data such as acceleration data, angle data, and angular velocity data of user terminal device 10 is calculated based on the output from these sensor elements. FIG. 7 shows an example of changes in acceleration detected by sensor 17 when the user moves from the front toward the back. At time t1, the user is stationary, and the acceleration is 0. When the user starts walking toward the back at time t2, a positive acceleration is output. When the user starts walking at a constant speed at time t3, the acceleration approaches 0. When the user stops walking, a negative acceleration is output at time t4. When the user resumes walking toward the back at time t5, a positive acceleration is output. When the user is walking at a constant speed at time t6, the acceleration approaches 0. When the user starts running toward the back at time t7, a positive acceleration higher than when the user started walking is output. When the user is running at a constant speed as at time t8, the acceleration approaches 0.
[0084] As described above, when the user moves by walking or running, the processing unit 29 of the user terminal device 10 can detect the direction and speed of movement from the acceleration, angle, and angular velocity measured by the sensor 17. As shown in Fig. 8, in step S31, the processing unit 29 determines whether the acceleration has changed, and in step S32, changes the avatar object 102 according to the acceleration. When walking or running, up-down and forward-backward movements occur, and when running, these changes are faster and larger than when walking.
[0085] Specifically, processing unit 29 runs pedometer program 28b linked to distribution program 28a. Processing unit 29 then determines whether the acceleration is greater than a first threshold, and if it is less than the first threshold, determines that the person is standing still, and if it is equal to or greater than the first threshold, determines that the person is walking. Processing unit 29 also determines whether the acceleration is greater than a second threshold, and if it is less than the second threshold, determines that the person is walking, and if it is equal to or greater than the second threshold, determines that the person is running.
[0086] The processing unit 29 moves the avatar object 102 up and down and forward and backward when the avatar object 102 is walking or running, and when the avatar object 102 is running, the amount of movement is greater than when the avatar object 102 is walking. When the entire body of the avatar object 102 is displayed on the display screen, the processing unit 29 moves the arms and legs as if the avatar object 102 is running or walking, even if the depth sensor 15 does not detect the arms and legs. Even when only the upper body of the avatar object 102 is displayed on the display screen, the processing unit 29 moves the avatar object 102 up and down.
[0087] In the example of FIG. 7 , the user is moving forward from the front to the back. Therefore, although the depth sensor 15 of the user terminal device 10 actually faces the user's face, the processing unit 29 displays the back of the head and back of the avatar object 102 as if the avatar object 102 is moving from the front to the back. When the user is stationary, such as from time t4 to time t5, the processing unit 29 reverses the orientation of the avatar object 102 so that its face faces forward. When explaining famous places, historical sites, etc., the user is often stationary, and it is easier for viewers to see the user's facial expressions, etc., through the facial expressions reproduced on the avatar object 102, if the face of the avatar object 102 is visible during the explanation. In the above example, the back of the avatar object 102 is displayed when the user is moving, but the avatar object 102 may always face forward.
[0088] The movement of the user may be detected not only by the sensor 17 but also by a combination of other means. For example, the sensor 17 may be used in combination with GNSS. Furthermore, the depth sensor 15 and SLMA (Simultaneous Localization and Mapping) may be used in combination. Furthermore, temporal changes in position information obtained by photogrammetry (Visual SLAM) may be used supplementarily to determine the walking direction.
[0089] [Other distribution processing examples] Furthermore, when processing unit 29 detects a specific action based on the movement information from depth sensor 15, it additionally displays an effect object 103 related to the action in addition to avatar object 102. As an example, when a person hits the palm of one hand with the fist of the other hand, it is a moment of inspiration. Therefore, as shown in FIG. 9( a), when processing unit 29 detects, based on the movement information from depth sensor 15, the action of the user hitting the palm of one hand with the fist of the other hand, it displays effect object 103, which is an additional virtual object resembling a light bulb, near avatar object 102. Effect object 103 may be a stationary object or a moving object, and may further be audio data such as a voice or music that evokes an inspiration.
[0090] 9(b), while processing is being performed in accordance with distribution program 28a to overlay avatar object 102 on image 101 captured by outer camera 11 and generate composite data 100, processing is also being performed in parallel to perform voice recognition processing to generate text data and add effect object 104 to image 101. Here, effect object 104 can be configured by adding a speech bubble object to avatar object 102 and displaying text within the speech bubble object. Alternatively, in a composite image based on composite data 100, text may be displayed in a field object at any of the top, bottom, left, or right portions.
[0091] Furthermore, when the processing unit 29 detects a specific action based on the movement information from the depth sensor 15, it may add a speech bubble object to the avatar object 102 and display text in the speech bubble object to form the effect object 104. In this case, the specific action may be, for example, a specific gesture used when introducing something, such as raising one palm up to shoulder height.
[0092] Furthermore, viewers watching on viewer terminal devices 50 can send comments to the user who is broadcasting. In server 30, a broadcast management database 31 manages viewer accounts in association with broadcaster accounts. When the control unit 34 of server 30 receives comment data associated with a broadcaster account or content ID from a viewer terminal device 50, it transmits the comment to the user terminal device 10 of the corresponding broadcaster. In such a case, the user terminal device 10 displays the comment from the viewer on the display screen of the display unit 19. For example, if a user is broadcasting while walking around a tourist spot, the viewer can send a comment requesting the user to go next from the viewer terminal device 50 to the user terminal device 10 via the server 30. In this way, viewers can send comments to users, allowing the user to shoot while communicating with the viewer.
[0093] Furthermore, the viewer can display a virtual object related to the viewer himself / herself in the image 101 through the viewer terminal device 50. For example, the viewer's own avatar object can be added to a composite image based on the composite data 100. The avatar object, which is the viewer's virtual object, changes based on movement information generated by the depth sensor 15 provided in the viewer terminal device 50. The avatar object is transmitted from the viewer terminal device 50 to the server 30 and then displayed on the user terminal device 10 via the server 30. In this case, composite data including the avatar object 102 of the broadcaster and the avatar object of the viewer may be generated by any of the viewer terminal device 50, the server 30, and the user terminal device 10. The broadcaster's avatar object 102 and the viewer's avatar object are displayed in the composite image based on the composite data 100. The viewer's avatar object to be displayed in the image 101 is also selected by the viewer terminal device 50, and the selected avatar object is displayed in a predetermined position in the image 101. Furthermore, instead of an avatar object, a virtual object such as a gift for the avatar object 102 may be displayed. For example, virtual objects include offerings, a microphone for broadcasting, bouquets of flowers, glasses, clothing, and other decorative items. Viewers can select such presents from a list of presents. The selected virtual object is then sent to the server 30. According to the distribution system 1, the following effects can be obtained.
[0094] (1) The avatar object 102 is not superimposed on the user's own user object, but is superimposed on the image 101 in place of the user object. This prevents the movement of the avatar object 102 from being unable to follow the movement of the user object, and prevents the user object from overflowing the avatar object 102. Also, even if the user object is smaller than the avatar object 102, it prevents the user object from overflowing the avatar object 102. In this way, it is possible to prevent the avatar object 102 from appearing unnatural in the composite image.
[0095] (2) The avatar object 102 can be automatically added to the image 101. Alternatively, the position where the avatar object 102 is to be displayed can be easily specified using the touch panel 20.
[0096] (3) The outer camera 11 provided on the user terminal device 10 can be used to capture an image of a background subject to generate an image 101, and the depth sensor 15 can then detect the user who captured the image and add an avatar object 102 to the image 101.
[0097] (4) Since the user terminal device 10 includes the sensor 17, the avatar object 102 can be changed using the detection data from the sensor 17 even in areas where the depth sensor 15 cannot detect the user's movement.
[0098] (5) The user terminal device 10 has a pedometer function. Therefore, it is possible to detect the movement of the user and the movement of the user terminal device 10 separately. In other words, it is possible to determine whether the user is standing still and moving only the user terminal device 10 with their hands, or whether the user is actually walking. When it is detected that the user is walking, it is also possible to change the avatar object 102 to appear to be walking. Furthermore, it is possible to detect whether the user is running, and when the user is running, it is possible to change the avatar object 102 to appear to be running. Furthermore, by using the depth sensor 15, outer camera 11, sensor 17, and pedometer functions, the broadcaster can walk ahead of the viewer without walking backward, creating the appearance of a tour guide.
[0099] (6) When the processing unit 29 detects a specific movement from the movement information, it adds an effect object 103. This diversifies the expression style of the avatar object 102, making it possible to further enhance entertainment value.
[0100] (7) The processing unit 29 recognizes the user's voice, converts it into text, and adds the text to the image 101. This makes it possible for the viewer to easily understand what the user is saying.
[0101] The above-described live distribution system can be modified as follows. The user terminal device 10 may be configured as shown in FIG. 10. That is, the user terminal device 10 shown in FIG. 10 includes an out-depth sensor 41 on the outer surface opposite the user holding the user terminal device 10, and the out-depth sensor 41 is connected to an out-camera IF 42. The out-depth sensor 41 is, for example, located near the out-camera 11. In FIG. 10, the depth sensor located on the inner surface facing the user is an in-depth sensor 15, which is connected to an in-depth sensor IF 16. The in-depth sensor 15 and the in-depth sensor IF 16 correspond to the depth sensor 15 and the depth sensor IF 16 in FIG. 2, respectively.
[0102] The out-depth sensor 41 has a similar configuration to the in-depth sensor 15, and is, for example, an infrared camera, and includes a light-projecting unit that projects infrared rays and an infrared detection unit that detects the infrared rays. Depth information, such as three-dimensional position information in real space, is obtained from the time it takes for the infrared pulse projected from the light-projecting unit to be reflected and returned.
[0103] The in-depth sensor 15 mainly detects a user holding the user terminal device 10 in their hand. That is, the in-depth sensor 15 detects objects that are relatively close to the in-depth sensor 15. In contrast, the out-depth sensor 41 detects objects that are farther away than the distance to the detection target of the in-depth sensor 15. For this reason, as an example, it is preferable that the light-emitting section of the out-depth sensor 41 has a higher output than that of the in-depth sensor 15, and the detection section has a higher sensitivity than that of the in-depth sensor 15.
[0104] The out-depth sensor 41 is disposed near the out-camera 11, and its detection range is approximately the same as the imaging range of the out-camera 11. The detection range of the out-depth sensor 41 may be wider or narrower than the imaging range of the out-camera 11. The out-depth sensor 41 detects depth information of objects present within the detection range. As an example, when a person is walking in the user's direction of travel, the out-depth sensor 41 detects the movement of that person. Furthermore, when there are stairs ahead, the out-depth sensor 41 detects each step of the stairs. Furthermore, when a person ahead is climbing the stairs, the out-depth sensor 41 detects their movement. Then, the out-camera IF 42, which serves as a depth information acquisition unit, acquires the depth information from the out-depth sensor 41.
[0105] When generating composite data 100 of image 101 overlaid with avatar object 102 and distributing composite data 100 via server 30, processing unit 29 utilizes depth information and movement information from out-depth sensor 41 as follows. For example, when a person ahead is climbing stairs, there is a high probability that the next user will also be climbing stairs. Out-depth sensor 41 detects the stairs and the movement of the person ahead. Processing unit 29 detects that the user has reached the stairs according to the depth information and movement information from out-depth sensor 41. When processing unit 29 detects that the user has transitioned to a state of climbing stairs based on movement information from in-depth sensor 15 and detection data from sensor 17, processing unit 29 applies the movement information of the person ahead climbing the stairs to avatar object 102, thereby changing the stairs so that avatar object 102 appears to be climbing the stairs. This allows avatar object 102 to appear as if it is climbing the stairs when stairs exist in real space, preventing unnatural changes to avatar object 102 relative to image 101.
[0106] Here, the case where the user climbs stairs to match the person in front has been described, but this is just one example. For example, if an obstacle such as a flower bed, a desk, or a chair is present ahead in the user's direction of travel, the user will take action to avoid the obstacle. Based on the depth information of the obstacle detected by the out-depth sensor 41, the processing unit 29 can change the avatar object 102 to avoid the obstacle. That is, the avatar object 102 can be displayed taking into account its front-to-back relationship with real objects such as a flower bed, a desk, or a chair. In this way, the user terminal device 10 equipped with the out-depth sensor 41 can change the avatar object 102 to match the real-space object captured by the out-camera 11.
[0107] Instead of using built-in outer and inner cameras 11 and 13, the user terminal device 10 may use a 360-degree camera device 201 paired by wire or wirelessly through a communication IF 21, which serves as a first and second acquisition unit. As shown in FIG. 11(a), the 360-degree camera device 201 includes a first camera 202 facing a first direction on a first surface, and a second camera 203 facing a second direction on a second surface opposite the first surface. Each of the first camera 202 and the second camera 203 includes an image receiving element. In this case, one of the first camera 202 and the second camera 203 corresponds to the outer camera 11, and the other corresponds to the inner camera 13. A depth sensor is disposed near the lens of at least the camera corresponding to the inner camera 13. This results in a user terminal device 10 similar to that shown in FIG. 2. If depth sensors are disposed near both cameras, the user terminal device 10 will be similar to that shown in FIG. 10.
[0108] 2 or 10, the 360-degree camera device 201 itself becomes the user terminal device 10 of FIG. 2 or 10.
[0109] The 360-degree camera device 201 having such a configuration is preferably used by attaching it to an attachment device that secures the camera to a helmet, an attachment device that attaches the camera attached to a wrist mount to the back of the hand, wrist, or arm, an attachment device that secures the camera by clamping it to the handlebars of a bicycle with a fixing clamp, an attachment device that secures the camera to the body, an attachment device that secures the camera with a suction plate, or a handy stick of fixed length.
[0110] As shown in FIG. 11(b), the first camera 206 and the second camera 207 may be separated. That is, two hemispherical cameras are used. In this case, a depth sensor is provided at least on the camera corresponding to the in-camera 13. This results in a user terminal device 10 similar to that shown in FIG. 2, and if depth sensors are provided on both cameras, the user terminal device 10 will be similar to that shown in FIG. 10. Because the first camera 206 and the second camera 207 are separated, the camera corresponding to the out-camera 11 and the camera corresponding to the in-camera 13 do not need to be located in the same place. As an example, by connecting the first camera 206 and the second camera 207 to the user terminal device 10 via a wired or wireless connection, the camera corresponding to the out-camera 11 and the camera corresponding to the in-camera 13 do not need to be located in the same place. A user can install the camera corresponding to the out-camera 11 in a location remote from the user's own home or other location and live stream the composite data 100 while at home or other such location.
[0111] The example in FIG. 11(c) uses one hemispherical camera 211. The hemispherical camera 211 is, for example, a wide-angle camera. The hemispherical camera 211 is paired via a wired or wireless connection through the communication IF 21. In this case, the angle of view is, for example, approximately 170 degrees to 185 degrees, which is a wide angle. An image incident from one lens unit 212 is received by one image receiving element. In a captured image 213, a user appears in a partial area 214, and the remaining area 215 is the area that becomes the image 101.
[0112] The hemispherical camera 211 does not have a depth sensor, and the depth information and movement information of the user are calculated by image analysis using photogrammetry technology or the like. The depth information and movement information of the user can be detected using AI that utilizes deep learning. The processing unit 29 overlays an avatar object 102, which changes in accordance with the user's movement information detected in the area 214, onto the image 101 in the area 215 to generate composite data 100. The user terminal device 10 distributes this composite data 100 to the viewer terminal device 50 via the server 30.
[0113] 11(d) is a modified example of the user terminal device 10 shown in FIG. 2. The user terminal device 10 is equipped with an outer camera 11 and an inner camera 13, but the image receiving element 230 is a single element that is divided into a first image receiving area 231 and a second image receiving area 232. The first image receiving area 231 receives an image from the outer camera 11, and the second image receiving area 232 receives an image from the inner camera 13. The first image receiving area 231 receives the image from the outer camera 11, and the processing unit 29 generates an image 101 that serves as a background image. The user is received in the second image receiving area 232. The processing unit 29 generates user movement information by performing image analysis on the image in the second image receiving area 232. The processing unit 29 superimposes the avatar object 102, which changes in accordance with the user's movement information detected in the second image receiving area 232, onto the image 101 generated from the image received in the first image receiving area 231, to generate composite data 100. The user terminal device 10 distributes this composite data 100 to the viewer terminal device 50 via the server 30.
[0114] The in-camera 13 may be replaced by the depth sensor 15. In this case, the second image receiving area 232 serves as an infrared detection unit that detects infrared rays. In this case, the number of light receiving elements is reduced to one, which allows the number of components in the user terminal device 10 to be reduced.
[0115] The function of overlaying the avatar object 102 on the image 101 to generate the composite data 100 may be provided by the server 30, rather than by the user terminal device 10. In this case, the user terminal device 10 transmits the image 101 of the subject captured by the outer camera 11 and the movement information detected by the depth sensor 15 to the server 30. The network IF 33 functions as a first acquisition unit and a second acquisition unit. The server 30 then overlays the avatar object 102, which moves based on the movement information, on the image 101 to generate the composite data 100, and transmits the composite data 100 to the user terminal device 10 and distributes it to the viewer terminal device 50.
[0116] The processing performed by the server 30 may be distributed to edge servers. The composite data 100 may be generated not only by the user terminal device 10 but also by the viewer terminal device 50 or the server 30.
[0117] The user terminal device 10 does not need to have the function of performing speech recognition processing to generate text data and add a text object to the image 101. This is because it is possible to communicate with the viewer using only speech.
[0118] When the user terminal device 10 detects a specific user action, it does not need to perform processing to add an effect object 103 related to that action in addition to the avatar object 102. This is because the user's intention can be communicated to the viewer even without adding the effect object 103.
[0119] The user terminal device 10 does not have to have a pedometer function. In this case, the movement of the user terminal device 10 is determined as the movement of the user based on the detection data detected by the sensor 17, and the avatar object 102 is changed accordingly.
[0120] The sensor 17 may not be provided. Even if the avatar object 102 does not change to appear as if it is walking, as long as the user's face is at least detected and the avatar object 102, which is the user's personification, changes in accordance with the user's facial expression, the user's facial expression will be conveyed to the viewer via the avatar object 102. Alternatively, the direction of the user's movement may be detected by changes in the image from the outer camera 11 or the inner camera 13, and the walking display of the avatar object 102 may be changed accordingly.
[0121] The sensor 17 should at least be equipped with an acceleration sensor to detect acceleration. The user terminal device 10 does not need to have an external camera 11, as long as it has an acquisition unit that acquires video. This is because an external camera can be connected as shown in Figures 11(a) and 11(b). In this case, the communication IF 21 functions as an acquisition unit such as a first acquisition unit that acquires the image 101 and a second acquisition unit that acquires movement information.
[0122] The display unit 19 does not necessarily have to include a touch panel 20. This is because, if a pointing device such as a mouse is connected to the user terminal device 10 by wire or wirelessly, a position can be input using a pointer displayed on the display surface.
[0123] The user terminal device 10 may be installed with a voice changer program that processes input voice and converts it to sound like a different voice. As an example, when distributing the composite data 100, the voice collected by an acoustic-electrical conversion device such as a microphone can be converted into a different voice by a voice changer, and the converted voice can be distributed. For example, a voice changer can change the sampling period of voice data digitized by PCM (pulse code modulation) to change the voice pitch, delay the voice data by adding a time delay, or combine the delayed voice data with the original voice data to create an echo.
[0124] The user terminal device 10 may have installed thereon a voice reading program for reading text data aloud. For example, when the user terminal device 10 receives a request comment from a viewer as text data from the viewer terminal device 50, the user terminal device 10 converts the text data into voice and reads the text data aloud. In this case, the voice may be, for example, a synthesized voice or the voice of a voice actor or actor. If the avatar object 102 is a character object that appears in an animation, the voice may be the voice of the character's voice actor. If the avatar object 102 is a person object that imitates a famous person, a famous person, or an ordinary person, the voice may be the voice of that person.
[0125] The composite data 100 is composed of an image 101 with an avatar object 102 superimposed thereon and sound. However, it is preferable that the user's motion information, the detection data from the sensor 17, the GNSS location information, and the sound at the location where the image 101 was captured, all acquired at the same time as the image 101, be perfectly synchronized. A slight time lag between the image 101 and the motion information is acceptable as long as it does not cause discomfort to the viewer. In such a case, the movement of the avatar object 102 will be slightly delayed relative to the image 101. For example, as shown in FIG. 11(b), when the user terminal device 10 is separated from the first camera 206 and the second camera 207, a situation may occur in which one of the image 101 and the motion information is delayed relative to the other, depending on the communication environment. In such a case, a time lag between the image 101 and the motion information is acceptable.
[0126] When live streaming is not performed, the time acquired from a GNSS, an NTP server, or the built-in clock of the user terminal device 10 or the server 30 may be used as a reference to acquire user movement information, detection data from the sensor 17, GNSS location information, and sound at the location where the image 101 was captured, at the same time as the reference time assigned to the image 101. As shown in FIG. 11(b), when the user terminal device 10, the first camera 206, and the second camera 207 are separate, the information and data from the same period may be acquired according to the built-in clock of each device, and the composite data 100 may be generated.
[0127] The user's motion information may be generated by calculating inter-frame differences and motion vectors based on the video captured by the in-camera 13. [Explanation of symbols]
[0128] 1...Distribution system, 2...Network, 3...Support member, 10...User terminal device, 11...Outer camera, 12...Outer camera IF, 13...Inner camera, 14...Inner camera IF, 15...Depth sensor, 16...Depth sensor IF, 17...Sensor, 18...Sensor IF, 19...Display unit, 20...Touch panel, 21...Communication IF, 22...Operation unit, 23...GNSS receiver, 24...Audio input device, 25...Audio output device, 26...Audio IF, 27...Network IF, 28...Memory unit, 28a...Distribution program, 28b...Pedometer program, 28c...Speech recognition program, 29...Processing unit, 30...Server, 31...Distribution management Database, 32...storage unit, 33...network IF, 32a...distribution management program, 34...control unit, 41...out-depth sensor, 42...out-camera IF, 50...viewer terminal device, 100...synthesized data, 101...image, 102...avatar object, 103...effect object, 104...effect object, 201...360-degree camera device, 202...first camera, 203...second camera, 206...first camera, 207...second camera, 211...semi-spherical camera, 212...lens unit, 213...captured image, 214...area, 215...area, 230...image receiving element, 231...first image receiving area, 232...second image receiving area.
Claims
1. a first acquisition unit for acquiring image information of a first subject; a second acquisition unit for acquiring movement information regarding a movement of a second subject different from the first subject; a third acquisition unit that acquires sound information related to a sound; a processing unit that displays, on a display surface, a first image based on the image information, a virtual object that changes based on the movement information, and a second image based on the sound information; A data processing device comprising:
2. 2. The data processing device according to claim 1, The processing unit is a data processing device that generates the second image based on the sound information.
3. 3. The data processing device according to claim 1, The processing unit is a data processing device that performs recognition processing to recognize the sound information.
4. 4. The data processing device according to claim 3, The processing unit is a data processing device that generates text information from waveform information based on the sound information in the recognition process, and generates the second image based on the text information.
5. 5. The data processing device according to claim 4, A data processing device wherein the second image is a character image.
6. 6. The data processing device according to claim 1, The processing unit is a data processing device that controls the timing of displaying the second image on the display surface based on the movement information.
7. 7. The data processing device according to claim 6, The processing unit is a data processing device that controls the timing of displaying the second image on the display surface based on movement information related to the mouth of the second subject.
8. 8. The data processing device according to claim 1, The processing unit associates a timing at which the virtual object, which changes based on the movement information, is changed with a timing at which the second image is displayed.
9. 9. The data processing device according to claim 8, The processing unit is a data processing device that synchronizes a timing of changing the virtual object that changes based on the movement information with a timing of displaying the second image.
10. 10. The data processing device according to claim 9, the virtual object has a first part that changes based on a movement of the mouth of the second subject; The processing unit is a data processing device that synchronizes the timing of changing the first region with the timing of displaying the second image.
11. 11. The data processing device according to claim 1, The processing unit is a data processing device that displays a predetermined third image on a display surface when the motion information is predetermined motion information.
12. 12. The data processing device according to claim 11, The processing unit is a data processing device that displays the second image together with the third image.
13. 13. The data processing device according to claim 12, The processing unit is a data processing device that displays the second image on the display surface so as to overlap the third image.
14. 14. The data processing device according to claim 13, A data processing device, wherein the third image is an image of a speech bubble.
15. 15. A data processing device according to any one of claims 1 to 14, a receiving unit that receives first data transmitted by operating the external device; The processing unit is a data processing device that displays the first image, the virtual object, the second image, and an image based on the first data on the display surface.
16. 16. The data processing device according to claim 15, The receiving unit receives, as the first data, data generated by operating the external device.
17. 17. The data processing device according to claim 15 or claim 16, the external device includes a first terminal device and a second terminal device; The receiving unit receives, as the first data, data transmitted by operating the first terminal device and data transmitted by operating the second terminal device.
18. 18. A data processing device according to any one of claims 15 to 17, The receiving unit is a data processing device that receives the first data transmitted from a server.
19. 18. A data processing device according to any one of claims 15 to 17, A data processing device comprising a transmission unit that transmits second data to be distributed to the external device.
20. 20. The data processing device according to claim 19, The transmission unit transmits, as the second data, data for displaying an image on the external device.
21. 21. The data processing device according to claim 20, The transmission unit is a data processing device that transmits, as the second data, data for displaying at least one of the first image, the virtual object, the second image, and an image based on the first data on the external device.
22. 22. A data processing device according to any one of claims 19 to 21, The transmitting unit is a data processing device that transmits data for reproducing sound on the external device based on the sound information.
23. A program that causes a data processing device to operate to acquire image information of a first subject, acquire movement information regarding the movement of a second subject different from the first subject, acquire sound information regarding sound, and display on a display surface a first image based on the image information, a virtual object that changes based on the movement information, and a second image based on the sound information.
Citation Information
Patent Citations
Moving image generation device and program
JP2015184689A