Computer program, server device and method

A system using sensor-based body movement analysis automates avatar expression control, addressing the inefficiencies of manual screen flicking in video distribution, ensuring accurate and effortless expression changes.

JP7726544B2Active Publication Date: 2025-08-20GLEE HOLDINGS CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023214338
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-08-20
Estimated Expiration
2039-12-27

AI Technical Summary

Technical Problem

Existing video distribution systems require performers to manually flick the smartphone screen to control avatar expressions, which is cumbersome and prone to errors.

Method used

A computer program and server device that analyze body movements using sensors to detect specific facial expressions or gestures, determining when predefined thresholds are exceeded to accurately control avatar expressions.

Benefits of technology

Facilitates easy and accurate expression control for avatars by automating the process, reducing performer effort and minimizing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726544000001
    Figure 0007726544000001
  • Figure 0007726544000002
    Figure 0007726544000002
  • Figure 0007726544000003
    Figure 0007726544000003
Patent Text Reader

Abstract

To provide a computer program, a server device, and a method for allowing a performer and the like to cause an avatar object to represent desired facial expressions or motions in an easy and accurate manner.SOLUTION: A computer program causes one or more processors to execute: retrieving a change amount of each of a plurality of specific portions of a body based on data regarding a motion of the body retrieved by a sensor; determining that a specific facial expression or motion is formed in a case where all change amounts for one or more specific portions previously specified, among the change amounts of the plurality of specific portions, exceed respective threshold values; and generating an image or movie in which a specific expression corresponding to the determined specific facial expression or motion is reflected on an avatar object corresponding to a performer.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology disclosed in the present application relates to computer programs and server equipment related to video distribution. It relates to the installation and method. [Background technology]

[0002] Conventionally, video distribution services that distribute videos to terminal devices via a network have been known. In this type of video distribution service, the distribution users (performers) who distribute the videos An environment is provided in which corresponding avatar objects are displayed.

[0003] In addition, in relation to video distribution services, the facial expressions and movements of avatar objects can be displayed in a way that matches the movements of actors, etc. This service, called "Custom Cast," uses technology that controls the content based on the user's actions. A service is known (Non-Patent Document 1). In this service, performers can For each of the multiple flick directions on the screen, a large number of facial expressions and actions are prepared. You can assign any of the facial expressions or movements to the camera in advance, and when streaming a video, you can select the desired facial expression or movement. The performer flicks the smartphone screen in the direction corresponding to the action. The avatar object displayed in the video can be made to express its facial expression or movement.

[0004] The above-mentioned Non-Patent Document 1 is incorporated herein by reference in its entirety. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] "Custom Cast", [online], Custom Cast Inc., [Retrieved December 10, 2019], Internet (URL: https: / / customcast.jp / ) Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the technology disclosed in Non-Patent Document 1, when distributing videos, The performer has to flick the smartphone screen while speaking, which is a hassle for the performer. It is difficult to perform the flick operation, and the flick operation is likely to be performed erroneously.

[0007] Therefore, some embodiments disclosed in the present application allow performers to easily and accurately A computer can make an avatar object express a desired facial expression or movement. A program, a server device, and a method are provided. [Means for solving the problem]

[0008] According to one aspect, a computer program is provided for execution by one or more processors. Based on the data on the body movement acquired by the sensor, The amount of change in each of the specific parts is acquired, and a predetermined amount of change in each of the multiple specific parts is selected. When all of the changes in at least one of the specific portions exceed the respective thresholds, It is determined that a specific facial expression or gesture has been formed, and a specific facial expression or gesture that has been determined is formed. An image or video in which a specific expression corresponding to the performer is reflected on an avatar object corresponding to the performer. The present invention causes the processor to function so as to generate an image.

[0009] A server device according to one aspect includes a processor, and the processor executes a The sensor executes instructions readable by the user to obtain information about the body's movements. Based on the data, the amount of change in each of the plurality of specific body parts is obtained, and the plurality of specific parts are Among the changes in each of the specific portions, each of at least one specific portion specified in advance is When all of the changes in the facial expression or gesture are greater than the respective thresholds, it is determined that a specific facial expression or gesture has been formed. The determined specific expression corresponding to the specific facial expression or gesture is displayed as an avatar of the performer. It generates an image or video that reflects the object.

[0010] In one aspect, a method includes one or more computer-readable instructions executing the instructions. A processor-implemented method for detecting body movements acquired by a sensor a change amount acquiring step of acquiring the change amount of each of the plurality of specific body parts based on the data; and at least one of the change amounts of the specific portions that is specified in advance. When all of the change amounts of the specific parts exceed the respective thresholds, a specific facial expression or behavior is formed. a determining step of determining that the specific facial expression or gesture has been made; An image or a video in which a specific expression corresponding to the performer is reflected on an avatar object corresponding to the performer. and a generating step of generating a moving image. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a communication system according to an embodiment. [Figure 2] FIG. 2 is a block diagram schematically illustrating an example of a hardware configuration of the terminal device (server device) illustrated in FIG. [Figure 3] FIG. 3 is a block diagram schematically illustrating an example of the functions of the studio unit shown in FIG. [Figure 4A] FIG. 4A is a diagram showing the relationship between a specific part identified in response to a specific facial expression "one eye closed (wink)" and its threshold value. [Figure 4B] FIG. 4B is a diagram showing the relationship between the specific part identified corresponding to the specific expression "smiling face" and its threshold value. [Figure 5] FIG. 5 is a diagram showing the relationship between a specific facial expression or gesture and a specific expression (specific action or facial expression). [Figure 6] FIG. 6 is a diagram schematically illustrating an example of a user interface unit. [Figure 7] FIG. 7 is a diagram schematically illustrating an example of a user interface unit. [Figure 8] FIG. 8 is a diagram schematically illustrating an example of a user interface unit. [Figure 9] FIG. 9 is a flow diagram showing an example of a part of the operations performed in the communication system shown in FIG. [Figure 10] FIG. 10 is a flow diagram showing an example of part of the operations performed in the communication system shown in FIG. [Figure 11] FIG. 11 is a diagram showing a modified example of the third user interface unit. DETAILED DESCRIPTION OF THE INVENTION

[0012] Various embodiments of the present invention will be described below with reference to the accompanying drawings. Components throughout the drawings are given the same reference numerals. Please note that some elements may be omitted in other drawings for clarity of illustration. Furthermore, the accompanying drawings are not necessarily drawn to scale. Furthermore, the term application refers to software or programs. It may also be called a program, which is an instruction to a computer that will produce some kind of result. It is sufficient that the combination be such that the desired effect can be obtained.

[0013] 1.Communication System Configuration FIG. 1 is a block diagram showing an example of the configuration of a communication system 1 according to an embodiment. As shown in FIG. 1, the communication system 1 includes one or more terminal devices connected to a communication network 10. 20 and one or more server devices 30 connected to the communication network 10. In addition, in FIG. 1, three terminal devices 20A to 20C are illustrated as examples of the terminal device 20. As an example of the server device 30, three server devices 30A to 30C are shown. As the terminal device 20, one or more terminal devices 20 other than these are connected to the communication network 10. Alternatively, one or more other server devices 30 may be used. The network 10 may be connected to the network.

[0014] The communication system 1 also includes one or more studio units connected to a communication network 10. 1 shows an example of the studio unit 40. Although studio units 40A and 40B are shown as examples, the studio unit 40 One or more other studio units 40 may be connected to the communication network 10 .

[0015] In the "first aspect," the communication system 1 shown in FIG. 1 is, for example, a studio room, etc. or the studio unit 40 installed in another location may be used in the studio room or other location. After acquiring physical data about the performers present, The amount of change in each of the plurality of body parts (specific parts) is acquired, and the change in each of the specific parts is calculated. When it is determined that all of the quantities exceed the respective thresholds, a predetermined specific expression is given to the performer. A video (or image) is generated that reflects the avatar object. The unit 40 transmits the generated video to the server device 30, and the server device 30 transmits the video to the studio unit. The video acquired (received) from the unit 40 is transmitted to one or more terminal devices via the communication network 10. 20, and a specific application (an application for watching videos) is executed and operated. The image can be distributed to the terminal device 20 that has transmitted a signal requesting image distribution.

[0016] Here, in the "first mode," the studio unit 40 instructs the performer to use a predetermined specific expression. A video that reflects the corresponding avatar object is generated and transmitted to the server device 30. Instead of the above configuration, the studio unit 40 collects data on the bodies of the performers and other Data on the amount of change in each of multiple specific parts of the performer's body based on the data (for the aforementioned judgment) The server device 30 transmits the data related to the studio unit 40 to the server device 30. According to the data received from the Alternatively, a rendering system may be used to generate a video that reflects the The studio unit 40 stores data on the body of the performer, etc., and the body of the performer, etc. based on the data. The data on the amount of change in each of the multiple specific parts of the body (the data on the aforementioned judgment) The server device 30 transmits the data received from the studio unit 40 to the terminal device 30. The terminal device 20 then sends the data received from the server device 30. Then, a video is generated in which a predetermined specific expression is reflected in an avatar object corresponding to the performer. A rendering method may be adopted.

[0017] In the "second embodiment," the communication system 1 shown in FIG. 1 is operated by, for example, a performer, etc. A terminal device 2 that executes a specific application (such as an application for video distribution) 0 (for example, terminal device 20A) acquires data relating to the body of a performer or the like facing terminal device 20A. Based on this data, each of multiple specific parts of the performer's body is and determining that all of the changes in the specific portions exceed the respective thresholds. Taking this as an opportunity, we reflected certain specific expressions in the avatar objects corresponding to the performers. Then, the terminal device 20A transmits the generated video (or image) to the server device 3. 0, and the server device 30 transmits the video acquired (received) from the terminal device 20A to the communication network 10. via one or more other terminal devices 20 to perform a specific application (video viewing) Terminal device 2 that executed the application for transmitting a signal requesting video distribution 0 (for example, terminal device 20C).

[0018] Here, in the "second mode," the terminal device 20 (terminal device 20A) is A video is generated in which the current state is reflected in an avatar object corresponding to the performer, and this is transmitted to a server device. Instead of transmitting data related to the body of a performer or the like to the terminal device 30, the terminal device 20 may transmit the data related to the body of a performer or the like to the terminal device 30. Data on the amount of change in each of multiple specific parts of the performer's body based on the data (the aforementioned judgment The server device 30 receives the data from the terminal device 20. According to the data, a specific expression is reflected in the avatar object corresponding to the performer. Alternatively, a rendering system may be employed to generate a moving image based on the video. 20 (terminal device 20A) acquires data on the body of a performer, etc., and Data on the amount of change in each of multiple specific parts of the body (data related to the aforementioned judgment) The server device 30 transmits the data received from the terminal device 20A to another One or more terminal devices 20 that execute a specific application and receive video distribution. The request signal is sent to the terminal device 20 (for example, terminal device 20C), and The terminal device 20C performs a predetermined specific expression according to the data received from the server device 30. A rendering method for generating video that reflects the user's avatar object may be adopted.

[0019] In the "third aspect," the communication system 1 shown in FIG. 1 is, for example, a studio room, etc. Alternatively, a server device 30 (for example, server device 30B) installed in another location may be used to After acquiring data on the bodies of performers in a room or other location, Based on the data, the amount of change in each of multiple parts (specific parts) of the performer's body is obtained, When it is determined that all of the changes in the specific portions exceed the respective thresholds, A video (or image) is generated in which a specific expression is reflected on an avatar object corresponding to the performer. Then, the server device 30B transmits the generated video to one or more The above terminal device 20 is configured to have a specific application (an application for watching videos) ) and deliver the video to the terminal device 20 that has transmitted the signal requesting the video delivery. In this "third aspect," as in the above, the server device 30 (server device 30 B) generates a video in which a specific expression is reflected in an avatar object corresponding to the performer. Instead of the configuration in which the server device 30 receives the information related to the body of the performer, etc., and transmits it to the terminal device 20, and data relating to the amount of change in each of a plurality of specific parts of the body of the performer, etc. based on the data. The terminal device 20 transmits the data (data related to the aforementioned determination) to the server. According to the data received from the avatar device 30, a predetermined specific expression is displayed on the avatar corresponding to the performer. A rendering system may be employed to generate a moving image that is reflected in the object.

[0020] The communication network 10 may be a mobile phone network, a wireless LAN, a fixed telephone network, the Internet, an intranet, or the like. may include, but are not limited to, Ethernet and / or It is something.

[0021] The aforementioned performers, etc., include not only performers but also those in studio rooms, etc. or other locations. This may include supporters who are with the performers, studio unit operators, etc.

[0022] The terminal device 20 executes the specific application installed thereon. After acquiring data on the physical condition of the performers, etc., The amount of change in each of a plurality of body parts (specific parts) is acquired, and the total amount of change in each of the specific parts is calculated. When it is determined that the performance exceeds each threshold, the system assigns a specific expression to the performer. Generate a video (or image) that reflects the butter object, and then use the generated video on the server. Alternatively, the terminal device 2 may perform an operation of transmitting the information to the terminal device 30. The web browser 0 executes the installed web browser, and the web server 30 accesses the web page. It can receive and display web pages, perform similar operations, and so on.

[0023] The terminal device 20 may be any terminal device capable of performing such operations. Smartphones, tablets, mobile phones (feature phones) and / or personal computers This may include, but is not limited to, a computer, etc.

[0024] In the "first aspect" and the "second aspect", the server device 30 It acts as an application server by running the application A predetermined specific expression is reflected on the avatar object from the geo unit 40 or the terminal device 20. The video is received via the communication network 10, and the received video (together with other videos) is transmitted via the communication network 10. 10 to each terminal device 20. The server device 30 also executes a specific application installed on it to access the web. By functioning as a server, the terminal device 20 can simultaneously It is possible to perform the same operations.

[0025] In the "third aspect," the server device 30 By executing the above and functioning as an application server, the server device 30 Acquiring physical data of performers in studio rooms or other locations Furthermore, based on this data, changes in multiple parts (specific parts) of the performer's body can be calculated. and determining that all of the change amounts of the specific portions exceed the respective thresholds. As a trigger, a video ( or images) can be generated, and the generated video (along with other videos) can be transmitted over a communication network 1 0 to each terminal device 20. In addition, the server device 30 executes a specific application installed therein to provide a website. By functioning as a server, the terminal device 20 can similarly communicate with the user via a web page sent to the terminal device 20. The above operations can be executed.

[0026] The studio unit 40 stores information to run the specific application installed. By functioning as a processing device, the studio unit 40 After acquiring data about the bodies of performers in a theater or other location, The amount of change in each of multiple parts (specific parts) of the performer's body is obtained based on the data, and the When it is determined that all of the change amounts of the specific portions exceed the respective thresholds, a predetermined Generate a video (or image) that reflects a specific expression on an avatar object corresponding to the performer. The generated video (together with other videos) can be transmitted to a server via a communication network 10. The operation can be performed such that the information is transmitted to the device 30.

[0027] 2. Hardware configuration of each device Next, the hardware included in each of the terminal device 20, the server device 30, and the studio unit 40 will be described. An example of the hardware configuration will be described below.

[0028] 2-1. Hardware configuration of terminal device 20 An example of the hardware configuration of each terminal device 20 will be described with reference to FIG. 2 is a block diagram schematically illustrating an example of a hardware configuration of the terminal device 20 shown in FIG. (Note that in FIG. 2, the reference numerals in parentheses indicate the respective server devices 30, as will be described later.) (These are the documents attached as a result of the application.)

[0029] As shown in FIG. 2, each terminal device 20 mainly comprises a central processing unit 21, a main memory device 22, and , an input / output interface 23, an input device 24, an auxiliary storage device 25, and an output device 26. These devices are connected to each other by a data bus and / or a control bus. It is being done.

[0030] The central processing unit 21 is called a "CPU" and is stored in a main memory 22. The result of the calculation is stored in the main memory device 22. Furthermore, the central processing unit 21 receives input data from an input device via an input / output interface 23. The terminal device 20 can control the computer 24, the auxiliary storage device 25, and the output device 26. , may include one or more such central processing units 21.

[0031] The main memory device 22 is referred to as a "memory" and is connected to the input device 24, the auxiliary memory device 2 5 and the communication network 10 (server device 30, etc.) via the input / output interface 23. It stores the instructions and data entered and the results of the calculations performed by the central processing unit 21. The storage device 22 includes a RAM (random access memory), a ROM (read only memory), and The memory may include, but is not limited to, a memory card and / or flash memory.

[0032] The auxiliary storage device 25 is a storage device having a larger capacity than the main storage device 22. Specific applications (video streaming applications, video viewing applications, etc.) It stores instructions and data (computer programs) that make up the browser, etc. These instructions and data (computer instructions) are stored in the memory and controlled by the central processing unit 21. A computer program is transmitted to the main memory 22 via the input / output interface 23. The auxiliary storage device 25 may be a magnetic disk device and / or an optical disk device. These may include, but are not limited to:

[0033] The input device 24 is a device for inputting data from the outside, and includes a touch panel, buttons, keys, etc. This includes, but is not limited to, a keyboard, a mouse, and / or a sensor. As will be described later, one or more cameras and / or one or more microphones The sensors may include, but are not limited to:

[0034] The output device 26 may include a display device, a touch panel, and / or a printer device. This can include, but is not limited to:

[0035] In such a hardware configuration, the central processing unit 21 stores the data in the auxiliary storage device 25. The instructions and data that make up a specific application stored in the are sequentially loaded into the main memory 22, and the loaded instructions and data are operated to Controlling an output device 26 via an output interface 23, or alternatively, an input / output interface via the device 23 and the communication network 10, other devices (for example, a server device 30, a studio unit 40 and other terminal devices 20, etc.).

[0036] As a result, the terminal device 20 executes the specific application that has been installed. By doing so, we can obtain data on the physical condition of the performers, etc., and then use this data to The amount of change in each of a plurality of parts (specific parts) of the body of the performer, etc. is acquired, and each of the specific parts is When all of the changes in the Generate a video (or image) that reflects the butter object, and then use the generated video on the server. Alternatively, the terminal device 2 may perform an operation of transmitting the information to the terminal device 30. The web browser 0 executes the installed web browser, and the web server 30 accesses the web page. It can receive and display web pages, perform similar operations, and so on.

[0037] The terminal device 20 may be used in place of the central processing unit 21 or in addition to the central processing unit 21. or higher microprocessor and / or graphics processing unit The device may include a graphics processing unit (GPU).

[0038] 2-2. Hardware configuration of server device 30 An example of the hardware configuration of each server device 30 will be described with reference to FIG. The hardware configuration of each server device 30 may be the same as that of each terminal device 20 described above. Therefore, the same hardware configuration as that of each server device 30 can be used. The reference numerals for the components included in the .

[0039] As shown in FIG. 2, each server device 30 mainly includes a central processing unit 31 and a main memory device 32. an input / output interface 33, an input device 34, an auxiliary storage device 35, and an output device 36 These devices are connected to each other by a data bus and / or a control bus. It has been done.

[0040] Central processing unit 31, main memory 32, input / output interface 33, input device 34, auxiliary The storage device 35 and the output device 36 are respectively included in the central A processing unit 21, a main memory unit 22, an input / output interface 23, an input unit 24, an auxiliary memory unit The input device 25 and the output device 26 can be substantially the same.

[0041] In such a hardware configuration, the central processing unit 31 stores the data in the auxiliary storage device 35. The instructions and data that make up a specific application stored in the are sequentially loaded into the main memory 32, and the loaded instructions and data are operated to Controlling an output device 36 via an output interface 33, or alternatively, an input / output interface The device 33 and the communication line 10 are connected to other devices (for example, each terminal device 20 and studio user It is possible to send and receive various information between the device and the network (such as the network unit 40).

[0042] As a result, the server device 30 can perform the installation in the "first mode" and the "second mode". It acts as an application server, running specific applications installed on it. As a result, a predetermined specific expression is displayed on the avatar object from the studio unit 40 or the terminal device 20. The video reflected in the project is received via the communication network 10, and the received video is displayed (in combination with other videos). In addition, the operation of distributing the information to each terminal device 20 via the communication network 10 can be performed. Alternatively, the server device 30 may execute a specific application that has been installed. By functioning as a web server, the web page to be transmitted to each terminal device 20 Similar operations can be performed via

[0043] In the "third aspect," the server device 30 also By executing the application and functioning as an application server, this server device 3 Acquire physical data of performers in studio rooms where 0 is installed or other locations Then, based on this data, each of the performer's body parts (specific parts) When all of the changes in the specific parts exceed the thresholds, Then, a video (or image) is created in which a predetermined specific expression is reflected in an avatar object corresponding to the performer. The generated moving image (together with other moving images) can be transmitted via the communication network 10. and distributes the data to each terminal device 20. The server device 30 runs a specific application installed on it and acts as a web server. By functioning as a web page, the same operation can be performed via a web page transmitted to each terminal device 20. etc. can be carried out.

[0044] The server device 30 may be used in place of or in addition to the central processing unit 31. One or more microprocessors and / or graphics processing units It may also include a graphics processing unit (GPU).

[0045] 2-3.Hardware configuration of Studio Unit 40 The studio unit 40 can be implemented by an information processing device such as a personal computer. Although not shown, the terminal device 20 and the server device 30 are the same as those described above. The main components are a central processing unit, a main memory, an input / output interface, an input device, and an auxiliary The device may include a storage device and an output device. These devices may be connected to each other via a data bus and / or are connected by a control bus.

[0046] The Studio Unit 40 runs certain installed applications to collect information. By functioning as a processing device, the studio unit 40 After acquiring data about the bodies of performers in a theater or other location, The amount of change in each of multiple parts (specific parts) of the performer's body is obtained based on the data, and the When all of the changes in the specific parts exceed the respective thresholds, a predetermined specific expression is executed. It is possible to generate a video (or image) that reflects the user's avatar object. The generated video (together with other videos) is transmitted to the server device 30 via the communication network 10. It is possible to perform operations such as:

[0047] 3. Functions of each device Next, the functions of the studio unit 40, the terminal device 20, and the server device 30 are as follows: An example of this will be described.

[0048] 3-1. Functions of Studio Unit 40 An example (one embodiment) of the function of the studio unit 40 will be described with reference to FIG. FIG. 3 is a block diagram showing an example of the functions of the studio unit 40 shown in FIG. (Note that in FIG. 3, the reference numerals in parentheses indicate the terminal device 20 and the server device 21, as will be described later.) (These are attached in relation to the server device 30.)

[0049] As shown in FIG. 3, the studio unit 40 collects data on the bodies of performers and the like from sensors. and a sensor unit 100 for acquiring data on the performer's body based on the data acquired from the sensor unit 100. a change amount acquiring unit 110 for acquiring the change amount of each of a plurality of specific body parts; The total of the change amounts of at least one specific portion specified in advance among the respective change amounts. If it is determined that the results exceed the thresholds, the performer will be given a special a determination unit 120 that determines that a specific facial expression has been formed; The specific expressions corresponding to the facial expressions of the performers are reflected in the avatar objects corresponding to the performers. A generating unit 130 that generates a video (or image) may be included.

[0050] Furthermore, the studio unit 40 allows the performer or the like to set each of the above thresholds as appropriate. The system may further include a user interface unit 140 that can

[0051] Furthermore, the studio unit 40 may also generate a video (or image) a display unit 150 that displays the moving image, and a storage unit 16 that stores the moving image generated by the generating unit 130. 0 and the video generated by the generation unit 130 is transmitted to the server device 30 via the communication network 10. and a communication unit 170 for communicating with the

[0052] (1) Sensor unit 100 The sensor unit 100 is disposed in, for example, a studio room (not shown). In this event, performers perform various performances, and the sensor unit 100 detects the movements and expressions of the performers. , and speech (including singing), etc.

[0053] The performers' movements, facial expressions, and speech (singing) are monitored by various sensors in the studio room. In this case, the recording will be recorded in the studio room. There may be one or more actors in the song.

[0054] The sensor unit 100 includes one or more sensors that acquire data about the performer's face, hands, feet, and other parts of the body. a first sensor (not shown) for detecting speech and / or singing uttered by the performer; and one or more second sensors (not shown) for acquiring data.

[0055] In a preferred embodiment, the first sensor is an RGB camera that captures visible light and a near-infrared camera. and a near-infrared camera for capturing an image of the line. , and can include a motion sensor and a tracking sensor, which will be described later. For example, the iPhone X (registered trademark) TrueDepth camera is used as a near-infrared camera. It is possible to use the one included in the True Depth camera. The sensor may include a microphone for recording sound.

[0056] Regarding the first sensor, the sensor unit 100 is a second sensor placed close to the face, hands, feet, etc. of the performer. The first sensor (the camera included in the first sensor) is used to capture images of the performer's face, hands, feet, etc. As a result, the sensor unit 100 can record the image acquired by the RGB camera as a time code (acquisition The data recorded over a unit time interval (e.g., M PEG file). Furthermore, the sensor unit 100 can generate a near-infrared camera. A number (e.g., a floating-point number) indicating the depth of a given number (e.g., 51) obtained by The data recorded over a unit time period (for example, a TSV file) is Generate a file (a file in which multiple data are recorded with tabs separating the data) It is possible.

[0057] Regarding near-infrared cameras, specifically, a dot projector projects a dot pattern. The infrared laser is emitted onto the face, limbs, etc. of the performer, and the near-infrared camera captures the It captures the infrared dots projected onto the surface and reflected by it, and generates an image of the infrared dots captured in this way. The sensor unit 100 detects dots emitted by a dot projector that has been registered in advance. The image of the pattern was compared with the image captured by the near-infrared camera. Positional deviation at each point (each feature point) (e.g., each of 51 points / feature points) The depth of each point (each feature point) (the distance between each point and the near-infrared camera) is calculated using The sensor unit 100 calculates the depth calculated in this way. As shown above, the numerical values are associated with the time code to generate data recorded over a unit time. It is possible.

[0058] The sensor unit 100 in the studio room detects the body of the performer (for example, the wrist, the instep, Various motion sensors (not shown) are attached to the performer's waist, top of the head, etc., and sensors held in the performer's hands are used. Furthermore, the studio room may have a controller (not shown) for controlling the In addition to the components described above, the system includes a plurality of base stations (not shown) and a tracking system. It may also have a sensor (not shown) or the like.

[0059] The motion sensor, in cooperation with the base station, determines the position and orientation of the performer. In one embodiment, multiple base stations are used in a multi-axis After emitting a flashing light for synchronization, one base station For example, one base station may scan a laser beam around a vertical axis, while another base station may scan a laser beam around a horizontal axis. The motion sensor is configured to scan the laser beam around the base station. It is equipped with multiple optical sensors that detect the incidence of flashing light and laser light from the time difference between the timing of the laser light being received and the timing of the laser light being incident, the time it takes for each optical sensor to receive the light, The motion sensor can detect the angle of incidence of the detected laser light, etc. For example, Vive Tracker provided by HTC CORPORATION. Alternatively, you can use the Xsens M provided by ZERO C SEVEN Inc. It may also be VN Analyze.

[0060] The sensor unit 100 detects the position and The motion sensors can detect the movements of the performer's wrists, insteps, and feet. By attaching it to the waist, top of the head, etc., the position and orientation of the motion sensor can be detected. The motion sensor can detect the movements of each part of the performer's body. The detected information indicating the position and orientation is the body of the performer in the video (in the virtual space included in the video). The position coordinates of each part are calculated in the XYZ coordinate system. For example, the X axis is The Y axis is the depth direction in the video, and the Z axis is the vertical direction in the video. Therefore, the movements of each part of the performer's body are all set to X It is detected as a position coordinate value in the YZ coordinate system.

[0061] In one embodiment, multiple motion sensors are equipped with multiple infrared LEDs. The light from the external LEDs is detected by infrared cameras installed on the floor and walls of the studio room. The position and orientation of the motion sensor may be detected by using a light source other than the infrared LED. A visible light LED is used, and the light from this visible light LED is detected by a visible light camera. The position and orientation of the motion sensor may be detected.

[0062] In one embodiment, the motion sensors are replaced by multiple reflective markers. Reflective markers can also be attached to the performer with adhesive tape. The performer with the car attached is photographed to generate photographic data, which is then processed into images. As a result, the position and orientation of the reflective marker (as mentioned above, the position coordinates in the XYZ coordinate system) The configuration may be such that the value is detected.

[0063] The controller outputs control signals according to the performer's finger movements, etc. , which is acquired by the generation unit 130.

[0064] The tracking sensor acquires the setting information of the virtual camera to construct the virtual space included in the video. The tracking information is used to determine the three-dimensional orthogonal coordinates. The tracking information is calculated as a position in the coordinate system and an angle around each axis. Get.

[0065] Next, regarding the second sensor, the sensor unit 100 detects the second sensor located close to the performer. The audio of the speech and / or singing uttered by the performer is acquired using a sensor. Therefore, the sensor unit 100 records data over a unit time in association with the time code. In one embodiment, the sensor unit 100 can generate a file (e.g., an MPEG file). The system uses a first sensor to acquire data about the performer's face and limbs, and simultaneously uses a second sensor to acquire data about the performer's face and limbs. The sensor acquires audio data relating to speech and / or singing uttered by a performer. In this case, the sensor unit 100 can detect the image captured by the RGB camera. and audio data relating to speech and / or singing uttered by the performer using a second sensor. and data recorded over a unit time period in association with the same time code (e.g. MP EG file).

[0066] The sensor unit 100 receives the motion data (M PEG files and TSV files, etc.), and data on the position and orientation of each part of the performer's body and audio data (MPEG files) relating to the speech and / or singing uttered by the performers. The generated image data (such as a rule) can be output to the generating unit 130 described later.

[0067] In this way, the sensor unit 100 detects M PEG files and other video files, as well as the positions (coordinates, etc.) of the performers' faces, hands, and feet, are stored as data related to the performers. It can be obtained as data.

[0068] According to such an embodiment, the sensor unit 100 detects, for example, the presence or absence of a subject on the face, hands, or feet of the performer. For each part, MPEG files captured for each unit time section and Specifically, the sensor unit 100 can acquire data including the position (coordinates). For each unit time interval, for example, the right eye includes information indicating the position (coordinates) of the right eye. For example, with respect to the upper lip, information indicating the position (coordinates) of the upper lip can be included.

[0069] In another preferred embodiment, the sensor unit 100 is It is possible to use a technique called Argumented Faces. is disclosed at https: / / developers.google.com / ar / develop / java / augmented-faces / are available and are incorporated herein by reference in their entirety.

[0070] The sensor unit 100 detects the facial and limbs of the performer, which are generated as described above. Among them, the motion data (MPEG files, TSV files, etc.) related to several specific parts, The image can be further output to the change amount acquisition unit 110, which will be described later. refers to any part of the body, such as the head, part of the face, shoulders (or clothing covering the shoulders) ), and limbs, etc. More specifically, parts of the face, such as the forehead, eyebrows, This may include, but is not limited to, the eyelids, cheeks, nose, ears, lips, mouth, tongue, and chin.

[0071] The sensor unit 100 detects the movements, facial expressions, speech, etc. of the performers present in the studio room. As explained above, in addition to this, there is a need to be present with the performers in the studio room. The motions and expressions of the supporters and the operators of the studio unit 40 are detected. In this case, the sensor unit 100 may detect the face, hands, feet, etc. of the supporter or operator. Data relating to multiple specific body parts (MPEG files, TSV files, etc.) ) may be output to the change amount acquisition unit 110 described later.

[0072] (2) Change Amount Acquisition Unit 110 The change amount acquisition unit 110 acquires the performance data of the performer (as described above, Based on data on the body movements of the performer (which may be a computer or operator), The change amount (displacement amount) of each of a plurality of specific body parts is acquired. Specifically, 110 is a position acquired in unit time section 1 for a specific part, for example, the right cheek. The difference between the position (coordinate) of the current position and the position (coordinate) acquired in unit time interval 2 is calculated. Therefore, the amount of change in the specific part, the right cheek, between unit time interval 1 and unit time interval 2 is The change amount acquiring unit 110 can similarly acquire the specific values for other specific portions. The amount of change in the part can be obtained.

[0073] The change amount acquiring unit 110 acquires the change amount of each specific portion in an arbitrary unit time. The position (coordinates) acquired in a unit time interval and the position (coordinates) acquired in another unit time interval are The difference between the coordinates can be used. , variable, or a combination thereof.

[0074] (3) Determination unit 120 Next, the determination unit 120 will be described with reference to FIGS. 4A and 4B. The relationship between the specific parts identified in response to the facial expression "Closing one eye (wink)" and their thresholds FIG. 4B shows a specific part identified in correspondence with a specific facial expression "smiling face" and the corresponding 10 is a diagram showing the relationship between the thresholds.

[0075] The determination unit 120 determines the change in each of the plurality of specific portions acquired by the change amount acquisition unit 110. The amount of change in each of at least one specific portion specified in advance is determined to be within each threshold. If it is determined that the value is exceeded, a specific facial expression is displayed by the performer. Specifically, the determination unit 120 determines that the specific facial expression is formed, for example, "Smiling face", "Closing one eye (wink)", "Surprised face", "Sad face", "Angry face", "Scheming face," "Bashful face," "Closing both eyes," "Stick out tongue," "Open mouth," "Cheeks Use facial expressions such as, but not limited to, "pushing one's head," "opening one's eyes," and "opening one's eyes" For example, gestures such as "shaking one's shoulders" or "shaking one's head" can be expressed as specific expressions. However, these specific facial expressions and specific The gestures are consciously performed by the performer (who, as mentioned above, may be a supporter or operator). It is preferable that the determination unit 120 determines only the facial expressions (or gestures) of the performer. In order to prevent erroneous judgments that were not intentionally made by the performers, It is important to select expressions that do not overlap with the various performances and facial expressions used during speech. preferable.

[0076] The determination unit 120 determines at least one location corresponding to each of the specific facial expressions (or specific gestures) described above. The amount of change in the specific portion is specified in advance. Specifically, as shown in FIG. 4A, for example, If the specified facial expression is "close one eye (wink)", the eyebrow (right eyebrow or left eyebrow), eyelid (right eyelid or left eyelid) eyelids), eyes (right or left), cheeks (right or left), and noses (right or left) For example, the amount of change can be obtained by using the following: The specific parts can be the eyebrows, right eyelid, right eye, right cheek, and nose. For example, if a specific facial expression is a "smiling face," the mouth (right or left), lips (right or left of the lower lip), The amount of change is obtained by specifying the area of the face (left side) and the inside of the eyebrows (or forehead).

[0077] Furthermore, as shown in FIGS. 4A and 4B, a predetermined facial expression corresponding to the specific facial expression is generated. A threshold is set for each of the changes in the specific parts. In the case of "Close eyes (wink)", the threshold of the eyebrow change (downward movement) is set to 0.7, and the eyelid change ( The threshold for the eye change (amount of narrowing) is 0.9, the threshold for the eye change (amount of narrowing) is 0.6, and the threshold for the cheek change The threshold for the amount of change (amount of rise) of the nose is set to 0.4, and the threshold for the amount of change (amount of rise) of the nose is set to 0.5. Similarly, if the specific facial expression is a "smile face," the threshold for the change in the mouth (lift) is set to 0.4, and the threshold for the change in the lower lip is set to 0. The threshold for the amount of change (downward movement) is 0.4, and the threshold for the amount of change (upward movement) inside the eyebrows is 0.1. The values of these thresholds are set via the user interface unit 140 as will be described later. The amount of narrowing of the eyes can be set appropriately by adjusting the amount of narrowing of the eyes. For example, it is the amount by which the distance between the upper and lower eyelids has decreased.

[0078] In addition, the specific parts corresponding to specific facial expressions can also be changed as appropriate. As shown in Figure 4A, when a specific facial expression is "close one eye (wink)", the eyebrows, eyelids, eyes, Five areas, including the cheeks and nose, may be specified in advance as specific areas. However, only the three areas of the face, the eyes, and the face may be specified in advance. Only facial expressions (or gestures) consciously performed by the person (who may be a supporter or operator) are recorded. It is preferable that the decision is made by the decision unit 120. Therefore, if the performer has consciously performed In order to prevent erroneous judgments, it is preferable to have a large number of specific parts corresponding to specific facial expressions. It's nice.

[0079] In this way, the determination unit 120 determines, for example, that "closing one eye (wink)" The changes in the eyebrows, eyelids, eyes, cheeks, and nose as specific parts acquired by the change amount acquisition unit 110 When all of these changes exceed the thresholds, the system will close one eye ( The "ink" is formed by the performer (who, as mentioned above, may be a supporter or operator). In this case, it is determined that all of the changes actually exceed the respective threshold values. Even if the determination unit 120 determines that "one eye closed (wink)" has been formed at the time of turning, Alternatively, the state in which all of the amounts of change actually exceed the respective thresholds described above may be maintained for a predetermined period of time (for example, 1 second or 2 seconds). The additional condition of "closing one eye (wink)" was added, and the condition of "closing one eye (wink)" was formed. By adopting the latter mode, the determination unit 120 may determine that the This makes it possible to efficiently avoid erroneous determinations due to the above.

[0080] In addition, when the determination unit 120 makes the above-mentioned determination, the determination unit 120 Judgment result (for example, the judgment result that "closing one eye (wink)" was formed by the performer) The determination unit 12 outputs information (signals) related to the result to the generation unit 130. The information of the determination result output from the generating unit 130 to the generating unit 130 includes, for example, the change in each specific part. Information showing the amount of change, and specific facial expressions formed when the amount of change in each specific part exceeds each threshold. or indicates that a decision has been made to reflect specific expressions corresponding to the actions in the avatar object. The avatar object then receives a specific expression corresponding to the cue and the specific facial expression or gesture formed. The ID of a specific expression (also called the "special expression ID") as information that requests that it be reflected in the ), and at least one of

[0081] Here, the relationship between a specific facial expression or gesture and a specific expression (specific action or facial expression) is shown in Figure 1. 5. FIG. 5 shows a specific facial expression or gesture and a specific expression (specific action or facial expression). ) is a diagram showing the relationship.

[0082] The relationship between a specific facial expression or gesture and a specific expression (specific action or facial expression) is the same relationship, similarity, The relationship can be appropriately selected from either a relationship that is related to the relationship or a relationship that is completely unrelated. For example, for a specific facial expression such as "close one eye (wink)" as in Specific Expression 1 in Figure 5, The corresponding specific expression may be the same as "close one eye (wink)." ,As shown in Figure 5, specific expression 2 corresponds to a specific facial expression "smiling face", and "Close your eyes (wink)" corresponds to "kick up your right leg", "Sad face" corresponds to " It can also be unrelated, such as "sleeping," "closing one eye," etc. It may be a "sad face" or similar. Furthermore, a "scheming face" may be used in response to a "smiling face." Furthermore, the same relationship, similar relationship, and completely unrelated relationship may be used. In this case, a cartoon-like image may be used as the specific expression. Use it as a trigger to reflect a specific expression on an avatar object. can be done.

[0083] The relationship between a specific facial expression or gesture and a specific expression (specific action or facial expression) will be explained later. The settings are changed appropriately via the user interface unit 140.

[0084] (4) Generation unit 130 The generating unit 130 receives motion data (MP) relating to the face, limbs, etc. of the performer from the sensor unit 100. EG files and TSV files, etc.), data on the position and orientation of each part of the performer's body, and audio data (such as MPEG files) relating to the speech and / or singing of the performers. ) based on the performance data, a video containing animations of the avatar object corresponding to the performer is generated. As for the animation of the avatar object itself, the generation unit 130 generates the animation as shown in FIG. Various information (e.g., geometry information, button information, texture information, shader information, blendshape information, etc.) By having the rendering part that does not use the avatar object perform the rendering, You can also generate videos.

[0085] Furthermore, when the generating unit 130 acquires the information on the aforementioned determination result from the determining unit 120, the generating unit 130 The specific expressions corresponding to the information of the determination results are generated as videos of the avatar object generated as described above. Specifically, for example, the determination unit 120 may determine whether the user has closed one eye ( A specific facial expression or gesture called "wink" is formed by the performer, and the corresponding "one-eyed" Generate an ID for the specific expression "close (wink)" (or information about the aforementioned cue). When the generating unit 130 receives the “close one eye (win A video (or image) in which a specific expression "" is reflected in an avatar object corresponding to the performer. image).

[0086] Incidentally, the generating unit 130 may perform the following operations regardless of whether or not the information on the determination result of the determining unit 120 is acquired. As mentioned above, the sensor unit 100 outputs motion data (MPEG files and TSV files), data on the position and orientation of each part of the performer's body, and Audio data (MPEG files, etc.) relating to the speech and / or singing of the performers Based on the obtained data, a video including animation of the avatar object corresponding to the performer is generated. (For convenience, this moving image is referred to as the "first moving image"). When obtaining the information on the aforementioned determination result, the generation unit 130 receives the face of the performer from the sensor unit 100. Motion data of the performer's hands and feet (MPEG files, TSV files, etc.) Data on the position and orientation of body parts, and sounds related to speech and / or singing uttered by the performer Based on the voice data (MPEG file, etc.) and the information of the determination result received from the determination unit 120, Generate a video (or image) that reflects a specific expression on an avatar object. (For convenience, we will refer to this video as "Video 2").

[0087] (5) User Interface Unit 140 Next, the user interface unit 140 will be described with reference to FIGS. 6 to 8 are diagrams showing an example of the user interface unit 140.

[0088] The user interface section 140 in the studio unit 40 is displayed on the display section 150. The above-mentioned video (or image) is transmitted to the server device 30, and various information related to the above-mentioned thresholds, etc. Various information can be input through the operations of the performers, etc., and various information can be shared visually with the performers, etc. It is possible.

[0089] For example, the user interface unit 140 may be configured to recognize a specific facial expression or gesture as shown in FIG. The corresponding threshold values for the specific parts can be set (changed). The user interface unit 140 is configured to provide a specific portion (for example, the right side of the mouth, the left side of the mouth in FIG. 6) , the right side of the lower lip, the left side of the lower lip, and the forehead. In FIG. 6, the display mode of these specific parts is The slider 141a of the display unit 150 is displayed in a manner that emphasizes the slider 141a by using a font, color, or the like. The threshold value can be adjusted appropriately based on the touch operation on the screen and can be changed to any value between 0 and 1. In FIG. 6, when a "smiling face" is set as a specific facial expression, In this case, the specific parts described in FIG. 4B, the right side of the mouth (rising), the left side of the mouth (rising), and the lower lip The thresholds for the right side (downward), left side of the lower lip (downward), and forehead (upward) were set to 0.4 or 0.1. These threshold values can be changed by operating the slider 141a. This slider 141a can be conveniently displayed on the first user interface 14. 1. In addition, in FIG. 6, the right side of the mouth (downward) and the left side of the mouth (downward) are targets for setting thresholds. Since the slider 141a is not displayed in these areas, the slider 141a is not displayed in these areas. In other words, when setting a threshold, it is necessary to identify a specific part that corresponds to a specific facial expression or behavior. After determining the amount of change, it is necessary to further identify the type of change (increase, decrease, etc.). As shown in FIG. 6, the user interface unit 140 has a right-hand side (lowering) and a left-hand side. On the side (downward), not only the slider 141a but also the right side (downward) and the left side (downward) ) tab itself is not displayed on the user interface unit 140 (display unit 150), Alternatively, a dedicated slider 141x may be provided separately. The specific portion and the corresponding threshold and slider 141a are displayed on the screen. Alternatively, a dedicated slider 141y may be provided separately to enable selection without using the slider. The riders 141x and 141y are an example of an operation unit that switches the display mode.

[0090] As mentioned above, the specific part corresponding to the specific facial expression is also input to the user interface unit 14. 0 (first user interface unit 141) can be changed as appropriate. For example, As shown in Figure 6, when the specific facial expression is a "smiling face," the specific parts are the right side of the mouth, the left side of the mouth, To change the location from five (right side of lower lip, left side of lower lip, and forehead) to four (with the forehead removed), click " By clicking on the "Forehead Raise" tab, you can change the specific parts that correspond to the "Smiling Face" It is possible.

[0091] The user interface unit 140 also determines the threshold values of the specific parts corresponding to the specific facial expressions. to a predetermined value automatically without operating the slider 141a. Specifically, for example, two modes may be prepared in advance. Then, the two modes are selected based on a selection operation on the user interface unit 140. When either of the above is selected, the thresholds (predetermined values) corresponding to the selected mode are automatically set. In this case, in FIG. 6, "easy to appear" and Two modes are prepared: "hard to get out" and "hard to get out." By touching the screen, you can select either "easy" or "hard" mode. It is possible to select "easy to appear" and "hard to appear" in Figure 6. The corresponding tab is referred to as a second user interface section 142 for convenience. The interface unit 142 can be thought of as a set menu in which the values of each threshold are predetermined. do.

[0092] By the way, in the aforementioned "easy to appear" mode, each threshold is generally low (for example, For example, the right side of the mouth, the left side of the mouth, the right side of the lower lip, and the left side of the lower lip are specific parts of a specific facial expression, "smiling face." Each threshold is set to a value less than 0.4 and the forehead threshold is set to a value less than 0.1 This allows the frequency with which the determination unit 120 determines that a "smiling face" has been formed by the performer, etc. On the other hand, it is possible to increase the number of times ... In the "good" mode, each threshold is set to a high value overall (for example, a specific facial expression "smiling face" is set to a high value). The specific parts of the right side of the mouth, the left side of the mouth, the right side of the lower lip, and the left side of the lower lip have thresholds greater than 0.4. The threshold value for the forehead is set to a value greater than 0.1. The frequency with which the determination unit 120 determines that a "smiling face" has been formed is reduced, or the determination unit 120 The determination can be made in a limited manner.

[0093] In addition, the threshold values (predetermined values) that are predetermined in the "likely to appear" mode are The value may be different for each minute, or may be the same for at least two specific parts. Specifically, for example, the right and left sides of the mouth, which are specific parts of a specific facial expression "smiling face", The thresholds for the right and left lower lip may be set to 0.2, and the threshold for the forehead may be set to 0.05. The threshold of the left side of the mouth is 0.3, the threshold of the right side of the lower lip is 0.01, and the threshold of the left side of the lower lip is The threshold for the upper and lower eyelids may be set to 0.2, and the threshold for the lower eyelids may be set to 0.05. than the default value at the time a specific application is installed on unit 40. is set small.

[0094] Similarly, the threshold values (predetermined values) that are predetermined in the "hard to come out" mode are also It may be a different value for each part, or it may be the same value for at least two specific parts. Specifically, for example, the right and left sides of the mouth, which are specific parts of a specific facial expression "smiling face," The thresholds for the right and left lower lip may be set to 0.7, and the threshold for the forehead may be set to 0.5. The threshold of the left side of the mouth is 0.7, the threshold of the right side of the lower lip is 0.8, the threshold of the left side of the lower lip is 0.6, 0.9, and the forehead threshold can be set to 0.3. Alternatively, from the "easy to appear" mode, To change to "difficult to come out" mode (or vice versa), use the right side of the mouth, the left side of the mouth, and the bottom A specific part of the right side of the lip, the left side of the lower lip, and a specific part of the forehead (for example, the left side of the lower lip and the forehead) The threshold for minutes is the specified value for the "easy" mode (or the specified value for the "hard" mode). ) can also be used as is.

[0095] As the second user interface section 142, referring to FIG. 6, As explained above, there are two modes (tabs) for "hard to appear" and "hard to appear", but For example, three or more modes (tabs) may be provided. For example, "Normal", Three modes may be provided: "easy to appear", "very easy to appear", "normal", Four modes may be provided: "easy to get", "very easy to get", and "extremely easy to get". In these cases, the value of each threshold may be adjusted to suit the particular application of the studio unit 40. It may be set to a value less than the default value when the application is installed, or it may be set to a value less than the default value. May be set higher than the default value.

[0096] Also, as the second user interface unit 142, A tab for disabling the operation by the unit 142 may be provided. In FIG. 6, a tab for "disable" is provided. When this tab is touched, the performer or other person will be able to access the first user interface. The threshold value is set appropriately using only the unit 141.

[0097] In addition, the first user interface unit 141 or the second user interface unit 142 The tab for resetting all the thresholds set in the above to the default values is available in the user interface. Alternatively, it may be provided separately in the surface unit 140.

[0098] The reason for appropriately setting (changing) the values of each threshold in this way is that Of course, there are individual differences among performers, and some people tend to form certain facial expressions (or certain expressions). On the other hand, another person may be more likely to be judged by the judgment unit 120 as having formed a relationship with the particular person. Therefore, no matter who the person is, In order for the determination unit 120 to accurately determine that a specific facial expression has been formed, It is preferable to reset each threshold value (preferably each time the person to be judged changes).

[0099] Furthermore, each time a person related to the performer to be judged changes, the threshold (amount of change) is reset to the initial value. As shown in FIG. 6, the threshold value for any particular part is preferably set to the value of the particular part. If there is no change in the part, the standard is 0, and the maximum change in the specific part is 1. In this case, the threshold is set appropriately between 0 and 1. Then, the threshold value for a certain person X is different from the standard value of 0 to 1. The range of the standard 0 to 1 for person Y is different (for example, in the case of person X, the standard 0 to 1 for person Y is different). (In this case, the maximum change amount for person Y may be equivalent to only 0.5 for person X.) Therefore, in order to express the amount of change in a specific part of all people as 0 to 1, It is preferable to set the width to the initial setting (multiplying by a predetermined magnification). The initial settings will be performed by touching the "rate" tab.

[0100] The user interface unit 140 inputs the values of the thresholds in the first user interface as described above. It can be set in both the first user interface section 141 and the second user interface section 142. This configuration allows, for example, fast video distribution without being concerned with detailed threshold settings. For performers who want to test their skills, the second user interface unit 142 can be used. On the other hand, performers who are particular about setting the thresholds in detail can set the first user interface corresponding to each threshold. By operating the slider 141a on the interface 141, you can customize the threshold value to your own specifications. By using such a user interface unit 140, it is possible to change the settings based on the preferences of the performer, etc. Each threshold can be set as desired, making it easy for performers to use. Furthermore, for example, the second user interface unit 142 can be used to select a predetermined mode (e.g., For example, after setting the mode of "easy to appear", The slider 141a can also be operated, so the user interface unit 140 This can also improve the variety of ways of use.

[0101] The user interface unit 140 can also be used to set various values and information other than the threshold values. For example, the user interface unit 140 may set or change the Regarding the determination operation by 120, all of the changes in the specific parts corresponding to the specific facial expressions exceed the respective thresholds. If the condition is that the actual exceedance continues for a specified period of time (for example, 1 or 2 seconds), , a user interface (not shown in FIG. 6) for setting the predetermined time However, for example, a slider may be separately included. The specific expressions corresponding to the specific facial expressions are recorded as videos of the avatar objects corresponding to the performers ( The user interface section also determines the time (for example, 5 seconds) to be reflected in the image. 140 (not shown in FIG. 6, for example, sliders 141x and 141y You can set (change) the value as needed using a different slider.

[0102] Furthermore, the user interface unit 140 may be configured to display the specific facial expression or The third method allows you to set or change the relationship between a gesture and a specific expression (specific action or facial expression). The third user interface unit 143 may be included. 43 is a special expression that is reflected in the avatar object for a "smiling face" as a specific facial expression. The expression is a "smiling face" that is the same as the specific facial expression, a completely unrelated "angry face" Select from multiple options such as "face" or "raise both hands" by touching (or flicking) the screen. (For convenience, in Figure 6, "smiling face" is selected as the specific expression. As shown in Figure 7 below, the candidate features are The specific expression may be an image of an avatar object that reflects the specific expression of the candidate. .

[0103] Furthermore, the user interface unit 140 may include a specific facial expression or gesture, a specific facial expression or gesture, The specific part corresponding to the gesture, the threshold value corresponding to the specific part, the specific expression or gesture and the specific When setting or changing the correspondence with the expression, the predetermined time, or the fixed time, It includes image information 144 and text information 145 relating to a specific facial expression or gesture. As shown in FIG. 7, the user interface unit 140 may display, as a specific facial expression, for example, " When setting "Stick out tongue", to easily let the target person know the face that "Sticks out tongue" (To instruct the target person), image information 144 as an illustration of "sticking out tongue" The text information includes "Stick out your tongue!!" etc., consider the image information 144 and the text information 145 (either one may be displayed). The user can set or change each piece of information. (Display unit 150) allows the user to select whether to display image information 144 (and character information 145). A dedicated slider 144x may be provided separately.

[0104] Furthermore, a specific facial expression or gesture, a specific part corresponding to a specific facial expression or gesture, The threshold values corresponding to the parts, the correspondence between a specific facial expression or gesture and a specific expression, a predetermined time, and A determination unit that determines whether a specific facial expression or gesture is formed when any of the predetermined time periods is set or changed. If the determination is made by the user interface unit 140, the specific facial expression or A first test video 147 (or Specifically, as shown in FIG. 7, for example, etc., based on the image information 144 and / or character information 145, As a result of the facial expression of "sticking out tongue" being made as a specific facial expression in If it is determined that a specific facial expression indicating "sticking out tongue" has been formed, an action that reflects the specific expression "sticking out tongue" is displayed. A first test video 147 (first test image 147) that is a butter object is displayed. This allows the performers to decide what kind of avatar they want to create for the specific facial expression or gesture they have created. It will be easier to recognize the image of the target object as an image or video is generated. .

[0105] Furthermore, a specific facial expression or gesture, a specific part corresponding to a specific facial expression or gesture, The threshold values corresponding to the parts, the correspondence between a specific facial expression or gesture and a specific expression, a predetermined time, and A determination unit that determines whether a specific facial expression or gesture is formed when any of the predetermined time periods is set or changed. If the determination is made by 120, the user interface unit 140 displays the predetermined time Even after the time has elapsed, the first test video 147 (first test image) is 147), and the same video (or image) as the first test video 147 (first test image 1 47) includes a second test video 148 (or a second test image 148) that is smaller in size than the first test video 148 (or the second test image 148). Specifically, as an example, when the performer or the like makes an expression indicating "sticking out tongue," the determination unit 12 0 determines that the specific facial expression of "sticking out tongue" has been formed, and the first test motion as shown in Figure 7 is executed. After the image 147 (first test image 147) is displayed, the judgment is released and the image is displayed for a certain period of time. After the time has passed, as shown in FIG. 8, the avatar object 1000 no longer has any specific expression. However, as shown in Figure 8, the first test A video (or image) with the same content as the first video 147 (first test image 147) is used as the second test video 148 (second test image 148) in the user interface section 140. For example, the performers can check the correspondence between a specific facial expression or gesture and a specific expression by using related images. You can set it slowly over time while watching. It may be a period between the predetermined times or a time different from the predetermined time.

[0106] As described above, the user interface unit 140 allows the performer to set various information. It is also possible to share various information visually with performers. For example, a specific facial expression or gesture, a specific part corresponding to a specific facial expression or gesture, the specific part the corresponding thresholds, the correspondence between a specific facial expression or gesture and a specific expression, the predetermined time, and the predetermined time The setting or change of the interval may be performed before (or after) the video (or image) distribution. Also, the user interface unit 140 in relation to FIGS. As an example, they may be displayed as separate pages on the display unit 150, each linked to the other. Alternatively, all of the images may be displayed on the same page, and the images may be scrolled vertically or horizontally on the display unit 150. It may be configured so that the performer can see it by clicking the In the data section 140, the various pieces of information shown in FIGS. 6 to 8 are arranged as shown in FIGS. It is not necessary to display the information in a specific form or combination. For example, instead of some of the information shown in FIG. 7 or 8 may be displayed on the same page.

[0107] (6) Display section 150 The display unit 150 displays the video generated by the generation unit 130 and the user interface unit 140. The screen for the studio unit 40 is displayed on the display (touch panel) and / or The display unit 150 can display the information on a display connected to the geo unit 40. The moving images generated by the generating unit 130 can be displayed sequentially, or can be stored in the storage unit 160. The stored video can also be displayed on a display or the like according to instructions from the performer or the like.

[0108] (7) Storage section 160 The storage unit 160 can store the video (or image) generated by the generation unit 130. The storage unit 160 can store the above-mentioned threshold value. 160 is a predefined default when a particular application is installed. The threshold values can be stored, or the threshold values set by the user interface unit 140 can be The value can also be stored.

[0109] (8) Communications Department 170 The communication unit 170 receives the information generated by the generation unit 130 (and stored in the storage unit 160). The video (or image) can be transmitted to the server device 30 via the communication network 10.

[0110] The operation of each of the above-mentioned parts depends on the specific application installed in the studio unit 40. Applications (for example, video distribution applications) are handled by this studio unit 40. Alternatively, the operations of the above-described parts can be performed by The browser installed in the studio unit 40 is provided by the server device 30. This can be performed by the studio unit 40 by accessing a website. As explained in the "first aspect" above, the studio unit 4 The generating unit 130 generates the above-mentioned moving images (first moving image and Instead of generating the video (video 2), the generating unit 130 is arranged in the server device 30, and the video is The unit 40 stores data about the body of the performer, etc., and the body of the performer, etc. based on the data. Data relating to the amount of change in each of the plurality of specific parts (including information on the determination result by the determination unit 120) The server device 30 transmits the studio unit 100 the studio unit 102 via the communication unit 170. According to the data received from the network 40, a predetermined specific expression is displayed on the avatar corresponding to the performer. The rendering method for generating the video (first video and second video) reflected in the object is Alternatively, the studio unit 40 may store data relating to the bodies of performers, etc. and data on the amount of change in each of a plurality of specific parts of the body of the performer, etc. based on the data. (including information on the determination result by the determination unit 120) via the communication unit 170 to the server device 30 The server device 30 transmits the data received from the studio unit 40 to the terminal device 20. The generating unit 130 provided in the terminal device 20 transmits the received data from the server device 30. According to the data, a predetermined specific expression is reflected in the avatar object corresponding to the performer. Alternatively, a rendering method may be employed to generate the moving images (first moving image and second moving image).

[0111] 3-2. Functions of the terminal device 20 A specific example of the functions of the terminal device 20 will be described with reference to FIG. As a function, for example, the function of the studio unit 40 described above can be used. Therefore, the reference numerals for the components of the terminal device 20 are shown in parentheses in FIG. is shown.

[0112] In the above-mentioned "second aspect," the terminal device 20 (for example, the terminal device 20A in FIG. 1) The sensor unit 200 to the communication unit 270 are respectively related to the studio unit 40. The sensor unit 100 to the communication unit 170 may be the same as those described above. The operations of the above-mentioned components are controlled by a specific application installed in the terminal device 20. An application (for example, an application for video distribution) is executed by this terminal device 20. This can be executed by the terminal device 20. As explained in the previous section, the terminal device 20 is provided with a generating unit 230. Instead of generating the above-mentioned moving image by the server device 30, the generating unit 230 is arranged in the server device 30. The terminal device 20 then acquires data relating to the body of the performer, etc., and the performer, etc., based on the data. Data on the amount of change in each of a plurality of specific body parts (information on the determination result by the determination unit 220) The server device 30 transmits the information (including the information) to the terminal device via the communication unit 270. According to the data received from the device 20, a predetermined specific expression is displayed as an avatar image corresponding to the performer. A configuration may be adopted in which moving images (first moving image and second moving image) are generated that reflect the object. Alternatively, the terminal device 20 may store data relating to the body of a performer or the like and a character of the performer based on the data. Data on the amount of change in each of a plurality of specific body parts (determination results by the determination unit 220) (including the information) is transmitted to the server device 30 via the communication unit 270, and the server device 30 The data received from the terminal device 20 is forwarded to another terminal device 20 (for example, the terminal device 2 in FIG. 1). 0C), and the generation unit 230 provided in the other terminal device 20 transmits the According to the data received from the Alternatively, a configuration may be adopted in which moving images (first moving image and second moving image) are generated that reflect the above.

[0113] On the other hand, in the "first mode" and the "third mode", for example, the terminal device 20 By having at least the communication unit 270 among the communication units 0 to 270, the studio unit The video ( or image) can be received via the communication network 10. 0 indicates that a specific application (e.g., a video viewing application) is installed. The user executes the above operation to send a signal (request) to the server device 30 requesting the delivery of the desired video. By transmitting a signal, the desired video is received from the server device 30 in response to the signal. It can be received via the specific application.

[0114] 3-3. Functions of the server device 30 A specific example of the functions of the server device 30 will be described with reference to FIG. As this function, for example, the function of the studio unit 40 described above can be used. Therefore, the reference numerals for the components of the server device 30 are used collectively in FIG. is shown in the arc.

[0115] In the above-mentioned "third aspect," the server device 30 includes the sensor unit 300 to the communication unit 370. The sensor unit 100 to the communication unit 170 described in relation to the studio unit 40 are The operations of the above-mentioned components can be controlled by the server. A specific application (for example, a video distribution application) installed on the device 30 The server device 30 can execute the above-mentioned application. In the "third aspect," the server device 30 is provided with a generating unit 330, and the generating unit 330 Instead of generating the above-mentioned moving image by the generating unit 330, the generating unit 330 is The server device 30 stores data relating to the body of the performer and the like and a performance based on the data. Data on the amount of change in each of a plurality of specific parts of the body of the person, etc. (determination result by the determination unit 320) The terminal device 20 transmits the result information (including the result information) to the terminal device 20 via the communication unit 370. According to the data received from the avatar device 30, a predetermined specific expression is displayed on the avatar corresponding to the performer. A configuration may be adopted in which moving images (first moving image and second moving image) are generated that are reflected in the object. stomach.

[0116] 4. Overall Operation of Communication System 1 Next, the overall operation of the communication system 1 having the above configuration will be described with reference to FIG. 9 and 10. In the communication system 1 shown in FIG. 10 is a flowchart showing an example of a part of the operation performed by the This shows the aforementioned "first aspect" as an example.

[0117] First, in step (hereinafter referred to as "ST") 500, performers etc. (as mentioned above, a user interface section 14 of the studio unit 40 0 to set a specific facial expression or gesture, as explained above. For example, "Smiling face", "Closing one eye (wink)", "Surprised face", "Sad face", "Angry face", "Scheming face," "Bashful face," "Closing both eyes," "Stick out tongue," "Open mouth," "Cheeks Facial expressions such as "inflating one's chest" and "opening one's eyes wide" and parts such as "shaking one's shoulders" and "shaking one's head" are also included. The action can be set as a specific facial expression or gesture, without being limited to these.

[0118] Next, in ST501, the performer or the like accesses the user interface section of the studio unit 40. 140 (first user interface unit 141) as described above with reference to FIG. As explained above, each specific facial expression (e.g., "close one eye (wink)" or "smile") Specific parts of the performer's body that correspond to the "face" (e.g., eyebrows, eyelids, eyes, cheeks, nose, mouth, lips, etc.) Set.

[0119] Next, in ST502, the performer or the like uses the user interface of the studio unit 40. As described above with reference to FIG. 6, the setting in ST501 is performed via the unit 140. In this case, each threshold value is set corresponding to the amount of change in each of the specific portions. As described above, the setting is performed for each specific part using the first user interface unit 141. It may be set to any value, or a predetermined mode may be selected using the second user interface unit 142. By selecting a mode (for example, a mode that is "easy to appear"), each threshold value is set to a predetermined value. Also, a predetermined mode may be set in the second user interface unit 142. After the selection, the first user interface section 141 is used to customize the threshold. Good too.

[0120] Next, in ST503, the performer or the like uses the user interface of the studio unit 40. 5 to 8 through the ST500. In this case, a correspondence relationship between the specific facial expression or gesture and the specific expression is set. As described above, the corresponding relationship is set using the third user interface unit 143. is executed.

[0121] Next, in ST504, the performer or the like uses the user interface of the studio unit 40. The predetermined time and the fixed time described above can be set to appropriate values via the unit 140. Cut.

[0122] ST500 to ST504 shown in FIG. 9 are settings in the overall operation of the communication system 1. Also, ST500 to ST504 are not necessarily limited to the order shown in FIG. For example, the order of ST502 and ST503 may be reversed, or The order of T501 and ST503 may be reversed. After the setting operation is performed (or after the video generation operation shown in FIG. 10 is performed), When changing only one of the values, some of ST500 to ST504 Specifically, only the setting operations in ST500 to ST504 may be executed. If you want to change only the threshold after the operation has been performed, you can execute only ST502. That's fine.

[0123] As described above, when the setting operation shown in FIG. 9 is completed, the video generation shown in FIG. 10 is performed. The following operations can be performed.

[0124] A request (operation) for generating a moving image is made by a performer or the like via the user interface unit 140. When the program is executed, first, in ST505, the sensor unit 100 of the studio unit 40 However, as mentioned above, data on the physical movements of performers, etc. is acquired.

[0125] Next, in ST506, the change amount acquisition section 110 of the studio unit 40 Based on the data on the body movements of the performers, etc., acquired by 100, The amount of change (amount of displacement) in each of a plurality of specific body parts is acquired.

[0126] Next, in ST507, the generation section 130 of the studio unit 40 receives the signal from the sensor section 100. generates the first moving image based on various information acquired by the

[0127] Next, in ST508, the determination section 120 of the studio unit 40 determines whether or not the All of the change amounts of the set specific parts exceed the respective thresholds set in ST502. If it is "exceeding," the decision unit 120 determines whether the ST If it is determined that the specific facial expression or gesture set in step 500 has been formed, the process proceeds to step ST520. On the other hand, if the result in ST508 is "not greater," proceed to ST509. .

[0128] Next, if it is not "exceeded" in ST508, in ST509, The communication section 170 of the unit 40 serves the first moving image generated by the generation section 130 in ST507. Then, in ST509, the communication unit 170 transmits the information to the server device 30. In ST510, the first moving image transmitted to the terminal device 30 is transmitted to the server device 30. The first moving image is transmitted to the server device 20. In ST530, the terminal device 20 displays the first moving image on the display unit 250. In this way, if the result in ST508 is "not greater," the series of steps ends.

[0129] On the other hand, if the ST508 is "exceeded," the ST520 will be a studio unit. The generation unit 130 of 40 determines the information of the determination result that a specific facial expression (or gesture) has been formed. and assigning a specific expression corresponding to the specific facial expression or gesture to the avatar object. At this time, the generation unit 130 generates a secondary moving image that reflects the above-mentioned project. By referring to the settings in the avatar object, specific expressions corresponding to specific facial expressions or actions can be assigned to the avatar object. This can be reflected in the project.

[0130] Then, in ST521, the communication section 170 transmits the second moving image generated in ST520. The second moving image is transmitted to the server device 30. The second moving image transmitted by the server device 30 is At T522, the server device 30 transmits the request to the terminal device 20. In ST530, terminal device 20 receives the second moving image transmitted from terminal device 30. The second moving image is displayed on the display unit 250. In this way, in ST508, If "YES", the series of steps ends.

[0131] A request (operation) regarding video generation (video distribution) is made via the user interface unit 140. When executed, the process for the series of steps of video generation (video distribution) shown in Figure 10 is executed. It is repeatedly performed. For example, a specific facial expression or gesture may be performed by a performer. is determined to be formed, and the series of steps shown in FIG. 10 (for convenience in this paragraph, While the processing related to the first processing is being performed, another specific table is If it is determined that the emotion or behavior has been formed, the process shown in FIG. 10 is repeated to follow the initial process. The avatar object has a separate process for the sequence of steps that are performed. Specific expressions corresponding to specific facial expressions or gestures formed by the user, etc., malfunction in real time. The intentions of the performers are accurately reflected without any distortion.

[0132] 9 and 10, the "first embodiment" has been described as an example. However, in the "second embodiment" and the "third embodiment", the same configuration as in FIGS. 9 and 10 is basically used. This is a series of steps. That is, the sensor unit 100 to the communication unit 170 in FIGS. 9 and 10 is replaced with a sensor unit 200 to a communication unit 270, or a sensor unit 300 to a communication unit 370.

[0133] As described above, according to various embodiments, a performer or the like can easily and accurately create an avatar object. A computer program, a server device, and a More specifically, according to various embodiments, a device and method may be provided. Even while speaking, you can make the avatar object express a specific expression (desired expression) by simply forming a specific facial expression. It is now possible to create videos that reflect the user's facial expressions and movements more accurately and easily than before, without any operational errors or misfires. Furthermore, the performer or the like can express a particular facial expression or The user sets (changes) the actions and the like as described above, and performs the various actions and the like as described above from the terminal device 20. Furthermore, when streaming video, the terminal device held by the performer, etc. The device 20 can capture changes in the performers (face and body changes) at any time and respond to those changes. This allows the avatar object to reflect a specific expression.

[0134] 5. Variations In the embodiment described above, the performer or the like uses the user interface unit 140 While operating the device, a specific facial expression or gesture may be formed by the user. However, the present invention is not limited to this. For example, a supporter or an operator operates the user interface unit 140 while a performer It may be possible to form a specific facial expression or gesture. The user sets the threshold value and the like while checking the user interface unit 140 shown in FIGS. At the same time, the sensor unit 100 can detect the movements, facial expressions, and speech (singing) of the performer. When it is determined that the performer has made a specific facial expression or gesture, the In this way, the user interface unit 140 displays an image of an avatar object that reflects a specific expression. An image or video is displayed.

[0135] Regarding the third user interface unit 143, with reference to FIGS. 6 to 8, As explained above, another embodiment may be used as shown in FIG. FIG. 11 shows a modified example of the third user interface unit 143. In this case, First, for each specific facial expression or gesture formed by a performer, etc., For example, for a specific facial expression such as "opening both eyes wide," The control number is "1" for the specific facial expression, and the control number is "2" for the specific facial expression, "close both eyes tightly." The control number for the specific facial expression "sticking out tongue" is "3," and the control number for the specific facial expression "making a face" is "3." The control number for the facial expression "4" is assigned, and the control number for the specific facial expression "puffing out cheeks" is assigned. For the specific facial expression, "5" is assigned to the control number "6" and for the specific facial expression, "smiling face" is assigned to the control number "6". The control number for the specific facial expression "surprise face" is "7", and the control number for the specific facial expression "surprise face" is "8". The control number is "8" for the specific action of "shaking shoulders," and the control number is "9" for the specific action of "shaking head." The control number "10" is set for each specific action.

[0136] Next, the performer or the like inputs a specific expression via the third user interface unit 143. A specific facial expression or gesture can be selected based on the aforementioned reference number. As shown in 11, the control number "1" is selected for the specific expression "open both eyes wide." In response to the specific expression "open both eyes", the specific expression "open both eyes" is displayed as an avatar object. For example, the control number " When "2" is selected, the specific expression "Close both eyes" corresponds to the specific expression "Close both eyes" "Open" is reflected in the avatar object. When the control number "8" is selected for the specific expression "make your mouth look like a smile," In response to the "surprised face," a specific expression "mouth clenched" is reflected in the avatar object. In this way, by managing various specific facial expressions or actions with control numbers, performers can It is now possible to more easily set or change the correspondence between a specific facial expression or gesture and a specific expression. become.

[0137] In this case, the specific facial expression or gesture and the associated management number are The correspondence relationship is stored in the storage unit 160 (storage unit 260, storage unit 360). The third user interface section 143 shown in FIG. 11 is not linked to FIGS. 6 to 8. It may be displayed as a separate page or may be displayed on the same page as the pages shown in FIGS. The display unit 150 is configured to be visible by scrolling vertically or horizontally. It may also be possible to use the following.

[0138] For example, when a specific facial expression is associated with a management number and stored in the storage unit 160, When the determination unit 120 determines that a specific facial expression or gesture has been formed by the performer, etc., The generation unit 130 outputs a control number corresponding to a specific facial expression or gesture. Based on the correspondence between the specific expression and the predetermined control number (specific facial expression or gesture), and a second image in which a specific expression corresponding to the specific facial expression or behavior is reflected in an avatar object. 2 You may generate videos.

[0139] 6. Various Aspects The computer program according to the first aspect is "executed by one or more processors" By this, based on the data on the movement of the body acquired by the sensor, Amounts of change in each of a plurality of specific portions are acquired, and a predicted amount of change in each of the plurality of specific portions is selected. All of the amounts of change in at least one of the specified portions specified for this purpose exceed the respective thresholds. If the facial expression or gesture is formed, it is determined that the facial expression or gesture has been formed. An image in which a specific expression corresponding to the work is reflected on an avatar object corresponding to the performer. or causes the processor to function to generate a video.

[0140] The computer program according to the second aspect is the computer program according to the first aspect, "The phrase 'includes a particular action or facial expression.'

[0141] A computer program according to a third aspect of the present invention is a computer program according to the first aspect or the second aspect of the present invention. In this case, "the body is the body of the performer."

[0142] A computer program according to a fourth aspect of the present invention is a computer program according to any one of the first to third aspects of the present invention. In either case, "the processor is configured to When all of the changes in the respective minutes exceed the respective thresholds for a predetermined time, the specific facial expression or behavior is formed. "It is determined that the

[0143] A computer program according to a fifth aspect of the present invention is a computer program according to any one of the first to fourth aspects of the present invention. In either case, "the processor determines the specific facial expression or gesture corresponding to the determined specific facial expression or gesture." A specific expression is reflected for a certain period of time on an avatar object corresponding to the performer. "generates images or videos."

[0144] A computer program according to a sixth aspect of the present invention is a computer program according to any one of the first to fifth aspects of the present invention. In either case, "the specific facial expression or gesture, the specific facial expression or gesture corresponding to the specific facial expression or gesture" the part, each of the threshold values, the correspondence between the specific facial expression or gesture and the specific expression, At least one of the time and the predetermined time is set via a user interface. or be changed."

[0145] The computer program according to the seventh aspect is the same as the sixth aspect, in which "each of the thresholds Each of the specific portions can be set or changed to an arbitrary value via the user interface. It is something that can be done.

[0146] The computer program according to the eighth aspect of the present invention is the same as the sixth aspect of the present invention, except that "each of the thresholds Each of the specific portions is configured to input a plurality of predetermined values via the user interface. The value is "set or changed to one of the following values."

[0147] The computer program according to the ninth aspect is the same as the sixth aspect in that "the user interface" is The interface includes a first user interface for setting each of the threshold values to an arbitrary value for each of the specific portions. and an interface, wherein each of the threshold values is set to any one of a plurality of predetermined values that are predetermined for each of the specific portions. a second user interface for setting the specific facial expression or gesture and the specific expression; and a third user interface for setting a correspondence relationship between the "It is something like that.

[0148] A computer program according to a tenth aspect of the present invention is a computer program according to any one of the sixth to ninth aspects of the present invention. In either case, "the specific facial expression or gesture, the specific facial expression or gesture corresponding to the specific facial expression or gesture" a predetermined part, each of the threshold values, a correspondence relationship between the specific facial expression or gesture and the specific expression, When at least one of the predetermined time and the predetermined time is set or changed, The interface displays at least image information and text information related to the specific facial expression or behavior. "Both of these are included."

[0149] A computer program according to an eleventh aspect of the present invention is a computer program according to any one of the sixth to tenth aspects of the present invention. In either case, "the specific facial expression or gesture, the specific facial expression or gesture corresponding to the specific facial expression or gesture" a specific part, each of the thresholds, a correspondence between the specific facial expression or gesture and the specific expression, When at least one of the predetermined time and the fixed time is set or changed, When it is determined that the facial expression or gesture of the particular A first test in which the specific expression identical to a specific facial expression or behavior is reflected in the avatar object. "These include stock images or the first test video."

[0150] The computer program according to the 12th aspect is the same as the computer program according to the 11th aspect, except that "the specified the specific part corresponding to the specific facial expression or gesture, the threshold value, the correspondence between the specific facial expression or gesture and the specific expression, the predetermined time, and the fixed time, When at least one of the above is set or changed, the specific facial expression or gesture is formed. If it is determined that the user interface is A second test image or a second test video identical to the first test image or the first test video is then generated. It includes videos of people stalking the camera.

[0151] The computer program according to the thirteenth aspect is the same as the sixth aspect in that "the specific The correspondence between the facial expression or the behavior and the specific expression is The same relationship, the relationship in which the specific facial expression or gesture is similar to the specific expression, and the relationship in which the specific expression is similar to the specific expression "Either the emotion or action is unrelated to the specific expression."

[0152] A computer program according to a fourteenth aspect of the present invention is a computer program according to any one of the sixth to thirteenth aspects of the present invention. In either case, "the specific facial expression or gesture, the specific facial expression or gesture corresponding to the specific facial expression or gesture" a specific part, each of the thresholds, a correspondence between the specific facial expression or gesture and the specific expression, At least one of the predetermined time and the fixed time is changed during the distribution of the image or video. It is something that is "changed."

[0153] A computer program according to a fifteenth aspect of the present invention is a computer program according to any one of the first to fourteenth aspects of the present invention. In any of the above, "the specific part is a part of the face."

[0154] The computer program according to the 16th aspect is the same as the computer program according to the 15th aspect, except that "the specified the part is selected from the group including eyebrows, eyes, eyelids, cheeks, nose, ears, lips, tongue, and chin .

[0155] A computer program according to a seventeenth aspect of the present invention is a computer program according to any one of the first to sixteenth aspects of the present invention. In any of the above, "the processor is a central processing unit (CPU), a microprocessor or a graphics processing unit (GPU)."

[0156] A computer program according to an eighteenth aspect of the present invention is a computer program according to any one of the first to seventeenth aspects of the present invention. "The processor is a smartphone, tablet, mobile phone, or is installed in a personal computer or a server device.

[0157] A server device according to a 19th aspect includes a processor, The instructions readable by the computer are executed to detect the movement of the body captured by the sensor. Based on the data on the movement, the amount of change in each of the plurality of specific parts of the body is obtained, and the plurality Among the changes in the respective specific portions of the When all of the changes in each part exceed the respective thresholds, it is determined that a specific facial expression or gesture has been formed. The specific expression corresponding to the determined specific facial expression or gesture is displayed on an avatar corresponding to the performer. It generates an image or video that reflects the target object.

[0158] A server device according to a twentieth aspect is the server device according to the nineteenth aspect, Central Processing Unit (CPU), microprocessor or graphics processing unit (GPU)".

[0159] The server device according to the 21st aspect is the server device according to the 19th aspect or the 20th aspect. It will be placed in the studio.

[0160] The method according to the twenty-second aspect is a method for performing one or more computer-readable instructions. A method executed by a plurality of processors, A change amount for each of the plurality of specific body parts is obtained based on the data relating to An acquisition step, and acquiring at least one of the amounts of change of the plurality of specific portions, which is specified in advance. When all of the change amounts of the specific parts exceed the respective thresholds, a specific facial expression or a specific part is detected. a determining step of determining that a specific facial expression has been formed; Or, a specific expression corresponding to the gesture is reflected in an avatar object corresponding to the performer. and generating an image or video.

[0161] The method according to the 23rd aspect is the same as the method according to the 22nd aspect except that "the change amount acquiring step, the determination step" The determining step and the generating step are carried out on a smartphone, a tablet, a mobile phone, and a personal computer. The program is executed by the processor installed in a terminal device selected from the group including a computer. It is something that ``happens.''

[0162] The method according to the 24th aspect is the same as the method according to the 22nd aspect except that "the change amount acquiring step, the determination step" The determining step and the generating step are executed by the processor installed in the server device. "It is something like that.

[0163] The method according to the 25th aspect is the method according to any one of the 22nd to 24th aspects. "The processor may be a central processing unit (CPU), a microprocessor, or a graphics It is a graphics processing unit (GPU).

[0164] A system according to a 26th aspect of the present invention includes a first device including a first processor, and a second processor. a second device including a processor and connectable to the first device via a communication line, and a system for detecting a movement of the body based on data relating to the movement of the body acquired by a sensor. A change amount acquisition process for acquiring a change amount for each of a plurality of specific portions; of the change in each of at least one specific portion specified in advance A determination process in which a specific facial expression or gesture is determined to have been formed when all of the values exceed the respective thresholds. A specific expression corresponding to the specific facial expression or gesture determined by the determination process is given to the performer. A generation process for generating an image or video reflected on a corresponding avatar object. wherein the first processor included in the first device reads the The change amount acquisition process, the determination process, and the generation process are performed by executing the available instructions. and performing at least one of the processes not performed by the first processor. If there is any remaining processing, the second processor included in the second device: The remaining steps are performed by executing computer readable instructions. It is something that ``happens.''

[0165] A system according to a 27th aspect is the system according to the 26th aspect, wherein "the processor Processing Unit (CPU), microprocessor or graphics processing unit ( It is a "GPU (Graphical Processing Unit)".

[0166] The system according to the 28th aspect is the same as the system according to the 26th or 27th aspect. The communication lines include the Internet.

[0167] The terminal device according to the 29th aspect of the present invention is configured to: and obtaining a change amount for each of the plurality of specific body parts based on the above. Among the respective change amounts, the change in each of the specific portions at least at one or more locations specified in advance If all of the amounts exceed the respective thresholds, it is determined that a specific facial expression or gesture has been formed. A specific expression corresponding to the specific facial expression or gesture is displayed on an avatar object corresponding to the performer. "It generates an image or video that reflects the image."

[0168] A terminal device according to a 30th aspect is the terminal device according to the 29th aspect, wherein "the processor Central Processing Unit (CPU), microprocessor or graphics processing unit (GPU)".

[0169] 7. Fields to which the technology disclosed in this application is applicable The technology disclosed in the present application can be applied, for example, in the following fields: It is something. (1) An application server that distributes live videos featuring avatar objects. bis (2) An app that can communicate using text and avatar objects Application services (chat applications, messengers, email applications) (communication, etc.) [Explanation of symbols]

[0170] 1. Communication Systems 10. Communication Network 20(20A~20C) Terminal equipment 30(30A~30C) Server equipment 40 (40A, 40B) Studio Unit 100 (200, 300) Sensor part 110 (210, 310) Change amount acquisition section 120(220, 320) Judgment section 130(230, 330) Generation part 140 (240, 340) User interface section 141 First user interface section 142 Second User Interface Section 143 Third User Interface Section 144 Image Information 145 Text Information 147 First test image (first test video) 148 Second test image (Second test video) 150(250, 350) Display section 160(260, 360) Storage section 170(270, 370) Communications Department

Claims

1. one or more processors, a first means for accepting selection of one of a plurality of modes in which thresholds for the amount of change in a plurality of body parts are preset; a second means for acquiring a change amount of the body part based on data relating to the body movement; and a third means for determining that a specific facial expression or gesture has been formed when the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode; The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; A computer program in which a threshold value for the amount of change of the second portion in the first mode is different from a threshold value for the amount of change of the second portion in the second mode.

2. the plurality of modes include a first mode and a second mode; The program according to claim 1 , wherein when the second mode is selected, it is more likely that the specific facial expression or gesture has been formed than when the first mode is selected.

3. causing the processor to function as fifth means for accepting settings of each of the thresholds regardless of selection of any of the plurality of modes; 3. The program according to claim 1, wherein the third means makes the determination using the threshold value set for the mode selected by the first means or the threshold value set by the fifth means.

4. The program according to claim 3 , further comprising causing the processor to function as sixth means for accepting a selection to invalidate a selection of any one of the plurality of modes.

5. one or more processors, a first means for accepting selection of one of a plurality of modes in which thresholds for the amount of change in a plurality of body parts are preset; a second means for acquiring a change amount of the body part based on data relating to the body movement; and a fourth means for generating an image or video in which a specific expression corresponding to a specific facial expression or gesture is reflected on an avatar object corresponding to the performer when the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode; The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; A program in which a threshold value of the amount of change of the second portion in the first mode is different from a threshold value of the amount of change of the second portion in the second mode.

6. The program according to claim 1 , wherein the data on the body movement is acquired by a sensor.

7. the plurality of modes include a first mode and a second mode; a threshold value set for a first portion of the plurality of portions is lower than a threshold value set in advance for the second mode; The program according to claim 1 , wherein the threshold value set for the second part of the plurality of parts is the same as the threshold value set in advance for the first mode and the threshold value set in advance for the second mode.

8. the plurality of modes include a first mode and a second mode; A default value is set for each of the thresholds, A threshold value smaller than the default value is preset in the first mode, The program according to claim 1 , wherein a threshold value greater than the default value is preset for the second mode.

9. a first means for accepting selection of one of a plurality of modes in which thresholds of amounts of change in a plurality of body parts are preset; a second means for acquiring a change amount of the body part based on data relating to the body movement; and a third means for determining that a specific facial expression or gesture has been made when the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode, The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; The system wherein the threshold for the amount of change of the second portion in the first mode is different from the threshold for the amount of change of the second portion in the second mode.

10. a first means for accepting selection of one of a plurality of modes in which thresholds for the amount of change in a plurality of body parts are preset; a second means for acquiring a change amount of the body part based on data relating to the body movement; and fourth means for generating an image or video in which a specific expression corresponding to a specific facial expression or gesture is reflected on an avatar object corresponding to the performer when the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode, The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; The system wherein the threshold for the amount of change of the second portion in the first mode is different from the threshold for the amount of change of the second portion in the second mode.

11. A method executed by one or more processors, comprising: Accepting a selection of one of a plurality of modes in which thresholds for the amount of change in a plurality of body parts are set in advance; obtaining a change amount of the body part based on data relating to the body movement; determining that a specific facial expression or gesture has been made when the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode; The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; The method, wherein the threshold amount of change of the second portion in the first mode is different from the threshold amount of change of the second portion in the second mode.

12. A method executed by one or more processors, comprising: Accepting a selection of one of a plurality of modes in which thresholds for the amount of change in a plurality of body parts are set in advance; obtaining a change amount of the body part based on data relating to the body movement; When the amount of change in the plurality of body parts exceeds each of the thresholds set for the selected mode, an image or video is generated in which a specific expression corresponding to a specific facial expression or gesture is reflected on an avatar object corresponding to the performer; The change amount is a displacement amount between a position of the corresponding body part at a certain time point and a position of the corresponding body part at another time point, the plurality of modes includes at least a first mode and a second mode; the plurality of body parts includes at least a first part and a second part; a threshold value of the amount of change of the first portion in the first mode is different from a threshold value of the amount of change of the first portion in the second mode; The method, wherein the threshold amount of change of the second portion in the first mode is different from the threshold amount of change of the second portion in the second mode.

Citation Information

Patent Citations

  • Image / Voice communication system and video telephone transmission / Reception method

    JP1998271470A

  • Information processing device, control method therefor, computer program, and memory medium

    JP2007087345A

  • Information processing device, control method therefor, computer program, and memory medium

    JP2007087346A

  • Imaging apparatus and its control method and program and storage medium

    JP2008131204A

  • Image processing device, method, and program

    JP2008299430A