System, program and method

The system uses a machine-learned prediction model to analyze lip movements and positions in video to accurately evaluate voice quality, addressing the limitations of existing voice evaluation systems and providing actionable feedback for users.

JP2025093715APending Publication Date: 2025-06-24OKUCHY INC

Patent Information

Application Number
JP2023209531
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing voice evaluation systems lack the ability to accurately assess a person's voice quality using video information, particularly focusing on lip movements and positions, which is crucial for evaluating speech fluency, pronunciation, and other vocal attributes.

Method used

A system utilizing a machine-learned prediction model that processes video information, including images of lips and their positional data, to evaluate a user's voice by specifying evaluation information based on these inputs, incorporating features like movement and speed of predetermined points on the lips.

Benefits of technology

The system provides accurate and detailed evaluation of voice quality, enabling users to understand their vocal performance and receive feedback for improvement, leveraging machine learning to enhance the precision of voice assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093715000001_ABST
    Figure 2025093715000001_ABST
Patent Text Reader

Abstract

To provide a novel system for evaluating human vocalization.SOLUTION: A system comprises at least one computer device. The system comprises specification means that uses a machine-learned prediction model that takes video information containing images of a person's lip during vocalization, or information regarding the position of a specified point on an image of a person's lip during vocalization as input information, and takes evaluation information related to the vocalization as output information so as to identify evaluation information related to the user's vocalization based on video information containing an image of the user's lip at the user's vocalization, or information related to the position of predetermined point on an image of the user's lip at the user's vocalization.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system, a program, and a method.

Background Art

[0002] As a method for evaluating a person's voice, a method of recording the voice during voice production and evaluating the recorded voice is disclosed. For example, Patent Document 1 discloses a system that can provide content for objectively judging and training the ability to respond to voice changes when a language is being vocalized.

[0003] The system of Patent Document 1 associates and stores a plurality of indexes that are factors in evaluating the ability to respond to voice changes when a language is being vocalized, with the portions in the character string of the language used for evaluation where the sounds corresponding to the indexes appear. The system receives, as voice information, an input of an answer content that the user attempts to reproduce the character string of the language used for evaluation. Then, the system compares the received answer content with the character string of the language used for evaluation, determines whether the portions where the indexes are associated and stored in the character string output as voice are correct, and calculates the correct rate for each index.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] An object of the present invention is to provide, for example, a novel system for evaluating a person's voice.

Means for Solving the Problems

[0006] The problems of the present invention are [1]A system comprising at least one computer device, which uses a prediction model that is machine-learned with video information including an image of a lip during a person's voice utterance or information regarding the position of a predetermined point of an image of a lip during a person's voice utterance as input information and evaluation information regarding the voice utterance as output information, and includes a specifying means for specifying evaluation information regarding a user's voice utterance based on video information including an image of a lip during the user's voice utterance or information regarding the position of a predetermined point of an image of a lip during the user's voice utterance; [2]The prediction model in the system according to [1] above, wherein the prediction model is further machine-learned with video information including an image of a lip during a person's voice utterance or information regarding the position of a predetermined point of an image of a lip during a person's voice utterance, and voice information during the person's voice utterance as input information and evaluation information regarding the voice utterance as output information, and the specifying means uses the prediction model to specify evaluation information regarding the user's voice utterance based on video information including an image of a lip during the user's voice utterance or information regarding the position of a predetermined point of an image of a lip during the user's voice utterance, and voice information during the user's voice utterance; [3]The system according to [1] or [2] above, wherein the information regarding the position of a predetermined point of an image of a lip during the person's voice utterance includes information regarding the amount of movement or speed of the predetermined point, or information regarding the area obtained by the predetermined point; [4]The system according to any one of [1] to [3] above, further including a position specifying means for specifying information regarding the position of a predetermined point of an image of a lip from video information including an image of a lip during the user's voice utterance, and the position specifying means specifies the amount of movement or speed of the predetermined point based on a reference point of the image; [5]The system according to any one of [1] to [4] above, wherein the evaluation information regarding the voice utterance includes information regarding whether the voice utterance is good or bad; [6]A program that functions as a specifying means for specifying evaluation information regarding a user's voice based on moving image information including an image of lips during a person's voice or information regarding the positions of predetermined points of an image of lips during a person's voice, using a prediction model that has been machine-learned with the moving image information including an image of lips during a person's voice or information regarding the positions of predetermined points of an image of lips during a person's voice as input information and the evaluation information regarding the voice as output information; [7]A method executed in at least one computer device, the method including a specifying step of specifying evaluation information regarding a user's voice based on moving image information including an image of lips during a person's voice or information regarding the positions of predetermined points of an image of lips during a person's voice, using a prediction model that has been machine-learned with the moving image information including an image of lips during a person's voice or information regarding the positions of predetermined points of an image of lips during a person's voice as input information and the evaluation information regarding the voice as output information; can be solved by. [Advantages of the Invention]

[0007] According to the present invention, for example, a novel system for evaluating a person's voice can be provided. [Brief Description of the Drawings]

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described. However, the present invention is not limited to the following embodiments as long as it does not contravene the gist of the present invention. The order of each process constituting the flowchart described below can be in any order as long as there are no contradictions or inconsistencies in the process content. Also, as long as there are no contradictions or inconsistencies in the process content, a part of each process constituting the flowchart can be omitted, or a new process can be added to each process constituting the flowchart. Further, the device that is the subject of executing each process constituting the flowchart can be changed to another device as long as it does not contravene the gist of the present invention. In that case, the process content can be changed so that there are no contradictions or inconsistencies in the process content.

[0010] [Configuration of the System] FIG. 1 is a block diagram showing the configuration of a system according to an embodiment of the present invention. The system includes at least one computer device. Specifically, the system includes a user terminal 1 and a server device 2. The user terminal 1 and the server device 2 are communicably connected to each other via a communication network 3. In the system, either the user terminal 1 or the server device 2 can function as an information processing device. When either the user terminal 1 or the server device 2 functions as an information processing device, information is transmitted and received between the user terminal 1 and the server device 2 as necessary.

[0011] The user terminal 1 is operated by the user. The server device 2 may be managed by an administrator who manages the system. Although not shown, the system may include an administrator terminal for the administrator who operates the server device 2, and a terminal for constructing a prediction model described later.

[0012] Note that the system may include only one user terminal 1. Further, the system may include two or more user terminals 1. Further, the system may include two or more server devices 2. For example, the system may include a first server device for constructing a prediction model, a second server device for receiving information from the user terminal 1, and / or a third server device for receiving information from the second server device and performing execution processing of a program. Further, the server device 2 may be distributed and function among a plurality of computer devices. For example, instead of the server device 2, a distributed ledger technology such as a blockchain may be used.

[0013] [User Terminal] The user terminal 1 is operated by a user who receives an evaluation related to voice, for example. The user is not particularly limited, but specifically, it is a learner or a student who learns a predetermined language, a trainee or a student aiming at a predetermined occupation, or the like. The user terminal 1 is not particularly limited as long as it is a computer device including an imaging device. The user terminal 1 is preferably a computer device including a display device and an imaging device. Examples of the user terminal 1 include a conventional mobile phone, a tablet terminal, a smartphone, and the like.

[0014] FIG. 2 is a diagram showing an example of the appearance of the user terminal according to the embodiment of the present invention. The user terminal 1 has a rectangular shape. The user terminal 1 includes a camera 5 on the surface provided with the display screen 4 (also referred to as the front surface of the user terminal 1). Further, the user terminal 1 may include a speaker and a microphone 6 on the front surface. The display device provided in the user terminal 1 has a rectangular display screen 4. In the embodiment of the present invention, the rectangle may include a substantially rectangular shape such as a shape in which the corners of the rectangle are rounded, an oval shape, or a shape in which a part of the rectangle is missing.

[0015] The camera 5 functions as an imaging device. When the longitudinal direction of the user terminal 1 is the vertical direction, the camera 5 is disposed near the upper side of the user terminal 1 in the normal orientation when the user operates the user terminal 1. It can also be said that the camera 5 is installed near the upper end in the longitudinal direction of the user terminal 1. That is, the imaging device is disposed near either one end in the longitudinal direction of the display screen. Note that the vicinity of one end in the longitudinal direction of the user terminal 1 is, for example, a portion from one end in the longitudinal direction of the user terminal 1 toward the center in the longitudinal direction of the user terminal 1 and up to 1 / 3 of the length in the longitudinal direction of the user terminal 1. Also, the arrangement position of the camera 5 in the left-right direction is not particularly limited. For example, the camera 5 may be disposed near the center of the user terminal 1, or may be disposed at a position on the right side or the left side of the user terminal 1.

[0016] Also, when the user exists in the space that can be imaged by the camera 5, the camera 5 is arranged so that the user can visually recognize the user imaged by the camera 5 on the display screen 4. For example, the camera 5 and the display screen 4 are provided on the same surface of the user terminal 1. Thereby, when the user turns the display screen 4 toward himself / herself while holding the user terminal 1, the camera 5 photographs the user himself / herself, and the user himself / herself photographed is displayed on the display screen 4. The user can visually recognize himself / herself displayed on the display screen 4 while photographing himself / herself with the camera 5.

[0017] Preferably, the optical axis of the camera 5 is provided to be perpendicular to the display screen 4. The camera 5 can image the space in the direction perpendicular to the display screen 4. The camera 5 can also be said to be a so-called in-camera or front camera of the user terminal 1. Note that the user terminal 1 may be provided with other cameras not only on the back surface of the surface on which the display screen 4 of the user terminal 1 is provided but also on the back surface of the camera 5.

[0018] FIG. 3 is a block diagram showing the hardware configuration of a user terminal according to an embodiment of the present invention. The user terminal 1 includes a control unit 11, a RAM 12, a storage unit 13, an input unit 14, a display unit 15, and a communication interface 16, which are connected by an internal bus. Further, the user terminal 1 may include an imaging unit 17 such as a camera 5, a recording unit 18 such as a microphone or a recorder, and an output sound unit 19 such as a speaker.

[0019] The control unit 11 is composed of a CPU and a ROM. The control unit 11 executes a program stored in the storage unit 13 to control the user terminal 1. The RAM 12 is a work area of the control unit 11. The storage unit 13 is a storage area for storing programs and data. That is, the storage unit 13 functions as a recording medium storing a program. The control unit 11 performs arithmetic processing based on the program and data read from the RAM 12, and the data input by the input unit 14, the imaging unit 17, or the recording unit 18.

[0020] The display unit 15 has a display screen 4. The control unit 11 outputs a video signal for displaying an image on the display screen 4 according to the result of the arithmetic processing. Here, the display screen 4 of the display unit 15 may be a touch panel having a touch sensor. In this case, the touch panel functions as the input unit 14. The recording unit 18 converts the sound around the user terminal 1 into an electrical signal and acquires it as sound information. The recording unit 18 is an acquisition device capable of acquiring sound information. Note that the recording unit 18 may be provided in the user terminal 1 by attaching an external microphone.

[0021] The communication interface 16 can be connected to the communication network 3 wirelessly or by wire, and can transmit and receive data to and from other computer devices via the communication network 3. The data received via the communication interface 16 is loaded into the RAM 12, and arithmetic processing is performed by the control unit 11.

[0022] Further, the user terminal 1 may have a sensor unit. The sensor unit may include at least one or more sensors selected from the group consisting of a depth sensor, an acceleration sensor, a gyro sensor, a GPS sensor, a fingerprint authentication sensor, a proximity sensor, a magnetic sensor, a luminance sensor, a GPS sensor, and a barometric pressure sensor.

[0023] [Server device] FIG. 3 is a block diagram showing the hardware configuration of the server device according to an embodiment of the present invention. The server device 2 includes at least a control unit 21, a RAM 22, a storage unit 23, and a communication interface 24, which are connected by an internal bus.

[0024] The control unit 21 is composed of a CPU and a ROM, executes a program stored in the storage unit 23, and controls the server device 2. The control unit 21 also includes an internal timer for measuring time. The RAM 22 is a work area of the control unit 21. The storage unit 23 is a storage area for storing programs and data. That is, the storage unit 23 functions as a recording medium storing a program. The control unit 21 reads a program and data from the RAM 22, and performs execution processing of the program based on information received from the user terminal 1 and the like.

[0025] Further, the program may be stored in a recording medium such as a CD-ROM. In this case, the program stored in the recording medium may be installed in the user terminal 1 or the server device 2 to execute a predetermined function.

[0026] Alternatively, the program may be distributed from a computer device external to the system. In this case, the program distributed from the computer device external to the system may be installed in the user terminal 1 or the server device 2 to execute a predetermined function.

[0027] [Evaluation process] The system according to an embodiment of the present invention includes an evaluation process for acquiring the state when the user speaks and evaluating the user's speech. FIG. 5 is a diagram showing a flowchart of the evaluation process according to an embodiment of the present invention. The user logs in to the system by starting the application program downloaded to the user terminal 1 and accessing the server device 2. The user may log in to the system by accessing the server device 2 from the user terminal 1 via a web browser.

[0028] When the user logs in to the system for the first time, the user may perform initial settings by operating the user terminal 1 and inputting user information regarding the user. As initial settings, the user inputs user information such as a user name, password, gender, date of birth, etc. Next, by transmitting the input user information to the server device 2, the user is registered in the server device 2. The registered user may be assigned user identification information (hereinafter also referred to as user ID) for identifying the user. The user ID is stored in association with the user information. Also, in the system, a page dedicated to the registered user is created.

[0029] When logging in to the system, it may be required to input a pre-registered user ID and password. Alternatively, when logging in to the system, it may be required to input a pre-registered email address and password.

[0030] Next, the user inputs the start of the evaluation process on the user terminal 1 (step S1). In step S1, the user selects an evaluation item. The information regarding the start of the input evaluation process is transmitted from the user terminal 1 to the server device 2 and received by the server device 2. Thereby, the server device 2 accepts the start of the evaluation process (step S2).

[0031] The evaluation items are items that indicate the content for evaluating the user's utterance. The user's utterance is the voice emitted by the user. For each evaluation item, the user's utterance that can be evaluated is different. The evaluation items are, for example, evaluation of fluency, evaluation of voice projection, evaluation of pronunciation, evaluation of volume, comprehensive evaluation, etc. The comprehensive evaluation evaluates at least two or more of fluency, voice projection, pronunciation, and volume respectively. The evaluation of fluency is an evaluation of the smoothness of pronunciation and voice projection, and includes, for example, an evaluation of whether a phrase of 1 can be spoken without hesitation. The evaluation of voice projection includes, for example, an evaluation of whether the voice is clear, an evaluation of whether abdominal breathing can be done, an evaluation of the skill and the number of times per unit time of techniques such as vibrato, fist, and sob. The evaluation of pronunciation is an evaluation of the way of making sounds, and includes, for example, an evaluation of whether a predetermined character or word can be pronounced appropriately. Information regarding the evaluation items is stored in the server device 2 or the user terminal 1.

[0032] The evaluation items are set for each language. The languages are, for example, Japanese, English, Chinese, Spanish, German, etc. Also, the evaluation items may be items for evaluating the user's utterance by occupation. The occupations are, for example, announcer, cabin attendant, receptionist, vocal music, etc.

[0033] One or more pieces of instruction information are associated with one evaluation item. When a plurality of pieces of instruction information are associated with one evaluation item, the order in which the instruction information is output is stored in the server device 2 or the user terminal 1.

[0034] For example, the evaluation item for evaluating the fluency of Japanese is associated with instruction information A. The evaluation item for evaluating the pronunciation of Japanese is associated with instruction information B. Similarly, the evaluation item for evaluating the pronunciation of "r" in English is associated with instruction information X. The evaluation item for evaluating the pronunciation of "l" in English is associated with instruction information Y. The evaluation item for evaluating the pronunciation of English may include instruction information X and instruction information Y.

[0035] Here, the instruction information is instruction information for requesting a user to perform a predetermined utterance, and includes information regarding an instruction for causing the user to perform the utterance. The predetermined utterance requested of the user is also referred to as the utterance requested of the user.

[0036] The utterance requested of the user refers to, for example, the utterance to be made by the user in order to acquire video information or sound information necessary for evaluating the user's utterance. Specifically, the utterance requested of the user may be a predetermined phrase, a predetermined song, etc. The predetermined phrase, the predetermined song, etc. are stored in the server device 2 in advance. Examples of the predetermined phrase include words, character strings, phrases, fast talk, etc. The utterance requested of the user may be not only Japanese but also a foreign language such as English or Chinese. The video information or sound information to be acquired includes, for example, video information while uttering a predetermined phrase, etc., sound information corresponding to the voice uttering a predetermined phrase, etc. Note that the video information includes one or more still images, and is information in which the still images and time information are associated.

[0037] Examples of the instruction information include text information regarding the utterance requested of the user, moving image information regarding the utterance, and / or sound information regarding the utterance. It may be any one of the text information, the moving image information, or the sound information, or may be a combination of two or more. The moving image information includes still images and videos. The instruction information may use, as the moving image information, a video of the person making the utterance to be performed or an illustration.

[0038] The instruction information may include information indicating the distance from a predetermined part of the user to the camera 5 or information indicating the volume of the voice to be uttered. Specifically, the instruction information may include information indicating the position or size of the user's lips in the image included in the video information to be acquired, or information indicating the angle formed by a predetermined part of the user and the camera 5 with the horizontal plane.

[0039] In step S13 described below, in order for the server device 2 to evaluate the user's voice, it is necessary to acquire uniform video information regarding the brightness of the image, the position within the screen, the size occupied within the screen, the orientation or angle, etc. of a predetermined part of the user being imaged. Alternatively, it is necessary to acquire uniform sound information regarding the content and volume. The instruction information is used to acquire uniform video information and / or sound information at the user terminal 1.

[0040] The instruction information is stored in the server device 2 in association with each other for each voice required of the user. In the present embodiment, the instruction information may include first instruction information and second instruction information. The first instruction information and the second instruction information will be described later.

[0041] Returning to the description of the flowchart of FIG. 5, when receiving the start of the evaluation process in step S2, adjustment information for adjusting the brightness of the display screen 4 of the user terminal 1 is transmitted from the server device 2 to the user terminal 1 (step S3). The adjustment information is received at the user terminal 1 (step S4). When the adjustment information is received at the user terminal 1, adjustment processing for adjusting the brightness of the display screen 4 is executed according to the received adjustment information (step S5).

[0042] In the adjustment process, for example, it may be to adjust the brightness of the display screen 4 of the display device, to output brightness instruction information regarding the brightness for which the user is requested to adjust the brightness of the display screen 4 of the display device, or to output adjustment information for adjusting the brightness of the display screen 4 of the display device. The brightness of the display screen 4 of the display device is preferably set to the maximum brightness that can be set at the user terminal 1.

[0043] Also, in the adjustment process, for example, brightness instruction information instructing to photograph in a room with a predetermined brightness may be output. Further, in step S1, image information may be acquired by the camera 5, and the brightness of the acquired image information may be determined at the user terminal 1 or the server device 2. The user terminal 1 or the server device 2 may specify the brightness from the image information, and when the brightness is equal to or less than a predetermined threshold value, the processes of steps S3 to S5 may be executed.

[0044] In the adjustment process, when adjusting the brightness of the display screen 4, the display screen 4 may be automatically adjusted to a predetermined brightness on the user terminal 1. Information regarding the predetermined brightness is registered in the server device 2 in advance. The predetermined brightness is a brightness determined in advance and is preferably the maximum brightness.

[0045] Alternatively, in the adjustment process, when outputting brightness instruction information regarding the brightness for which the user is requested to adjust the brightness of the display screen 4, text information, moving image information, and / or sound information may be output as the brightness instruction information regarding the brightness. The user can adjust the brightness of the display screen 4 according to the brightness instruction information regarding the brightness.

[0046] Alternatively, in the adjustment process, when outputting adjustment information for adjusting the brightness of the display screen 4, a brightness adjustment screen or a brightness adjustment bar may be displayed on the display screen 4 as the adjustment information. The user can adjust the brightness of the display screen 4 according to the adjustment information.

[0047] Alternatively, in the adjustment process, when adjusting the brightness of the display screen 4 of the display device, the processes of steps S3 to S5 may be executed simultaneously with steps S9 and 10 described later. For example, in steps S9 and S10, the display screen 4 of the user terminal 1 may be automatically adjusted to a predetermined brightness only while acquiring moving image information. The brightness adjustment process of the display screen 4 in steps S9 and S10 may be executed in response to the voice requested by the user.

[0048] In the following step S13, in order to identify evaluation information regarding the user's voice from the acquired video information, it is necessary to acquire video information including an image of the user's lips. In one aspect of the present invention, when acquiring an image of the user, the camera 5 provided on the surface where the display screen 4 is provided is used. Generally, while a light is provided near the camera on the back of the user terminal 1, there may be no light provided near the camera 5 on the front of the user terminal 1. When using the camera on the back of the user terminal 1, the light provided on the back can be used, but when using the camera 5 on the front of the user terminal 1, there may be no light available. Therefore, in one aspect of the present invention, even when using the camera 5 on the front of the user terminal 1, by adjusting the brightness of the display screen 4, the user terminal 1 can acquire a clear image of the user's oral cavity and the back of the throat.

[0049] Also, when the distance between the user and the camera 5 is at a predetermined distance and the brightness of the display screen 4 is set to a predetermined brightness, video information including a clear image of the user's lips can be acquired using the brightness of the display screen 4. Also, since the display screen 4 is bright, the user can easily visually recognize the user's figure and the instruction information displayed on the display screen 4. Therefore, the user can appropriately acquire video information including the images necessary for evaluating the user's voice.

[0050] When the adjustment information is transmitted, instruction information is transmitted from the server device 2 to the user terminal 1 (step S6). At the user terminal 1, the instruction information is received (step S7). When the instruction information is received at the user terminal 1 (step S7) and the adjustment of the brightness of the display screen 4 in step S5 is completed, the first instruction information is output at the user terminal 1 (step S8). The first instruction information is instruction information for presenting to the user the information necessary for acquiring the video information or audio information before acquiring the video information or audio information necessary for evaluating the user's voice. Note that outputting includes not only displaying text information and / or moving image information on the display screen 4 but also outputting audio information from the speaker of the user terminal 1.

[0051] The first instruction information may be information regarding the specific words uttered by the user. For example, on the user terminal 1, first instruction information such as "Please pronounce 'a e i o u' three times." is output. Also, the first instruction information may be information regarding the specific singing uttered by the user. For example, on the user terminal 1, first instruction information such as information regarding the pitch like lyrics and musical scores is output.

[0052] When the first instruction information is output in step S8, on the user terminal 1, acquisition of the user's video information is started (step S9). The video information acquired in step S9 includes an image of the user's lips when speaking. In step S9, the sound information when the user is speaking may be acquired. Step S9 can be said to be a process of imaging the space by the imaging device, or a process of recording the sound by the recording device, or a process of imaging the space by the imaging device after the adjustment of the brightness of the display screen is completed, or a process of recording the sound by the recording device after the adjustment of the brightness of the display screen is completed.

[0053] Note that the imaging of the space by the imaging device and the output of the captured image to the display screen 4 may be executed simultaneously with step S8, and video information recording may be started in step S9. Also, the acquisition of the video information in step S9 may be started in response to an operation input to the user terminal 1.

[0054] When the acquisition of the user's video information is started, the second instruction information is output (step S10). Also, in step S10, while the acquisition of the user's video information continues on the user terminal 1, the images included in the simultaneously acquired video information are output. Note that the output has the same meaning as above. When acquiring the user's voice, the user's voice may not be output.

[0055] The second instruction information is instruction information for presenting to the user the information necessary to acquire video information or audio information required to evaluate the user's state during the acquisition of the video information or audio information. In step S10, the second instruction information is output as the instruction information. Note that depending on the utterance required of the user, the first instruction information and the second instruction information may be the same information.

[0056] In step S10, the image included in the video information obtained by imaging and the second instruction information which is the instruction information can be displayed on the display screen 4 of the display device. Thereby, since the user can simultaneously confirm the content of the utterance required of the user and their own imaged appearance, the user can acquire the video information of the user imaged within the preset position, size, and range.

[0057] As the first instruction information and / or the second instruction information, it is also possible to display a guide corresponding to at least a part of a human face in the area for displaying the image obtained by the imaging in step S9. The guide is, for example, a point, a line, or a figure combining these displayed on the display screen 4 for the user to confirm whether the position and angle of the imaging device with respect to the user are appropriate when imaging an image including the user's lips.

[0058] Examples of the types of guides include a solid line, a dotted line, a marker, an illustration, etc. Also, the guide may be one that partially varies in color and type within the guide. For example, the outline of the face may be displayed as a blue dotted line, and the nose part may be displayed as a red solid line. The guide is displayed superimposed on the captured image. The guide may be displayed in a transparent manner.

[0059] Examples of at least a part of a human face include the outline of a human face, predetermined parts such as the jaw and nose, etc. By displaying a guide corresponding to at least a part of a human face on the display screen 4, the user can confirm the captured image of themselves and adjust the position of the user terminal 1 and their own position so that their face fits within the dotted line of the guide.

[0060] Also, according to the instruction information, by changing at least a part of the type, position, size of the guide, and the face of the person corresponding to the guide, for each voice corresponding to the instruction information, the distance between the camera 5 and the user and the angle at which the camera 5 is directed at the user can be specified.

[0061] Next, by operating the user terminal 1, the acquisition of video information is terminated, and the video information obtained by imaging is transmitted from the user terminal 1 to the server device 2 (step S11). When the sound information is acquired in step S10, the sound information is transmitted in step S11. The server device 2 receives the video information (step S12). Note that the video information may be stored in the user terminal 1 and / or the server device 2.

[0062] For one type of voice, the processes of steps S6 to S12 are executed. Therefore, the processes of steps S6 to S12 are repeatedly executed until the output of all the instruction information included in the evaluation items is completed. In the repeated processes of steps S6 to S12, according to the voice corresponding to the instruction information, the text information regarding the voice, the moving image information regarding the voice, and / or the sound information regarding the voice displayed by the first instruction information and the second instruction information change.

[0063] When the output of all the instruction information is completed, the server device 2 analyzes the video information (step S13). When the video information is analyzed in step S13, the evaluation information regarding the user's voice is specified. Hereinafter, the evaluation information regarding the user's voice may be referred to as evaluation information. The evaluation information will be described later.

[0064] The evaluation information is transmitted from the server device 2 to the user terminal 1 (step S14), and the user terminal 1 receives the evaluation information (step S15). The evaluation information is displayed on the user terminal 1 (step S16). The user can grasp the evaluation regarding his / her own voice by checking the evaluation information.

[0065] In addition, by storing the evaluation information in the server device 2 in association with the user identification information, the evaluation information is registered (step S17). The evaluation information may also be stored in association with the time information. The time information may be any of the time when the evaluation process is received in step S2, the time when the instruction information is transmitted in step S6, the time when the video information is received in step S12, or the time when the evaluation of the video information is executed in step S13. In step S17, the video information or audio information evaluated in association with the evaluation information may be stored. By executing the processes of S1 to S17 above, the evaluation process ends.

[0066] In addition, in the above, the configuration in which the video information at the time of the user's voice is acquired by the user terminal 1 operated by the user has been described, but it is not limited to this. For example, the video information at the time of the user's voice may be acquired by the user terminal 1 operated by a second user different from the user. In this case, the video information at the time of the user's voice is acquired by the camera provided on the back surface of the user terminal 1.

[0067] [Analysis Process] Next, the analysis process regarding the evaluation of the video information executed in step S13 will be described. The analysis process is executed after all the video information necessary for evaluating the evaluation items has been received by the repetitive process of steps S6 to S12. Alternatively, the analysis process may be executed each time the video information is received in step S12, or may be executed in response to an operation input to the user terminal 1 after the video information is received in step S12.

[0068] By the analysis process in the server device 2, the evaluation information regarding the user's voice can be specified. The evaluation information of the user specified by the analysis process includes information regarding the quality of the voice. The evaluation information of the user corresponds to the evaluation item selected in step S1.

[0069] FIG. 6 is a diagram showing a flowchart of the analysis process according to an embodiment of the present invention. In the server device 2, video information is received (step S21). Next, in the server device 2, information regarding the position of a predetermined point of the lip image at the time of the user's voice (hereinafter, the information regarding the position of the predetermined point of the lip image is also referred to as lip information) is specified from the image included in the received video information (step S22).

[0070] The process of step S22 is a process of specifying information regarding the position of a predetermined point of the lip image from the video information. Then, in the server device 2, evaluation information is specified using the lip information as input data (step S23). Through the processes of steps S21 to S23 above, the analysis process ends. The processes of steps S21 to S23 are repeatedly executed for each piece of video information received in step S12.

[0071] In step S23, when the server device 2 receives the input of the lip information, the control unit 21 of the server device 2 uses the machine-learned prediction model to specify evaluation information regarding the user's voice based on the information regarding the position of a predetermined point of the lip image at the time of the user's voice. The information regarding the position of a predetermined point of the lip image at the time of the user's voice is the lip information specified in step S22. Thereby, it is possible to evaluate the user's voice corresponding to the video information acquired in step S10.

[0072] The prediction model used in step S23 is machine-learned with information regarding the position of a predetermined point of the lip image at the time of a person's voice as input information and evaluation information regarding the voice as output information. The machine learning algorithm is not particularly limited, and a known one can be used. As the input information, information regarding the position of a predetermined point of the lip image at the time of a person's voice (lip information) is used. On the other hand, as the output information, evaluation information regarding the person's voice is stored. The output information will be described later.

[0073] A process of identifying lip information from the images included in the video information in step S22 will be described. The lip information preferably changes over time. The lip information includes, for example, information regarding the coordinates of a predetermined point, information regarding the movement amount or speed of a predetermined point, or information regarding the area obtained by a predetermined point. Note that what lip information the server device 2 identifies changes according to the information input to the prediction model described later.

[0074] First, a process of identifying the coordinates of a predetermined point from the images included in the video information will be described. In the present embodiment, the predetermined point is an arbitrary point selected from the feature points of the face. For example, as the predetermined point, a point forming the contour of the lips, a point located in the cheek area, or a point located in the iris is selected. Information regarding the selected predetermined point is stored in the server device 2 in advance.

[0075] The process of step S22 is executed for each frame (image) included in the video information. First, the server device 2 identifies a face area corresponding to the user's face from one image. Next, the server device 2 identifies the feature points of each face from the identified face area. Then, from the identified feature points, the position of the predetermined point for identifying the information regarding the position in step S23 is identified. Note that the method for identifying the feature points of the face in step S22 is not particularly limited, and known methods can be utilized.

[0076] The identification of the predetermined point from the feature points of the face is executed as follows. Different numbers are assigned to the respective feature points of the face. Also, the assignment of the numbers of the feature points of the face is executed according to the same rule for each frame. For example, the 10th feature point of the face in the first frame and the 10th feature point of the face in the second frame indicate the same part of the face. Therefore, the predetermined point identified in step S22 is identified by presetting a specific number.

[0077] Note that which predetermined point is specified from the facial feature points can be set as appropriate. Also, the number of predetermined points specified from the facial feature points can be set as appropriate. The numbers and the number of the specified predetermined points are stored in the server device 2.

[0078] FIG. 7 is a diagram for explaining the predetermined points specified in step S22 according to the present embodiment. In FIG. 7, in one frame (image 100), a plurality of facial feature points are specified. Also, in FIG. 7, some of the specified facial feature points are superimposed and displayed on the image 100. The number of facial feature points and the number of predetermined points specified in step S22 are not particularly limited and can be increased or decreased as appropriate according to need. In FIG. 7, eight predetermined points 110 are specified in the lip region, two in the iris, two in the cheek portion, and two at the inflection point of the jaw.

[0079] The moving image information is information in which a plurality of still images are arranged in time series. Therefore, in the specific process of the coordinates of the predetermined points in step S22, the coordinates of the predetermined points are specified for each frame. The specified coordinates of the predetermined points are stored in the server device 2 in association with time or frame. The coordinates of the predetermined points and the information associated with time or frame can be input information for a prediction model described later. Also, it can be input information for the moving image information to be evaluated in step S23.

[0080] Next, a process of specifying information regarding the positions of other predetermined points from the coordinates of the specified predetermined points will be described.

[0081] For example, in step S22, the server device 2 can identify the movement amount or speed of a predetermined point based on the reference point of the image included in the video information. The reference point can be identified from the feature points of the face. As described above, since numbers are assigned to the feature points of the face, the reference point is identified by presetting a specific number. Which face feature point is used as the reference point can be set as appropriate. For example, the iris, the tip of the nose, the center point of the forehead, etc. can be used as the reference point. The movement amount of a predetermined point refers to the total movement distance of the predetermined point in a predetermined time. The speed of a predetermined point refers to the value obtained by dividing the total movement distance of the predetermined point in a predetermined time by the predetermined time.

[0082] The server device 2 can identify the movement amount or speed by obtaining the distance from the reference point to the predetermined point. Specifically, the server device 2 obtains the distance from the reference point to the predetermined point in a plurality of frames. Next, the server device 2 can identify the amount of change in distance by comparing the distance in one frame with the distance in the next frame. The server device 2 can identify the movement amount of the predetermined point by identifying the amount of change in the distance from the reference point to the predetermined point for each frame in which the predetermined point is identified. Also, the server device 2 can identify the speed of the predetermined point from the movement amount of the predetermined point and the time information corresponding to the frame in which the movement amount is identified. In this embodiment, the server device 2 can assume that the width of the black eye is a predetermined distance and convert the distance in the image into an actual length (for example, a length in meters).

[0083] Since the server device 2 identifies the movement amount or speed of a predetermined point based on the reference point, it can identify the actual movement amount or speed of the predetermined point excluding the shaking of the face, etc. For example, even when the user is speaking while shaking the face, the shaking of the user's face in the video information can be corrected to identify the movement amount or speed of the predetermined point.

[0084] Also, for example, in step S22, the server device 2 can identify the amount of change in the area at a predetermined point on a predetermined part of the user. For example, the server device 2 can obtain the area of the user's jaw. In the process of obtaining the area of the jaw, the server device 2 obtains a predetermined point that is a predetermined point on the face contour and is an inflection point in the lower right area and the lower left area of the face. Then, the server device 2 can obtain, as the area of the jaw, the portion surrounded by the straight line connecting the two inflection points and the predetermined points that form the face contour located below the inflection points. The inflection points and the predetermined points that form the face contour may be identified by setting specific numbers in advance. The server device 2 can compare the area in one frame with the area in the next frame to identify the amount of change in the area. The server device 2 can identify the amount of change in the area in the video information by comparing the areas for each frame in which the predetermined points are identified. For example, the server device 2 can identify the amount of change in the area in the video information by adding the absolute values of the amount of change in the area for each frame in which the predetermined points are identified.

[0085] Next, another example of the analysis process executed in step S13 will be described. FIG. 8 is a diagram showing a flowchart of another example of the analysis process according to an embodiment of the present invention. The analysis process is executed by the server device 2.

[0086] The server device 2 receives video information (step S31). In the server device 2, evaluation information is identified using the received video information (step S32). By steps S31 and S32, the analysis process ends.

[0087] In step S32, when the server device 2 receives the input of video information, the control unit 21 of the server device 2 uses the machine-learned prediction model to identify evaluation information regarding the user's voice based on the video information including an image of the user's lips when speaking. Thereby, the voice of the user corresponding to the video information acquired in step S10 can be evaluated.

[0088] The prediction model used in step S32 is machine-learned with video information including an image of the lips during a person's vocalization as input information and evaluation information regarding the vocalization as output information. The machine learning algorithm is not particularly limited, and known algorithms can be used. As the input information, video information including an image of the lips during a person's vocalization is used. On the other hand, as the output information, evaluation information regarding the person's vocalization is stored. The output information will be described later.

[0089] The prediction model used in step S23 or S32 is machine-learned with video information including an image of the lips during a person's vocalization or information regarding the position of a predetermined point of the image of the lips during a person's vocalization as input information and evaluation information regarding the vocalization as output information. The prediction model, input information, and evaluation information regarding the vocalization, which is the output information, in the system according to the present embodiment will be described in detail in the following Embodiments 1 to 3.

[0090] 〔Embodiment 1〕 The evaluation information regarding a person's vocalization includes, for example, information regarding the evaluation of smoothness of speech, pronunciation, voice projection, or the appropriateness of volume for the voice corresponding to the vocalization. Further, the evaluation information may be information comprehensively evaluated based on at least two or more evaluation items among smoothness of speech, pronunciation, voice projection, volume, etc. in a person's vocalization. In Embodiment 1, the evaluation information regarding the vocalization, which is the output information, is, for example, information obtained by an evaluator listening to the person's vocalization corresponding to the video information or the information regarding the position of a predetermined point of the image of the lips, which is the input information of the prediction model, and comprehensively evaluating the person's vocalization for each of a plurality of items.

[0091] In Embodiment 1, the input information and output information used in the prediction model can be obtained as follows. For example, the evaluated person A makes a sound for a predetermined word. An evaluator evaluates the voice during the sound production according to a predetermined standard for items related to the sound production (for example, smoothness of speech, voice projection, pronunciation, and volume). This evaluation may be performed by one or more evaluators, or a known evaluation device for smoothness of speech, voice projection, pronunciation, or volume may be used. The result of this evaluation represented as a numerical value, or the result ranked step by step using symbols, which is converted into data, becomes the output information used in the prediction model. Also, video information of the evaluated person A during the above-mentioned sound production by the evaluated person A is imaged, and information regarding the position of a predetermined point of this video information or an image of the lips obtained from the video information can be used as the input information for the prediction model.

[0092] In this way, regarding the information on the position of a predetermined point of the video information or the image of the lips during one sound production by the evaluated person A, it is used as the input information for the prediction model, and regarding the information on the evaluation result of the sound production during that sound production, it is used as the output information for the prediction model. For a plurality of other evaluated persons different from the evaluated person A, in the same way, input information and output information are obtained, and machine learning is performed based on these plurality of input information and output information.

[0093] Note that the items related to the sound production to be evaluated may be all the set items, may be one item, or may be any combination of items.

[0094] The evaluator may be, for example, an expert in the items related to the sound production to be evaluated. Also, the evaluation information may be for evaluating a person's voice by language. It is preferable that the evaluation of the items related to the sound production is performed for each language. For example, when evaluating a person's voice in English, the evaluator is a person whose native language is English.

[0095] In addition, the evaluation of items related to voice may be performed for each occupation. Since the evaluation criteria for items related to voice differ depending on the occupation, even when evaluating the voice of a person corresponding to information on a predetermined point position of the same video information or lip image, the scores or ranks used as output information will be different. For example, when the occupation is an announcer, the evaluator is an announcer, and when the occupation is a stage actor, the evaluator is a stage actor. The evaluation of items related to voice being performed for each language or occupation is the same in the following Embodiment 2 and Embodiment 3.

[0096] 〔Embodiment 2〕 The evaluation information related to a person's voice includes, for example, information related to the evaluation of a specific smoothness of speech, a specific pronunciation, or a specific voice. In Embodiment 2, the evaluation information related to voice as output information is, for example, obtained by an evaluator listening to the voice of a person corresponding to information on a predetermined point position of video information or lip image which is the input information of a prediction model, and evaluating one evaluation item for the voice of the person.

[0097] In Embodiment 2, the input information and output information used in the prediction model can be obtained as follows. The descriptions similar to those in Embodiment 1 are omitted. For example, the evaluated person A makes a voice for a predetermined word. The evaluator evaluates, according to a predetermined criterion, what kind of voice it was specifically for one item related to voice (for example, smoothness of speech, voice projection, pronunciation, or volume) during the voice. The evaluator performs an evaluation, for example, regarding one item, such as what pronunciation the voice during the voice corresponds to, or whether the voice of the voice is good or bad, correct or incorrect, or preferable, ordinary, or not preferable. The data obtained from this evaluation becomes the output information used in the prediction model. In addition, the video information of the evaluated person A during the above voice is imaged, and the video information or information on the position of a predetermined point of the lip image obtained from the video information can be used as the input information used in the prediction model.

[0098] In this way, information regarding the position of a predetermined point in video information or an image of lips at the time of the utterance of "1" by the evaluated person A is used as input information, and information regarding the evaluation result of the utterance at the time of the utterance is used as output information for the prediction model. For a plurality of other evaluated persons different from the evaluated person A, input information and output information are obtained in the same way, and machine learning is performed based on these plural pieces of input information and output information.

[0099] As specific output information, for example, regarding the voice of a person at the time of utterance corresponding to the information regarding the position of a predetermined point in the video information or the image of the lips, which is the input information, the evaluator evaluates whether the person's pronunciation is the pronunciation of "r" or the pronunciation of "l", and the result of determining which pronunciation it is can be used as the output information.

[0100] In addition, in Embodiments 1 and 2, for a plurality of evaluated persons, a prediction model was created based on a plurality of combinations of input information and output information when evaluating the utterance of each. However, for example, the same evaluated person may utter a sound multiple times, and a prediction model may be constructed based on a plurality of combinations of input information and output information when evaluating the utterance of each time.

[0101] 〔Embodiment 3〕 The evaluation information regarding a person's utterance includes, for example, information regarding the evaluation of a specific occupation or a specific attribute. In Embodiment 3, the evaluation information regarding the utterance, which is the output information, may be, for example, the one obtained by judging a person's utterance according to the attribute of the person corresponding to the information regarding the position of a predetermined point in the video information or the image of the lips, which is the input information of the prediction model.

[0102] In Embodiment 3, the input information and output information used in the prediction model can be obtained as follows. Descriptions similar to those in Embodiment 1 are omitted. For example, the evaluated person A makes a voice for a predetermined phrase. Information regarding the attributes of the evaluated person A who made the voice becomes output information used in the prediction model as evaluation information. Also, the video information of the evaluated person A during the above-mentioned voice is imaged, and information regarding the position of a predetermined point of this video information or the lip image obtained from the video information can be used as the input information for the prediction model. In this way, regarding the information regarding the position of a predetermined point of the video information or the lip image at the time of one voice by the evaluated person A as the input information, and the information regarding the attributes of the evaluated person who made the voice as the output information, it is used in the prediction model. For a plurality of evaluated persons different from the evaluated person A, in the same way, input information and output information are obtained, and machine learning is performed based on these plurality of input information and output information.

[0103] As specific output information, for example, when constructing a prediction model for evaluating a person's voice with the occupation of an announcer, for the person corresponding to the information regarding the position of a predetermined point of the video information or the lip image which is the input information, it is specified whether the person is an announcer with a work history of 10 years or more, an announcer with a work history of less than 3 years, or an ordinary person, and the information regarding the specified attributes is used as the output information.

[0104] In this case, in step S13, the server device 2 can specify which of an announcer with a work history of 10 years or more, an announcer with a work history of less than 3 years, or an ordinary person the voice of the user corresponding to the input video information is closest to.

[0105] In addition, in the present embodiment, examples of the machine learning algorithm include linear regression, multiple regression analysis, support vector machine, decision tree, random forest, and deep learning using a multi-layer neural network. The multi-layer neural network has an input layer, an output layer, and a plurality of intermediate layers. Weights are set on the edges connecting the nodes of each layer. Weights corresponding to each input to the node are set on the edge, and the weights corresponding to each input to the node are multiplied, and the value obtained by multiplying these weights and the bias are added. The value obtained by the addition is non-linearly transformed using an activation function to calculate an activation value. The calculated activation value becomes the input value passed to the nodes of the next layer. The number of intermediate layers can be appropriately designed. The weights are optimized by the above-described teacher data.

[0106] Note that the acquired video information and the output evaluation information differ depending on the evaluation item selected by the user. According to the evaluation item received in step S1, the server device 2 selects a prediction model used to execute the evaluation process in step S13. The acquired video information is information corresponding to the input information of the prediction model, and the output evaluation information is information corresponding to the output information of the prediction model.

[0107] The evaluation information specified in step S23 or S32 and displayed in step S16 is evaluation information for evaluating the user's voice. The evaluation information may be displayed, for example, on the user terminal 1 as a numerical value such as 1 to 100 points for the user's voice, or may be displayed as a hierarchical evaluation such as ABC. The evaluation information may display, for example, on the user terminal 1 what pronunciation the user's voice corresponds to, or what attributes the user's voice corresponds to. By displaying these, the user can recognize the quality of his or her own voice.

[0108] In step S16, information regarding advice may be displayed together with the evaluation information. The information regarding advice is specified by the server device 2 based on the video information acquired in step S10. For example, the information regarding advice includes information for bringing the video information at the time of the voice of a person who has been evaluated as good closer to the video information of the user by comparing the video information at the time of the voice of a person who has been evaluated as good with the video information of the user. Specifically, the user terminal 1 may display specifically which part of the user's lips should be moved and how.

[0109] Also, in step S16, the user terminal 1 can reproduce the model video information or sound information and the video information or sound information acquired in step S10. The model video information or sound information may be stored in the server device 2 in advance, or may be video information or sound information determined to have a high evaluation in the output information. Thereby, the user can qualitatively understand the difference between the model and his / her own voice.

[0110] Also, in step S16, the user terminal 1 may compare and display the amount of movement, speed, and change in area of a predetermined point corresponding to the video information acquired in step S10 and the model video information. The model video information may be stored in the server device 2 in advance, or may be video information determined to have a high evaluation in the output information. Thereby, the user can quantitatively understand the difference between the model and his / her own voice.

[0111] Also, in step S16, the user terminal 1 may display the video information of the user acquired in step S10 and the video information of the user acquired in the past, and may display the evaluation information of the user specified in step S13 and the evaluation information of the user registered in the past. The past evaluation information or past video information to be displayed may be selected according to an operation on the user terminal 1, or may be associated with the time closest to the time associated with the evaluation information in step S17. That is, the evaluation information or video information acquired one time before and the evaluation information or video information acquired this time may be compared.

[0112] In step S16, when the user terminal 1 displays the user's video information acquired in step S10 and the user's video information acquired in the past, the user's video information acquired in step S10 may be displayed so as to superimpose the user's video information acquired in the past. When two or more pieces of video information are superimposed and displayed, the user terminal 1 may display by changing the transparency of either piece of video information. Further, when two or more pieces of video information are superimposed and displayed, a face region may be specified from one piece of video information, the specified face region may be extracted and superimposed on the other piece of video information for display. Thereby, the user can recognize whether the current or past lip movement represents the user's lip movement.

[0113] Each piece of video information can be superimposed and displayed based on the reference points included in the video information. The reference points can be specified by the method described in the process of step S22. Thereby, the user can qualitatively understand the difference between the lip movement when the user was speaking in the past and the lip movement when the user is speaking currently.

[0114] Further, in step S16, the user terminal 1 can reproduce the user's video information or audio information acquired in step S10 and the user's video information or audio information acquired in the past, respectively. Also, the user terminal 1 can reproduce the audio information acquired in step S10 or the audio information of the user acquired in the past simultaneously with the reproduction of the video information, with the user's video information acquired in step S10 in a state where the user's video information acquired in the past is superimposed. Thereby, the user can qualitatively understand the difference between the user's past and current voices.

[0115] In step S16, when the user terminal 1 displays the evaluation information of the user specified in step S13 and the evaluation information of the user registered in the past, the evaluation information of the user specified in step S13 and the evaluation information of the user registered in the past are displayed so that they can be compared. The user terminal 1 may simultaneously display the past score or rank and the current score or rank for each evaluation item, or may compare and display the movement amount, speed, and change amount of a predetermined point corresponding to the video information acquired in step S10 and the user's own past video information. Thereby, the user can quantitatively understand the difference between his / her past and current voices.

[0116] Also, in step S16, the user terminal 1 may be able to play back the video information or audio information acquired by another user terminal operated by another user and the video information or audio information acquired by its own user terminal 1. The video information or audio information acquired by the other user terminal and the user terminal 1 is stored in the server device 2. The video information or audio information acquired by another user terminal played back by the user terminal 1 may be the video information or audio information corresponding to a pre-registered other user, or may be the video information or audio information corresponding to a randomly specified other user.

[0117] Also, it may be possible to register a group of users who can display video information or audio information in the user terminal 1 in advance by operating the user terminal 1. The users registered in the group are stored in the server device 2 by associating information that can identify the group with the user ID. The users registered in the group can play back each other's video information or audio information on their respective user terminals. The user can qualitatively understand the difference in voices from others, which will motivate the user to acquire video information or audio information.

[0118] In addition, when video information acquired by the user terminal 1 or another user terminal is played back on the user terminal 1, video information processed so that the actual face of the user included in the video information cannot be recognized may be played back. For example, on the user terminal 1, video information subjected to processing such as replacing the user's face area with an avatar, blurring the user's face area, or displaying only the feature points of the face in the user's face area may be played back.

[0119] Also, in step S16, the user terminal 1 may compare and display evaluation information corresponding to the video information acquired in step S10 and the video information acquired by another user terminal, the movement amount of a predetermined point, the speed, and / or the change amount of the area. The evaluation information etc. corresponding to the video information acquired by another user terminal displayed on the user terminal 1 may be the evaluation information etc. corresponding to a pre-registered other user, or the evaluation information etc. corresponding to a randomly specified other user. Similarly to the above, it may be possible to register in advance a group of users for whom the user terminal 1 can display evaluation information etc. by operating the user terminal 1. Users registered in the group can display each other's evaluation information etc. on their respective user terminals. Thereby, the user can quantitatively understand the difference in voices from others, and it becomes a motivation for the user to practice so as to make a voice with higher evaluation.

[0120] Note that the above prediction model may be further machine-learned using, as input information, video information including an image of the lips when a person speaks or information regarding the position of a predetermined point of the image of the lips when a person speaks, and sound information when the person speaks, and using, as output information, evaluation information regarding the speech. In this case, in step S23 or S32, the server device 2 uses the above prediction model to specify evaluation information regarding the user's speech based on video information including an image of the user's lips when the user speaks or information regarding the position of a predetermined point of the image of the user's lips when the user speaks, and the sound information when the user speaks. The sound information when the user speaks is the sound information acquired in step S10.

[0121] Note that, in the prediction model, the video information including the image of the lips when a person speaks, which is input as input information, may be subjected to the same processing as that in step S22 in the server device 2 when constructing the prediction model. The process of specifying information on the position of a predetermined point in step S22 is similarly executed as the process of specifying information on the position of a predetermined point from the video information of the evaluation target acquired in step S10. The process executed on the video information to obtain the input information of the prediction model is the same as the process executed on the video information acquired in step S10.

[0122] Note that, for all frames included in the video information, the process of specifying the coordinates of a predetermined point may be executed, or for the frames selected at a predetermined interval, the process of specifying the coordinates of a predetermined point may be executed. Also, the video information for which the process of specifying the coordinates of a predetermined point is executed may be the entire video information received in step S21 or a part of the video information. A part of the video information is, for example, video information corresponding to the part where the user is speaking. A part of the video information may be specified in correspondence with sound information.

[0123] Note that if the face region identification process and the face feature point identification process in step S22 are executed in the first frame, in the frames after the second frame, it may be possible to use the face region and face feature points identified in the first frame for identification.

[0124] Note that in the present embodiment, the lip information may be any one or more pieces of information among information on the coordinates of a predetermined point, information on the movement amount or speed of a predetermined point, and information on the area obtained by a predetermined point. For example, the system according to the present embodiment may use a prediction model that is machine-learned with the movement amount of a predetermined point in the image of the lips when a person speaks as input information and the evaluation information related to the voice as output information, and based on the movement amount of a predetermined point in the image of the lips when the user speaks, specify the evaluation information related to the user's voice.

[0125] In the above-described embodiment, a mode has been described in which video information is transmitted to the server device 2 and the server device 2 executes analysis processing of the video information. However, the user terminal 1 may execute the analysis processing of the video information. In this case, the processes of transmitting and receiving the video information in steps S11 and S12, and the processes of transmitting, receiving, and displaying the evaluation information in steps S14 to S16 can be omitted. Note that when the analysis processing is executed in the user terminal 1, necessary programs are stored in the user terminal 1 in advance. The programs may be transmitted from the server device 2.

[0126] In the above-described embodiment, a mode has been described in which the processes of acquiring and transmitting the video information in steps S1 to S11 and the process of displaying the evaluation information in steps S15 and S16 are executed by the same user terminal 1. However, the processes of steps S1 to S11 and steps S15 and S16 may be executed by different user terminals 1. In this case, the user terminal 1 that executes steps S1 to S11 only needs to include the imaging unit 17. Also, the user terminal 1 that executes steps S15 and S16 only needs to include the display unit 15.

[0127] Note that the processes of steps S3 to S5 may be omitted. Also, the processes of steps S6 to S12 do not necessarily have to be repeatedly executed. Also, the processes after step S13 may be executed in response to an operation input of the user terminal 1. For example, the processes from steps S1 to S12 and the processes after step S13 may be executed on different days. Also, the process of displaying the evaluation information in steps S15 and S16 may be executed at an arbitrary timing of the user. For example, when the user logs in to the system, the evaluation information may be transmitted from the server device 2 and the evaluation information may be displayed on the user terminal 1.

[0128] Note that the start of obtaining the video information in step S9 may be determined by the user terminal 1 or the server device 2 as to whether the user is in focus, and may be executed when it is determined that the user is in focus. When the acquisition of the video information is started by the user terminal 1, the user terminal 1 may output information for notifying that the acquisition of the video information has started.

[0129] Note that the sound information acquired in step S10 is information obtained by converting the sound around the user terminal 1 acquired by the recording unit 18 of the user terminal 1 into an electrical signal. The sound around the user terminal 1 includes not only the voice uttered by the user but also noise other than the voice uttered by the user. In step S23 or S32, when the sound information is input, the server device 2 may input the sound information from which the noise has been removed. As a method for removing the noise, a method for removing known noise components can be adopted. For example, the server device 2 adopts methods such as filtering using the spectral subtraction method, spectral restoration, and human voice models to remove the ambient sound from the sound information.

[0130] According to one aspect of the present invention, in order to specify evaluation information regarding the user's voice based on video information including an image of the user's lips during voice utterance or information regarding the position of a predetermined point of the image of the user's lips during voice utterance, by inputting the video information, the evaluation information regarding the user's voice can be specified.

[0131] According to one aspect of the present invention, since the prediction model is further machine-learned with video information including an image of a person's lips during voice utterance or position information of a predetermined point of the image of a person's lips during voice utterance, and sound information during the person's voice utterance as input information, and evaluation information regarding the voice as output information, the evaluation information regarding the user's voice can be specified with higher accuracy.

[0132] According to one aspect of the present invention, since information regarding the position of a predetermined point in an image of lips during human vocalization includes information regarding the movement amount or speed of the predetermined point, or information regarding the area obtained by the predetermined point, evaluation information regarding the user's vocalization can be specified from video information.

[0133] According to one aspect of the present invention, in order to specify the movement amount or speed of a predetermined point based on a reference point of an image, even when the user's face in video information moves independently of vocalization, the movement amount or speed of the predetermined point derived from vocalization can be specified.

[0134] According to one aspect of the present invention, since the evaluation information regarding vocalization includes information regarding the quality of vocalization, the user can understand whether their vocalization is good or bad.

Explanation of Reference Numerals

[0135] 1 User terminal, 2 Server device, 3 Communication network, 11 Control unit, 12 RAM, 13 Storage unit, 14 Input unit, 15 Display unit, 16 Communication interface, 17 Imaging unit, 18 Recording unit, 19 Sound output unit, 21 Control unit, 22 RAM, 23 Storage unit, 24 Communication interface 100 Image, 110 Predetermined point

Claims

1. A system comprising at least one computer device, wherein: Using a prediction model that has been machine-learned with video information including an image of lips during a person's speech or information regarding the position of a predetermined point on the image of the lips during a person's speech as input information, and evaluation information regarding the speech as output information, based on video information including an image of the user's lips during the user's speech or information regarding the position of a predetermined point on the image of the user's lips during the user's speech, a specifying means for specifying evaluation information regarding the user's speech. A system comprising the above.

2. The prediction model is further machine-learned with video information including an image of lips during a person's speech or information regarding the position of a predetermined point on the image of the lips during a person's speech, and audio information during the person's speech as input information, and evaluation information regarding the speech as output information. The specifying means uses the prediction model to specify evaluation information regarding the user's speech based on video information including an image of the user's lips during the user's speech or information regarding the position of a predetermined point on the image of the user's lips during the user's speech, and audio information during the user's speech. The system according to claim 1.

3. The information regarding the position of a predetermined point on the image of the lips during the person's speech includes information regarding the amount of movement or speed of the predetermined point, or information regarding the area obtained by the predetermined point. The system according to claim 1 or 2.

4. A position specifying means for specifying information regarding the position of a predetermined point on the image of the lips from video information including an image of the user's lips during the user's speech. Comprising: The position specifying means specifies the amount of movement or speed of the predetermined point based on a reference point of the image. The system according to claim 1 or 2.

5. The evaluation information regarding the speech includes information regarding whether the speech is good or bad. The system according to claim 1 or 2.

6. A program that causes a computer device to function as: Using a prediction model that has been machine-learned with video information including an image of lips during a person's speech or information regarding the position of a predetermined point on the image of the lips during a person's speech as input information, and evaluation information regarding the speech as output information, based on video information including an image of the user's lips during the user's speech or information regarding the position of a predetermined point on the image of the user's lips during the user's speech, a specifying means for specifying evaluation information regarding the user's speech. A program.

7. A method executed in at least one computer device, using a prediction model that is machine-learned with video information including an image of lips during human speech or information regarding the positions of predetermined points of an image of lips during human speech as input information and evaluation information regarding the speech as output information, to identify evaluation information regarding a user's speech based on video information including an image of the user's lips during speech or information regarding the positions of predetermined points of an image of the user's lips during speech; a specifying step The method comprising the steps above.

Citation Information

Patent Citations

  • Language training system, language training terminal and language training program

    JP2017083658A

Cited By

  • Apparatus for monitoring operation of disconnecting switch

    KR1020250060852A