Image processing device, image processing method, learning model production method, and program

The image processing device uses a learning model trained on motion and non-motion sections of body parts to recognize actions like speech without specifying the exact time frame, enhancing the accuracy of speech recognition in moving images.

JP7752479B2Active Publication Date: 2025-10-10GLORY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021046762
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-22
Publication Date
2025-10-10
Estimated Expiration
2041-03-22

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to accurately identify the vocalization period of each character in a moving image, making it difficult to recognize the content of a person's speech without specifying the exact time frame of the action.

Method used

An image processing device generates a learning model using video data that includes both motion and non-motion sections of specific body parts, such as lip movements, to identify actions like speech without specifying the exact time frame, utilizing machine learning to recognize movements like turning the face, moving eyeballs, blinking, rotating the head, or waving hands.

Benefits of technology

Enables the recognition of the content of a person's actions in a moving image without needing to specify the exact time of the action, improving the accuracy of speech recognition in image processing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752479000001
    Figure 0007752479000001
  • Figure 0007752479000002
    Figure 0007752479000002
  • Figure 0007752479000003
    Figure 0007752479000003
Patent Text Reader

Abstract

To provide a technique for recognizing contents of a motion of a person without specifying an execution period of the motion in a video image.SOLUTION: An image processing apparatus 30 generates a learning model which receives, as input, a video image 330 obtained by imaging a subject person and outputs information on a motion of the subject person. The image processing apparatus 30 executes machine learning of the learning model using multiple pieces of training data that receive, as input, the video image 330 obtained by imaging a person and including a non-motion section (non-speech section, etc.) B of a specific section (lip, etc.) of the person, and output information 350 (speech word, etc.) on the motion (speech motion, etc.) of the specific section. The image processing apparatus 30 specifies the motion of the subject person on the basis of the learning model.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing technique for processing an image to recognize an action, and to techniques related thereto. [Background technology]

[0002] There are technologies that process images to recognize (identify) actions (action details). For example, in the technology of Patent Document 1, when person authentication is performed by image recognition processing using a person's lip image (mainly a still image), the content of the person's utterance (utterance character string) is recognized as a result.

[0003] In the technology of Patent Document 1, facial image templates corresponding to the pronunciation of the letters "a," "i," "u," "e," and "o" are pre-registered for multiple users (individuals). During authentication, information about the person to be authenticated (such as name and ID number) is sent to an authentication server, and the user corresponding to the information (such as name and ID number) is identified. Next, a random character string (e.g., "a," "e," "i," and "u") is generated, and the user is requested to pronounce the character string. A video of the user pronouncing the character string is captured. The authentication server (authentication device) acquires a captured image (a facial image of the user) from a terminal when the character string is pronounced. The authentication server then compares the features of the user's pre-registered facial image template (a speech pattern image of a specific individual) for each character with the features of a facial image extracted from the captured image at the time of authentication (when the character is spoken). Based on the comparison result, the authentication server authenticates whether the person in the captured image is a genuine user. In detail, if it is determined that the person is speaking the specified character string, the person is determined to be a genuine user. In other words, a combination of multiple facial image features corresponding to a random character string is used as an authentication key that changes each time, and the person's authentication process is executed. As a result of this process, the content of the person's speech (the spoken character string) is also recognized. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-306090 Summary of the Invention [Problem to be solved by the invention]

[0005] The technology of Patent Document 1 uses facial images (still images) of characters (such as "a" and "e") being spoken, so it is necessary to identify the period during which each character is spoken. For example, when extracting facial images of a person speaking each character from a moving image (photographed image), it is necessary to identify the period during which each character is spoken (the execution period (action period) of the speaking action of each character) within the moving image.

[0006] However, it is not necessarily easy to identify the vocalization period of each character in a moving image (the period during which each character's vocalization action is performed).

[0007] Therefore, an object of the present invention is to provide a technique that can recognize the content of a person's action without specifying the period during which the action is performed in a moving image. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems, the image processing device according to the present invention is an image processing device that generates a learning model by inputting a video image of a subject person and outputting information about the motion of the subject person. The image processing device includes a control unit that executes machine learning of the learning model using a plurality of training data that receives as input a video image of a subject person, the video image including a motion section of a specific body part of the subject person and a non-motion section of the specific body part both before and after the motion section, and outputs information about the motion of the specific body part, and the video image used as training data includes the motion section in which the motion of the specific body part was performed within a specified period. The action is characterized in that it is a speech action. To solve the above problem, an image processing device according to the present invention is an image processing device that receives a video of a subject person and generates a learning model that outputs information about the subject person's movements. The image processing device includes a control unit that executes machine learning of the learning model using a plurality of training data that receives a video of a subject person, the video including a section in which a specific body part of the subject is in motion and a section in which the specific body part is not in motion both before and after the section in which the specific body part is not in motion, and outputs information about the movements of the specific body part, the video used as training data includes the section in which the movement of the specific body part was performed within a specified period, and the movements are any of a movement of turning a face in a specific direction, a movement of moving eyeballs in a specific direction, a movement of blinking at least one of both eyes, a movement of rotating a head in a specific direction, a movement of raising at least one of both hands, and a movement of waving at least one of both hands in a specific direction.

[0010] The video used as training data may be a video captured of a person's actions in response to an action instruction to complete a specified action related to the specific body part within a specified period of time.

[0011] In order to solve the above-described problems, the image processing device according to the present invention includes a control unit that identifies the movements of a subject person based on a learning model that has been machine-learned using a plurality of training data that receives as input moving images of a person, the moving images including movement sections of specific parts of the subject person and non-movement sections of the specific parts both before and after the movement sections, and outputs information about the movements of the specific parts, the moving images used as training data include the movement sections in which the movements of the specific parts were performed within a specified period, and the control unit identifies the movements of the subject person based on the information about the movements of the specific parts that is output from the learning model in response to inputting moving images of the movements of the subject person, the moving images including non-movement sections of the specific parts of the subject person, into the learning model. The action is a speech action. It is characterized by: In order to solve the above problem, the image processing device of the present invention includes a control unit that identifies the movements of a subject person based on a learning model that has been machine-learned using multiple training data that takes as input moving images of a person, the moving images including movement sections of specific parts of the subject person and non-movement sections of the specific parts both before and after the movement sections, and outputs information about the movements of the specific parts, the moving images used as training data include the movement sections in which the movements of the specific parts were performed within a specified period, and the control unit identifies the movements of the subject person based on information about the movements of the specific parts that is output from the learning model in response to inputting moving images of the movements of the subject person, the moving images including non-movement sections of the specific parts of the subject person, into the learning model, and the movements are any of a movement of turning the face in a specific direction, a movement of moving the eyeballs in a specific direction, a movement of blinking at least one of both eyes, a movement of rotating the head in a specific direction, a movement of raising at least one of both hands, and a movement of waving at least one of both hands in a specific direction.

[0012] The specific area includes the lips. That's fine.

[0013] The subject person is a person who is the target of the authentication process, and the control unit may determine that the subject person is a genuine person with respect to the authentication process, on the condition that it is determined that the subject person has performed the specified action based on the output obtained by inputting the video image of the subject person into the learning model.

[0014] The specific part may include lips, and the movement of the specific part may be a speech movement.

[0015] The movement of the specific body part may be a speaking movement of one word, and the information relating to the movement of the specific body part may be information indicating the spoken word.

[0016] In order to solve the above-mentioned problems, the present invention provides a learning model production method for producing a learning model that receives as input a video of a subject person and outputs information about the subject person's movements, the method comprising the steps of: a) executing machine learning of the learning model using a plurality of training data in which a video of a subject person is input, the video including a section in which a specific part of the subject person moves and a section in which the specific part is not moving both before and after the section in which the specific part moves, and outputting information about the movements of the specific part; and and the action is a speech action. It is characterized by: In order to solve the above problem, the learning model production method of the present invention is a learning model production method that produces a learning model that inputs video footage of a subject person and outputs information about the subject person's movements, and includes the steps of: a) performing machine learning of the learning model using a plurality of training data that inputs video footage of a subject person, the video footage including movement sections of specific parts of the subject person and non-movement sections of the specific parts both before and after the movement sections, and outputting information about the movements of the specific parts, wherein the video footage used as training data includes the movement sections in which the movements of the specific parts were performed within a specified period, and the movements are any of the following: turning the face in a specific direction, moving the eyes in a specific direction, blinking at least one of both eyes, rotating the head in a specific direction, raising at least one of both hands, and waving at least one of both hands in a specific direction. In order to solve the above problem, the program according to the present invention is characterized in that it is a program for causing a computer to execute the above learning model production method.

[0017] In order to solve the above problem, the image processing method of the present invention comprises the steps of: a) identifying the movement of a subject person based on a learning model machine-learned using a plurality of training data in which moving images of a person, the moving images including movement sections of specific parts of the subject person and non-movement sections of the specific parts both before and after the movement sections, are input, and information on the movement of the specific parts is output, the moving images used as training data include the movement sections in which the movement of the specific parts was performed within a specified period, and in step a), the movement of the subject person is identified based on information on the movement of the specific parts output from the learning model in response to inputting moving images of the movement of the subject person, the moving images including non-movement sections of the specific parts of the subject person, into the learning model. , the action is a speech action It is characterized by: In order to solve the above problem, the image processing method of the present invention comprises: a) a step of identifying the movements of the subject person based on a learning model machine-learned using a plurality of training data that input moving images of a person, the moving images including movement sections of specific parts of the subject person and non-movement sections of the specific parts both before and after the movement sections, and output information regarding the movements of the specific parts, wherein the moving images used as training data include the movement sections in which the movements of the specific parts were performed within a specified period, and in step a), the movements of the subject person are identified based on information regarding the movements of the specific parts output from the learning model in response to inputting moving images of the movements of the subject person, the moving images including non-movement sections of the specific parts of the subject person, into the learning model, and the movements are characterized in that the movements are any of a movement of turning the face in a specific direction, a movement of moving the eyeballs in a specific direction, a movement of blinking at least one of both eyes, a movement of turning the head in a specific direction, a movement of raising at least one of both hands, and a movement of waving at least one of both hands in a specific direction. In order to solve the above problem, the present invention provides a program for causing a computer to execute the above image processing method. [Effects of the Invention]

[0018] According to the present invention, it is possible to recognize the content of a person's movement without specifying the period during which the movement is performed in a moving image. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a schematic diagram illustrating an image processing system. [Figure 2] FIG. 2 is a diagram illustrating functional blocks of the image processing device. [Figure 3] FIG. 2 is a diagram illustrating functional blocks of a terminal device. [Figure 4] FIG. 1 is a conceptual diagram illustrating the processing in the learning stage of a learning model. [Figure 5] FIG. 1 is a conceptual diagram showing the processing of the inference stage of a learning model. [Figure 6] 10 is a flowchart showing the process of the learning stage. [Figure 7] 10 is a flowchart showing the processing at the inference stage. [Figure 8] 10 is a flowchart illustrating an authentication process using the inference results of a learning model. [Figure 9] FIG. 1 is a conceptual diagram showing how input video to a learning model is generated. [Figure 10] FIG. 2 is a diagram showing a plurality of feature points in a face image. [Figure 11] FIG. 10 is a diagram showing speech periods and non-speech periods for each number. [Figure 12] FIG. 2 is a diagram showing a display screen of a terminal device. [Figure 13] FIG. 10 is a diagram showing a display screen (screen for speaking the first number) on the terminal device. [Figure 14] FIG. 10 is a diagram showing a screen for speaking the second number. [Figure 15] FIG. 10 is a diagram showing a screen for speaking the third number. [Figure 16] FIG. 10 is a diagram showing a screen for speaking the fourth number. [Figure 17] FIG. 10 is a diagram showing a display screen including a determination result and the like. [Figure 18] FIG. 2 is a diagram showing original video data. [Figure 19] FIG. 10 is a diagram showing evaluation results. DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0021] <1. System Overview> Fig. 1 is a schematic diagram showing an image processing system 1. As shown in Fig. 1, the image processing system 1 includes an image processing device 30 and a plurality of terminal devices 70 (see also Figs. 2 and 3). The image processing device 30 is a device that performs image processing and the like on images acquired by the plurality of terminal devices 70. Note that Fig. 2 is a diagram showing the functional blocks of the image processing device 30, and Fig. 3 is a diagram showing the functional blocks of the terminal device 70.

[0022] The image processing device 30 and each terminal device 70 can communicate with each other via a network 108 (including a LAN (Local Area Network) and the Internet, etc.). The image processing device 30 is configured as a server device (such as a cloud server), and the terminal devices 70 are configured as client devices. In other words, the image processing system 1 is configured as a client-server system. The connection to the network 108 may be a wired connection or a wireless connection. For example, the image processing device 30 is connected to the network 108 by wire, and each terminal device 70 is connected to the network 108 wirelessly. Alternatively, all devices 10 and 70 may be connected to the network 108 wirelessly.

[0023] As will be described later, the image processing device 30 and the image processing system 1 are devices that detect information (such as characteristics of the movements) related to the movements of a person in a captured image, and are therefore also called a movement detection device and a movement detection system.

[0024] The image processing device 30 includes a learning model 410 (see FIG. 2). The learning model 410 after being trained by machine learning is also referred to as a trained model 420. Specifically, the learning parameters of the learning model 410 (learner) are adjusted using a predetermined machine learning method, and a trained learning model 410 (trained model 420) is generated (see FIG. 2).

[0025] For example, a neural network model consisting of multiple layers is used as the learning model 410. Then, weighting coefficients (learning parameters) and the like between the multiple layers (input layer, (one or more) intermediate layers, and output layer) in the neural network model are adjusted by a predetermined machine learning method such as deep learning.

[0026] In this embodiment, a video 310 (see FIG. 1) captured by a terminal device 70 is transmitted from the terminal device 70 to the image processing device 30. Then, the video 310 is subjected to a predetermined process by the image processing device 30 to generate a video 330 (a video for input to the learning model 410), and the video 330 is input to the learning model 410.

[0027] The image processing device 30 executes a learning stage process in machine learning. Note that Fig. 4 is a conceptual diagram showing the learning stage process.

[0028] The learning model 410 is machine-learned using a plurality of training data. Each of the plurality of training data is training data that takes as input a video 330 (also referred to as 330a) of a person, the video 330 (330a) including a non-movement section B (described later) of a specific body part of the person, and outputs information 350 (350a) about the movement of the specific body part. In other words, the learning model 410 is a model that learns the characteristics of the movement of a specific body part of a person (such as classification information of the movement).

[0029] Here, a description will be given of a mode in which a captured image (moving image 330) including a person's lips is used as input information, and information 350 relating to the movement of the lips (speech movement) is used as output information. Here, speech movement of a human's lips (more specifically, the movement of one word) is mainly exemplified as the movement of a specific part, and information indicating the content of the speech (the word spoken) is exemplified as information 350.

[0030] By executing the learning stage processing in the image processing device 30, machine learning of the learning model 410 is performed.

[0031] Specifically, the learning model 410 is trained so that when a video 330 (also referred to as 330b) of a subject person (target person) is input, the learning model 410 outputs information 350 (350b) about the movement (speech movement, etc.) of a specific part (lips, etc.) of the subject person. Examples of the information 350 include information about the classification of the movement (for example, words spoken in the speech movement). Through such machine learning, a trained model of the learning model 410 (trained model 420) is generated.

[0032] Furthermore, the image processing device 30 executes processing at the inference stage in machine learning (see FIG. 5). The image processing device 30 estimates information 350 relating to the movement of a subject person using a trained model 420. Specifically, when a captured image (video) 330b of a subject person (a person whose movement is to be determined) is input to the trained model 420, information 350 (350b) relating to the movement (speech movement, etc.) of a specific part (lips, etc.) of the subject person (information on the classification of the movement, etc.) is output from the trained model 420. That is, the image processing device 30 acquires the information 350b output from the trained model 420. Note that FIG. 5 is a conceptual diagram showing processing at the inference stage.

[0033] In this embodiment, a video 310 (see FIG. 1) captured by a terminal device 70 is transmitted from the terminal device 70 to the image processing device 30. Then, the image processing device 30 performs predetermined processing on the video 310 to generate a video 330 (video for input to the learning model 410), and the video 330 (330a, 330b) is input to the learning model 410. Note that the video 330a is used as training data (input data) in the learning stage, and the video 330b is used as input data in the inference stage.

[0034] <2. Image processing device 30> As shown in FIG. 2, the image processing device 30 includes a controller 31 (also referred to as a control unit), a storage unit 32, a communication unit 34, and an operation unit 35.

[0035] The controller 31 is a control device that is built into the image processing device 30 and controls the operation of the image processing device 30 .

[0036] The controller 31 is configured as a computer system including one or more hardware processors (for example, a central processing unit (CPU) and a graphics processing unit (GPU)). The controller 31 performs various processes by executing, in the CPU or the like, a predetermined software program (hereinafter also simply referred to as a program) stored in a storage unit (a non-volatile storage unit such as a ROM and / or a hard disk) 32. The program (more specifically, a group of program modules) may be recorded on a portable recording medium such as a USB memory, read from the recording medium, and installed in the image processing device 30. Alternatively, the program may be downloaded via a communication network or the like and installed in the image processing device 30.

[0037] The controller 31 executes processing related to the learning stage of machine learning. Specifically, the controller 31 first executes processing to acquire multiple pieces of teacher data. Next, the controller 31 executes machine learning processing based on the multiple pieces of teacher data. In particular, the controller 31 executes processing to optimize learning parameters related to the learner (learning model 410) based on the multiple pieces of teacher data, and generates a trained model 420.

[0038] Furthermore, the controller 31 executes processing related to the inference stage of machine learning. Specifically, the controller 31 uses the learning model 410 (trained model 420) in which the learning parameters have been adjusted to detect information 350 (350a) related to the movement of a specific part of the subject person based on the moving image 330 of the subject person, etc.

[0039] The controller 31 also executes face recognition processing based on a face image (still image). In the face recognition processing, a detection processing using the trained model 420 (a motion detection processing (motion recognition processing) based on a moving image of a person) is used.

[0040] The storage unit 32 is configured with a storage device such as a hard disk drive (HDD) and / or a solid state drive (SSD). The storage unit 32 stores a learning model 410 (including learning parameters and programs related to the learning model) (and thus a trained model 420) and the like. The storage unit 32 also stores a registration database 450 used for authentication processing by the image processing device 30. Various information related to registered users (user IDs, names, facial image information, etc.) is registered in advance in the registration database 450.

[0041] The communication unit 34 is capable of performing network communication via a network. In this network communication, various protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol) are used. By using this network communication, the image processing device 30 can exchange various data with a desired destination (for example, a terminal device 70).

[0042] The operation unit 35 includes an operation input unit 35a that accepts operation inputs to the image processing device 30, and a display unit 35b that displays and outputs various information. A mouse, a keyboard, or the like is used as the operation input unit 35a, and a display (such as a liquid crystal display) is used as the display unit 35b. A touch panel that functions as both a part of the operation input unit 35a and a part of the display unit 35b may also be provided.

[0043] Note that the image processing device 30 also executes the learning stage processing for the learning model 410 (generation processing of the learning model 410), and is therefore also referred to as a learning model generation device.

[0044] <3. Terminal Device 70> Each terminal device 70 is an information input / output terminal device (information processing device) capable of network communication with the image processing device 30. Each terminal device 70 is configured as a smartphone, a tablet terminal, or a personal computer (which may be either a fixed (desktop) or portable type), etc. FIG. 1 illustrates the terminal device 70 configured as a smartphone.

[0045] The terminal device 70 exchanges various types of information with the image processing device 30. The terminal device 70 generates a captured image (image data) 310 of a person present in the vicinity of the terminal device 70, and transmits the captured image 310 to the image processing device 30. The terminal device 70 also receives information from the image processing device 30 and displays the information on its display unit 75b (touch panel 75c or the like) (see FIGS. 1 and 3).

[0046] FIG. 3 is a functional block diagram showing a schematic configuration of the terminal device 70. As shown in FIG.

[0047] As shown in the functional block diagram of Figure 3, the terminal device 70 includes a controller 71, a memory unit 72, a communication unit 74, an operation unit 75, and an imaging unit 76, and various functions are realized by operating these units in combination.

[0048] The controller (control unit) 71 is a control device that controls the terminal device 70 .

[0049] The controller 71 has the same hardware configuration as the controller 31. The controller 71 executes, in a CPU or the like, a predetermined software program stored in a storage unit (a non-volatile storage unit such as a ROM and / or a hard disk) 72, thereby realizing various processes. The program (more specifically, a group of program modules) may be recorded on a portable recording medium such as a USB memory, read from the recording medium, and installed in the terminal device 70. Alternatively, the program may be downloaded via a communication network or the like and installed in the terminal device 70.

[0050] The controller 71 executes the program and performs the following various processes.

[0051] Specifically, during the learning stage of the learning model 410, the controller 71 performs a photographing process of a moving image 300 (310, etc.) to be used as training data for learning, specifically, a photographed image including a specific part of a person (here, a facial image including the lips).

[0052] Furthermore, the controller 71 performs a process of capturing a captured image (here, a facial image (moving image) including the lips) including a specific part of the subject person (a person whose speech behavior is to be detected) in an inference stage using the learning model 410 (more specifically, the trained model 420). The controller 71 also transmits a captured image 310 relating to the subject person to the image processing device 30. The image processing device 30 inputs the captured image 310 (more specifically, a captured image 330 obtained by adjusting the captured image 310) to the trained model 420, and obtains information 350b (speech content) relating to the subject person's behavior as an output result from the trained model 420. The image processing device 30 then transmits the information 350b (speech content) and the like to the terminal device 70. The terminal device 70 (controller 71) displays the information 350b (speech content) and the like received from the image processing device 30 on the display unit 75b.

[0053] The storage unit 72 has the same hardware configuration as the storage unit 32 .

[0054] The communication unit 74 has the same hardware configuration as the communication unit 34. By utilizing network communication via the communication unit 74, the terminal device 70 can exchange various data with a desired destination (for example, the image processing device 30).

[0055] The operation unit 75 includes an operation input unit 75a that accepts operation input to the terminal device 70, and a display unit 75b that displays and outputs various information. A mouse, a keyboard, etc. are used as the operation input unit 75a, and a display (such as a liquid crystal display) is used as the display unit 75b. A touch panel 75c that functions as both a part of the operation input unit 75a and a part of the display unit 75b may also be provided.

[0056] The imaging unit 76 is configured with a camera 76c (see FIG. 1) (including a lens and an imaging element (CCD, etc.)). The imaging unit 76 is capable of generating an image 310 (more specifically, moving image data) captured by the camera 76c. Specifically, the camera 76c has an RGB image sensor or the like that captures visible light images (color images, etc.), and is capable of capturing color moving images.

[0057] <4. Processing Overview> In this embodiment, a detection process (a spoken word detection process using the trained model 420) relating to the movement (speech movement) of a subject person in a captured image is mainly used for authentication of the subject person. Note that the authentication process is mainly executed by the image processing device 30.

[0058] Here, face recognition processing is executed as the authentication processing. However, in order to prevent "impersonation" using a face photograph (still image) or the like in the face recognition processing, it is also determined whether the subject person (the person to be authenticated) can perform the action as instructed. Specifically, it is further determined using the trained model 420 whether the action (the action of speaking the specified words) in response to the action instruction from the authentication device (terminal device 70 and / or image processing device 30) is actually performed. In other words, the person to be authenticated is determined to be a genuine specific person not only on the condition that the face recognition processing for the person to be authenticated is authenticated as a specific person, but also on the condition that the content of the speech of the person to be authenticated in response to the speech instruction matches the content of the speech instruction.

[0059] In detail, after the face authentication process, an instruction to speak a word is given to the user (subject person (person to be authenticated)). Here, the instruction to speak is an instruction to speak one word selected (designated) (by the authentication device) from 10 words ("zero," "one," "two," "three," "four," "five," "six," "seven," "eight," and "nine") representing each single digit. The instruction to speak is given to the user from the terminal device 70 (and / or the image processing device 30).

[0060] In response to the speech instruction, the user speaks the word specified in the speech instruction (for example, "3 (san)"). The terminal device 70 captures an image (video) 310 of the person to be authenticated (the person to be authenticated) and transmits the video 310 to the image processing device 30. The image processing device 30 inputs the video 310 (more specifically, a video 330 obtained by processing (adjusting) the video 310) into the trained model 420, and obtains information 350b (the content of the speech) relating to the movements of the subject person (the person to be authenticated) as an output result from the trained model 420. Note that the trained model 420 has been generated in advance by machine learning.

[0061] Then, the image processing device 30 transmits the information 350b (utterance content) to the terminal device 70. Furthermore, the terminal device 70 displays the information 350b (utterance content) received from the image processing device 30 on the display unit 75b.

[0062] Furthermore, the image processing device 30 transmits a collation result (i.e., whether or not the content as instructed has been spoken) of comparing the information 350b (the content of the utterance) with the content of the instruction to the terminal device 70. Furthermore, the terminal device 70 displays the collation result on the display unit 75b.

[0063] Furthermore, the image processing device 30 executes authentication processing based on the matching result. Specifically, the person to be authenticated is determined to be a genuine specific person not only if the person to be authenticated is authenticated as a specific person by the face authentication processing on the person to be authenticated, but also if the content of the utterance of the person to be authenticated in response to the speech instruction matches the content of the speech instruction.

[0064] The movement detection process (detection process related to the movement of the photographed person) used in such authentication process is executed based on the trained model 420.

[0065] Below, we will first explain the generation process of the trained model 420 (in other words, the learning stage process of the training model 410), and then explain the inference stage process using the trained training model 410 (trained model 420).

[0066] <5. Learning Stage Processing> First, the learning stage processing will be described.

[0067] Fig. 6 is a flowchart showing the processing in the learning stage. The processing shown in Fig. 6 is mainly executed by the controller 31 of the image processing device 30. However, part of the processing in Fig. 6 (step S11, etc.) is executed in cooperation with the terminal device 70.

[0068] 6 is also a diagram showing a method for generating a trained model. In this application, generating a trained model 420 means producing (manufacturing) the trained model 420, and the "method for generating a trained model" means the "method for producing a trained model."

[0069] The learning stage processing in FIG. 6 is roughly divided into preparatory processing of teacher data (steps S11, S12) and machine learning processing of the learning model 410 using the teacher data (steps S13, S14).

[0070] In step S11, the controller 31 receives and acquires a moving image 310 of a person (a person to be photographed) from the terminal device 70, and generates a moving image 330 based on the moving image 310.

[0071] First, the processing of the terminal device 70 will be described.

[0072] 12 is displayed on the display unit 75b of the terminal device 70. The screen 110 has an image display area 111, a plurality of number display fields 112, a plurality of reading display fields 113, an explanation field 114 regarding speaking instructions, and a start button 115.

[0073] The image display area 111 is an area where a moving image 310 captured by the terminal device 70 is displayed in real time.

[0074] The multiple (four in this example) number display fields 112 are fields that display multiple (four) numbers (one by one). Also, below each number display field 112, there is provided a reading display field 113 corresponding to each number display field 112. Each reading display field 113 displays the reading of the number displayed in the corresponding (directly above) number display field 112.

[0075] The explanation column 114 is a column where explanatory text regarding the speaking instruction is displayed. The explanation column 114 displays the following text: "A total of four numbers will be displayed in order from left to right in the upper column. For each number that is displayed, please pronounce that number clearly within 1.5 seconds."

[0076] After reading the explanatory text in the explanation column 114, the user presses the start button 115 to start the input process.

[0077] In response to pressing the start button 115, the terminal device 70 sequentially displays a plurality of numbers (four in this example) randomly selected from ten numbers (0, 1 to 9) at predetermined time intervals (for example, every three seconds) (or at random time intervals). In detail, the first number (for example, "2") is displayed in the leftmost number display field 112, the second number (for example, "4") is displayed in the second number display field 112 from the left, the third number (for example, "6") is displayed in the third number display field 112 from the left, and the fourth number (for example, "7") is displayed in the rightmost number display field 112 (fourth from the left).

[0078] The user speaks (utters) the indicated numbers (words) in response to each speech instruction. Specifically, the user speaks the first number in response to the start of the first number's display (speech instruction). Next, the user speaks the second number in response to the start of the second number's display (speech instruction). Thereafter, the user similarly speaks the third number in response to the start of the third number's display (speech instruction), and speaks the fourth number in response to the start of the fourth number's display (speech instruction). In this way, the user performs each specified action (speaking each number) in response to each action instruction that specifies that the specified action (speaking a number using the lips) related to a specific part of the person should be completed within the specified period D1 (here, 1.5 seconds).

[0079] FIG. 11 is a diagram showing an utterance section (period) F relating to one of four numbers, and non-utterance sections B (B1, B3) before and after the utterance section F. In FIG.

[0080] For example, if a speech instruction is given at time T1, the user responds to the speech instruction and starts speaking at time T3. Generally, it takes a human being about 100 ms (milliseconds) to receive and react to an external stimulus, and about 300 ms (milliseconds) to recognize it and begin to act. Therefore, time T3 is often several hundred ms (milliseconds) after time T1.

[0081] In particular, since the user is instructed to speak within a short time (here, within 1.5 seconds), the user should complete the speech in as short a time as possible and start speaking as early as possible.

[0082] As a result, the user starts speaking around time T3 after the speech instruction time T1 and finishes speaking around time T5 (for example, 1.2 seconds after time T1). After that, no speech is made until time T9, 1.5 seconds after the time T5, (i.e., the period from time T5 to time T9 is a non-speech period).

[0083] In this way, a non-speech section B1 (time points T1 to T3) having a length corresponding to the human reaction time (to an instruction), a speech section F (time points T3 to T5) in response to the instruction, and the remaining non-speech section B3 (time points T5 to T9) from the end of the speech to the end of the specified period occur in this order.

[0084] The terminal device 70 captures such speech actions of the user to generate the moving image 310. In detail, the moving image 310 capturing the movement of the human lips (speech actions) and the like over a period from time T1 to time T9 is generated by the imaging unit 76 and the like.

[0085] Then, the terminal device 70 transmits the moving images 310 to the image processing device 30. Specifically, a total of four moving images 310 relating to each number are transmitted sequentially.

[0086] The image processing device 30 receives and acquires each video 310 transmitted from the terminal device 70.

[0087] 9, the image processing device 30 generates a moving image 320 (moving image after region adjustment) by extracting the region near the lips from a moving image 310 capturing the entire face of a person. Then, a moving image 330 for input to the learning model 410 is generated by performing adjustment processing on the moving image 320. That is, the moving image 330 is generated by performing various adjustment processing on the moving image 310.

[0088] Specifically, first, the controller 31 executes image processing to extract a plurality of feature points as shown in Fig. 10 from each frame image of the video 310. Fig. 10 is a diagram showing a plurality of feature points (68 in this case) in a face image. Fig. 10 shows 17 feature points P1 to P17 related to the facial contour (jaw, etc.), 10 feature points P18 to P27 related to the eyebrows, 9 feature points P28 to P36 related to the nose, 12 feature points P37 to P48 related to the eyes, and 20 feature points P49 to P68 related to the lips (lips). In particular, with regard to the lips, several points on the contours of the upper and lower lips are extracted as feature points.

[0089] Based on these extracted feature points P1 to P68, an affine transformation (a normalization process of two-dimensional positions based on three-dimensional positions) is performed on each frame image so that each frame image becomes an image of the face viewed from the front. Note that this affine transformation transforms a face image that is not facing directly ahead (such as a face image facing an angle) into a face image that faces forward (more specifically, a lip image).

[0090] Then, in each frame image of the affine-transformed video 310 (310a), an image of a predetermined region (an image of a predetermined region including the periphery of the lips) including feature points near the lips (20 feature points P49 to P68 related to the lips) is extracted as a lip image. The lip image can also be expressed as an image including feature points (facial feature points) around the lips. A collection of the frame images extracted in this way (images after affine transformation and region extraction) is generated as video 320 (320a) (see FIG. 9).

[0091] Furthermore, moving image 320 is resized to a predetermined size (for example, 128 pixels x 64 pixels), and the pixel value of each pixel (0 to 255) is converted to a value (0 to 1) divided by 255 and standardized. By performing this standardization process on moving image 320, moving image 330 (330a) is generated.

[0092] In this embodiment, based on four videos 310 relating to the speech actions of four numbers (words), four videos 330 relating to the speech actions of the four numbers are sequentially acquired. Specifically, a first video 330 relating to the speech action in response to the speech instruction for the first number and a second video 330 relating to the speech action in response to the speech instruction for the second number are sequentially acquired. Furthermore, a third video 330 relating to the speech action in response to the speech instruction for the third number and a fourth video 330 relating to the speech action in response to the speech instruction for the fourth number are sequentially acquired.

[0093] Each video 330 (330a) can also be expressed as a video capturing a person's movement (speaking movement) in response to an instruction to complete a specified movement (speaking movement using lips) related to a specific part of the person within a specified period D1. Each video 330a includes a movement section in which the movement of the specific part is performed within the specified period D1 (e.g., 1.5 seconds). The same applies to video 310 (310a).

[0094] Each video 330 comprises, in this order, a non-speech section B1 until speech, a speech section F in which one word (number) is spoken, and a non-speech section B3 after the speech section F (such as a non-speech section until the next utterance).

[0095] Each video 330 has the length of the designated period D1 (the length from time points T1 to T9) (for example, 1.5 seconds). However, without being limited to this, the video 330 may have a predetermined length (2 seconds) that is, for example, the designated period D1 plus a margin of time (0.5 seconds). Furthermore, if there is a slight error in the duration of the multiple videos 330 (especially if they are shorter than the predetermined length), the frame images corresponding to the missing time may be filled in with zero-padding images (black images).

[0096] In the next step S12 (FIG. 6), training data is generated using each video 330 (330a).

[0097] Specifically, information (spoken words) relating to the classification of the movement of a specific part (lips) of the person in each video 330 is identified. Specifically, each number designated in the screen 110 (particularly the number display field 112 (and the reading display field 113)) in step S11 is identified as information (information indicating spoken words) 350 relating to the speech movement. Then, each corresponding number (number designated in each speech instruction) is assigned to each video 330 as a label (correct answer data). Note that each number designated in each number display field 112 in step S11 is assumed to have been transmitted from the terminal device 70 to the image processing device 30 in step S11 together with the corresponding video 310.

[0098] For example, for video 330 in which a speech action corresponding to "2 (ni)" specified in the leftmost number display field 112 and reading display field 113 on screen 110 is captured, the specified number "2" is assigned as information 350 (correct answer data) related to the speech action. Also, for video 330 in which a speech action corresponding to "4 (yon)" specified in the second number display field 112 and reading display field 113 from the left on screen 110 is captured, the specified number "4" is assigned as information 350 related to the speech action.

[0099] In this way, each video 330 is associated with information 350 (correct answer data) indicating the specified spoken word (number), and teacher data is generated. In other words, teacher data is generated by inputting video 330 (330a) of a person that includes a lip non-movement section (non-speech section) B (B1, B3) of the person, and outputting information 350 (350a) related to lip movement. Here, four teacher data are generated, each associated with the speech content of four videos 330.

[0100] 6, the processes of steps S11 and S12 are repeatedly executed, thereby generating (obtaining) a plurality (a large number) of training data.

[0101] In this embodiment, four pieces of teacher data are generated using the screen 110, but this is not limiting. For example, a relatively small number of pieces of teacher data (e.g., one piece each) may be generated, or a relatively large number of pieces of teacher data (e.g., ten pieces each) may be generated.

[0102] Furthermore, multiple pieces of teacher data relating to different speech contents of the same person may be acquired using one terminal device 70, or multiple terminal devices 70. Multiple pieces of teacher data relating to multiple people may be acquired using multiple terminal devices 70, or one terminal device 70.

[0103] In the next step S13, the controller 31 uses a plurality of training data to perform machine learning of the relationship between the moving image (lip image) and the action content, and obtains a learning model 410 (trained model 420).

[0104] The learning model 410 is constructed using a convolutional neural network (CNN) 200. The learning model 410 has a hierarchical structure in which a plurality of layers (hierarchies) are hierarchically connected. Specifically, the learning model 410 includes an input layer, intermediate layers (such as a convolutional layer and a pooling layer), and an output layer.

[0105] As the learning model 410, for example, various convolutional neural networks such as ResNet (specifically, ResNet34-3D) are used. ResNet (Residual Network) is a convolutional neural network that includes adding up residuals between layers. ResNet is configured with a plurality of residual blocks, each of which is configured by a combination of a convolutional layer, an activation function, and a skip connection (shortcut connection). ResNet34-3D is a convolutional neural network in which the dimension of the filters in the convolutional layer and pooling layer of ResNet34 is expanded from two dimensions to three dimensions (three dimensions including not only two dimensions within an image but also the time axis).

[0106] In step S13, machine learning of the learning model 410 is performed using a plurality of training data sets, in which the video 330a of a person is used as input and information related to the movement of the specific body part is used as output. Here, only the video of the video 330a is used (the audio of the video 330a is not used).

[0107] In this machine learning, by using the above-mentioned convolutional neural network equipped with a three-dimensional filter (a filter extended in the time axis direction), it is possible to analyze motion (movement) including information in the time axis direction. Note that the neural network used in the learning model 410 is not limited to ResNet34-3D (a ResNet extended in the time axis direction). For example, other neural networks obtained by extending a recurrent neural network (RNN) such as a long short-term memory (LSTM) in the time axis direction may also be used in the learning model 410.

[0108] Then, the learning model 410 is machine-learned using a plurality of training data in step S13, and a trained model 420 is generated (step S14).

[0109] Furthermore, by using this trained model 420, when a video 330b of the person to be determined is input, information 350 (content of spoken words) related to the speech behavior of the person to be determined is output, as will be described below. In short, the information 350 (content of the person's speech, etc.) is identified based on the video 330b. Then, based on the information 350, it is possible to identify the behavior of the person to be determined (content of the speech behavior (spoken words)).

[0110] Here, in step S13, machine learning of the learning model 410 is performed using a plurality of training data sets in which a moving image 330a of a person is input and information on the movement of the specific part is output. This makes it possible to generate a learning model 410 (trained model 420) that can recognize the content of the person's movement based on the moving image 330.

[0111] In particular, the video 330a is a video that includes a non-movement section B of a specific part of the person. When using the video 330a in machine learning, it is not necessary to exclude the non-movement section B from the video 330a. Therefore, it is possible to generate a learning model (trained model) that can recognize the content of a person's movement without specifying the period F during which the movement is performed in the video. Furthermore, it is possible to recognize the content of a person's movement (such as spoken words) without specifying the period during which the movement is performed in the video. In other words, the technology according to this embodiment makes it possible to recognize the content of a person's movement by using machine learning or the like without specifying the period during which the movement is performed in the video.

[0112] Furthermore, each video 330a used as training data is a video captured of a person's action (speaking action) in response to an action instruction to complete a specified action (speaking action of a specific word) related to a specific part of the person's body ("lips") within a specified period D1.

[0113] The speech actions in such moving images 330a are generally performed within the designated period D1 in response to the action instruction. That is, the execution period (speech period) F of the designated action is performed within the designated period D1. In other words, each moving image 330a includes a movement period (speech period) F in which the movement of a specific body part (speech action using the lips) is performed within the designated period D1. This makes it possible to prevent the length of the speech period F from becoming longer than the period D1, and makes it possible to make the lengths of the speech periods F uniform (aligned) compared to when the designated period D1 is not specified.

[0114] Furthermore, by executing (completing) the speech action within the designated period D1 in response to the action instruction, the start time of the speech section F is prevented from being shifted later than the end time T9 of the designated period D1 (for example, the speech section F starts 5 seconds after time T1). In short, it is possible to align the position of the action section F within the video.

[0115] In this way, it is possible to align the length of the movement section F and the position of the movement section F within the video, which in turn improves the learning accuracy of the learning model 410, in other words, improves the estimation accuracy (accuracy of identifying the movement of the subject person) of the learning model 410.

[0116] In particular, the action instruction instructs that a single word should be spoken. By speaking word by word in accordance with the action instruction, it is possible to make the length of the speech section F more uniform (uniform) than when a combination of multiple words is spoken. This in turn improves the learning accuracy of the learning model 410, or in other words, improves the estimation accuracy (accuracy of identifying the subject person's action) of the learning model 410.

[0117] In particular, it is preferable that the specified period D1 be approximately the same as the required time for the specified action (e.g., 0.3 to 1 second) (more specifically, a period with a small margin added to the required time) (e.g., 1.5 seconds). That is, it is preferable that a very short period (approximately 1.5 to 3 seconds) be specified as the specified period D1. This makes it possible to further suppress variations in the length of the speech sections F and achieve greater uniformity compared to when the specified period D1 is relatively long (e.g., 5 seconds or more). For example, the length of the speech sections is prevented from varying over a relatively wide range (e.g., 6 seconds), and the length of the speech sections F is uniformed to around 1 second. It is also possible to align the positions of the speech sections F within each video 330a. Specifically, it is possible to align the start points of the speech sections F within each video 330a to approximately the same time (near time T3, which is a predetermined period (time required for reaction, etc.) after the start point T1 of the video). More specifically, the speech section F (e.g., approximately 0.8 seconds) is located between the immediately preceding non-motion section B1 (e.g., approximately 0.3 seconds) and the immediately following non-motion section B3 (e.g., approximately 0.4 seconds), in short, near the center of the video 330b.

[0118] In this way, with the above-described video acquisition technique, the length of the speech intervals (movement intervals) F is uniformed, and the appearance positions of the speech intervals F (appearance positions within the video) are also uniformed. This makes it possible to easily generate multiple pieces of training data in which the start times and lengths of the speech intervals F are roughly uniform. In other words, when generating training data for learning, it is not necessary to perform a process to identify the start times of the speech intervals F, etc. That is, it is possible to generate a learning model that can recognize the content of a person's movement without identifying the period during which the movement is performed within the video (movement interval F). Ultimately, it is possible to recognize the content of a person's movement without identifying the period during which the movement is performed within the video.

[0119] <6. Processing of the inference stage of the trained model 420> Next, the inference stage process using the trained model 420 and the authentication process using the inference results will be described.

[0120] Fig. 8 is a flowchart showing authentication processing using the inference results of the trained model 420. The processing shown in Fig. 8 is mainly executed by the controller 31 of the image processing device 30. However, part of the processing in Fig. 8 is executed in cooperation with the terminal device 70.

[0121] First, in step S51, authentication processing (face authentication processing) is performed based on the face image of the person to be logged in. Specifically, a face image (still image) of the person to be logged in is captured using the imaging unit 76 of the terminal device 70, and the face image is transmitted from the terminal device 70 to the image processing device 30. The image processing device 30 then performs face authentication processing based on the face image (step S51). In the face authentication processing, a process of matching the face with authentic information (authentic face image information, etc.) registered in the registration database 450 (see FIG. 2) is performed. However, face authentication processing alone may result in "impersonation" using a photograph, etc. Therefore, in this embodiment, a confirmation process for preventing impersonation is performed in the next step S52.

[0122] Specifically, in step S52, it is further determined using the trained model 420 whether or not an action (utterance of the specified word) in response to the action instruction from the authentication device is actually being performed.

[0123] In step S52, the process shown in Fig. 7 (steps S31 to S33) is executed. Fig. 7 is a flowchart showing the process of the inference stage using the trained model 420. The process shown in Fig. 7 is executed by the controller 31, etc. The process of the inference stage using the trained model 420 determines whether or not the speech action of the specified word is actually being performed.

[0124] First, in step S31, a moving image (lip image) of a person to be authenticated (photographed person) is acquired. In step S31, similarly to step S11 (see FIG. 6), moving images 310, 320, and 330 (specifically, 310b, 320b, and 330b) are acquired. The terminal device 70 is used for the process of acquiring the moving image 310.

[0125] Specifically, first, the person to be authenticated reads the explanatory text in the explanatory field 114 displayed on the display unit 75b of the terminal device 70, and then presses the start button 115 to start the input process.

[0126] In response to pressing the start button 115, the terminal device 70 sequentially displays a plurality of numbers (four in this case) randomly selected from ten numbers (0, 1 to 9) at predetermined time intervals (for example, every three seconds) (or at random time intervals). Then, the person to be authenticated speaks (vocals) the instructed numbers in response to the speaking instruction. The terminal device 70 captures a video 310 (310b) related to the user's speaking action in response to the speaking instruction, and transmits the video 310 (310b) to the image processing device 30. The image processing device 30 generates a video 320 (320b) and a video 330 (330b) based on the received video 310b.

[0127] In detail, first, the first number (for example, "5") is displayed in the number display field 112 on the left side (see FIG. 13). The person to be authenticated speaks the first number in response to the display of the first number (speaking instruction). The terminal device 70 captures a video 310 (310b) related to the user's facial image and transmits the video 310 (310b) to the image processing device 30. The image processing device 30 generates a video 320 (320b) and a video 330 (330b) based on the received video 310b. As a result, a video 330b (captured image) including a speaking action in response to the display of the first number (for example, "5") (speaking instruction) is acquired.

[0128] Next, the second number (for example, "3") is displayed in the second number display field 112 from the left (see FIG. 14). The person to be authenticated speaks the second number in response to the display of the second number (speaking instruction). The terminal device 70 captures a video 310 (310b) related to the user's face image and transmits the video 310 (310b) to the image processing device 30. The image processing device 30 generates a video 320 (320b) and a video 330 (330b) based on the received video 310b. As a result, a video 330b (captured image) including a speaking action in response to the display of the second number (for example, "3") (speaking instruction) is acquired.

[0129] Furthermore, the third number (for example, "9") is displayed in the third number display field 112 from the left (see FIG. 15). The person to be authenticated speaks the third number in response to the display of the third number (speaking instruction). The terminal device 70 captures a video 310 (310b) related to the user's facial image and transmits the video 310 (310b) to the image processing device 30. The image processing device 30 generates a video 320 (320b) and a video 330 (330b) based on the received video 310b. As a result, a video 330b (captured image) including a speaking action in response to the display of the third number (for example, "9") (speaking instruction) is acquired.

[0130] Furthermore, a fourth number (e.g., "8") is displayed in the rightmost (fourth from the leftmost) number display field 112 (see FIG. 16). The person to be authenticated speaks the fourth number in response to the display of the fourth number (speaking instruction). The terminal device 70 captures a video 310 (310b) related to the user's facial image and transmits the video 310 (310b) to the image processing device 30. The image processing device 30 generates a video 320 (320b) and a video 330 (330b) based on the received video 310b. As a result, a video 330b (captured image) including a speaking action in response to the display of the fourth number (e.g., "8") (speaking instruction) is acquired.

[0131] In this way, the image processing device 30 acquires the moving image 330 (330b). The moving image 330b captures the subject's actions in response to an instruction to perform a specified action related to a specific body part within a specified period D1. The moving image 330b also includes a non-motion section B (see FIG. 11) of the specific body part of the subject. No processing to eliminate the non-motion section B (excluding processing including processing to identify the non-motion section B) has been performed on the moving image 330b.

[0132] The terminal device 70 also transmits four numbers ("5," "3," "9," and "8") randomly selected by the terminal device 70 to the image processing device 30. In other words, the terminal device 70 also transmits the words (spoken instruction contents) specified in each spoken instruction to the image processing device 30. Note that, although the terminal device 70 determines the spoken words (words) in each spoken instruction here, this is not limiting. For example, the image processing device 30 may determine the spoken words (words) in each spoken instruction and transmit them to the terminal device 70.

[0133] Next, in step S32, the image processing device 30 inputs the moving image 330 (330b) of the subject person into the learning model 410 (trained model 420) and obtains a movement estimation result from the learning model 410. That is, the image processing device 30 obtains information 350b (utterance content) regarding the movement of the subject person as an output result from the trained model 420. Then, the information 350b regarding the movement of the specific body part output from the learning model 410 (trained model 420) is identified (recognized) as it is as a recognition result (recognition result of the spoken word) of the movement of the subject person (utterance movement of each word).

[0134] Specifically, first, a video 330b (captured image) including a speech action corresponding to the display (speech instruction) of the first number (e.g., "5") is input to the trained model 420, and output information 350b is obtained from the trained model 420. The recognition result of the spoken word related to the first instruction (e.g., "5") is obtained as the output information 350b.

[0135] Next, a video 330b (captured image) including a speech action in response to the display (speech instruction) of the second number (e.g., "3") is input into the trained model 420, and output information 350b (recognition result of the spoken word (e.g., "3")) is obtained from the trained model 420.

[0136] In addition, a video 330b (captured image) including a speech action in response to the display (speech instruction) of a third number (e.g., "9") is input into the trained model 420, and output information 350b (recognition result of the spoken word (e.g., "9")) is obtained from the trained model 420.

[0137] Furthermore, a video 330b (captured image) including a speech action in response to the display (speech instruction) of the fourth number (e.g., "8") is input into the trained model 420, and output information 350b (recognition result of the spoken word (e.g., "8")) is obtained from the trained model 420.

[0138] Then, in step S33, additional authentication processing is performed based on the inference result in step S32.

[0139] Specifically, it is determined for each of the four speech instructions whether the words specified in each speech instruction match the recognition results for the actual spoken words (the recognized content of the speech action).

[0140] Specifically, it is determined whether or not the word specified in the first spoken instruction (e.g., "5") matches the recognition result of the spoken words for the first instruction (e.g., "5"). Next, it is determined whether or not the word specified in the second spoken instruction (e.g., "3") matches the recognition result of the spoken words for the second instruction (e.g., "3"). It is also determined whether or not the word specified in the third spoken instruction (e.g., "9") matches the recognition result of the spoken words for the third instruction (e.g., "9"). It is also determined whether or not the word specified in the fourth spoken instruction (e.g., "8") matches the recognition result of the spoken words for the fourth instruction (e.g., "8").

[0141] Only when the words specified in each speech instruction and the recognition results of the spoken words corresponding to each of the four speech instructions match, is it determined that the action corresponding to the action instruction (the speech action of the specified words) is actually being performed. In other words, the person to be authenticated is determined to be a genuine specific person not only on the condition that the person to be authenticated is authenticated as a specific person by the face recognition process for the person to be authenticated, but also on the condition that the content of the speech of the person to be authenticated in response to the speech instruction matches the content of the speech instruction.

[0142] The image processing device 30 also transmits the information 350b (recognized spoken words) to the terminal device 70, and the terminal device 70 displays the information 350b received from the image processing device 30 in each recognized word display field 116 of the display unit 75b (see FIG. 17). In FIG. 17, the leftmost recognized word display field 116 displays the recognition result (spoken content) "5" regarding the speech action corresponding to the first speech instruction. The remaining three recognized word display fields 116 display the respective recognition results ("3," "9," and "8") regarding the speech actions corresponding to the second to fourth speech instructions.

[0143] Furthermore, the image processing device 30 transmits the collation result (i.e., whether or not the content (four words) was spoken as instructed) of comparing the information 350b (utterance content) with the utterance instruction content to the terminal device 70. Then, the terminal device 70 displays the collation result in the result display field 117 of the display unit 75b (see FIG. 17). In FIG. 17, the result display field 117 displays that the collation result is good (the judgment result "GOOD" indicating that the four words were recognized as actually spoken).

[0144] In step S52, the above-mentioned confirmation operation is performed.

[0145] In the next step S53, the person to be authenticated is determined to be a genuine specific person not only on the condition that the person to be authenticated is authenticated as a specific person by the face authentication process for the person to be authenticated in step S51, but also on the condition that it is determined in step S52 that the content of the speech of the person to be authenticated in response to each speech instruction matches the instruction content of each speech instruction. In other words, the person to be authenticated is determined to be a genuine specific person on the condition that it is determined that the person to be authenticated (subject person) has performed a specified action. This prevents "impersonation" using a photograph or the like in the face authentication process.

[0146] According to the above process, the motion of the subject person can be identified (recognized) based on information 350b about the motion of a specific body part output from the learning model 410 in response to inputting the video 330b to the learning model 410 (trained model 420). In particular, the video 330b is a video that includes a non-motion section B of a specific body part of the subject person. There is no need to exclude the non-motion section B from the video 330b that is to be input to the trained model 420. Therefore, it is possible to recognize the content of the person's motion without identifying the period during which the motion is performed in the video.

[0147] Furthermore, video 330b (video of the subject person) is a video of the subject person's actions in response to an instruction to perform a specified action related to a specific body part within a specified period. Therefore, it is possible to align the length of the action section F and the position of the action section F within the video. Specifically, it is possible to align the length of the speech section F within video 330b to be inferred with the length of the speech section F within other videos 330b and / or the length of the speech section within training video 330a. It is also possible to align the start time of the speech section F within video 330b to be inferred with the start time of the speech section F within other videos 330b and / or the start time of the speech section F within training video 330a (a period close to time T3, a predetermined period (time required for reaction, etc.) from start time T1 of the video). This ultimately improves the accuracy of identifying (estimating) the subject person's actions.

[0148] Furthermore, if the subject person (person to be authenticated) is judged to have performed a specified action based on the output obtained by inputting video 330b of the subject person into the learning model 410 (trained model 420), the subject person is judged to be a genuine person in the authentication process. Therefore, this is useful for preventing impersonation, etc.

[0149] Furthermore, since the speech movement, which is a highly advanced biological movement, is used as the movement of a specific part of a human body, it is possible to appropriately determine the existence of a human. In particular, the speech movement here is the movement of recognizing a designated word and then uttering the designated word (movement involving the recognition of the designated word, etc.), which makes it possible to extremely appropriately determine the existence of a human.

[0150] In the above-described step S52, a single estimation result (only one number that is most likely to be the spoken word (in short, the top estimation result)) is output as the output result from the trained model 420, but this is not limited to this. For example, numbers indicating the top three estimation results (three numbers: a number corresponding to the highest probability, a number corresponding to the next highest probability, and a number corresponding to the probability after that) may be output as the output result (action recognition result) from the trained model 420. If the number of the spoken instruction content (the word that the person is instructed to speak) is included in these top three numbers, it may be determined that the spoken content of the person to be authenticated in response to the spoken instruction matches the instruction content of the spoken instruction (the person to be authenticated has performed the specified action).

[0151] Furthermore, in the above-described step S52, first, captured images (four videos) of speech actions related to four numbers (words) are acquired in step S31, and then an estimation process of the speech content related to the four numbers is performed (four estimation results are acquired) in step S32. However, this is not limited to this. For example, a video acquisition process and an estimation process based on the video may be performed for each number (word). In particular, a video based on a first speech instruction may be acquired and an estimation process based on the video may be performed, and then a video based on a second speech instruction may be acquired and an estimation process based on the video may be performed. Furthermore, a video acquisition process based on a third speech instruction and an estimation process based on the video may be performed, and then a video acquisition process based on a fourth speech instruction and an estimation process based on the video may be performed. Furthermore, the estimation process based on a certain video (e.g., a video based on the first speech instruction) and the acquisition process of the next video (e.g., a video based on the second speech instruction) may be performed in parallel.

[0152] <7. Verification of Learning Model 410> Next, a verification example of the learning model 410 will be described.

[0153] In this verification example, the test data (learning data and evaluation data) used was not the data (video 330) produced by the above-described method itself, but data similar to it. This verification example employs a method different from that of the above-described embodiment, in that, as will be described later, the speech section F and the non-speech section B in the video are identified in advance (the start point of the speech section F in the video is consequently identified). Simply put, in this verification example, the above-described video 330 is artificially (artificially) generated. Therefore, this verification example cannot be considered a rigorous verification example. However, this verification example makes it possible to grasp a general trend.

[0154] Specifically, in this verification example (experimental example), certain video data (300 people, 10 digits spoken per video, a total of 319,350 digits) was used as learning data. Furthermore, other video data (10 people, 3-5 characters spoken per video, a total of 1,795 digits) was used as evaluation data. These data were not collected using the method of the above-described embodiment (see FIG. 12, etc.), but were video data of a person's face uttering multiple digits one by one with approximately regular silent intervals in between (speaking intermittently). Among these data, data deemed unsuitable for learning or evaluation was omitted.

[0155] More specifically, in each video data, a speech section F and a non-speech section B are specified in advance. Then, video 330 for each word (each number) is extracted from each video data using the following method. Each video data has a frame rate of 30 fps (frames per second).

[0156] Specifically, a video 330 of the "45-frame model" (video of 1.5 seconds in length) will be described. In this 45-frame model, a number of frames of a specific number of speech intervals F are first extracted from the numerous frame images (also simply referred to as frames) that make up each facial image data. Then, the number of frames in the extracted frames is subtracted from 45 frames, and half of the remaining frames are extracted from each of the frames before and after the extracted frames. This (artificially) generates a video having a speech interval F sandwiched between two non-action intervals B (B1, B3) before and after it (a video in which the speech interval F is located near the center).

[0157] Fig. 18 is a diagram showing the original (before extraction) video data. Each line in Fig. 18 lists the start frame number, end frame number, and status type (speech / non-speech, etc.) in that order. The status type indicates the number (such as "7") actually spoken in each frame section (the section from the start frame number to the end frame number), or "pause," which indicates non-speech.

[0158] In the video data, for example, frames 0 to 33 are a non-speech section B ("pause"), and frames 34 to 48 are a speech section F for the number "9." Also, frames 49 to 66 are a non-speech section B ("pause"). Furthermore, frames 67 to 84 are a speech section F for the number "8," and frames 85 to 101 are a non-speech section B ("pause").

[0159] A case will be described in which video 330 relating to the number "8" is extracted from such original video data. In this case, a total of 18 frames are extracted from the section from frame 67 to frame 84, which is the speech section F of the number "8." Then, of the remaining 27 frames, obtained by subtracting 18 frames from the total of 45 frames, approximately half (13 frames and 14 frames) are extracted from before and after the speech section F of the number "8." Specifically, a total of 13 frames, from frame 54 to frame 66, are extracted as frame images of the non-speech section B1 (see FIG. 11), and a total of 14 frames, from frame 85 to frame 98, are extracted as frame images of the non-speech section B3. Then, a total of 45 frames, from frame 54 to frame 98, are extracted as video 330 relating to the number "8."

[0160] In this way, video images 330 for learning and evaluation are extracted (generated) including video images relating to other numbers.

[0161] Then, using a plurality of learning videos 330a as training data, the learning model 410 undergoes machine learning to generate a trained model 420. Furthermore, each evaluation video 330b is input to the trained model 420, and an output result is obtained from the trained model 420. Furthermore, the obtained output result is compared with the actual spoken digits and evaluated.

[0162] FIG. 19 shows the evaluation results. The "Top-1 recognition rate" of the 45-frame model was 94.47%, and the "Top-3 recognition rate" of the 45-frame model was 99.85%. Thus, good recognition results were obtained. Furthermore, the recognition time (processing time required for recognition) was approximately 0.05 seconds (when using a Core-i7-7820HQ 2.90GHz CPU), which means that the recognition process is extremely fast.

[0163] Here, the "Top-1 recognition rate" indicates the rate at which an actually spoken digit was correctly recognized as the top (top one) estimation result from the trained model 420. The "Top-3 recognition rate" indicates the rate at which an actually spoken digit (correct answer) was correctly recognized as one of the top three estimation results from the trained model 420 (the rate at which the correct answer was recognized as being included in the top three estimation results).

[0164] FIG. 19 also shows the evaluation results for another frame model (110 frame model).

[0165] We will now explain the video 330 of the "110-frame model" (video of approximately 3.7 seconds in length). In this 110-frame model, from the many frames that make up each face image data, multiple frames of a specific number of speech intervals F and the non-speech intervals B before and after it (the immediately preceding "pause" interval and the immediately following "pause" interval) are extracted. The remaining frames (those at the back) that do not reach 110 frames are zero-padded (filled with black images). This (artificially) generates a video having a speech interval F sandwiched between the preceding and following non-motion intervals B (B1, B3) (a video in which the speech interval F is located near the center).

[0166] A case will be described where a video 330 relating to the number "8" is extracted as a 110-frame model from such original video data (see FIG. 18). In this case, a total of 18 frames of the speech section F of the number "8" (the section from frame 67 to frame 84) are first extracted. Furthermore, a total of 18 frames of the non-speech section B (the section from frame 49 to frame 66) immediately before the speech section F and a total of 17 frames of the non-speech section B (the section from frame 85 to frame 101) immediately after the speech section F are further extracted. That is, a total of 53 (=18+18+17) frames from frame 49 to frame 101 are extracted. Furthermore, the remaining frames after these 53 frames (the remaining 57 (=110-53) frames of the total 110 frames) are padded using zero padding.

[0167] Videos relating to other numbers are generated in the same manner, thereby generating a plurality of video images 330 for learning and evaluation.

[0168] FIG. 19 also shows the evaluation results for the 110-frame model. The "Top-1 recognition rate" for the 110-frame model was 94.19%, and the "Top-3 recognition rate" for the 110-frame model was 99.75%. As shown, good recognition results were obtained. Furthermore, the recognition time (processing time required for recognition) was approximately 0.5 seconds (when using a Core-i7-7820HQ 2.90GHz CPU), which is sufficiently fast.

[0169] <8.Other> Although the embodiment of the present invention has been described above, the present invention is not limited to the above-described contents.

[0170] For example, in the above-described embodiments, the movement of a specific part of a person (human) is exemplified by a speech movement of the person's lips, but is not limited thereto. Specifically, the movement of a specific part of a person (human) may be a movement of turning the face in a specific direction (turning left, turning right, turning up, turning down, etc.). Alternatively, the movement of the specific part may be a movement of moving the eyeballs (pupils) in a specific direction (moving left and right, moving up and down, etc.), or a blinking movement of at least one of both eyes (specific eyes) (wink with the right eye, wink with the left eye, wink with both eyes, etc.). Alternatively, the movement of the specific part may be a movement of rotating the head (neck) in a specific direction (clockwise, counterclockwise, etc.). Alternatively, the movement of the specific part may be a movement of raising at least one of both hands (specific hands) high (raising the right hand, raising the left hand, raising both hands). Alternatively, the movement of the specific part may be a movement of waving at least one of both hands (a specific hand) in a specific direction (waving the right hand up and down, waving the left hand up and down, waving both hands up and down, waving the right hand from side to side, waving the left hand from side to side, waving both hands from side to side), etc.

[0171] Furthermore, in the above-described embodiments, each video 330 (310) is a video capturing a person's actions in response to an action instruction to complete a specified action related to a specific body part within a specified period D1, but the present invention is not limited to this. For example, the action instruction does not necessarily specify (explicitly indicate) the specified period D1. More specifically, the action instruction may be an instruction to perform a specified action (utterance of one word) in as short a time as possible, and / or an instruction to respond as quickly as possible and perform the specified action.

[0172] Furthermore, each of the moving images 330 (310) does not necessarily have to be images captured at times T1 to T9.

[0173] For example, each of the videos 330 may be images captured from a specific time point related to the issuance of an instruction to start an action (such as the issuance of an instruction to start speaking). Specifically, each of the videos 330 may be images captured from a specific time point near time point T1 (the time point at which the instruction to start an action is issued). More specifically, each of the videos 330 does not have to start being captured from time point T1, and may start being captured from time point T2, etc.

[0174] Furthermore, each of the videos 330 may be images captured up to a specific point in time related to the action completion deadline (such as the speech completion deadline). Specifically, each of the videos 330 may be images captured up to a specific point in time near time T9 (the point in time when the action completion deadline arrives). More specifically, each of the videos 330 (310) does not have to end capturing at time T9, and may also include, for example, a period from time T9 to time T10, when a predetermined margin period has elapsed.

[0175] In the above-described embodiments, the image processing device 30 executes both the learning process and the inference process, but the present invention is not limited to this. The learning process and the inference process may be performed by separate devices (30a, 30b). In this case, the device 30b that executes the inference process may acquire, for example, information (such as trained learning parameters) related to the trained model 420 generated by the device 30a that executed the learning process from the device 30a, and use the trained model 420.

[0176] Furthermore, in the above-described embodiment, the actions of the person are captured by the terminal device 70, but this is not limiting. For example, the image processing device 30 may capture the actions of the person and obtain the captured image 310. In this case, the screen 110 may be displayed on the display unit 35b of the image processing device 30. In other words, the processes of both devices may be executed by a single device (such as the image processing device 30). [Explanation of symbols]

[0177] 1. Image processing system 30 Image processing device 70 Terminal Equipment 75c touch panel 76c camera 110 screens 111 Image display area 112 Number display field 113 Reading column 114 Description 115 Start button 116 Recognized word display field 117 Result display field 300, 310, 320, 330, 330a, 330b Video 410 Learning Model 420 trained models B, B1, B3 Non-speech section (non-movement section) F Speech section (action section) P1~P68 feature points

Claims

1. An image processing device that receives a moving image of a subject person and generates a learning model that outputs information about the subject person's movements, a control unit that executes machine learning of the learning model using a plurality of training data in which a moving image of a person is input, the moving image including a section in which a specific part of the person is in motion and a section in which the specific part is not in motion both before and after the section in which the specific part is in motion, and information regarding the movement of the specific part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, The image processing device is characterized in that the action is a speaking action.

2. An image processing device that receives a moving image of a subject person and generates a learning model that outputs information about the subject person's movements, a control unit that executes machine learning of the learning model using a plurality of training data in which a moving image of a person is input, the moving image including a section in which a specific part of the person is in motion and a section in which the specific part is not in motion both before and after the section in which the specific part is in motion, and information regarding the movement of the specific part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, The operation is The movement of turning your face in a specific direction, The action of moving the eyeball in a specific direction, A blinking action of at least one of the eyes; The movement of turning the head in a specific direction, A gesture of raising at least one hand; A motion of waving at least one hand in a specific direction 1. An image processing device comprising:

3. 3. The image processing device according to claim 1, wherein the video used as training data is a video captured of a person's actions in response to an action instruction to complete a specified action related to the specific part within a specified period of time.

4. a control unit that identifies the motion of the subject person based on a learning model that is machine-learned using a plurality of training data, the learning model having as input a video of a person, the video including a motion section of a specific part of the person and a non-motion section of the specific part both before and after the motion section, and the output being information about the motion of the specific part; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, the control unit identifies the motion of the subject person based on information on the motion of the specific part output from the learning model in response to inputting a video captured of the motion of the subject person, the video including a non-motion section of the specific part of the subject person, into the learning model; The image processing device is characterized in that the action is a speaking action.

5. a control unit that identifies the motion of the subject person based on a learning model that is machine-learned using a plurality of training data, the learning model having as input a video of a person, the video including a motion section of a specific part of the person and a non-motion section of the specific part both before and after the motion section, and the output being information about the motion of the specific part; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, the control unit identifies the motion of the subject person based on information on the motion of the specific part output from the learning model in response to inputting a video captured of the motion of the subject person, the video including a non-motion section of the specific part of the subject person, into the learning model; The operation is The movement of turning your face in a specific direction, The action of moving the eyeball in a specific direction, A blinking action of at least one of the eyes; The movement of turning the head in a specific direction, A gesture of raising at least one hand; A motion of waving at least one hand in a specific direction 1. An image processing device comprising:

6. 6. The image processing device according to claim 4, wherein the moving image of the subject person is a moving image of the subject person's actions in response to an execution instruction to perform a specified action related to the specific part within a specified period.

7. The photographed person is a person who is a target of authentication processing, The image processing device of claim 6, characterized in that the control unit determines that the subject person is a genuine person with respect to the authentication process, on the condition that it is determined that the subject person has performed the specified action based on the output obtained by inputting the video image of the subject person into the learning model.

8. The image processing device according to claim 1 or 4, wherein the specific part includes lips.

9. the movement of the specific body part is a speech movement of one word, The image processing device according to claim 8 , wherein the information relating to the movement of the specific body part is information indicating a spoken word.

10. A learning model production method for producing a learning model that receives a moving image of a subject person as input and outputs information about the subject person's movements, comprising: a) executing machine learning of the learning model using a plurality of training data in which a video of a person is input, the video including a section in which a specific body part of the person is in motion and a section in which the specific body part is not in motion both before and after the section in which the specific body part is not in motion, and information regarding the motion of the specific body part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, A learning model production method, wherein the action is a speech action.

11. A learning model production method for producing a learning model that receives a moving image of a subject person as input and outputs information about the subject person's movements, comprising: a) executing machine learning of the learning model using a plurality of training data in which a video of a person is input, the video including a section in which a specific body part of the person is in motion and a section in which the specific body part is not in motion both before and after the section in which the specific body part is not in motion, and information regarding the motion of the specific body part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, The operation is The action of turning your face in a specific direction, The action of moving the eyeball in a specific direction, A blinking action of at least one of the eyes; The movement of turning the head in a specific direction, A gesture of raising at least one hand; A motion of waving at least one hand in a specific direction A learning model production method characterized by being any one of the above.

12. a) identifying the motion of the subject person based on a learning model machine-learned using a plurality of training data sets, in which a moving image of a person is taken, the moving image including a section in which a specific part of the person is in motion and a section in which the specific part is not in motion both before and after the moving section, and in which information regarding the motion of the specific part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, In step a), a motion of the subject person is identified based on information about the motion of the specific body part output from the learning model in response to inputting a motion image of the subject person, the motion image including a non-motion section of the specific body part of the subject person, into the learning model; The image processing method, wherein the action is a speaking action.

13. a) identifying the motion of the subject person based on a learning model machine-learned using a plurality of training data sets, in which a moving image of a person is taken, the moving image including a section in which a specific part of the person is in motion and a section in which the specific part is not in motion both before and after the moving section, and in which information regarding the motion of the specific part is output; Equipped with the video used as training data includes the movement section in which the movement of the specific body part is performed within a specified period, In step a), a motion of the subject person is identified based on information about the motion of the specific body part output from the learning model in response to inputting a motion image of the subject person, the motion image including a non-motion section of the specific body part of the subject person, into the learning model; The operation is The action of turning your face in a specific direction, The action of moving the eyeball in a specific direction, A blinking action of at least one of the eyes; The movement of turning the head in a specific direction, A gesture of raising at least one hand; A motion of waving at least one hand in a specific direction An image processing method characterized in that:

14. A program for causing a computer to execute the learning model production method according to claim 10 or 11.

15. A program for causing a computer to execute the image processing method according to claim 12 or 13.

Citation Information

Patent Citations

  • Device and method for authenticating individual and recording medium

    JP2000306090A

  • Electronic device, control device, control program, and operating method of electronic device

    JP2019079449A