Program in which behavior changes according to expression of user
By developing an AI program that can recognize user expressions and change game processing, the lack of flexible interfaces in existing game AI technologies has been solved, and new input methods and richer player experiences have been achieved.
Patent Information
- Application Number
- CN202411517607.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-20
AI Technical Summary
Among the existing game AI technology, there are fewer examples of flexibly utilizing AI in the interface of gamers, and there is a lack of new input units to enhance the player experience.
By combining a computer, an image display device, an input device and a photography device, an AI program is developed that can recognize the user's expression and process it according to the recognition results, thereby changing the input method of the game program.
New input methods are implemented, enhanced player experience, and provided unprecedented freshness and interactivity.
Smart Images

Figure CN120020680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a program in which the behavior using AI technology changes according to the user's expression. Background Art
[0002] In recent years, due to the emergence of ChatGPT, the attention to AI has increased sharply. Along with this, the related investment in AI has increased, and thus it has evolved, and the applications in various fields have been continuously expanding.
[0003] In the fields of games and entertainment, AI is also flexibly used in various forms. Especially at the game development site, it plays a very important role in improving work efficiency. For example, an AI generator is also used in which a game developer inputs a text prompt (a simple description of the content to be visualized) and can immediately generate an image.
[0004] Regarding the actual game content, there are also technologies in which AI is directly involved. For example, character AI in which characters in the game act autonomously, and Meta AI that grasps and controls the overall situation of the game from an overhead viewpoint such as the appearance of enemy characters or changes in the game environment according to the player's actions.
[0005] Prior Art Documents Patent Documents Patent Document 1: Japanese Patent Application Laid-Open No. 2012-516507 In such conventional game AI technologies, there are few examples of flexibly using AI in the game player's interface. Most game players use conventional controllers such as joysticks or game pads to play games.
[0006] In addition, Patent Document 1 discloses a system that takes in the user's gesture as an input to an application program such as a game. This system analyzes the gesture by obtaining depth information from scattered light pulse imaging and requires special equipment.
[0007] Therefore, an object of the present invention is to provide a new input unit for a program executed by a computer using AI technology. Summary of the Invention
[0008] To solve the above problems, a program according to one aspect of the present invention is a program executed by a system including a computer, an image display device, an input device, and a photographing device that photographs a user's face. The program is characterized in that the computer is connected to the image display device, the input device, and the photographing device, and based on information from the input device and the photographing device, a predetermined image is displayed on the image display device. The program causes the computer to execute an identification step of identifying the user's expression based on information from the photographing device, and the processing performed by the program changes according to the user's expression identified in the identification step.
[0009] Furthermore, in one embodiment, the program according to claim 1 is characterized in that in the identification step, the user's expression is classified into multiple categories by a convolutional neural network.
[0010] Furthermore, in one embodiment, the program according to claim 2 is characterized in that the program is a game program, and in addition to a controller operated by the user's hand, a change in the expression of the user's face also serves as an input unit for the game program.
[0011] According to the program of the present invention, a new input method is installed, enabling an unprecedented fresh experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 FIG. 1 is a perspective view of a flight simulator 1 incorporating a program according to an embodiment of the present invention.
[0013] Figure 2 FIG. 2 is a diagram for explaining the principle that the display screen of a curved surface display 20 is imaged as an aerial image G through an optical plate 40 in a flight simulator incorporating a program according to an embodiment of the present invention. The figure shows only the optical plate 40, the display screen of the curved surface display 20, and the aerial image G of the flight simulator 1 as viewed from the left side of the flight simulator 1. Figure 1 of the flight simulator 1.
[0014] Figure 3 FIG. 3 is a diagram for explaining the principle that the display screen of a curved surface display 20 is imaged as an aerial image G through an optical plate 40 in a flight simulator incorporating a program according to an embodiment of the present invention. The figure shows only the optical plate 40, the display screen of the curved surface display 20, and the aerial image G of the flight simulator 1 as viewed from directly above the flight simulator 1. Figure 1 of the flight simulator 1.
[0015] Figure 4This is a diagram showing the principle by which, in a flight simulator using the program of an embodiment of the present invention, the display screen of the curved display 20 is imaged as an aerial image G through the optical plate 40. The diagram showing the flight simulator 1 as viewed from the front only shows Figure 1 the optical plate 40, the display screen of the curved display 20, and the aerial image G of the flight simulator 1.
[0016] Figure 5 This is a block diagram showing the architecture of a convolutional neural network implemented in the program of an embodiment of the present invention.
[0017] Figure 6 This shows the screen image when flying with a happy expression in a flight simulator using the program of an embodiment of the present invention.
[0018] Figure 7 This shows the screen image when flying with a sad expression in a flight simulator using the program of an embodiment of the present invention. Detailed Embodiments
[0019] Hereinafter, embodiments of the program of the present invention will be described with reference to the drawings. Here, it is applied to a flight simulator that can simulate actual operation while sitting in a cockpit that mimics an actual cockpit and operating the aircraft while watching an image.
[0020] That is, as Figure 1 shown, the flight simulator 1 serves as the cockpit control seat of a simulated aircraft and is composed of a simulator main body 10 and a seat 17 for the operator who operates the flight simulator 1. Moreover, the simulator main body 10 is equipped with a control stick 12, control pedals 14, instruments 16, a camera 18, a speaker 19, etc.
[0021] Inside the simulator main body 10, a curved display 20, a control device 30, and an optical plate 40 are installed. The control device 30 is connected to the control stick 12, control pedals 14, instruments 16, camera 18, and curved display 20, etc. via internal wiring (omitted in the figure), exchanges signals with them, and simulates the flight situation for the boarding experience.
[0022] The incident surface of the optical plate 40 faces downward and is opposite to the display surface of the curved display 20 at a certain angle (for example, 45 degrees). Moreover, the image on the display surface of the curved display 20 is refocused as an aerial image G at a symmetric position on the opposite side of the optical plate 40, forming the same image as the original. That is, it can be considered that the imaging position in the air constitutes the aerial display G.
[0023] Of course, the images displayed on the curved display 20 are flight images (background, actual aircraft, etc.) generated by the flight simulator 1. In addition, the specific implementation of flight operation simulation and the like performed by the flight simulator 1 here, except for the control based on the images of the camera 18 described later, is the same as that of conventional flight simulators, so detailed description is omitted here.
[0024] The curved display 20 is a convex curved liquid crystal display device that is placed horizontally with the display surface facing upward. Here, the curvature of the convex surface is, for example, 1000R. Instead of the curved liquid crystal display device, a flexible display composed of an organic EL display, an electronic paper with a backlight, or the like can also be bent at a desired curvature for use. In short, the display surface faces upward to form a convex surface.
[0025] Here, the position of the curved display 20 is fixed, but a support structure can also be designed to adjust the position in the vertical direction. In this case, the condensing position of the aerial image G can be adjusted to a position that is easy for the operator to see.
[0026] The control device 30 is essentially a small computer, which is composed of a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), a storage device for storing various programs and data, an input / output interface, etc. As the input / output interface, for example, a USB port, a wireless LAN such as WIFI, etc. are installed. Through this input / output interface, various data and programs related to flight simulation can be updated, and the model of the learned convolutional neural network described later can be downloaded. Moreover, the control device 30 outputs an image signal to the curved display 20 to perform the display that forms the basis of the aerial image, and outputs a drive signal to the speaker 19 to play a sound field synchronized with the image on the curved display 20.
[0027] In addition, as the optical plate 40, for example, the optical imaging element (two-sided orthogonal reflector) described in Japanese Patent Application Laid-Open No. 2011-175297 can be used. This optical imaging element is realized by arranging a plurality of mutually orthogonal planar light reflection parts at a certain interval. In addition to this, a structure such as a dihedral reflector in which a reflecting surface is formed on the side surface of a quadrangular hole described in Japanese Patent No. 4900618 can also be used.
[0028] Figure 2 It is a diagram for explaining the principle that the display screen of the curved display 20 is imaged as the aerial image G through the optical plate 40. For the sake of simplicity of explanation, it is shown only as a view observed from the left side of the flight simulator 1. Figure 1The optical plate 40 of the flight simulator 1, the display screen of the curved display 20, and the aerial image G.
[0029] In addition, Figure 3 This is also a diagram illustrating the principle by which the display screen of the curved display 20 forms an aerial image G through the optical plate 40. For simplicity of explanation, only the Figure 1 optical plate 40 of the flight simulator 1, the display screen of the curved display 20, and the aerial image G are shown as a view from directly above the flight simulator 1.
[0030] Furthermore, Figure 4 This is also a diagram illustrating the principle by which the display screen of the curved display 20 forms an aerial image G through the optical plate 40. For simplicity of explanation, only the Figure 1 optical plate 40 of the flight simulator 1, the display screen of the curved display 20, and the aerial image G are shown as a view from the front of the flight simulator 1 ( Figure 1 behind the seat 17).
[0031] Under certain conditions of incident light, the optical plate 40 with a double - reflection structure such as a two - plane orthogonal reflector or a two - plane angular reflector does not change the component of the incident light orthogonal to the panel plane, but recursively reflects the incident light in the direction of the panel plane (recursive transmission).
[0032] Refer to Figure 2 for an explanation. As the coordinate system (x, y, z), the x - direction is the paper - plane direction along the surface of the optical plate 40 (i.e., the 45 - degree direction connecting the upper - left and lower - right corners of the paper), the z - direction is the paper - plane direction perpendicular to the surface of the optical plate 40 (i.e., the 45 - degree direction connecting the upper - right and lower - left corners of the paper), and the y - direction is the direction perpendicular to the paper (i.e., the direction passing through the paper at a right angle). Here, the incident - light vector (v x 、v y 、v z ) is recursively transmitted through the optical plate 40 and becomes the outgoing - light vector (-v x 、-v y 、v z ).
[0033] Therefore, the real image of the mirror image is formed on the opposite side of the optical plate 40. As a result, the display screen of the curved display 20 and the aerial image G are mirror - symmetric with respect to the optical plate 40.
[0034] That is, the position L1 (the most bulging position) at the horizontal center of the curved display 20 is closest to the optical plate 40, and the light emitted from the position L1 converges at the position M1 closest to the optical plate 40. That is, the light emitted from the position L1 is reflected at any positions R1, R1 on the optical plate 40 and converges at the position M1 that is the same distance away from the optical plate 40 as the distance between the position L1 and the optical plate 40 on the opposite side of the optical plate 40.
[0035] Similarly, the position L2 on the lateral outside of the curved display 20 is far from the optical plate 40, and the light emitted from the position L2 converges at the position M2 away from the optical plate 40. That is, the light emitted from the position L2 is reflected at any positions R2, R2 on the optical plate 40 and converges at the position M2 that is the same distance away from the optical plate 40 as the distance between the position L2 and the optical plate 40 on the opposite side of the optical plate 40.
[0036] As a result, when observed by the operator sitting on the seat 17, the aerial image G displayed in front appears to be curved into a concave surface, and the concave curved aerial display is installed to float in the air.
[0037] In a flight simulator using a conventional flat display, it is felt that the left and right ends of the display are far from the eyes and difficult to see clearly, and the image is distorted at the ends, increasing the difference from the original field of view. In contrast, in this concave curved aerial display, the field of view of the operator actually seen during flight can be directly reproduced without a sense of incongruity.
[0038] In addition, in the aerial display, since there is no physical display itself, a more immersive and realistic operation experience with a three-dimensional effect can be achieved. Furthermore, in a physical display, the reflection of external light on the display surface always obstructs the operator's field of view anyway, but in the aerial display, there is no such reflection of external light, and the operator can experience a deep sense of immersion in the simulated operation experience without being hindered by external light.
[0039] Furthermore, in the present invention, the camera 18 plays an important role. The camera 18 is provided in front of the flight simulator 1 and captures the face of the operator sitting on the seat 17. The image of the operator's face captured by the camera 18 is sent to the control device 30 in real time during the operation of the flight simulator 1.
[0040] In this flight simulator 1, the weather changes according to the expression of the operator's face captured by the camera 18. By consciously making an expression, the operator can also make the weather change in accordance with the intention. In addition, when the operator is operating, emotions and consciousness are sometimes inadvertently expressed on the face, and thus the weather also changes.
[0041] The image data of the face of the operator captured by the camera 18 is sent to the control device 30. An expression recognition system for recognizing the received facial expressions is installed in the control device 30, which recognizes expressions such as anger, disgust, fear, joy, sadness, panic, and expressionless.
[0042] This expression recognition system is implemented by a convolutional neural network (CNN), which is a representative method of deep learning. Figure 5 Represents the architecture of the CNN. That is, the CNN consists of a first convolutional layer C1, a first pooling layer P1 that receives the output of the first convolutional layer C1, a second convolutional layer C2 that receives the output of the first pooling layer P1, a second pooling layer P2 that receives the output of the second convolutional layer C2, and a fully connected layer F that receives the output of the second pooling layer P2. Here, the convolutional layer and the pooling layer are connected in two levels, but it can also be more levels.
[0043] The image data of the face of the operator is input to the first convolutional layer C1. However, the image data is resized to match the first convolutional layer and transformed into a grayscale image. The image data is two-dimensional arranged data I, and n-dimensional arranged data is sometimes called a tensor.
[0044] In the first convolutional layer C1, the inner product of a kernel, which is a two-dimensional arranged data of a small size (e.g., 9×9), and a partial area of the same size (window: e.g., 9×9) of the input image data I is calculated to obtain an output. This inner product is performed over the entire input image data I while shifting the window, and as a result, the output O becomes two-dimensional arranged data. Each parameter of the kernel is optimized in advance through learning to extract the features of the input image data I. In addition, multiple (e.g., 16) kernels are used, and multiple sets of outputs are also obtained. However, instead of directly outputting the inner product calculation result, it is non-linearized through an activation function. Here, the activation function ReLU() that sets the negative value of the inner product calculation result to 0 is used.
[0045] The output of the first convolutional layer C1 is input to the first pooling layer P1. In the first pooling layer P1, a downsampling operation that only retains the representative value of each local spatial range (e.g., 2×2) is performed. Specifically, only the maximum value is retained. That is, by performing interval elimination and compression on the output result of the first convolutional layer C1, it has the advantages of coping with fine offsets and noises in the image and significantly reducing the computational amount.
[0046] The output of the first pooling layer P1 is input to the second convolutional layer C2, passed through the activation function ReLU(), and further input to the second pooling layer P2. The second convolutional layer C2 and the second pooling layer P2 have the same architecture as the first convolutional layer C1 and the first pooling layer P1, but different sizes, and the kernels, etc. are also optimized separately.
[0047] The output of the second pooling layer P2 is input to the fully connected layer F. As described later, in this embodiment, facial expressions are recognized in seven categories. Therefore, the fully connected layer F as the output layer has seven units (neurons). Each neuron is connected to all the outputs of the second pooling layer P2. Let the output of the second pooling layer P2 be pj (j = 1 to M). The output Ni (i = 1, 2,..., 7) of each neuron is calculated as follows.
[0048] Ni = bi + w i1 ·p 1 + w i2 ·p 2 +…+ w iM ·p M In the above formula, bi represents the bias, and w j represents the weight. Furthermore, in order to make the output Ni the probability of each category, that is, to make the sum of the outputs Ni equal to 1, the following softmax function is calculated as the final probability value Ei.
[0049] Ei = e Ni / (e N1 + e N2 +…+ e N7 ) In addition, the values of the above-mentioned parameters are pre-adjusted to the optimal values through learning using the error backpropagation method, etc. Each parameter and the learned model can be appropriately downloaded from the Internet to update the expression recognition system.
[0050] In this embodiment, facial expressions are recognized in seven categories. That is, anger, disgust, fear, joy, sadness, surprise, and expressionless. In addition, the recognition results are obtained as the probabilities of each category. Depending on the content of the program applying the present invention, all seven categories or their probabilities, or only a part thereof, are used.
[0051] In the installation of the flight simulator 1, the weather is calculated based on the seven categories and their probabilities. As the weather, for example, it includes cloud cover, precipitation, snowfall, wind direction, wind speed, humidity, etc. In addition, it may also include the change of wind, the gustiness of wind, turbulence (the degree of air flow disorder), and the setting of updrafts. In addition, it may also include the setting of cumulonimbus clouds, storms, etc. Appropriate selection is made from these items for expression-based control.
[0052] As an example, taking the output (probability) of the expression recognition system as anger E1, disgust E2, fear E3, joy E4, sadness E5, surprise E6, and expressionless E7, the parameters of the weather are calculated as follows.
[0053] Cloud cover (%) = MIN(100, 10 + 150×(1 - E4 - E7)) Precipitation (mm) = MAX(0, 10×(2×E1 + 3×E2 + E3 + 2×E5) - 2) Wind change (%) = 100×(E1 + E2 + E3 + E5 + E6) Wind speed (m / s) = 10×E1 + 10×E2 + 15×E3 + E4 + 2×E6 Turbulence (%) = 10×(5×E1 + 5×E2 + 10×E3 + 5×E6) Use this setting as an environment variable, which determines the weather when simulating and running the program of the flight simulator 1. Thus, the operator's emotions are reflected in the weather during the simulation. Or, by changing the facial expressions, the operator can change the weather for flying to a preferred one. For example, when wanting to enjoy flying comfortably under a sunny sky, fly with a happy expression ( Figure 6 ), and when wanting to feel tension in bad weather, fly with a sad expression ( Figure 7 ).
[0054] Industrial applicability According to the program of the present invention, no special input device is required. By using AI technology, an application program with a new input method for the program executed by a computer is implemented.
[0055] As described above, the program of the present invention has been described based on the embodiments, but the present invention is not limited thereto. Additions and changes can be made without departing from the gist of the present invention. If possible, the technologies described in each embodiment can also be combined, or known technologies can be combined, etc.
[0056] In the above embodiment, it is shown applied to a manipulation simulator that simulates a cockpit, but the application of the present invention is not limited thereto. For example, it can also be applied to a game program executed by a general personal computer with a web camera. As an example, in a fighting game, if an "angry" expression is recognized, the attack power is increased. In this case, in addition to the skills of operating the controller in the past, the competition in facial expressiveness is added, doubling the fun.
[0057] In addition, in the above embodiment, as the convolutional neural network for recognizing facial expressions, two sets of convolutional layers and pooling layers are alternately installed, but the present invention is not limited thereto. For example, the convolutional layer can be set to three layers, and the pooling layer can be combined with only a part. Furthermore, multiple fully connected layers can also be stacked.
[0058] Furthermore, the gist of the present invention can also be further developed by providing sensors for detecting information on changes in a living body (such as the heart, body temperature, muscles, brain waves, respiration, etc.), and changing the processing performed by the program according to their changes. Of course, it can also be a combination with the above-described change based on an expression. For example, the following situation can be considered: in a fighting game, the energy ball is inflated due to the stiffness of the arm muscles and the degree of effort on the face.
[0059] Description of Reference Numerals 1 Flight simulator 10 Simulator main body 12 Handle 14 Control pedal 16 Instruments 17 Seat 18 Camera 19 Speaker 20 Curved display 30 Control device 40 Optical plate C1, C2 Convolutional layer F Fully connected layer G Aerial image P1 Pooling layer P2 Pooling layer
Claims
1. A program executed by a system, wherein the system is composed of a computer, an image display device, an input device, and a photographic device for photographing a user's face, wherein the computer is connected to the image display device, the input device, and the photographic device, and a prescribed image is displayed on the image display device based on information from the input device and the photographic device, and the program causes the computer to execute a recognition step of recognizing the user's expression based on the information from the photographic device, and the processing performed by the program varies according to the expression of the user recognized in the recognition step.
2. The program according to claim 1, characterized in that In the recognition step, the user's expression is classified into multiple categories through a convolutional neural network.
3. The program according to claim 2, characterized in that The program is a game program, and in addition to the controllers operated by the user's hands, changes in the user's facial expressions also serve as input means for the game program.
Citation Information
Patent Citations
JP1974000618A
Method of manufacturing light control panel for use in optical imaging device
JP2011175297A
Standard gestures
JP2012516507A