Method and apparatus for realizing avatar facial expressions based on voice

The method employs a neural network to process voice signals and predict facial expression coefficients, allowing for realistic voice-based avatar facial expressions in the Metaverse without the need for camera tracking.

JP7699850B2Active Publication Date: 2025-06-30FLUENTT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023187984
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2023-11-01
Publication Date
2025-06-30
Estimated Expiration
2043-11-01

AI Technical Summary

Technical Problem

In the context of the Metaverse, there is a challenge in realistically embodying a user's facial expression in a 3D avatar without the use of a camera, especially when the user is wearing VR or AR devices.

Method used

A method and apparatus that utilize a neural network to process voice signals, dividing them into chunks, downsampling, and then upsampling to predict facial expression coefficients, which are used to realize an avatar's facial expression.

Benefits of technology

This approach enables the successful realization of voice-based avatar facial expressions by accurately predicting facial expression coefficients from voice signals, thereby overcoming the limitations of camera-based tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699850000002
    Figure 0007699850000002
  • Figure 0007699850000003
    Figure 0007699850000003
  • Figure 0007699850000004
    Figure 0007699850000004
Patent Text Reader

Abstract

To provide a method and an apparatus for realizing face expression of an avatar on an audio basis which realize face expression of an avatar based on an audio signal.SOLUTION: A method according to the present invention has a step of acquiring a plurality of chunks including amplitude information of an audio signal during a predetermined period of time, a step of inputting a plurality of input signals corresponding to the plurality of chunks into a neural network, a step of calculating, by at least one layer included in the neural network, at least one value reflecting relationship among at least two or more input signals among the plurality of input signals based on the plurality of input signals, a step of predicting a plurality of face expression coefficients by outputting a plurality of output signals based on at least one value by an output layer included in the neural network, and a step of realizing previously designated expression of an avatar based on the plurality of face expression coefficients.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for realizing an avatar facial expression based on voice, and more particularly, to a method and apparatus for realizing an avatar facial expression based on voice by applying a voice signal to a neural network.

Background Art

[0002] Metaverse means a network of 3D virtual worlds realized using virtual reality (VR) devices or augmented reality (AR) devices.

[0003] In order to generate a 3D avatar realized in a 3D virtual world in the same way as an actual user, a camera is required. However, when a user wears a VR device or an AR device, the camera cannot easily track the user's facial expression. Therefore, there is a problem that the user's facial expression cannot be realistically embodied.

[0004] Therefore, there is a need for a technology that can generate an expression similar to the actual expression of a user with only the user's voice without a camera for an avatar.

Prior Art Documents

Patent Documents

[0005] Patent Document 1: Korean Registered Patent Publication No. 10-2390781 (April 21, 2022)

Summary of the Invention

Problems to be Solved by the Invention

[0006] The technical problem to be achieved by the present invention is to provide a method and apparatus for realizing an avatar facial expression based on a voice signal for realizing an avatar facial expression.

Means for Solving the Problems

[0007] According to an embodiment of the present disclosure, at least one processor included in a computing device obtains, based on an audio signal, a plurality of chunks including amplitude information of the audio signal for a specific time; inputs a plurality of input signals corresponding to the plurality of chunks into a neural network; calculates, by at least one layer included in the neural network, at least one value reflecting a relationship between at least two or more of the plurality of input signals based on the plurality of input signals; predicts a plurality of facial expression coefficients by outputting a plurality of output signals based on the at least one value by an output layer included in the neural network; and realizes an expression of a pre-specified avatar based on the plurality of facial expression coefficients. At least two or more of the plurality of output signals included in the plurality of output signals are generated by reflecting a relationship between at least two or more of the plurality of input signals among the plurality of input signals. A method for processing an audio signal can be provided.

[0008] The plurality of facial expression coefficients can include at least one facial expression coefficient corresponding to the cheek of the avatar, at least one facial expression coefficient corresponding to the jaw of the avatar, or at least one expression coefficient corresponding to the mouth of the avatar.

[0009] The plurality of facial expression coefficients are obtained by performing an upsampling operation based on the plurality of output signals, and the upsampling operation may include a deconvolution operation of the output signal of the neural network.

[0010] The plurality of input signals may be obtained by downsampling the chunked audio signal.

[0011] The neural network is trained based on pre-stored learning conditions, and the pre-stored learning conditions may be set based on the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth.

[0012] When the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is greater than an arbitrary threshold, the learning condition is set to the first condition, and when the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is smaller than the arbitrary threshold, the learning condition may be set to a second condition different from the first condition.

Advantages of the Invention

[0013] The method and apparatus for realizing voice-based avatar facial expressions according to the embodiments of the present invention have the effect that voice-based avatar facial expressions can be successfully realized by applying a voice signal to a neural network to predict a facial expression coefficient.

Brief Description of the Drawings

[0014] To more fully understand the drawings cited in the detailed description of the present invention, a detailed description of each drawing is provided.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0015] Specific structural or functional descriptions of embodiments according to the concept of the present invention disclosed in this specification are merely exemplified for the purpose of explaining embodiments according to the concept of the present invention. Embodiments according to the concept of the present invention can be implemented in various forms and are not limited to the embodiments described in this specification.

[0016] Embodiments according to the concept of the present invention can be modified in various ways and can have various forms. Therefore, the embodiments are shown in the drawings and described in detail in this specification. However, this is not intended to limit embodiments according to the concept of the present invention to a specific disclosed form, and includes all modifications, equivalents, or replacements included in the spirit and technical scope of the present invention.

[0017] Terms such as first or second can be used to describe various components, but the components should not be limited by the terms. The terms are only for the purpose of distinguishing one component from another. For example, without departing from the scope of rights according to the concept of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component.

[0018] The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "including" or "having" are used to specify the presence of the described features, numbers, steps, operations, components, parts, or combinations thereof, and it should be understood that the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0019] Unless otherwise defined, all terms used in this specification, including technical or scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms such as those defined in advance and commonly used shall be construed to have a meaning consistent with the meaning in the context of the relevant art and shall not be construed in an idealized or overly formal sense unless clearly defined in this specification. Each block of the process flowchart in the drawings and combinations of the flowchart can be executed by computer program instructions. These computer program instructions can be implemented in the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices. Therefore, the instructions executed via the processor of a computer or other programmable data processing device generate means for performing the functions described in the flowchart blocks. These computer program instructions can also be stored in a computer-usable or computer-readable memory that can be directed to a computer or other programmable data processing device to perform functions in a specific manner. Therefore, the instructions stored in the computer-usable or computer-readable memory can also manufacture an article of manufacture that includes instruction means for performing the functions described in the flowchart blocks. Since the computer program instructions can also be implemented in a computer or other programmable data processing device, a series of operational steps are executed on the computer or other programmable data processing device, generating a process executed by the computer. Instructions for executing the computer or other programmable data processing device can also provide steps for performing the functions described in the flowchart blocks.

[0020] Also, a machine-readable storage medium can be provided in the form of a non-transitory storage medium. Here, "non-transitory" only means that the storage medium is a tangible device and does not include signals (such as electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and temporarily in the storage medium.

[0021] Also, each block can represent a module, segment, or part of code that contains one or more executable instructions for performing a specific logical function. Note that in some alternative execution examples, the functions referred to in the blocks can occur out of order. For example, two blocks shown in sequence may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order according to the functions to which they pertain. For example, operations performed by a module, program, or other component can be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations can be executed in a different order, omitted, or one or more other operations can be added.

[0022] As used herein, the term "unit" refers to a hardware component such as software or a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The "unit" serves a specific role, but is not meant to be limited to software or hardware. The "unit" may be configured to be in an addressable storage medium or may be configured to reproduce one or more processors. Thus, according to some embodiments, the "unit" includes components such as software components, object-oriented software components, class components, and task components, and processes, functions, attributes, processors, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided by the components and the "unit" can be combined into fewer components and "units" or further separated into additional components and "units". Further, the components and the "unit" can be implemented to reproduce one or more CPUs within a device or a security multimedia card. Also, according to various embodiments of the present disclosure, the "unit" can include one or more processors.

[0023] Hereinafter, the present invention will be described in detail by explaining preferred embodiments of the present invention with reference to the accompanying drawings.

[0024] FIG. 1 shows a block diagram of a voice-based avatar facial expression realization device according to an embodiment of the present invention.

[0025] Referring to FIG. 1, the voice-based avatar facial expression realization device 10 is a computing device. The computing device means an electronic device such as a notebook, a server, a smartphone, or a tablet PC.

[0026] The voice-based avatar facial expression realization device 10 includes a processor 11 and a memory 13. Although not shown in FIG. 1, the voice-based avatar facial expression realization device 10 can further include a general electronic configuration (for example, a communication circuit) included in an electronic device.

[0027] Processor 11 can include at least one processor implemented to provide at least partially different functions respectively. For example, it can execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the electronic device connected to Processor 11, and can execute various data processing or operations. According to one embodiment, as at least part of the data processing or operation, Processor 11 can store instructions or data received from other components in Memory 13 (e.g., volatile memory), process the instructions or data stored in the volatile memory, and store the result data in non-volatile memory. According to one embodiment, Processor 11 can include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with it. For example, when the electronic device includes a main processor and an auxiliary processor, the auxiliary processor can use less power than the main processor or can be set to specialize in a specified function. The auxiliary processor can be implemented separately from or as part of the main processor. The auxiliary processor can, for example, control at least part of the functions or states related to at least one of the components (e.g., a communication circuit) of the electronic device instead of the main processor while the main processor is in an inactive (e.g., sleep) state or together with the main processor while the main processor is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (e.g., an image signal processor or a communication processor) can be implemented as part of another functionally related component (e.g., a communication circuit). According to one embodiment, the auxiliary processor (e.g., a neural processing unit) can include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model can be generated through machine learning.Such learning may be performed, for example, on the electronic device itself on which the artificial intelligence model is executed, or may be performed via a separate server. The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the above examples. The artificial intelligence model can include a plurality of artificial neural network layers. The artificial neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, or one of combinations of two or more of the above, but is not limited to the above examples. In addition to the hardware structure, the artificial intelligence model can include, additionally or alternatively, a software structure. On the other hand, the operation of the electronic device (or computing device) described below can be understood as the operation of the processor 11. In the present disclosure, based on the functions of the processor, it is expressed by distinguishing into a plurality of units (for example, "~ unit"), but this is for convenience of explanation, and each unit is not necessarily realized by separate hardware. That is, the plurality of units included in the processor may be realized by separate hardware from each other, or may be realized by one hardware.

[0028] Also, the memory 13 can store at least one component of the electronic device (e.g., various data output by the processor 11). The data can include, for example, input data or output data for software (e.g., a program) and related instructions. The memory 13 can include a volatile memory or a non-volatile memory. The memory 13 can be implemented to store an operating system, middleware or an application, and / or the artificial intelligence model described above.

[0029] Also, the communication circuit can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device and an external electronic device (e.g., a user device), and the execution of communication via the established communication channel.

[0030] The processor 11 executes instructions for realizing a voice-based avatar facial expression.

[0031] The memory 13 stores instructions executed by the processor 11.

[0032] Hereinafter, it should be understood that the operations for realizing a voice-based avatar facial expression are executed by the processor 11.

[0033] FIG. 2 shows a block diagram for explaining the chunking operation of an audio signal according to an embodiment of the present invention.

[0034] Referring to FIGS. 1 and 2, the computing device 10 receives the voice signal 17 of the user 15 in order to realize a corresponding avatar facial expression based on the voice signal 17 of the user 15.

[0035] In the voice signal 17, the X-axis represents time and the Y-axis represents the amplitude of the signal. The amplitude of the signal is the amplitude generated at a sample rate of 16,000 times per second.

[0036] Processor 11 divides the audio signal 17 into a plurality of chunks (19-1 to 19-N; N is a natural number). Processor 11 divides the audio signal 17 at regular time intervals (for example, 10 seconds) to generate a chunked audio signal 19.

[0037] In FIG. 2, the interiors of the plurality of chunks 19-1 to 19-N are shown graphically, but in reality they are represented by amplitude values. The amplitude values have values within the 16-bit range. That is, the amplitude values have values between -32,768 and +32,767.

[0038] The reference symbol "S" indicates the total time length of the chunked audio signal 19.

[0039] FIG. 3 shows a block diagram of a neural network structure executed by the processor shown in FIG. 1.

[0040] Referring to FIGS. 1 to 3, processor 11 can preprocess the received audio signal to generate a neural network input signal 40. Specifically, processor 11 can downsample the chunked audio signal 19 to generate the neural network input signal 40.

[0041] Downsampling is performed by a downsampling module 30.

[0042] Processor 11 applies the chunked audio signal 19 to the downsampling module 30 and outputs the neural network input signal 40.

[0043] The chunked audio signal 19 includes a plurality of chunks 19-1 to 19-N.

[0044] The downsampling module 30 includes a plurality of downsampling layers (not shown). The plurality of downsampling layers are convolution layers. For example, the downsampling module 30 can include seven convolution layers. A 1D convolution operation is performed in the convolution layer. Each of the convolution layers can include a plurality of kernels. For example, the number of the plurality of kernels included in each of the convolution layers can be 512.

[0045] Set the interval (stride) of the kernels moving on the audio time axis to be greater than 1 so that the total time length S' of the neural network input signal 40 is generated to be shorter than the total time length S of the chunked audio signal 19.

[0046] The total time length S' of the neural network input signal 40 is shorter than the total time length S of the chunked audio signal 19. For example, assuming that the total time length S of the chunked audio signal 19 is 160,000 seconds, the total time length S' of the neural network input signal 40 is 499 seconds.

[0047] Since a plurality of kernels are used in the process of performing the 1D convolution operation, the neural network input signal 40 includes a plurality of dimensions (dimensions, D). For example, the neural network input signal 40 can include 512 dimensions. That is, the dimension D of the neural network input signal 40 is 512. The dimension means the number of channels used in the 1D convolution operation.

[0048] The neural network input signal 40 includes a plurality of input signals (IS-1, IS-2,..., and IS-K; K is a natural number) each having a total time length S' of 499 seconds. When the dimension D of the neural network input signal 40 is 512, K can be 512.

[0049] The number (N) of chunks (40-1 to 40-N; N is a natural number) of each of the plurality of input signals (IS-1, IS-2, ..., and IS-K) is equal to the number (N) of chunks 19-1 to 19-N of the chunked audio signal 19. Here, a chunk includes amplitude information of the audio signal 17 for a specific time of the user 15.

[0050] The processor 11 applies the neural network input signal 40 to the neural network 50 to generate a neural network output signal 60 in order to enhance the feature extraction of the downsampled audio signal 40.

[0051] The neural network 50 can include, but is not limited to, a transformer model.

[0052] The transformer model is a deep learning model mainly used in natural language processing. Since the transformer model is a widely known deep learning model, its specific internal structure is omitted.

[0053] The processor 11 can calculate at least one value that reflects the relationship between at least two or more of the plurality of input signals based on the plurality of input signals by at least one layer included in the neural network 50. Specifically, the processor 11 can obtain an output signal that reflects the relationship between the input signals by calculating the relationship (or correlation) between a plurality of chunks based on the input signals realized by the chunked time-series data.

[0054] In this case, the processor 11 can output a plurality of output signals based on at least one value reflecting the relationship between the input signals by the output layer included in the neural network 50. At this time, at least two or more of the output signals included in the plurality of output signals may be generated by reflecting the relationship between at least two or more of the plurality of input signals among the plurality of input signals.

[0055] Also, the processor 11 can predict a plurality of facial expression coefficients based on the plurality of output signals.

[0056] FIG. 4 shows a conceptual diagram for explaining the operation of the neural network shown in FIG. 3.

[0057] FIG. 4 is a conceptual diagram for explaining the operation of the neural network, and the operations in the actual neural network are calculated numerically.

[0058] The input signal IS-1 and the output signal OS-1 shown in FIG. 4 are any of the input signals IS-1 to IS-K shown in FIG. 3 and any of the output signals OS-1 to OS-K shown in FIG. 3.

[0059] Referring to FIGS. 1 to 4, the input signal IS-1 is applied to the neural network 50.

[0060] The input signal IS-1 includes a plurality of chunks 40-1 to 40-9.

[0061] The output signal OS-1 includes a plurality of chunks 60-1 to 60-9.

[0062] In FIG. 4, for convenience of explanation, it is shown that the number of chunks of the input signal IS-1 and the output signal OS-1 is nine, but the number of chunks may vary depending on the embodiment.

[0063] In FIG. 4, the initial weights of the neural network 50 are initialized with weights pre-trained on large-capacity audio data.

[0064] In FIG. 4, the operation of the neural network 50 when the second chunk 40-2 is applied to the neural network 50 will be described.

[0065] The second chunk 40-2 is a query vector.

[0066] The remaining chunks 40-1, 40-3 to 40-9 are key vectors. The remaining chunks 40-1, 40-3 to 40-9 are related to the second chunk 40-2.

[0067] The value vector is the amplitude value representing the characteristics of the actual voice signal 15.

[0068] The dot product of the query vector which is the second chunk 40-2 and the key vector which is the first chunk 40-1 is executed. The dot product of the query vector which is the second chunk 40-2 and the key vector which is the third chunk 40-3 is executed. The dot product of the query vector which is the second chunk 40-2 and the key vector which is the fourth chunk 40-4 is executed. Similarly, dot products are also executed for the remaining chunks 40-5 to 40-9.

[0069] When the dot product is executed, a score is generated. The score indicates how well the query vector and the key vector match.

[0070] The processor 11 multiplies the value vector by the score generated by the dot product and adds the multiplication values to generate the second chunk 60-2.

[0071] The second chunk 60-2, which is the neural network output signal 60, represents the relationship of the voice signal 15 better than the second chunk 40-2, which is the neural network input signal 40. For example, if a slight amplitude occurs in the first chunk 40-1, it is represented such that the amplitude becomes larger or smaller in the second chunk 40-2.

[0072] That is, the second chunk 60-2 is generated in consideration of the relationship between the chunks 40-1, 40-3 to 40-9 with different amplitudes of the voice signal 15 included in the second chunk 40-2 via the neural network 50.

[0073] In FIG. 4, the second chunk 60-2 is shown in the same way as the second chunk 40-2, but actually the second chunk 60-2 is different from the second chunk 40-2. The second chunk 40-2 and the second chunk 60-2 are actually represented by numbers.

[0074] Referring to FIG. 3, the feature extraction of the downsampled voice signal 40 via the neural network 50 can be enhanced.

[0075] The neural network output signal 60 includes a plurality of output signals (OS-1, OS-2,..., and OS-K; K is a natural number), each having a total time length S' of 499 seconds. When the dimension D' of the neural network output signal 60 is 512, K can be 512.

[0076] The processor 11 can post-process the neural network output signal 60 to predict a plurality of facial expression coefficients. Specifically, the processor 11 can upsample the neural network output signal 60.

[0077] The upsampling is performed by the upsampling module 70.

[0078] The upsampling module 70 includes a plurality of upsampling layers (not shown). The plurality of upsampling layers are deconvolution layers. For example, the downsampling module 30 can include seven convolution layers. In the convolution layer, a 1D deconvolution operation is performed. Each of the deconvolution layers can include a plurality of kernels.

[0079] The processor 11 applies the neural network output signal 60 to the upsampling module 70 to output a plurality of facial expression coefficients. The plurality of facial expression coefficients have values between 0 and 1. According to the plurality of facial expression coefficients, the meshes constituting the 3D avatar can be deformed. Therefore, the expressions of the 3D avatar may be different depending on the plurality of facial expression coefficients. The plurality of facial expression coefficients mean coefficients representing the avatar facial expression in relation to the movement of specific avatar face features. The plurality of facial expression coefficients can be called Shape Keys, Vertex Keys, or Morphing. At this time, the facial expression coefficients can correspond to the lower part of the avatar face. Specifically, the facial expression coefficients can include, but are not limited to, at least one facial expression coefficient corresponding to the mouth of the avatar, at least one facial expression coefficient corresponding to the jaw of the avatar, or at least one facial expression coefficient corresponding to the mouth of the avatar.

[0080] The upsampling output signal 80 includes a plurality of chunks (80-1 to 80-M; M is a natural number). One chunk has a time length of 10 seconds.

[0081] The facial expression coefficients can be generated at a rate of 60 per second. Assuming that the facial expression coefficients are generated at a rate of 60 per second, one chunk 80-1 can include 600 facial expression coefficients.

[0082] The total time length T of the upsampling output signal 80 can be 600 seconds. That is, since the time length of one chunk is 10 seconds, the number of chunks can be 60.

[0083] The total time length T of the upsampling output signal 80 is longer than the total time length S’ of the neural network output signal 60. For example, assuming that the total time length S’ of the neural network output signal 80 is 499 seconds, the total time length T of the upsampling output signal 80 is 600 seconds.

[0084] The upsampling output signal 80 includes a plurality of channels R. The plurality of channels R correspond to facial expression coefficients. For example, the first channel FE-1 can be a facial expression coefficient representing the size of the mouth shape. The second channel FE-2 can be a facial expression coefficient indicating the jaw opening of the avatar face. The number of the plurality of channels R can be 31. The 31 are facial expression coefficients at the lower part of the avatar face. The lower part of the face means the cheeks, jaw, or mouth of the avatar face.

[0085] The total time length S’ of the neural network output signal 60 and the total length of the facial expression coefficients are different. However, through the upsampling operation, the total time length S’ of the neural network output signal 60 and the total time length T of the upsampling output signal 80 can be matched.

[0086] The processor 11 can train an audio signal processing engine including the neural network 50 based on pre-stored training conditions. Specifically, the processor 11 can adjust at least one weight corresponding to the downsampling module 30, at least one weight corresponding to the upsampling module 70, and at least one weight corresponding to the neural network 50 based on pre-stored training conditions.

[0087] At this time, the pre-stored learning conditions may be set based on the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth. Specifically, when the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is greater than an arbitrary threshold value, the learning conditions can be set to the first condition. Also, when the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is smaller than the arbitrary threshold value, the learning conditions can be set to a second condition different from the first condition. The specific content regarding the above-described learning conditions will be described with reference to FIG. 5.

[0088] In order to train the weights included in the downsampling module 30, the weights included in the neural network 50, and the weights included in the upsampling module 70, the loss function can be defined as follows.

[0089] [Equation 1] JPEG0007699850000001.jpg17170

[0090] The said loss n represents the value of the loss function, the said y n represents the facial expression coefficient value of the n-th ground truth, the said y' n represents the n-th predicted facial expression coefficient value, and the said alpha represents a constant.

[0091] The loss function will be further described with reference to FIG. 5.

[0092] The processor 11 generates a facial expression coefficient value that matches the voice signal 17 of the user 15 in order to generate the ground truth. At this time, the facial expression coefficient value is the facial expression coefficient at the lower part of the avatar face.

[0093] The processor 11 extracts an arbitrary number (e.g., 100) of the generated facial expression coefficient values, an arbitrary time (e.g., 600 seconds), and the number of channels (e.g., 31) to generate a ground truth.

[0094] That is, the ground truth has 31 channels and 600 seconds, and the data per channel has 100.

[0095] The upsampling output signal 80 has 31 channels and 600 seconds. That is, the ground truth and the upsampling output signal 80 correspond to each other.

[0096] The channel values of the ground truth are normalized to have values between -1.0 and 1.0.

[0097] FIG. 5 shows graphs of facial expression coefficients according to parameter changes of a loss function to update the weights included in the downsampling module shown in FIG. 3, the weights included in the neural network, and the weights included in the upsampling module. (a) of FIG. 5 shows a graph of facial expression coefficients when the alpha value, which is a parameter of the loss function, is 0.1. (b) of FIG. 5 shows a graph of facial expression coefficients when the alpha value, which is a parameter of the loss function, is 1.0. (c) of FIG. 5 shows a graph of facial expression coefficients when the alpha value, which is a parameter of the loss function, is 10. In the graphs shown in FIGS. 5(a) to 5(c), the X-axis represents time, and the Y-axis represents the values of the facial expression coefficients.

[0098] In the graphs shown in FIGS. 5(a) to 5(c), "pred" represents the facial expression coefficients output via the downsampling module 30, the neural network 50, and the upsampling module 70, and "true" represents the facial expression coefficients of the ground truth.

[0099] The time unit of the X-axis is 1 / 60 second. Therefore, 600 on the X-axis means 10 seconds. The value of the facial expression coefficient has a value between 0 and 1.

[0100] Referring to FIG. 3, one chunk (e.g., 80-1) included in the upsampling output signal 80 has a time length of 10 seconds. The graphs shown in FIGS. 5(a), 5(b), or 5(c) correspond to chunks 80-1, 80-2, or 80-3 included in the upsampling output signal 80. In the graphs shown in FIGS. 5(a) to 5(c), the ground truth facial expression coefficients ("true" graphs) are the same. Therefore, it can be assumed that the graphs shown in FIGS. 5(a) to 5(c) correspond to chunk 80-1 included in the upsampling output signal 80.

[0101] Referring to FIG. 5(a), when the alpha value of the loss function is 0.1, the difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph) at points P1 and P2 is large. When the alpha value of the loss function is 0.1, the processor 11 determines noise at points P1 and P2 to exclude the value of the facial expression coefficient. The alpha value means the alpha value in the above formula (1).

[0102] The values of the ground truth facial expression coefficients (e.g., 0.3, 0.1) at points P1 and P2 are relatively larger or smaller than the value of the ground truth average facial expression coefficient (e.g., 0.2). The value of the ground truth facial expression coefficient (e.g., 0.3) at point P1 is relatively larger than the value of the ground truth average facial expression coefficient (e.g., 0.2). The value of the ground truth facial expression coefficient (e.g., 0.1) at point P2 is relatively smaller than the value of the ground truth average facial expression coefficient (e.g., 0.2).

[0103] On the one hand, there is little difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph) at point P3. The value of the ground truth facial expression coefficient at point P3 (for example, 0.2) is almost the same as the value of the average ground truth facial expression coefficient (for example, 0.2). The value of the average ground truth facial expression coefficient (for example, 0.2) can be calculated by taking the average of the values of the ground truth facial expression coefficients.

[0104] That is, when the alpha value of the loss function is 0.1, the closer the value of the ground truth facial expression coefficient is to the value of the average ground truth facial expression coefficient, the less difference there is between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph).

[0105] On the other hand, the greater the value of the ground truth facial expression coefficient deviates from the value of the average ground truth facial expression coefficient, the greater the difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph).

[0106] When the alpha value of the loss function is 0.1, the values of the facial expression coefficients at points P1 and P2 are not well predicted.

[0107] However, it is important to accurately predict large values of the facial expression coefficient. The greater the value of the facial expression coefficient, the more the facial expression of the avatar will deform. Therefore, in order to make the differences in avatar facial expressions more distinct, it is necessary to accurately predict large values of the facial expression coefficient.

[0108] Referring to Fig. 5(b), when the alpha value of the loss function is 1.0, the predicted facial expression coefficient does not differ significantly from the case when the alpha value of the loss function is 0.1 (the case of Fig. 5(a)).

[0109] At points P1 and P2, the difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph) is large. At point P3, there is little difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph).

[0110] Referring to Fig. 5(c), when the alpha value of the loss function is 10, the difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph) at points P1 and P2 is not significantly large compared to Fig. 5(a). This is because when the alpha value of the loss function is 10, the processor 11 does not determine the values of the facial expression coefficients at points P1 and P2 as noise.

[0111] On the other hand, the difference between the ground truth facial expression coefficient ("true" graph) and the predicted facial expression coefficient ("pred" graph) at point P3 is different from Fig. 5(a).

[0112] When the alpha value of the loss function is small as in Fig. 5(a) and Fig. 5(b), the facial expression coefficient values of the ground truth similar to the average facial expression coefficient value of the ground truth (point P3) are well predicted. On the other hand, the facial expression coefficient values of the ground truth with a large difference from the average facial expression coefficient value of the ground truth (points P1, P2) are not well predicted.

[0113] When the alpha value of the loss function is large as in Fig. 5(c), the facial expression coefficient values of the ground truth similar to the average facial expression coefficient value of the ground truth (point P3) are not well predicted. On the other hand, the facial expression coefficient values of the ground truth with a large difference from the average facial expression coefficient value of the ground truth (points P1, P2) are well predicted.

[0114] The processor 11 increases the alpha value to better predict the values of the ground truth facial expression coefficients (at points P1 and P2) that have a large difference from the value of the average facial expression coefficient of the ground truth, and trains the weights included in the downsampling module 30, the weights included in the neural network 50, and the weights included in the upsampling module 70.

[0115] Also, the processor 11 decreases the alpha value to better predict the value of the ground truth facial expression coefficient (at point P3) that is similar to the value of the average facial expression coefficient of the ground truth, and trains the weights included in the downsampling module 30, the weights included in the neural network 50, and the weights included in the upsampling module 70.

[0116] That is, the processor 11 calculates the value of the average facial expression coefficient of the ground truth. The ground truth means the facial expression coefficient value that matches the voice signal 17 of the user 15.

[0117] The processor 11 calculates the absolute difference value between the value of the average facial expression coefficient of the ground truth and the channel value of the upsampling output signal 90.

[0118] When the absolute value is greater than an arbitrary threshold, the processor 11 sets the alpha value of the loss function to a first value. The first value is greater than or equal to a first arbitrary value (for example, 5). For example, the arbitrary threshold can be 0.1. The first value can be 10. The first value has a range that is greater than or equal to a first arbitrary value (for example, 5) and less than or equal to a second arbitrary value (for example, 20). The second arbitrary value is greater than the first arbitrary value.

[0119] When the absolute value is less than an arbitrary threshold, the processor 11 sets the alpha value of the loss function to a second value. The second value is less than the first arbitrary value (for example, 5). For example, the arbitrary threshold can be 0.1. The second value can be 0.1. The second value has a range that is greater than or equal to 0 and less than the first arbitrary value (for example, 5).

[0120] According to an embodiment, the processor 11 can set the first arbitrary value, the second arbitrary value, the first value, or the second value to be different.

[0121] By setting the alpha value to be different according to the difference between the value of the average facial expression coefficient of the ground truth and the channel value of the upsampling output signal 80 as in the present invention, it is possible to better predict the values of the facial expression coefficients (at points P1 and P2) of the ground truth with a large difference from the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of the ground truth (at point P3) similar to the value of the average facial expression coefficient of the ground truth.

[0122] According to an embodiment, the processor 11 counts the number of epochs for the loss function. When the number of epochs for the loss function is a certain number of times (for example, 20) or more, the processor 11 can set the alpha for the loss function to the second value. This is because when the number of epochs for the loss function is a certain number of times (for example, 20) or more, the deviation of the error of the loss function is not large. According to an embodiment, the processor 11 can set the certain number of times (for example, 20) to be different.

[0123] FIG. 6 shows a flowchart of a method for realizing an audio-based avatar facial expression according to an embodiment of the present invention.

[0124] Referring to FIGS. 1 to 6, the processor 11 divides the audio signal 17 into a plurality of chunks 19-1 to 19-N (S10).

[0125] The processor 11 downsamples the chunked audio signal 19 to generate a neural network input signal 40 (S20).

[0126] Processor 11 applies the neural network input signal 40 to the neural network 50 to generate a neural network output signal 60 in order to enhance the feature extraction of the downsampled audio signal 40. The neural network input signal 40 and the downsampled audio signal 40 are the same signal (S30).

[0127] Processor 11 upsamples the neural network output signal 60 to predict a plurality of facial expression coefficients (S40). The plurality of facial expression coefficients are included in the upsampled output signal 80.

[0128] Processor 11 realizes the expression of the avatar face according to the predicted plurality of facial expression coefficients (S50).

[0129] FIG. 7 shows a 3D avatar image realized according to the voice according to an embodiment of the present invention. FIGS. 7(a) to 7(d) show images of avatars with different expressions realized over time.

[0130] Referring to FIGS. 1, 2, and 7, processor 11 realizes the expression of the avatar face according to the predicted plurality of facial expression coefficients.

[0131] The avatar can be realized in 3D. The reference image of the avatar can be the user 15 pre-photographed by a camera (not shown).

[0132] The present invention has been described with reference to one embodiment shown in the drawings, which is merely exemplary, and those of ordinary skill in the art should understand that various modifications and equivalent other embodiments are possible hereinafter. Therefore, the true technical protection scope of the present invention should be determined by the technical idea of the appended registered claims.

Explanation of Reference Numerals

[0133] 10 Device for realizing voice-based avatar facial expression 11 Processor 13 Memory 30 Downsampling Module 50 Neural Network 70 Upsampling Module

Claims

1. In a voice signal processing method, by at least one processor included in a computing device, acquiring a plurality of chunks including amplitude information of the voice signal for a specific time based on the voice signal; inputting a plurality of input signals corresponding to the plurality of chunks into a neural network; calculating at least one value reflecting a relationship between at least two or more of the plurality of input signals based on the plurality of input signals by at least one layer included in the neural network; predicting a plurality of facial expression coefficients by outputting a plurality of output signals based on the at least one value by an output layer included in the neural network; realizing an expression of a pre-specified avatar based on the plurality of facial expression coefficients, at least two or more of the output signals included in the plurality of output signals are generated by reflecting a relationship between at least two or more of the plurality of input signals among the plurality of input signals, the neural network is trained based on pre-stored training conditions, the pre-stored training conditions are set based on a difference between an average facial expression coefficient value of ground truth and a facial expression coefficient value of any ground truth, when an absolute value of a difference between the average facial expression coefficient value of the ground truth and the facial expression coefficient value of any ground truth is greater than an arbitrary threshold, the training condition is set to a first condition, A method for realizing a voice-based avatar facial expression, characterized in that when an absolute value of a difference between the average facial expression coefficient value of the ground truth and the facial expression coefficient value of any ground truth is smaller than the arbitrary threshold, the training condition is set to a second condition different from the first condition.

2. The plurality of facial expression coefficients include at least one facial expression coefficient corresponding to the cheek of the avatar, at least one facial expression coefficient corresponding to the jaw of the avatar, or at least one expression coefficient corresponding to the mouth of the avatar. The method for realizing a voice-based avatar facial expression according to claim 1.

3. The plurality of facial expression coefficients are obtained by performing an upsampling operation based on the plurality of output signals, The upsampling operation includes a deconvolution operation of an output signal of the neural network. The method for realizing a voice-based avatar facial expression according to claim 1.

4. The method for realizing an avatar facial expression based on voice according to claim 1, wherein the plurality of input signals are obtained by downsampling the chunked voice signal.

5. A processor for executing instructions, A memory for storing the instructions, and The processor Obtains a plurality of chunks including amplitude information of the voice signal for a specific time based on the voice signal, Inputs a plurality of input signals corresponding to the plurality of chunks into a neural network, Calculates at least one value reflecting the relationship between at least two or more of the plurality of input signals based on the plurality of input signals by at least one layer included in the neural network, Predicts a plurality of facial expression coefficients by outputting a plurality of output signals based on the at least one value by an output layer included in the neural network, Is set to realize an expression of a pre-specified avatar based on the plurality of facial expression coefficients, At least two or more of the output signals included in the plurality of output signals are generated by reflecting the relationship between at least two or more of the plurality of input signals, The neural network is trained based on pre-stored training conditions, The pre-stored training conditions are set based on the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth, When the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is greater than an arbitrary threshold value, the training condition is set to a first condition, An apparatus for realizing an avatar facial expression based on voice, characterized in that when the absolute value of the difference between the value of the average facial expression coefficient of the ground truth and the value of the facial expression coefficient of any ground truth is smaller than the arbitrary threshold value, the training condition is set to a second condition different from the first condition.

6. The plurality of facial expression coefficients The apparatus for realizing an avatar facial expression based on voice according to claim 5, characterized in that it includes at least one facial expression coefficient corresponding to the cheek of the avatar, at least one facial expression coefficient corresponding to the jaw of the avatar, or at least one expression coefficient corresponding to the mouth of the avatar.

7. The plurality of facial expression coefficients are obtained by performing an upsampling operation based on the plurality of output signals. The voice-based avatar facial expression realization device according to claim 5, wherein the upsampling operation includes a deconvolution operation of an output signal of the neural network.

8. The voice-based avatar facial expression realization device according to claim 5, wherein the plurality of input signals are obtained by downsampling the chunked voice signal.

Citation Information

Patent Citations

  • Learning method and learning device for generating training data acquired from virtual data on virtual world by using generative adversarial network (GAN), to thereby reduce annotation cost required in learning processes of neural network for autonomous driving, and testing method and testing device using the same

    JP2020123345A

  • Artificial intelligence-based voice-driven animation method and apparatus, device and computer program

    JP2022537011A