Training data construction method and expression coefficient extraction method

By collecting expression coefficients from the front perspective and combining multi-view image acquisition to construct training data, the problem of inaccurate expression coefficients at the non-frontal perspective is solved, and the quality of the training data and the expression coefficient extraction performance of the model are improved.

WO2025175892A1PCT designated stage Publication Date: 2025-08-28BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/140260
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2024-12-18
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

In the prior art, the expression coefficients collected by the facial expression data acquisition device at the non-facial perspective are inaccurate, resulting in a decrease in the quality of the training data, which in turn affects the model's expression coefficient extraction performance at the non-facial perspective.

Method used

The training data is constructed using expression coefficient acquisition sequence and multi-view image acquisition sequence to ensure that the expression coefficient acquisition device is collected from the front view angle, the image acquisition device is collected from different view angles, and the training data is constructed through frame alignment processing to maintain the consistency of expression changes and the balance of perspective angles.

Benefits of technology

The quality of the training data and the expression coefficient extraction performance of the model at different perspectives are improved, the proportion imbalance of front-face images and non-frontal images is avoided, and the expression coefficient extraction effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140260_28082025_PF_FP_ABST
    Figure CN2024140260_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a training data construction method and an expression coefficient extraction method. The training data construction method comprises: firstly, obtaining an expression coefficient acquisition sequence and at least two image acquisition sequences; and then constructing training data on the basis of the expression coefficient acquisition sequence and the at least two image acquisition sequences. In this way, model updating processing can be carried out subsequently on the basis of the training data, and the updated model is used to perform expression coefficient extraction processing.
Need to check novelty before this filing date? Find Prior Art

Description

A training data construction method and expression coefficient extraction method

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 202410190491.4, filed on February 20, 2024, entitled “A training data construction method and expression coefficient extraction method”. The entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present application relates to the field of data processing technology, and in particular to a training data construction method, an expression coefficient extraction method, an apparatus, a device, a medium, and a product. Background Art

[0004] In some application scenarios, such as face-driven scenarios, the model can be trained based on some training data first, so that the trained model has the ability to extract expression coefficients; then the image is processed with the help of the model to obtain the expression coefficients to complete the expression coefficient extraction task. Summary of the Invention

[0005] The present application provides a training data construction method, expression coefficient extraction method, device, equipment, medium, and product, which are conducive to improving the quality of training data.

[0006] In order to achieve the above objectives, the technical solutions provided by this application are as follows:

[0007] The present application provides a method for constructing training data, the method comprising:

[0008] Acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; the expression coefficient acquisition sequence and each of the image acquisition sequences are acquired for the same face; the acquisition viewing angle of the acquisition device for the expression coefficient acquisition sequence is a frontal viewing angle of the face; the acquisition viewing angles of the acquisition devices for different image acquisition sequences are different;

[0009] Training data is constructed based on the expression coefficient acquisition sequence and the at least two image acquisition sequences.

[0010] In a possible implementation, the at least two image acquisition sequences include at least one first image sequence; each of the first image sequences corresponds to an image acquisition device; and each of the image acquisition devices meets a synchronization condition.

[0011] In a possible implementation manner, the at least two image acquisition sequences include a second image sequence; and the expression coefficient acquisition sequence is determined based on the second image sequence.

[0012] In a possible implementation, the at least two image acquisition sequences include at least one first image sequence; and the training data is determined based on a frame alignment relationship between the expression coefficient acquisition sequence and each of the first image sequences.

[0013] In one possible implementation, the at least one first image sequence is in a frame alignment state; the frame alignment relationship is obtained by performing frame alignment processing on the expression coefficient acquisition sequence and the target image sequence in the at least one first image sequence; the difference between the acquisition perspective of the acquisition device of the target image sequence and the frontal perspective is not greater than the difference between the acquisition perspective of the acquisition device of any other first image sequence in the at least one first image sequence except the target image sequence and the frontal perspective.

[0014] In one possible implementation, the process of determining the frame alignment relationship includes:

[0015] Determining a coefficient acquisition sequence of a target dimension from the expression coefficient acquisition sequence, and determining a coefficient extraction sequence of the target dimension from an expression coefficient extraction sequence corresponding to the target image sequence; the expression coefficient extraction sequence is obtained by performing expression coefficient extraction processing on the target image sequence;

[0016] The frame alignment relationship is determined based on frame misalignment representation data between the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension.

[0017] In one possible implementation, the process of determining the frame misalignment characterization data includes:

[0018] determining a first data signal and a second data signal according to a coefficient acquisition sequence of the target dimension and a coefficient extraction sequence of the target dimension;

[0019] Updating the second data signal according to the relative frame number distance between the signals; the initial value of the relative frame number distance between the signals is a preset distance;

[0020] performing similarity determination processing on the first data signal and the second data signal according to the frame number intersection of the first data signal and the second data signal to obtain a similarity corresponding to the frame number distance;

[0021] The relative frame number distance between the signals is updated, and the step of updating the second data signal according to the relative frame number distance between the signals is continued until a preset stop condition is reached, and the frame misalignment characterization data is determined according to the similarity.

[0022] In one possible implementation, the process of constructing the training data includes:

[0023] For any of the first image sequences, if the frame alignment relationship indicates that there is an alignment relationship between a first frame number in the first image sequence and a second frame number in the expression coefficient acquisition sequence, a sample image is determined based on the image corresponding to the first frame number in the first image sequence, and an expression coefficient label corresponding to the sample image is determined based on the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence;

[0024] The training data is constructed based on the sample images and the expression coefficient labels corresponding to the sample images.

[0025] In one possible implementation manner, constructing training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences includes:

[0026] Determining an expression coefficient label sequence corresponding to each of the image acquisition sequences from the expression coefficient acquisition sequence;

[0027] The training data is constructed based on the at least two image acquisition sequences and the expression coefficient label sequences corresponding to the at least two image acquisition sequences.

[0028] In one possible implementation, the at least two image acquisition sequences include a second image sequence and / or at least one first image sequence; the at least one first image sequence is in a frame-aligned state; the expression coefficient label sequences corresponding to the at least one first image sequence are the same; and the expression coefficient label sequence corresponding to the second image sequence is the expression coefficient acquisition sequence.

[0029] The present application provides a method for extracting expression coefficients, the method comprising:

[0030] Get facial image;

[0031] The facial image is subjected to expression coefficient extraction processing using an expression coefficient extraction model to obtain an expression coefficient extraction result; the training data of the expression coefficient extraction model is constructed using the training data construction method provided in this application.

[0032] This application provides a training data construction device, comprising:

[0033] A first acquisition unit is configured to acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; the expression coefficient acquisition sequence and each of the image acquisition sequences are acquired from the same face; the acquisition viewing angle of the acquisition device for the expression coefficient acquisition sequence is a frontal viewing angle of the face; and the acquisition viewing angles of the acquisition devices for different image acquisition sequences are different;

[0034] A data construction unit is used to construct training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences.

[0035] The present application provides an expression coefficient extraction device, comprising:

[0036] a second acquiring unit, configured to acquire a facial image;

[0037] The expression extraction unit is used to perform expression coefficient extraction processing on the facial image using an expression coefficient extraction model to obtain an expression coefficient extraction result; the training data of the expression coefficient extraction model is constructed using the training data construction method provided in this application.

[0038] The present application provides an electronic device, the device comprising: a processor and a memory;

[0039] The memory is used to store instructions or computer programs;

[0040] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the training data construction method or expression coefficient extraction method provided in this application.

[0041] The present application provides a computer-readable medium, which stores instructions or computer programs. When the instructions or computer programs are executed on a device, the device executes the training data construction method or expression coefficient extraction method provided in the present application.

[0042] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the training data construction method or expression coefficient extraction method provided by the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] FIG1 is a flow chart of a training data construction method provided in an embodiment of the present application;

[0045] FIG2 is a schematic diagram of an alignment process provided in an embodiment of the present application;

[0046] FIG3 is a schematic diagram of a signal state before alignment provided by an embodiment of the present application;

[0047] FIG4 is a schematic diagram of a signal state after alignment provided by an embodiment of the present application;

[0048] FIG5 is a flow chart of an expression coefficient extraction method provided in an embodiment of the present application;

[0049] FIG6 is a schematic diagram of the structure of a training data construction device provided in an embodiment of the present application;

[0050] FIG7 is a schematic structural diagram of an expression coefficient extraction device provided in an embodiment of the present application;

[0051] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] Research has found that for some related schemes for implementing the process of acquiring training data, a facial expression data acquisition device can be used to collect data from the object to be collected, and obtain a facial image and an expression coefficient corresponding to the facial image; then, based on the facial image and the expression coefficient corresponding to the facial image, training data can be constructed.

[0053] Research has also found that for the solution shown in the previous paragraph, because the facial expression data acquisition equipment collects inaccurate expression coefficients from some non-frontal perspectives, such as a wide profile or a bird's-eye view, the quality of the training data constructed based on these expression coefficients is reduced, resulting in poor expression coefficient extraction performance for the model derived from this training data. Furthermore, because the facial expression data acquisition equipment is typically used to collect a large amount of data from a frontal perspective, the resulting training data contains a large number of frontal images and a small number of, or even no, non-frontal images. This leads to an excessive imbalance in the ratio between frontal and non-frontal images in the training data. This results in poor expression coefficient extraction performance for the model derived from this training data from some non-frontal perspectives, such as a wide profile or a bird's-eye view.

[0054] Based on the above research, in order to better improve the performance of expression coefficient extraction, the present application provides a training data construction method, which includes: first obtaining an expression coefficient acquisition sequence and at least two image acquisition sequences; and then constructing training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences. Since the expression coefficient acquisition sequence and each image acquisition sequence are all acquired for the same face, the expression changes described by the expression coefficient acquisition sequence are consistent with the expression changes described by each image acquisition sequence. This allows the expression coefficient acquisition sequence to represent the expression coefficient corresponding to each frame image in each image acquisition sequence, thereby ensuring that the training data constructed based on the expression coefficient acquisition sequence and these image acquisition sequences has good quality. Furthermore, since the acquisition device for the expression coefficient acquisition sequence acquires the expression coefficient acquisition sequence from a frontal perspective of the face, the expression coefficient acquisition sequence can accurately describe the expression changes of the face, thereby making the expression coefficient guidance information in the training data constructed based on the expression coefficient acquisition sequence more accurate. This effectively avoids defects caused by inaccurate expression coefficient guidance information in the training data, thereby improving the quality of the training data. Because the acquisition perspectives of the acquisition devices of different image acquisition sequences are different, these image acquisition sequences can represent facial images from various perspectives with a similar number of views. This is conducive to improving the proportional balance between facial images from various perspectives in the training data, thereby effectively avoiding defects caused by excessive imbalance in the proportions between frontal face images and non-frontal face images in the training data, and thus helping to improve the quality of the training data. In this way, the model determined based on the training data has better expression coefficient extraction performance, which is conducive to improving the expression coefficient extraction effect.

[0055] In addition, the present application does not limit the execution subject of the training data construction method provided in the embodiment of the present application. For example, the training data construction method provided in the embodiment of the present application can be applied to a terminal device or a server. For another example, the training data construction method provided in the embodiment of the present application can also be implemented with the help of the data interaction process between the terminal device and the server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.

[0056] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0057] To better understand the technical solution provided by this application, the training data construction method provided by this application is described below with reference to some figures. As shown in Figure 1, the training data construction method provided by the embodiment of this application includes the following steps S101-S102. Figure 1 is a flow chart of a training data construction method provided by the embodiment of this application.

[0058] S101: Acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; the expression coefficient acquisition sequence and each image acquisition sequence are all acquired for the same face; the acquisition perspective of the acquisition device of the expression coefficient acquisition sequence is the frontal perspective of the face; the acquisition perspectives of the acquisition devices of different image acquisition sequences are different.

[0059] The expression coefficient acquisition sequence is obtained by performing expression coefficient acquisition processing on the face of the collected object, so that the expression coefficient acquisition sequence can represent the expression change state of the face.

[0060] In addition, for the expression coefficient acquisition sequence described above, the expression coefficient acquisition sequence can include multiple frames of expression coefficients, so that the expression coefficient acquisition sequence can describe the facial expression state at different acquisition times; and for each frame of expression coefficients in the expression coefficient acquisition sequence, the expression coefficient can be multi-dimensional, such as 52 dimensions, so that the expression coefficient can better describe the facial expression state. It should be noted that this application is not limited to the 52-dimensional implementation method.

[0061] In addition, for the above-mentioned expression coefficient acquisition sequence, the acquisition perspective of the acquisition device of the expression coefficient acquisition sequence is the frontal perspective of the above-mentioned face, so that the expression coefficient acquisition sequence is used to describe the frontal state of the face at different acquisition moments, thereby enabling the expression coefficient acquisition sequence to more accurately represent the expression state of the face at different acquisition moments, thereby making the expression coefficient guidance information in the training data constructed based on the expression coefficient acquisition sequence more accurate, which is conducive to improving the quality of the training data. It can be seen that the acquisition device of the expression coefficient acquisition sequence can be used to perform expression coefficient acquisition and processing on the face of the collected object from a frontal perspective to obtain the expression coefficient acquisition sequence. It should be noted that the present application does not limit the implementation method of the frontal perspective. For example, the frontal perspective can be implemented using a 90° or a perspective close to 90°.

[0062] Furthermore, the present application does not limit the implementation of the device for collecting the expression coefficient collection sequence described above. For example, it can be implemented using any existing or future device capable of collecting and processing expression coefficients, such as the facial expression collection device or mobile phone shown in FIG2 . It should be noted that the present application does not limit the implementation of the facial expression collection device. For example, the frame rate of the facial expression collection device is 30 frames.

[0063] For any one of the at least two image acquisition sequences mentioned above, the image acquisition sequence is obtained by performing image acquisition processing on the face of the captured object at a certain viewing angle, so that the image acquisition sequence can represent the state changes of the face at the viewing angle. In this way, the at least two image acquisition sequences can represent the state changes of the face at various viewing angles as much as possible, so that the training data constructed based on the at least two image acquisition sequences can include facial images at various viewing angles with similar data volumes, which is conducive to improving the proportional balance between the facial images at various viewing angles in the training data, thereby helping to improve the quality of the training data.

[0064] It should be noted that, for the expression coefficient acquisition sequence and the at least two image acquisition sequences described above, both the expression coefficient acquisition sequence and each image acquisition sequence are acquired from the same face. Furthermore, the acquisition time periods of the expression coefficient acquisition sequence and the acquisition time periods of each image acquisition sequence intersect, and the duration of this intersection is greater than a preset duration. Furthermore, the acquisition time periods of each image acquisition sequence intersect, and the duration of this intersection is greater than a preset duration.

[0065] In addition, for any one of the at least two image acquisition sequences mentioned above, the image acquisition sequence may include multiple frames of images, so that the image acquisition sequence can describe the state of the face at different acquisition moments from the same viewing angle.

[0066] In addition, the present application does not limit the implementation of the at least two image acquisition sequences mentioned above. For example, the at least two image acquisition sequences may include at least one first image sequence. The at least one first image sequence refers to images acquired by the same image acquisition device at at least one viewing angle. Moreover, the present application does not limit the implementation of the at least one first image sequence. For example, the at least one first image sequence may be acquired by a synchronized surround multi-view acquisition camera, so that the at least one first image sequence can capture the state of the face at multiple viewing angles. It can be seen that in one possible implementation, the acquisition perspectives of the acquisition devices of different first image sequences are different, so that each first image sequence is used to describe the state of the face at different viewing angles, so that all first image sequences can describe the changes in the state of the face as comprehensively as possible, which is conducive to improving the quality of the training data.

[0067] Furthermore, for the at least one first image sequence described above, each first image sequence corresponds to an image acquisition device, such that a many-to-one or one-to-one relationship exists between these first image sequences and the image acquisition devices. For example, for any first image sequence in the at least one first image sequence, such as sequence 1, if the first image sequence is acquired by image acquisition device 1, then another first image sequence, such as sequence 2, may be acquired by image acquisition device 2 or by image acquisition device 1.

[0068] It can be seen that in one possible implementation, for at least one first image sequence above, different first image sequences may be acquired by different devices, that is, different first image sequences correspond to different image acquisition devices, so that a one-to-one relationship can be achieved. Among them, the image acquisition device corresponding to the first image sequence refers to the acquisition device of the first image sequence. In addition, the image acquisition device corresponding to each first image sequence is different from the acquisition device of the expression coefficient acquisition sequence above, and the present application does not limit the implementation methods of these acquisition devices. For example, the image acquisition devices corresponding to different first image sequences can be implemented using cameras deployed at different viewing angles; and the acquisition device of the expression coefficient acquisition sequence can be implemented using a facial expression acquisition device deployed at a frontal viewing angle.

[0069] Furthermore, for the at least one first image sequence mentioned above, if each first image sequence corresponds to an image acquisition device, then in order to better improve the construction efficiency, each image acquisition device satisfies a synchronization condition so that the acquisition devices of the at least one first image sequence achieve frame synchronization or approximate frame synchronization, so that the at least one first image sequence acquired by these devices is in a frame-aligned state, thereby effectively overcoming the defects caused when image sequences from different perspectives are in a frame-misaligned state, such as the need to consume a lot of time and resources for frame alignment processing. This not only improves construction efficiency but also saves resources. Among them, the synchronization condition refers to the condition that needs to be met when achieving multi-device frame synchronization or approximate frame synchronization; and this application does not limit the synchronization condition. For example, when device 1 and device 2 both meet the synchronization condition, if the i-th frame image acquired by device 1 corresponds to the i-th frame image acquired by device 2, then the difference between the timestamp corresponding to the i-th frame image acquired by device 1 and the timestamp corresponding to the i-th frame image acquired by device 2 is lower than a preset difference threshold. Among them, the timestamp corresponding to an image is used to describe the acquisition time of the image. Wherein, i is a positive integer.

[0070] It can be seen that in a possible implementation manner, at least one of the first image sequences is in a frame alignment state, so that images corresponding to the same frame in different first image sequences can describe the facial states of each viewing angle at the same moment.

[0071] In fact, in some application scenarios, when the acquisition device of the above expression coefficient acquisition sequence is different from the acquisition device of each first image sequence, if it is difficult for these two devices to achieve frame synchronization, then the expression coefficient acquisition sequence and each first image sequence are not in a frame alignment state, which requires subsequent frame alignment processing for the expression coefficient acquisition sequence and each first image sequence.

[0072] It should be noted that the present application does not limit the implementation method of the acquisition device of at least one first image sequence mentioned above. For example, it may include acquisition cameras that can surround the captured object from various perspectives, and these acquisition cameras are strictly frame synchronized, and the frame rate is set to 30 frames.

[0073] In fact, in some application scenarios, in order to better increase the amount of training data, the present application also provides a possible implementation of the at least two image acquisition sequences mentioned above. In this implementation, the at least two image acquisition sequences may include a second image sequence and at least one first image sequence mentioned above, and the expression coefficient acquisition sequence mentioned above is determined based on the second image sequence, so that the expression coefficient acquisition sequence includes the expression coefficients of some or all images in the second image sequence. The second image sequence is used to record some images acquired from a frontal perspective, and the acquisition device of the second image sequence is the same as the acquisition device of the expression coefficient acquisition sequence mentioned above. It can be seen that the second image sequence refers to the image sequence acquired when the acquisition device of the expression coefficient acquisition sequence acquires expression coefficients, so that the expression coefficients of each frame in the expression coefficient acquisition sequence are respectively used to describe the facial expression state shown in the image corresponding to the corresponding frame in the second image sequence, so that the second image sequence and the expression coefficient acquisition sequence are in a frame-aligned state.

[0074] Based on the content of the previous paragraph, it can be seen that for the acquisition device of the above-mentioned expression coefficient acquisition sequence, such as the facial expression acquisition device shown in Figure 2, because the acquisition device can perform facial image acquisition during the expression coefficient acquisition process, the acquisition device can be used not only to provide expression coefficients, but also to provide facial images; and because the acquisition perspective of the acquisition device is a frontal perspective, the facial image acquired by the acquisition device is used to describe the state of the above-mentioned face under the frontal perspective. Therefore, in order to increase the amount of training data as much as possible, training data can be constructed based on the facial images acquired by the acquisition device and the expression coefficients corresponding to the facial images, so that the training data can describe the facial state from as many perspectives as possible, which is conducive to improving the quality of the training data.

[0075] Based on the relevant content of the at least two image acquisition sequences above, it can be known that in one possible implementation manner, the acquisition perspectives of the acquisition devices of different image acquisition sequences are different, such as the acquisition perspectives of the second image sequence above and the acquisition perspectives of the acquisition devices of each first image sequence are different, so that the at least two image acquisition sequences can describe the facial state from as many perspectives as possible, thereby making the training data constructed based on the at least two image acquisition sequences have higher quality.

[0076] Based on the relevant content of S101 above, it can be known that in some application scenarios, for the collected object, a facial expression collection device can be deployed at the frontal perspective of the face of the collected object, so that the facial expression collection device can collect the facial image at the frontal perspective and the expression coefficient corresponding to the facial image, so that the expression coefficient can accurately represent the expression state of the face; and, collection cameras can be deployed at various perspectives around the collected object, so that these collection cameras can collect facial images at various perspectives, so that training data can be constructed subsequently based on the expression coefficients collected by the facial expression collection device, the facial images collected by the facial expression collection device, and the facial images collected by these collection cameras.

[0077] S102: Constructing training data based on the expression coefficient acquisition sequence and at least two image acquisition sequences.

[0078] The training data refers to the data required for model updating.

[0079] In addition, this application does not limit the implementation of the training data described above. For example, the training data may include at least one binary data set. For any binary data set, the binary data set may include a sample image and an expression coefficient label corresponding to the sample image. The sample image is a facial image; and the expression coefficient label corresponding to the sample image is used to indicate guidance information about the expression coefficient of the sample image, so that the model can subsequently learn how to better extract the expression coefficient of the sample image based on the expression coefficient label.

[0080] In addition, the present application does not limit the implementation of the above S102. For ease of understanding, two possible implementations are described below.

[0081] In one possible implementation, if the above expression coefficient acquisition sequence and each image acquisition sequence are in a frame-aligned state, then the above S102 can specifically be: first, the j-th frame image in the i-th image acquisition sequence is used as the sample image, and the j-th frame expression coefficient in the expression coefficient acquisition sequence is used as the expression coefficient label corresponding to the sample image, j is a positive integer, j≤the total number of frames in a sequence, i is a positive integer, i≤the number of sequences in the above at least two image acquisition sequences; then, based on each sample image and the expression coefficient label corresponding to each sample image, the training data is determined so that the training data includes these sample images and their corresponding expression coefficient labels.

[0082] In another possible implementation, when the at least two image acquisition sequences above include at least one first image sequence, and for any first image sequence, the first image sequence is not in a frame-aligned state with the expression coefficient acquisition sequence above, S102 above may specifically be: constructing training data based on the frame alignment relationship between the expression coefficient acquisition sequence and each first image sequence. The frame alignment relationship between the expression coefficient acquisition sequence and the kth first image sequence is used to indicate which frame of expression coefficients in the expression coefficient acquisition sequence is aligned with which frame of image in the kth first image sequence. k is a positive integer, and k≤the number of sequences in the at least one first image sequence.

[0083] In addition, this application does not limit the method for obtaining the frame alignment relationship in the above paragraph. For example, it can be implemented with the help of a pre-built frame alignment model. The frame alignment model is used to perform frame alignment processing on the input data of the frame alignment model; and this application does not limit the implementation method of the frame alignment model.

[0084] In addition, to better improve efficiency, the present application provides a possible implementation method of the above-mentioned frame alignment relationship determination process. Under this implementation method, when the above-mentioned at least one first image sequence is in a frame alignment state, the frame alignment relationship can be obtained by performing frame alignment processing on the above-mentioned expression coefficient acquisition sequence and the target image sequence in the at least one first image sequence. The target image sequence refers to a sequence in the at least one first image sequence whose acquisition perspective is closest to the frontal perspective, so that the target image sequence can describe the expression state of the face of the captured object as accurately as possible when each frame image in these first image sequences is acquired, which is conducive to improving the alignment accuracy. It can be seen that the difference between the acquisition perspective of the acquisition device of the target image sequence and the above-mentioned frontal perspective is not greater than the difference between the acquisition perspective of the acquisition device of any other first image sequence in the at least one first image sequence except the target image sequence and the frontal perspective, so that the target image sequence can represent the sequence in the at least one first image sequence whose acquisition perspective is closest to the frontal perspective.

[0085] In addition, the present application does not limit the implementation method of the frame alignment processing in the above paragraph. For example, it can be implemented using a pre-built frame alignment model.

[0086] In fact, for each frame's expression coefficient, since the expression coefficient is multidimensional, and different dimensions have different degrees of obviousness in expression changes, in order to better improve the alignment effect, the alignment process can be performed with the help of a dimension with a relatively high degree of obviousness, such as a left blink or a right blink. Based on this, the present application also provides a possible implementation method of the above-mentioned frame alignment relationship determination process. Under this implementation method, when at least one of the above-mentioned first image sequences is in a frame alignment state, and the acquisition perspective of the target image sequence in the at least one first image sequence is closest to the frontal perspective, the frame alignment relationship determination process includes the following steps 11 to 13.

[0087] Step 11: Determine the coefficient acquisition sequence of the target dimension from the expression coefficient acquisition sequence.

[0088] The target dimension refers to a dimension in the expression coefficients of each frame that changes significantly with the expression; this application does not limit the target dimension; for example, the target dimension can be implemented using a left blink or a right blink. Because left blinks or right blinks are more obvious, the recognition accuracy of the expression coefficients in these dimensions can be guaranteed to a certain extent in non-normal viewing angles, such as at various profile angles, effectively ensuring the accuracy of determining frame alignment relationships.

[0089] In addition, for the above-mentioned target dimension, the coefficient acquisition sequence of the target dimension refers to the data extracted from the above-mentioned expression coefficient acquisition sequence, which is used to describe the changes under the target dimension; and the coefficient acquisition sequence can include the numerical values ​​of the expression coefficients of each frame in the expression coefficient acquisition sequence under the target dimension.

[0090] Step 12: Determine a coefficient extraction sequence of a target dimension from the expression coefficient extraction sequence corresponding to the target image sequence; the expression coefficient extraction sequence is obtained by performing expression coefficient extraction processing on the target image sequence.

[0091] The expression coefficient extraction sequence corresponding to the target image sequence is obtained by performing expression coefficient extraction processing on the target image sequence, so that the expression coefficient extraction sequence can represent the expression state shown in each frame image in the target image sequence; and the present application does not limit the process of determining the expression coefficient extraction sequence. For example, the expression coefficient extraction sequence can be obtained by performing expression coefficient extraction processing on each frame image in the target image sequence using an existing expression coefficient extraction model. The expression coefficient extraction model refers to a model with expression coefficient extraction capabilities; and the expression coefficient extraction model needs to be optimized based on the training data constructed in this application.

[0092] The coefficient extraction sequence of the target dimension refers to the data extracted from the expression coefficient extraction sequence corresponding to the target image sequence above, which is used to describe the changes in the target dimension; and the coefficient extraction sequence can include the numerical values ​​of the expression coefficients of each frame in the expression coefficient extraction sequence under the target dimension.

[0093] It should be noted that this application does not limit the relationship between the execution time of step 12 above and the execution time of step 11 above. For example, the former may be earlier than the latter, or the latter may be earlier than the former, or both may be the same.

[0094] Step 13: Determine the frame alignment relationship based on the frame misalignment representation data between the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension.

[0095] Among them, the frame misalignment characterization data is used to represent the frame misalignment between the coefficient acquisition sequence of the target dimension above and the coefficient extraction sequence of the target dimension, so that the frame misalignment characterization data can represent the number of frames staggered between the coefficient acquisition sequence and the coefficient extraction sequence. It should be noted that frame misalignment is used to describe the temporal misalignment of the two sequences. For example, for acquisition device 1 and acquisition device 2, if the acquisition device 1 acquires the expression coefficient of the i-th frame in the expression coefficient sequence at the t-th moment, but the acquisition moment corresponding to the j-th frame image in the image sequence acquired by the acquisition device 2 is the moment t or the moment closest to the moment t, then the difference between i and j is the frame misalignment, and the absolute value of ij is the number of staggered frames. Among them, i is a positive integer, and j is a positive integer.

[0096] In addition, the present application does not limit the implementation method of the above frame misalignment characterization data. For example, the frame misalignment characterization data may include the misalignment between the coefficients of each frame in the coefficient acquisition sequence of the target dimension and the coefficients of the corresponding frame in the coefficient extraction sequence of the target dimension, so that frame alignment processing can be performed based on these misalignments. For another example, the frame misalignment characterization data may include the misalignment between the first frame in the coefficient acquisition sequence and the first frame in the coefficient extraction sequence, so that the coefficient acquisition sequence and the coefficient extraction sequence can be frame aligned to a certain extent based on the misalignment between the two first frames and the curves corresponding to the two sequences, so that the curves corresponding to the two aligned sequences are completely or approximately aligned, such as two frames with corresponding relationships in the two curves are located at the same coordinate on the horizontal axis or the coordinates are closest, which is conducive to improving the frame alignment effect.

[0097] In addition, this application does not limit the process of determining the frame misalignment characterization data described above. For example, it can be implemented using a pre-built frame misalignment recognition model. The frame misalignment recognition model is used to perform frame misalignment recognition processing on the input data of the frame misalignment recognition model; and this application does not limit the implementation method of the frame misalignment recognition model.

[0098] Furthermore, in order to better improve the alignment effect, the present application also provides a possible implementation of the above-mentioned process of determining the frame misalignment characterization data, which may specifically include the following steps 131 to 134.

[0099] Step 131: Determine a first data signal and a second data signal according to the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension.

[0100] Among them, the first data signal refers to a data signal that does not need to be updated, such as the signal represented by the blue curve shown in Figure 3 or Figure 4; and the present application does not limit the implementation method of the first data signal. For example, the first data signal can be implemented using a data sequence or a data vector.

[0101] In addition, the present application does not limit the implementation method of the above-mentioned first data signal. For example, the first data signal can be determined based on the coefficient acquisition sequence of the above-mentioned target dimension, so that the signal values ​​corresponding to each frame number in the first data signal respectively represent the coefficients under the corresponding frame number in the coefficient acquisition sequence of the target dimension, so that the first data signal can represent the data changes described by the coefficient acquisition sequence of the target dimension. For another example, the first data signal can be determined based on the coefficient extraction sequence of the target dimension, so that the signal values ​​corresponding to each frame number in the first data signal respectively represent the coefficients under the corresponding frame number in the coefficient extraction sequence of the target dimension, so that the first data signal can represent the data changes described by the coefficient extraction sequence of the target dimension.

[0102] The second data signal refers to a data signal that needs to be updated, such as the signal represented by the yellow curve shown in Figure 3 or Figure 4; and the present application does not limit the implementation method of the second data signal. For example, the second data signal can be implemented using a data sequence or a data vector.

[0103] In addition, the present application does not limit the implementation method of the second data signal. For example, if the first data signal is determined based on the coefficient acquisition sequence of the target dimension, the second data signal is determined based on the coefficient extraction sequence of the target dimension. For another example, if the first data signal is determined based on the coefficient extraction sequence of the target dimension, the second data signal is determined based on the coefficient acquisition sequence of the target dimension.

[0104] Step 132: Update the second data signal according to the relative frame distance between signals; the initial value of the relative frame distance between signals is a preset distance.

[0105] The relative frame distance between signals is used to represent the relative distance between the next signal in the current round and another signal, so that the relative frame distance between signals can represent the relative distance between the starting frame of one signal and the starting frame of another signal; and the initial value of the relative frame distance between signals is a preset distance. The preset distance can be pre-set, for example, the preset distance can be 0 or 1 frame.

[0106] Based on the relevant content of step 132 above, it can be known that for the current round, the second data signal is updated according to the relative frame distance between signals, so that the relative distance between the updated second data signal and the above first data signal reaches the relative frame distance between the signals, such as the position of the updated second data signal in the coordinate system is moved to the right by the relative frame distance between the signals relative to the position of the first data signal in the coordinate system.

[0107] Step 133: performing similarity determination processing on the first data signal and the second data signal based on the frame number intersection of the first data signal and the second data signal to obtain a similarity corresponding to a relative frame number distance between the signals.

[0108] Among them, the frame number intersection refers to the frame number range where the overlapping part between the first data signal and the second data signal is located, so that the first data signal has a signal value on each frame number in the frame number intersection, and the second data signal also has a signal value on each frame number in the frame number intersection.

[0109] In addition, for the similarity corresponding to the relative frame number distance between the above signals, the similarity is used to describe the similarity of the overlapping part between the above first data signal and the above second data signal, so that the similarity can represent the similarity between the two signals when the first data signal and the second data signal are aligned according to the relative frame number distance between the signals, thereby enabling the similarity to represent the possibility of aligning the first data signal and the second data signal according to the relative frame number distance between the signals.

[0110] In addition, the present application does not limit the implementation method of the above step 133. For example, it can be specifically as follows: first, the signal value corresponding to the above frame number intersection is extracted from the above first data signal, and the signal value corresponding to the frame number intersection is extracted from the above second data signal; then, the two extracted signal values ​​are subjected to similarity calculation processing, such as the cosine similarity calculation processing shown in formula (1) below, to obtain the similarity corresponding to the relative frame number distance between the signals, so that the similarity can represent the possibility of aligning the first data signal and the second data signal according to the relative frame number distance between the signals.

[0111] A signal;q represents the qth component in vector A; B q Represents the qth component in vector B; q is a positive integer, q≤Q, Q is a positive integer, and Q represents the number of components in a vector.

[0112] Step 134: Update the relative frame number distance between the signals, and return to continue executing the above step 132 and subsequent steps until the preset stop condition is reached, and determine the frame misalignment characterization data based on the above similarity.

[0113] It should be noted that the present application does not limit the implementation method of the above “updating the relative frame number distance between signals”. For example, it can be implemented using the following formula (2). D'=D+1 (2)

[0114] Where D' represents the relative frame distance between signals after the update; D represents the relative frame distance between signals before the update.

[0115] In addition, the preset stopping condition mentioned above refers to the condition that must be met when the iterative process stops, and this application does not limit the preset stopping condition. For example, the preset stopping condition may specifically be: the relative frame distance between the above signals is greater than two-thirds of the length of the above first data signal, and / or the relative frame distance between the above signals is greater than two-thirds of the length of the above second data signal. For another example, the preset stopping condition may specifically be: the length of the intersection of the above frames is less than one-third of the length of the first data signal, and / or the length of the intersection of the above frames is less than one-third of the length of the second data signal.

[0116] In addition, the present application does not limit the implementation method of the step of "determining the frame misalignment characterization data based on the above similarity" in the above step 134. For example, it can be specifically: after obtaining the similarity corresponding to the relative frame number distance between each signal, select the relative frame number distance between the signals with the highest similarity, and determine the relative frame number distance between the signals with the highest similarity as the above frame misalignment characterization data, so that the alignment effect between the above first data signal and the above second data signal achieved based on the frame misalignment characterization data is optimized, so that the frame numbers corresponding to the two data with an alignment relationship in the frame alignment relationship determined based on the frame misalignment characterization data differ by the relative frame number distance between the signals.

[0117] Based on the relevant contents of steps 131 to 134 above, it can be known that for some application scenarios, after obtaining the coefficient acquisition sequence of the target dimension above and the coefficient extraction sequence of the target dimension, the first data signal and the second data signal are first determined based on the two sequences; the first data signal is then fixedly processed; then, the second data signal is moved to the right by one frame each time, and the cosine similarity of the overlapping part between the two signals is calculated, so that the moving distance with the highest cosine similarity can be obtained, and the relative frame number distance between the signals is used as the frame misalignment representation data above, so that the alignment effect achieved based on the frame misalignment representation data is optimized, so that the expression coefficient acquisition sequence can be aligned with each first image sequence based on the frame misalignment representation data to obtain a frame alignment relationship, so that the frame alignment relationship can indicate which frame in the expression coefficient acquisition sequence has an alignment relationship with which frame in each first image sequence.

[0118] Based on the relevant contents of steps 11 to 13 above, it can be known that in one possible implementation, when at least one of the above first image sequences is in a frame alignment state, it can be determined that the acquisition time corresponding to the same frame image in these first image sequences is the same. Therefore, in order to improve efficiency, a one-time alignment process can be used to determine the frame alignment relationship between each first image sequence and the above expression coefficient acquisition sequence. This can effectively reduce the number of alignments, thereby effectively improving efficiency. Among them, for any first image sequence, the closer the acquisition perspective of the first image sequence is to the frontal perspective, the more accurate the facial expression described by the first image sequence. Therefore, in order to better improve the accuracy, this alignment process can be implemented based on the expression coefficient acquisition sequence and the first image sequence whose acquisition perspective is closest to the frontal perspective. This is conducive to improving the alignment accuracy, thereby improving the quality of the training data. In addition, for any frame expression coefficient, the expression coefficient is multidimensional, and because the coefficients of different dimensions show different degrees of change with the change of expression, the coefficients of different dimensions have different effects on the frame alignment process. Therefore, in order to better improve the alignment effect, only the dimensions with a relatively high degree of change, such as the left blink or right blink, can be used for alignment processing. This can effectively improve the alignment effect, such as alignment efficiency and alignment accuracy, etc., which is conducive to improving the quality of training data. In addition, the present application can achieve alignment processing through a step-by-step moving method, such as the moving method shown in Figures 3 and 4, which is conducive to accurately determining the number of frames staggered between the two sequences, thereby helping to improve alignment accuracy, and then helping to improve the quality of training data.

[0119] In addition, in order to better improve the quality of training data, the present application also provides a possible implementation method of the training data construction process. Under this implementation method, when the above at least two image acquisition sequences include at least one first image sequence, and each first image sequence is not in a frame alignment state with the above expression coefficient acquisition sequence, the training data construction process may include the following steps 21-22.

[0120] Step 21: For any first image sequence, after obtaining the frame alignment relationship between the above expression coefficient acquisition sequence and the first image sequence, if the frame alignment relationship indicates that there is an alignment relationship between the first frame number in the first image sequence and the second frame number in the expression coefficient acquisition sequence, then determine the sample image based on the image corresponding to the first frame number in the first image sequence, and determine the expression coefficient label corresponding to the sample image based on the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence.

[0121] In the present application, for the kth first image sequence, after obtaining the frame alignment relationship between the above-mentioned expression coefficient acquisition sequence and the kth first image sequence, if the frame alignment relationship indicates that there is an alignment relationship between the first frame number in the kth first image sequence and the second frame number in the expression coefficient acquisition sequence, it can be determined that the acquisition time represented by the first frame number in the kth first image sequence and the acquisition time represented by the second frame number in the expression coefficient acquisition sequence are the same time. Therefore, it can be determined that the image corresponding to the first frame number in the kth first image sequence and the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence were acquired at the same time. Therefore, a sample image can be determined based on the image corresponding to the first frame number in the kth first image sequence, and the expression coefficient label corresponding to the sample image can be determined based on the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence. The first frame number is used to represent an arrangement position in the kth first image sequence. The second frame number is used to represent an arrangement position in the expression coefficient acquisition sequence. k is a positive integer, k≤the number of sequences in the at least one first image sequence above.

[0122] Step 22: Construct training data based on the above sample image and the expression coefficient label corresponding to the sample image.

[0123] In the present application, after obtaining the above sample image and the expression coefficient label corresponding to the sample image, the binary data of (the sample image, the expression coefficient label corresponding to the sample image) can be constructed first; and then the training data can be determined based on the binary data so that the training data includes the binary data.

[0124] Based on the relevant contents of steps 21 to 22 above, it can be known that for some application scenarios, when the at least two image acquisition sequences above include at least one first image sequence, and each first image sequence is not in a frame alignment state with the expression coefficient acquisition sequence above, the frame alignment relationship between the expression coefficient acquisition sequence and each first image sequence can be determined first, so that the frame alignment relationship can indicate which frame in the expression coefficient acquisition sequence has an alignment relationship with which frame in each first image sequence; then, a pair of data is constructed based on each alignment relationship described by the frame alignment relationship; and then, based on these pair of data, the training data is determined, so that the training data includes these pair of data, so that the model can be updated based on the training data in the future.

[0125] In addition, the present application does not limit the implementation of the above S102. For example, in order to better improve the quality of training data, the S102 may specifically include the following steps 31 and 32.

[0126] Step 31: Determine the expression coefficient label sequence corresponding to each image acquisition sequence from the above expression coefficient acquisition sequence.

[0127] The expression coefficient label sequence corresponding to the i-th image acquisition sequence refers to the expression coefficient guidance information required for reference when constructing training data based on the i-th image acquisition sequence. i is a positive integer, i≤the number of sequences in the at least two image acquisition sequences above.

[0128] In addition, the present application does not limit the implementation method of the expression coefficient label sequence corresponding to the above i-th image acquisition sequence. For example, if the i-th image acquisition sequence is the above second image sequence, then because the i-th image acquisition sequence and the above expression coefficient acquisition sequence are different data acquired by the same device, the i-th image acquisition sequence and the expression coefficient acquisition sequence are in a frame alignment state, so that the expression coefficient acquisition sequence can accurately describe the changes in facial expressions described by the second image sequence. Therefore, the expression coefficient acquisition sequence can be directly determined as the expression coefficient label sequence corresponding to the i-th image acquisition sequence.

[0129] For another example, if the i-th image acquisition sequence above is any first image sequence, and the i-th image acquisition sequence is not in a frame-aligned state with the above expression coefficient acquisition sequence, then the expression coefficient label sequence corresponding to the i-th image acquisition sequence can be determined based on the frame alignment relationship between the expression coefficient acquisition sequence and the i-th image acquisition sequence, so that the expression changes described by the expression coefficient label sequence are consistent with the facial expression changes described by the i-th image acquisition sequence.

[0130] Based on the content of the previous paragraph, it can be known that in one possible implementation, for the at least two image acquisition sequences above, when the at least two image acquisition sequences include at least one first image sequence, if the at least one first image sequence is in a frame alignment state, it can be determined that the frame alignment relationship between each first image sequence and the above expression coefficient acquisition sequence is the same, so that it can be determined that the expression coefficient label sequence corresponding to the at least one first image sequence is the same, that is, the same frame image in these first image sequences corresponds to the same expression coefficient label.

[0131] Based on the relevant content of step 31 above, it can be known that for some application scenarios, after obtaining the above expression coefficient acquisition sequence and the above at least two image acquisition sequences, the expression coefficient label sequence corresponding to each image acquisition sequence can be determined from the expression coefficient acquisition sequence, so that the expression coefficient label sequence corresponding to each image acquisition sequence can accurately represent the facial expression state described by the corresponding image acquisition sequence, which is conducive to improving the accuracy of the expression coefficient label, thereby helping to improve the quality of the training data.

[0132] Step 32: Construct training data based on the above at least two image acquisition sequences and the expression coefficient label sequences corresponding to the at least two image acquisition sequences.

[0133] It should be noted that the present application does not limit the implementation method of the above step 32. For example, it can be specifically as follows: first, based on the t-th frame image in the i-th image acquisition sequence and the t-th frame expression coefficient in the expression coefficient label sequence corresponding to the i-th image acquisition sequence, construct a binary data, i is a positive integer, i≤the number of sequences in the above at least two image acquisition sequences, t is a positive integer, t≤the number of frames in a sequence; then, based on these binary data, construct training data so that the training data includes these binary data.

[0134] Based on the relevant contents of steps 31 to 32 above, it can be known that for some application scenarios, after obtaining at least two image acquisition sequences above, the expression coefficient label sequence corresponding to each image acquisition sequence is first determined from the expression coefficient acquisition sequence above, so that the images corresponding to the same acquisition moment in the image acquisition sequence have the same expression coefficient label, and the expression coefficient label is determined based on the expression coefficient recorded in the expression coefficient acquisition sequence at the corresponding acquisition moment; and then based on these two image acquisition sequences and their corresponding expression coefficient label sequences, training data is constructed so that the training data includes some binary data similar to (image, expression coefficient label corresponding to the image).

[0135] It can be seen that because the expression coefficient label sequences corresponding to the at least two image acquisition sequences described above can relatively accurately represent the changes in facial expressions described by these image acquisition sequences, the expression coefficient guidance information in the training data constructed based on the expression coefficient label sequences is relatively accurate. This can effectively avoid defects caused by inaccurate expression coefficient guidance information in the training data, such as poor expression coefficient extraction performance of the model. Furthermore, because the at least two image acquisition sequences can describe facial states from as many perspectives as possible, the training data constructed based on these image acquisition sequences includes facial images from as many perspectives as possible. This can effectively avoid defects caused by facial images in the training data covering fewer perspectives, such as poor expression coefficient extraction performance of the model for facial images from certain perspectives. Because different image acquisition sequences describe different acquisition perspectives, and the number of image frames in different image acquisition sequences is almost the same, the training data constructed based on these image acquisition sequences includes facial images from various perspectives with a similar number. This is conducive to improving the proportional balance between facial images from various perspectives in the training data, thereby effectively avoiding defects caused by excessive imbalance in the proportion between frontal face images and non-frontal face images in the training data, such as the model can only present better expression coefficient extraction performance for facial images from a certain perspective.

[0136] Based on the relevant contents of S101 to S102 above, it can be seen that for the training data construction method provided in the embodiment of the present application, an expression coefficient acquisition sequence and at least two image acquisition sequences are first obtained; then, training data is constructed based on the expression coefficient acquisition sequence and the at least two image acquisition sequences. Since the expression coefficient acquisition sequence and each image acquisition sequence are all acquired for the same face, the expression changes described by the expression coefficient acquisition sequence are consistent with the expression changes described by each image acquisition sequence, so that the expression coefficient acquisition sequence can represent the expression coefficient corresponding to each frame image in each image acquisition sequence, thereby making the training data constructed based on the expression coefficient acquisition sequence and these image acquisition sequences have better quality. Furthermore, since the acquisition perspective of the acquisition device of the expression coefficient acquisition sequence is a frontal perspective of the face, the expression coefficient acquisition sequence can accurately describe the expression changes of the face, thereby making the expression coefficient guidance information in the training data constructed based on the expression coefficient acquisition sequence more accurate. This can effectively avoid defects caused by inaccurate expression coefficient guidance information in the training data, thereby helping to improve the quality of the training data. Because the acquisition perspectives of the acquisition devices of different image acquisition sequences are different, these image acquisition sequences can represent facial images from various perspectives with a similar number of views. This is conducive to improving the proportional balance between facial images from various perspectives in the training data, thereby effectively avoiding defects caused by excessive imbalance in the proportions between frontal face images and non-frontal face images in the training data, and thus helping to improve the quality of the training data. In this way, the model determined based on the training data has better expression coefficient extraction performance, which is conducive to improving the expression coefficient extraction effect.

[0137] Based on the training data construction method provided in the embodiment of the present application, the embodiment of the present application further provides an expression coefficient extraction method, as shown in Figure 5, which includes the following S501-S502. Figure 5 is a flowchart of an expression coefficient extraction method provided in the embodiment of the present application.

[0138] S501: Acquire a facial image.

[0139] The facial image refers to an image that needs to be processed for expression coefficient extraction; and this application does not limit the method for obtaining the facial image.

[0140] S502: Perform expression coefficient extraction processing on the facial image using the expression coefficient extraction model to obtain an expression coefficient extraction result; the training data of the expression coefficient extraction model is constructed using any implementation of the training data construction method provided in this application.

[0141] The expression coefficient extraction model is used to perform expression coefficient extraction processing on input data of the expression coefficient extraction model.

[0142] Furthermore, the expression coefficient extraction model described above is obtained by updating the expression coefficient extraction model using the training data, enabling the expression coefficient extraction model to learn from the training data how to better extract expression coefficients. Because the training data is constructed using any embodiment of the training data construction method provided in this application, the training data is of high quality. Furthermore, because the expression coefficient guidance information in the training data is relatively accurate, the expression coefficient extraction model updated based on the training data exhibits better expression coefficient extraction performance. Furthermore, because the training data includes facial images from various perspectives, the expression coefficient extraction model updated based on the training data is suitable for extracting expression coefficients from facial images from various perspectives, thereby improving the application range of the expression coefficient extraction model. Furthermore, because the facial images from various perspectives in the training data are relatively well-balanced in proportion, the expression coefficient extraction model updated based on the training data exhibits better expression coefficient extraction performance for facial images from various perspectives, thereby improving the expression coefficient extraction effect of the expression coefficient extraction model.

[0143] The above expression coefficient extraction result refers to the result obtained by performing expression coefficient extraction processing on the facial image by the above expression coefficient extraction model, so that the expression coefficient extraction result can represent the expression state shown in the facial image.

[0144] Based on the relevant contents of S501 to S502 above, it can be seen that for the expression coefficient extraction method provided in the embodiments of the present application, after obtaining a facial image, the expression coefficient extraction model is used to perform expression coefficient extraction processing on the facial image to obtain an expression coefficient extraction result. In particular, because the training data of the expression coefficient extraction model is constructed using any embodiment of the training data construction method provided in the present application, the expression coefficient extraction model updated based on the training data has better expression coefficient extraction performance, and thus the expression coefficient extraction result determined based on the expression coefficient extraction model can more accurately represent the facial expression state shown in the facial image, which is conducive to improving the expression coefficient extraction effect.

[0145] In addition, the present application does not limit the execution entity of the expression coefficient extraction method provided in the embodiments of the present application. For example, the expression coefficient extraction method provided in the embodiments of the present application can be applied to a terminal device or a server. For another example, the expression coefficient extraction method provided in the embodiments of the present application can also be implemented through the data exchange process between the terminal device and the server.

[0146] Based on the training data construction method provided in the embodiments of the present application, the embodiments of the present application also provide a training data construction device, which will be explained and illustrated below in conjunction with Figure 6. Figure 6 is a schematic diagram of the structure of a training data construction device provided in the embodiments of the present application. It should be noted that for the technical details of the training data construction device provided in the embodiments of the present application, please refer to the relevant content of the training data construction method above.

[0147] As shown in FIG6 , the training data construction apparatus 600 provided in an embodiment of the present application includes:

[0148] The first acquisition unit 601 is configured to acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; the expression coefficient acquisition sequence and each of the image acquisition sequences are acquired from the same face; the acquisition viewing angle of the acquisition device for the expression coefficient acquisition sequence is a frontal viewing angle of the face; and the acquisition viewing angles of the acquisition devices for different image acquisition sequences are different;

[0149] The data construction unit 602 is configured to construct training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences.

[0150] In a possible implementation, the at least two image acquisition sequences include at least one first image sequence; each of the first image sequences corresponds to an image acquisition device; and each of the image acquisition devices meets a synchronization condition.

[0151] In a possible implementation manner, the at least two image acquisition sequences include a second image sequence; and the expression coefficient acquisition sequence is determined based on the second image sequence.

[0152] In a possible implementation, the at least two image acquisition sequences include at least one first image sequence; and the training data is determined based on a frame alignment relationship between the expression coefficient acquisition sequence and each of the first image sequences.

[0153] In one possible implementation, the at least one first image sequence is in a frame alignment state; the frame alignment relationship is obtained by performing frame alignment processing on the expression coefficient acquisition sequence and the target image sequence in the at least one first image sequence; the difference between the acquisition perspective of the acquisition device of the target image sequence and the frontal perspective is not greater than the difference between the acquisition perspective of the acquisition device of any other first image sequence in the at least one first image sequence except the target image sequence and the frontal perspective.

[0154] In one possible implementation, the process of determining the frame alignment relationship includes: determining the coefficient acquisition sequence of the target dimension from the expression coefficient acquisition sequence, and determining the coefficient extraction sequence of the target dimension from the expression coefficient extraction sequence corresponding to the target image sequence; the expression coefficient extraction sequence is obtained by performing expression coefficient extraction processing on the target image sequence; and determining the frame alignment relationship based on the frame misalignment representation data between the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension.

[0155] In one possible implementation, the process of determining the frame misalignment characterization data includes: determining the first data signal and the second data signal based on the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension; updating the second data signal based on the relative frame number distance between the signals; the initial value of the relative frame number distance between the signals is a preset distance; performing similarity determination processing on the first data signal and the second data signal based on the frame number intersection of the first data signal and the second data signal to obtain the similarity corresponding to the frame number distance; updating the relative frame number distance between the signals, and continuing to execute the step of updating the second data signal based on the relative frame number distance between the signals, until the preset stop condition is reached, and the frame misalignment characterization data is determined based on the similarity.

[0156] In one possible implementation, the process of constructing the training data includes: for any of the first image sequences, if the frame alignment relationship indicates that there is an alignment relationship between the first frame number in the first image sequence and the second frame number in the expression coefficient acquisition sequence, then determining a sample image based on the image corresponding to the first frame number in the first image sequence, and determining an expression coefficient label corresponding to the sample image based on the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence; and constructing the training data based on the sample image and the expression coefficient label corresponding to the sample image.

[0157] In one possible implementation, the data construction unit 602 is specifically used to: determine the expression coefficient label sequence corresponding to each of the image acquisition sequences from the expression coefficient acquisition sequence; and construct the training data based on the at least two image acquisition sequences and the expression coefficient label sequence corresponding to the at least two image acquisition sequences.

[0158] In one possible implementation, the at least two image acquisition sequences include a second image sequence and / or at least one first image sequence; the at least one first image sequence is in a frame-aligned state; the expression coefficient label sequences corresponding to the at least one first image sequence are the same; and the expression coefficient label sequence corresponding to the second image sequence is the expression coefficient acquisition sequence.

[0159] Based on the relevant content of the above-mentioned training data construction device 600, it can be seen that for the training data construction device 600 provided in the embodiment of the present application, an expression coefficient acquisition sequence and at least two image acquisition sequences are first obtained; then, training data is constructed based on the expression coefficient acquisition sequence and the at least two image acquisition sequences. Among them, because the expression coefficient acquisition sequence and each image acquisition sequence are all acquired for the same face, the expression changes described by the expression coefficient acquisition sequence are consistent with the expression changes described by each image acquisition sequence, so that the expression coefficient acquisition sequence can represent the expression coefficient corresponding to each frame image in each image acquisition sequence, thereby making the training data constructed based on the expression coefficient acquisition sequence and these image acquisition sequences have better quality. In addition, because the acquisition perspective of the acquisition device of the expression coefficient acquisition sequence is a frontal perspective of the face, the expression coefficient acquisition sequence can accurately describe the expression changes of the face, thereby making the expression coefficient guidance information in the training data constructed based on the expression coefficient acquisition sequence more accurate. This can effectively avoid defects caused by inaccurate expression coefficient guidance information in the training data, thereby helping to improve the quality of the training data. Because the acquisition perspectives of the acquisition devices of different image acquisition sequences are different, these image acquisition sequences can represent facial images from various perspectives with a similar number of views. This is conducive to improving the proportional balance between facial images from various perspectives in the training data, thereby effectively avoiding defects caused by excessive imbalance in the proportions between frontal face images and non-frontal face images in the training data, and thus helping to improve the quality of the training data. In this way, the model determined based on the training data has better expression coefficient extraction performance, which is conducive to improving the expression coefficient extraction effect.

[0160] Based on the expression coefficient extraction method provided in the embodiments of this application, the embodiments of this application also provide an expression coefficient extraction device, which is explained and illustrated below in conjunction with Figure 7. Figure 7 is a schematic diagram of the structure of the expression coefficient extraction device provided in the embodiments of this application. It should be noted that for technical details of the expression coefficient extraction device provided in the embodiments of this application, please refer to the relevant content of the expression coefficient extraction method above.

[0161] As shown in FIG7 , the expression coefficient extraction device 700 provided in an embodiment of the present application includes:

[0162] A second acquiring unit 701 is configured to acquire a facial image;

[0163] The expression extraction unit 702 is used to perform expression coefficient extraction processing on the facial image using an expression coefficient extraction model to obtain an expression coefficient extraction result; the training data of the expression coefficient extraction model is constructed using the training data construction method described in any one of claims 1-10.

[0164] Based on the above description of the expression coefficient extraction device 700, it can be seen that after acquiring a facial image, the expression coefficient extraction device 700 uses an expression coefficient extraction model to perform expression coefficient extraction processing on the facial image to obtain an expression coefficient extraction result. In particular, because the training data for the expression coefficient extraction model is constructed using any embodiment of the training data construction method provided in this application, the expression coefficient extraction model updated based on the training data has better expression coefficient extraction performance. Therefore, the expression coefficient extraction result determined based on the expression coefficient extraction model can more accurately represent the facial expression state shown in the facial image, thereby improving the expression coefficient extraction effect.

[0165] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the training data construction method provided in the embodiment of the present application, or executes any implementation of the expression coefficient extraction method provided in the embodiment of the present application.

[0166] Referring to FIG8 , a schematic diagram of the structure of an electronic device 800 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG8 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.

[0167] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0168] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0169] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0170] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0171] An embodiment of the present application also provides a computer-readable medium, which stores instructions or computer programs. When the instructions or computer programs are executed on a device, the device executes any implementation of the training data construction method provided in the embodiment of the present application, or executes any implementation of the expression coefficient extraction method provided in the embodiment of the present application.

[0172] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0173] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0174] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0175] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.

[0176] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0177] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0178] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.

[0179] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0180] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.

[0182] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0183] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0184] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0185] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing training data, comprising: Acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; The expression coefficient acquisition sequence and each of the image acquisition sequences are all acquired from the same face; The acquisition angle of the acquisition device of the expression coefficient acquisition sequence is the frontal angle of the face; The acquisition viewing angles of the acquisition devices of different image acquisition sequences are different; Training data is constructed based on the expression coefficient acquisition sequence and the at least two image acquisition sequences.

2. The method according to claim 1, wherein the at least two image acquisition sequences include at least one first image sequence; Each of the first image sequences corresponds to an image acquisition device; Each of the image acquisition devices meets synchronization conditions.

3. The method of claim 1 , wherein the at least two image acquisition sequences include a second image sequence; The expression coefficient acquisition sequence is determined based on the second image sequence.

4. The method according to claim 1, wherein the at least two image acquisition sequences include at least one first image sequence; The training data is determined based on a frame alignment relationship between the expression coefficient acquisition sequence and each of the first image sequences.

5. The method according to claim 4, wherein the at least one first image sequence is in a frame-aligned state; The frame alignment relationship is obtained by performing frame alignment processing on the expression coefficient acquisition sequence and the target image sequence in the at least one first image sequence; The difference between the capture angle of view of the capture device of the target image sequence and the frontal perspective is no greater than the difference between the capture angle of view of the capture device of any other first image sequence in the at least one first image sequence except the target image sequence and the frontal perspective.

6. The method according to claim 5, wherein the process of determining the frame alignment relationship comprises: Determining a coefficient acquisition sequence of a target dimension from the expression coefficient acquisition sequence, and determining a coefficient extraction sequence of the target dimension from an expression coefficient extraction sequence corresponding to the target image sequence; the expression coefficient extraction sequence is obtained by performing expression coefficient extraction processing on the target image sequence; The frame alignment relationship is determined based on frame misalignment representation data between the coefficient acquisition sequence of the target dimension and the coefficient extraction sequence of the target dimension.

7. The method according to claim 6, wherein the process of determining the frame misalignment characterization data comprises: determining a first data signal and a second data signal according to a coefficient acquisition sequence of the target dimension and a coefficient extraction sequence of the target dimension; Updating the second data signal according to the relative frame number distance between the signals; the initial value of the relative frame number distance between the signals is a preset distance; performing similarity determination processing on the first data signal and the second data signal according to the frame number intersection of the first data signal and the second data signal to obtain a similarity corresponding to the frame number distance; The relative frame number distance between the signals is updated, and the step of updating the second data signal according to the relative frame number distance between the signals is continued until a preset stop condition is reached, and the frame misalignment characterization data is determined according to the similarity.

8. The method according to claim 4, wherein the process of constructing the training data comprises: For any of the first image sequences, if the frame alignment relationship indicates that there is an alignment relationship between a first frame number in the first image sequence and a second frame number in the expression coefficient acquisition sequence, a sample image is determined based on the image corresponding to the first frame number in the first image sequence, and an expression coefficient label corresponding to the sample image is determined based on the expression coefficient corresponding to the second frame number in the expression coefficient acquisition sequence; The training data is constructed based on the sample images and the expression coefficient labels corresponding to the sample images.

9. The method according to claim 1, wherein constructing training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences comprises: Determining, from the expression coefficient acquisition sequence, an expression coefficient label sequence corresponding to each of the image acquisition sequences; The training data is constructed based on the at least two image acquisition sequences and the expression coefficient label sequences corresponding to the at least two image acquisition sequences.

10. The method according to claim 9, wherein the at least two image acquisition sequences include a second image sequence and / or at least one first image sequence; The at least one first image sequence is in a frame alignment state; the expression coefficient label sequences corresponding to the at least one first image sequence are the same; The expression coefficient label sequence corresponding to the second image sequence is the expression coefficient acquisition sequence.

11. A method for extracting expression coefficients, comprising: Get facial image; Performing expression coefficient extraction processing on the facial image using an expression coefficient extraction model to obtain an expression coefficient extraction result; The training data of the expression coefficient extraction model is constructed using the training data construction method described in any one of claims 1-10.

12. A training data construction device, comprising: A first acquisition unit is used to acquire an expression coefficient acquisition sequence and at least two image acquisition sequences; The expression coefficient acquisition sequence and each of the image acquisition sequences are all acquired from the same face; the acquisition viewing angle of the acquisition device of the expression coefficient acquisition sequence is the frontal viewing angle of the face; The acquisition viewing angles of the acquisition devices of different image acquisition sequences are different; A data construction unit is used to construct training data based on the expression coefficient acquisition sequence and the at least two image acquisition sequences.

13. An expression coefficient extraction device, comprising: a second acquiring unit, configured to acquire a facial image; An expression extraction unit is used to perform expression coefficient extraction processing on the facial image using an expression coefficient extraction model to obtain an expression coefficient extraction result; the training data of the expression coefficient extraction model is constructed using the training data construction method described in any one of claims 1-10.

14. An electronic device comprising: processor and memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 11.

15. A computer-readable medium storing instructions or a computer program, wherein when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 11.

16. A computer program product, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Three-dimensional face reconstruction system based on deep learning

    CN113971715A

  • Multi-view teaching expression recognition method and system in classroom environment

    CN114120403A

  • Expression tracking method and device, equipment and storage medium

    CN116228808A

Cited By

  • Multi-modal large model-based cue word automatic generation model training method and system

    CN121600511A