Method and apparatus for determining expression coefficient, device, medium, and product
By acquiring face images and reference images, the problem of low accuracy of expression coefficients in virtual reality scenes is solved, and the image sequence of the same object is used to improve the adaptability and accuracy of expression coefficients.
Patent Information
- Application Number
- PCT/CN2024/140098
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-19
- Filing Date
- 2024-12-17
- Publication Date
- 2025-08-28
AI Technical Summary
In virtual reality scenarios, single-frame images carry less effective information, resulting in low accuracy of expression coefficients and it is difficult to accurately represent the expression state of the object.
By obtaining face images and reference images, both are used to describe the same object and determine from the same image sequence. The reference images are used to supplement face characteristics and improve the accuracy of expression coefficients.
It improves the accuracy of the expression coefficient, avoids ambiguity problems caused by different objects having different facial features, and enhances the adaptability and accuracy of the expression coefficient.
Smart Images

Figure CN2024140098_28082025_PF_FP_ABST
Abstract
Description
A method, device, equipment, medium, and product for determining expression coefficient
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed on February 19, 2024, with application number 202410185900.1 and invention name “A method, device, equipment, medium, and product for determining an expression coefficient”. The entire contents of that application are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of image processing technology, and in particular to a method, device, equipment, medium, and product for determining an expression coefficient. Background Art
[0004] Some application scenarios, such as expression capture, require the following: After acquiring a facial image, it is necessary to extract expression coefficients from the facial image so that other tasks, such as facial reconstruction, can be performed based on the expression coefficients. The expression coefficients are used to represent the facial expression state described by the facial image. Summary of the Invention
[0005] The present application provides a method, apparatus, device, medium, and product for determining an expression coefficient, which are conducive to improving the accuracy of the expression coefficient.
[0006] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0007] The present application provides a method for determining an expression coefficient, the method comprising:
[0008] Acquire a facial image and a reference image corresponding to the facial image; the facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence;
[0009] Determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image.
[0010] In a possible implementation, the reference image is used to represent the expressionless state of the subject.
[0011] In one possible implementation, the expression coefficient determination method is applied to a model update scenario; the image sequence is determined based on a sample video;
[0012] The process of determining the reference image includes:
[0013] The reference image is determined from the image sequence based on the expression coefficient labels corresponding to each frame image in the image sequence, and the expression coefficient label corresponding to the reference image is not greater than the expression coefficient label corresponding to any other image in the image sequence except the reference image.
[0014] In one possible implementation, the expression coefficient determination method is used to determine a predicted expression coefficient corresponding to a current frame image in an image stream;
[0015] The image stream includes at least one frame of image to be selected; the acquisition time of the image to be selected is earlier than the acquisition time of the current frame of image, and the predicted expression coefficient corresponding to the image to be selected is lower than a preset coefficient threshold;
[0016] The number of image frames in the at least one frame to be selected is not less than a preset frame number threshold;
[0017] The process of determining the reference image includes:
[0018] Based on the predicted expression coefficient corresponding to each of the candidate images, the reference image is determined from the at least one frame of candidate images, and the predicted expression coefficient corresponding to the reference image is not greater than the predicted expression coefficient corresponding to any other candidate image in the at least one frame of candidate images except the reference image.
[0019] In a possible implementation manner, the facial image is a current frame image in an image stream;
[0020] The number of frames of the image to be selected in the image stream is lower than a preset frame number threshold; the acquisition time of the image to be selected is earlier than the acquisition time of the current frame image, and the predicted expression coefficient corresponding to the image to be selected is lower than a preset coefficient threshold;
[0021] The reference image is determined based on the current frame image.
[0022] In a possible implementation manner, the facial image is a current frame image in an image stream;
[0023] If the acquisition time of the reference image is earlier than the acquisition time of the current frame image, then after determining the predicted expression coefficient corresponding to the facial image, the method further includes:
[0024] If the predicted expression coefficient corresponding to the facial image is lower than the predicted expression coefficient corresponding to the reference image, the reference image is updated using the facial image.
[0025] In one possible implementation, the predicted expression coefficient is determined using an expression coefficient determination model;
[0026] The method further comprises:
[0027] Determining loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image;
[0028] The expression coefficient determination model is updated according to the loss representation data of the expression coefficient determination model.
[0029] In a possible implementation manner, there are multiple facial images;
[0030] The process of determining the loss representation data of the expression coefficient determination model includes:
[0031] For any of the facial images, determining an expression distribution range to which the facial image belongs from at least two expression distribution ranges based on a predicted expression coefficient corresponding to the facial image and an expression coefficient label corresponding to the facial image;
[0032] Determining loss representation data for the at least two expression distribution ranges based on the predicted expression coefficients corresponding to the plurality of facial images, the expression coefficient labels corresponding to the plurality of facial images, and the expression distribution ranges to which the plurality of facial images belong;
[0033] The loss characterization data of the expression coefficient determination model is determined based on the loss characterization data of the at least two expression distribution ranges.
[0034] In a possible implementation manner, the plurality of facial images include at least one image to be used, and the expression distribution range to which each of the images to be used belongs is a target range among the at least two expression distribution ranges;
[0035] The process of determining the loss characterization data of the target range includes:
[0036] For any of the images to be used, determining loss representation data of the image to be used according to a difference between a predicted expression coefficient corresponding to the image to be used and an expression coefficient label corresponding to the image to be used;
[0037] The loss characterization data of the target range is determined according to an average value of the loss characterization data of the at least one image to be used.
[0038] In one possible implementation, the image sequence includes an expressionless image, and the expressionless image is used to represent an expressionless state of the subject;
[0039] The method further comprises:
[0040] determining an image distinction prediction result based on the image features of the facial image and the image features of the reference image, wherein the image distinction prediction result is used to indicate whether the facial image and the reference image are the same image;
[0041] The step of determining loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image includes:
[0042] Determine the loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image, the expression coefficient label corresponding to the facial image, the image distinction prediction result, and the image distinction label; the image distinction label is determined based on the selection process of the reference image.
[0043] In one possible implementation, the reference image is determined using a process randomly selected from at least one candidate process;
[0044] The random selection process is: determining the expressionless image as the reference image, or determining the facial image as the reference image.
[0045] In one possible implementation, the process of determining the predicted expression coefficient includes: performing feature fusion processing on the facial image and the reference image to obtain a fusion feature; and performing expression coefficient regression processing based on the fusion feature to obtain the predicted expression coefficient corresponding to the facial image.
[0046] In one possible implementation, the expression coefficient determination method is applied to a virtual reality device; the virtual reality device is used to collect at least two image streams;
[0047] The acquiring of the facial image comprises:
[0048] For any of the image streams, determining a current frame image in the image stream as the facial image;
[0049] The method further comprises:
[0050] The current frame expression coefficient of the virtual reality device is determined according to the predicted expression coefficient corresponding to the current frame image in the at least two image streams.
[0051] The present application provides a device for determining an expression coefficient, comprising:
[0052] An image acquisition unit, configured to acquire a facial image and a reference image corresponding to the facial image; the facial image and the reference image are both used to describe the same object; and the facial image and the reference image are both determined from the same image sequence;
[0053] A coefficient determination unit is used to determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image.
[0054] The present application provides an electronic device, the device comprising: a processor and a memory;
[0055] The memory is used to store instructions or computer programs;
[0056] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the expression coefficient determination method provided in this application.
[0057] The present application provides a computer-readable medium having instructions or a computer program stored therein. When the instructions or the computer program are executed on a device, the device executes the expression coefficient determination method provided in the present application.
[0058] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the expression coefficient determination method provided by the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0060] FIG1 is a flow chart of a method for determining an expression coefficient provided in an embodiment of the present application;
[0061] FIG2 is a schematic diagram of an image processing process provided in an embodiment of the present application;
[0062] FIG3 is a schematic diagram of another image processing process provided in an embodiment of the present application;
[0063] FIG4 is a schematic structural diagram of an expression coefficient determination device provided in an embodiment of the present application;
[0064] FIG5 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] Research has found that in some application scenarios, such as virtual reality (VR) scenarios, the above expression coefficient acquisition scheme can be specifically as follows: after obtaining the current frame image in the image stream, the expression coefficient extraction process is directly performed on the current frame image to obtain the expression coefficient of the current frame image.
[0066] Research has also found that the acquisition scheme shown in the previous paragraph has the following defects: in VR scenarios, because each frame of the image may carry relatively little effective information, a single frame of the image cannot accurately represent the facial expression state of the captured object, resulting in a relatively low accuracy of the expression coefficient determined based on the acquisition scheme.
[0067] Based on the above research, in order to better improve the accuracy of the expression coefficient, the present application provides a method for determining an expression coefficient, which includes: first obtaining a facial image and a reference image corresponding to the facial image; then, based on the facial image and the reference image, determining a predicted expression coefficient corresponding to the facial image, so that the predicted expression coefficient can represent the expression state presented by the facial image. In this case, because the reference image and the facial image are both used to describe the same object, the facial features described by the reference image and the facial features described by the facial image are both facial features possessed by the object, so that the reference image can supplement some facial features for the facial image, and then, when performing expression coefficient determination processing on the facial image based on the reference image, as many facial features as possible can be used, so that the predicted expression coefficient determined for the facial image can more accurately represent the expression state presented by the facial image, thereby helping to improve the accuracy of the expression coefficient. Also, because the facial image and the reference image are both determined from the same image sequence, the acquisition configuration of the facial image is consistent with the acquisition configuration of the reference image, so that the facial features described by the reference image are more adaptable to the facial features described by the facial image, and thus the reference image can provide as much positive influence as possible for the expression coefficient prediction process of the facial image, which is conducive to improving the accuracy of the expression coefficient.
[0068] Compared with the related art, this application has at least the following advantages:
[0069] In the technical solution provided by the present application, a facial image and a reference image corresponding to the facial image are first obtained; then, based on the facial image and the reference image, a predicted expression coefficient corresponding to the facial image is determined, so that the predicted expression coefficient can represent the expression state presented by the facial image. Since the reference image and the facial image are both used to describe the same object, the facial features described by the reference image and the facial features described by the facial image are both facial features possessed by the object, thereby enabling the reference image to supplement some facial features for the facial image, and thus enabling as many facial features as possible to be used when determining the expression coefficient of the facial image based on the reference image. This allows the predicted expression coefficient determined for the facial image to more accurately represent the expression state presented by the facial image, thereby facilitating improved accuracy of the expression coefficient. Also, because the facial image and the reference image are both determined from the same image sequence, the acquisition configuration of the facial image is consistent with the acquisition configuration of the reference image, so that the facial features described by the reference image are more adaptable to the facial features described by the facial image, and thus the reference image can provide as much positive influence as possible for the expression coefficient prediction process of the facial image, which is conducive to improving the accuracy of the expression coefficient.
[0070] In addition, for the reference image corresponding to the above facial image, the reference image can be used to represent the expressionless state of the object described by the facial image, so that the reference image can better express the facial features of the object, such as the facial features in the expressionless state, etc., so that when determining the predicted expression coefficient corresponding to the above facial image based on the reference image, the facial features can be referred to, and then the predicted expression coefficient finally determined for the facial image can more accurately express the expression state of the object presented in the facial image, thereby effectively avoiding defects caused by different objects having different facial features, such as ambiguity problems, etc., which is conducive to improving the accuracy of the expression coefficient.
[0071] In addition, the present application does not limit the execution entity of the expression coefficient determination method provided in the embodiments of the present application. For example, the expression coefficient determination method provided in the embodiments of the present application can be applied to a terminal device or a server. For another example, the expression coefficient determination method provided in the embodiments of the present application can also be implemented through the data exchange process between a terminal device and a server. The terminal device can be a smartphone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.
[0072] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0073] To better understand the technical solutions provided by this application, the following describes the method for determining an expression coefficient provided by this application in conjunction with some accompanying figures. As shown in Figure 1 , the method for determining an expression coefficient provided by this embodiment of the application includes steps S1-S2 below. Figure 1 is a flow chart of a method for determining an expression coefficient provided by this embodiment of the application.
[0074] S1: Acquire a facial image and a reference image corresponding to the facial image; the facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence.
[0075] The facial image refers to an image that needs to be processed for expression coefficient determination. For example, the facial image can be implemented using the image 1 shown in FIG. 2 or FIG. 3 .
[0076] In addition, this application is not limited to the above implementation method of facial images. For better understanding, the following description is combined with two scenarios.
[0077] In scenario 1, when the expression coefficient determination method provided in this application is applied to a model update scenario, the facial image mentioned above may refer to a sample image in the training data. The training data refers to the data required for use in the model update process; this application does not limit the training data. For example, the training data may include the sample image and the expression coefficient label corresponding to the sample image, so that subsequent model update processing can be performed based on the sample image and the expression coefficient label corresponding to the sample image.
[0078] Regarding the sample images mentioned above, the sample images refer to images involved in the model update process that require expression coefficient determination processing. Furthermore, this application does not limit the method for obtaining the sample images. For example, when the expression coefficient determination method provided in this application is applied to a model update scenario of an expression coefficient determination model, the sample images can be extracted from a sample video so that the sample images can represent a frame of the sample video. The expression coefficient determination model refers to a model whose expression coefficient determination capabilities need to be further optimized through a model update process. The sample video refers to a video required for use in the model update process of the expression coefficient determination model. Furthermore, the sample video is used to describe the changes in facial expressions of an object over a period of time. Furthermore, this application does not limit the sample videos. For example, the sample videos can be implemented using videos captured by VR devices. It should be noted that this application does not limit the implementation of the object. For example, it can be implemented using any existing or future object with a face, such as a digital human or animal.
[0079] For the expression coefficient label corresponding to the sample image above, the expression coefficient label refers to the information involved in the model update process and used to guide the model update process; and the expression coefficient label is used to describe the expression state actually presented in the sample image. In addition, the present application does not limit the method for obtaining the expression coefficient label. For example, it can be implemented by manual annotation. For example, in order to better reduce the difficulty of obtaining training data, when the sample video and the reference video above are collected in the same time period, for the same object, and using different devices, and the sample video and the reference video are aligned in time, if the sample image is a frame image extracted from the sample video, then the process of obtaining the expression coefficient label corresponding to the sample image can be: first, the frame number corresponding to the sample image in the sample video is used as the target frame number; then, the image corresponding to the target frame number is extracted from the reference video as the label image, so that the label image can accurately represent the expression state shown in the sample image; then, the label image is subjected to expression coefficient extraction processing to obtain the expression coefficient label corresponding to the sample image. The reference video is used to provide expression coefficient guidance information for each frame of the sample video; and the present application does not limit the reference video. For example, the reference video is obtained by using a preset image acquisition device to capture the image of the object from a frontal angle, so that each frame of the reference video can accurately represent the expression state of the object at different acquisition times. In addition, the present application does not limit the preset image acquisition device. For example, it can be implemented using a terminal device such as a mobile phone. In addition, the present application does not limit the implementation method of the expression coefficient extraction process. For example, it can be implemented with the help of a pre-constructed model that can perform expression coefficient extraction processing on a single frame image, such as a single-frame facial drive model.
[0080] Based on the above three paragraphs, it can be seen that when the expression coefficient determination method provided in this application is applied to the model update scenario of the expression coefficient determination model, after obtaining the sample video and the reference video corresponding to the sample video, the sample video and the reference video can be first subjected to image extraction processing according to the same number of frames; then the image extracted from the sample video is used as a sample image, and the expression coefficient extraction result of the image extracted from the reference video is used as the expression coefficient label corresponding to the sample image, so that the model update processing can be performed based on the sample image and the expression coefficient label corresponding to the sample image.
[0081] Based on the relevant content of scenario one above, it can be seen that in one possible implementation, when the expression coefficient determination method provided in this application is applied to the model update scenario of the expression coefficient determination model, the facial image above can refer to any frame image extracted from the sample video; and the expression coefficient label corresponding to the facial image can refer to the expression coefficient extraction result of the image that exists in the reference video corresponding to the sample video and is temporally aligned with the facial image.
[0082] Scenario 2: When the expression coefficient determination method provided in this application is used to determine the predicted expression coefficient corresponding to the current frame image in an image stream, the facial image mentioned above may refer to the current frame image in the image stream. The image stream refers to a sequence of images captured in real time by a certain image acquisition device, such as a sequence of images captured in real time by a VR device. The current frame image refers to the latest frame image captured in the image stream, so that the current frame image can be used to represent the image captured in real time by the image acquisition device.
[0083] For the facial image mentioned above, the reference image corresponding to the facial image refers to the image required for reference when performing expression coefficient determination processing on the facial image, such as image 2 shown in Figure 2 or Figure 3, so that the reference image can provide some reference-worthy information for the expression coefficient determination process of the facial image, such as information such as facial features in a neutral state.
[0084] In addition, in order to better improve the reference value of the reference image corresponding to the facial image above, the reference image and the facial image can satisfy the following constraints: the facial image and the reference image are both used to describe the same object, and the facial image and the reference image are both determined from the same image sequence. It should be noted that in the implementation method of the image sequence of the present application, for example, when the expression coefficient determination method provided in the present application is applied to the model update scenario of the expression coefficient determination model, the image sequence is determined based on the sample video, so that the image sequence includes part or all of the images in the sample video. For example, when the expression coefficient determination method provided in the present application is used to determine the predicted expression coefficient corresponding to the current frame image in the image stream, the image sequence refers to the image stream.
[0085] Furthermore, this application does not limit the implementation of the above paragraph regarding "the image sequence is determined based on the sample video." For example, the sample video may be directly used as the image sequence, such that the image sequence includes each frame of the sample video. Another example may be extracting a number of images from the sample video as the image sequence, such that the image sequence includes these extracted images.
[0086] Research has found that there are differences in facial features between different objects, and this difference may cause ambiguity problems. For example, object 1 in a certain facial state is expressionless, but object 2 in the same facial state is expressive, etc., which affects the accuracy of the expression coefficient.
[0087] Based on the content of the previous paragraph, it can be seen that in order to better improve the accuracy, the present application also provides a possible implementation method of the reference image corresponding to the above facial image. Under this implementation method, the reference image is used to represent the expressionless state of the object described by the facial image, so that the reference image can express the facial features of the object to a certain extent, so that when the expression coefficient of the facial image is determined based on the reference image, it is possible to effectively avoid defects caused by different objects having different facial features, such as ambiguity problems, which is conducive to improving the accuracy of the expression coefficient.
[0088] In addition, this application does not limit the method for obtaining the reference image in the above paragraph. For example, the reference image can be manually specified by relevant personnel. For another example, to further improve flexibility, the reference image can be automatically obtained. For ease of understanding, the following two scenarios are combined for explanation.
[0089] Scenario 1. When the expression coefficient determination method provided in the present application is applied to a model update scenario, and the above facial image and the reference image corresponding to the facial image are both determined from the same image sequence, if the image sequence is determined based on a sample video, then the reference image determination process may include: determining a reference image from the image sequence based on the expression coefficient label corresponding to each frame image in the image sequence, the expression coefficient label corresponding to the reference image being no greater than the expression coefficient label corresponding to any other image in the image sequence except the reference image, so that the reference image can represent the image in the image sequence that is closest to the expressionless state.
[0090] It can be seen that in one possible implementation, if the above facial image and the reference image corresponding to the facial image are both obtained from a sample video, the process of determining the reference image may include: determining the reference image from the sample video based on the expression coefficient label corresponding to each frame image in the sample video, so that the expression coefficient label corresponding to the reference image is not greater than the expression coefficient label corresponding to any other image in the sample video except the reference image, so that the reference image can represent the image with the smallest expression coefficient label in the sample video, and further, the reference image can represent the image in the sample video that is closest to the expressionless state, so that the reference image can express the expressionless state of the object described by the facial image as accurately as possible, so as to realize the automatic determination of the reference image corresponding to each sample image in the model update scenario.
[0091] In addition, the present application does not limit the process of determining the reference image in the above paragraph. For example, when the expression coefficient label corresponding to a frame image is multi-dimensional, such as 52 dimensions, the process of determining the reference image can be specifically as follows: first, all dimensions of the expression coefficient label corresponding to the i-th frame image in the above sample video are summed up to obtain the coefficient and value corresponding to the i-th frame image, where i is a positive integer and i≤the total number of frames of images in the sample video; then, based on the coefficients and values corresponding to each frame image in the sample video, the image with the smallest coefficient and value is determined from the sample video as the reference image, so that the reference image can represent the image with the smallest coefficient and value existing in the sample video, thereby enabling the reference image to represent the image in the sample video that is closest to the expressionless state.
[0092] Scenario 2: When the expression coefficient determination method provided in this application is used to determine the predicted expression coefficient corresponding to the current frame image in the image stream, the process of determining the reference image corresponding to the above facial image may include the following steps 11-12.
[0093] Step 11: When the above facial image is the current frame image in the above image stream, if the frame number of the to-be-selected image in the image stream is lower than the preset frame number threshold, the current frame image is directly determined as the reference image corresponding to the facial image; the acquisition time of the to-be-selected image is earlier than the acquisition time of the current frame image, and the predicted expression coefficient corresponding to the to-be-selected image is lower than the preset coefficient threshold.
[0094] The candidate image refers to an image in the above image stream that is relatively close to a neutral state, so that the candidate image can represent an image that has been collected in a historical time period and is relatively close to a neutral state.
[0095] In addition, for the above-mentioned image to be selected, the image to be selected can meet the following constraints: the acquisition time of the image to be selected is earlier than the acquisition time of the current frame image in the above-mentioned image stream, and the predicted expression coefficient corresponding to the image to be selected is lower than the preset coefficient threshold. The preset coefficient threshold refers to a pre-set threshold value required for reference when screening images close to a neutral state. It should be noted that the present application does not limit the method for determining the conclusion that "the predicted expression coefficient corresponding to the image to be selected is lower than the preset coefficient threshold". For example, when the predicted expression coefficient is multi-dimensional, such as 52 dimensions, the process for determining the conclusion can be: if the sum of all dimensions of the predicted expression coefficient corresponding to the image to be selected is less than the preset coefficient threshold, then it can be determined that the predicted expression coefficient corresponding to the image to be selected is lower than the preset coefficient threshold. For another example, when the predicted expression coefficient is multi-dimensional, the process for determining the conclusion can be: if each dimension of the predicted expression coefficient corresponding to the image to be selected is less than the preset coefficient threshold, then it can be determined that the predicted expression coefficient corresponding to the image to be selected is lower than the preset coefficient threshold.
[0096] The preset frame number threshold refers to a predetermined minimum number of frames of the candidate image required to determine the expression coefficient with reference to the neutral state; and the preset frame number threshold can be set in advance based on the actual application scenario. For example, the preset frame number threshold can be implemented using K. K is a positive integer.
[0097] Based on the relevant content of step 11 above, it can be known that for the image stream collected in real time, the current frame image in the image stream can be regarded as the facial image mentioned above; and if the frame number of the selected image in the image stream is lower than the preset frame number threshold, it can be determined that the image stream is in the initial collection stage, so that it can be determined that the information carried by the image stream is relatively small, and then it can be determined that the expressionless state cannot be accurately analyzed from the image stream. Therefore, in order to avoid the impact caused by selecting the wrong expressionless state image, the current frame image in the image stream can be directly determined as the reference image corresponding to the current frame image, so that the expression coefficient can be determined based on the single frame image in the initial collection stage of the image stream, which is conducive to improving the accuracy of the expression coefficient.
[0098] Step 12: When the above facial image is the current frame image in the above image stream, if the image stream includes at least one frame of candidate image, and the number of image frames in the at least one frame of candidate image is not less than a preset frame number threshold, then based on the predicted expression coefficient corresponding to each candidate image, a reference image corresponding to the facial image is determined from the at least one frame of candidate image, and the predicted expression coefficient corresponding to the reference image is not greater than the predicted expression coefficient corresponding to any other candidate image in the at least one frame of candidate image except the reference image.
[0099] The predicted expression coefficient corresponding to the jth candidate image refers to the result obtained by performing expression coefficient determination processing on the jth candidate image, so that the predicted expression coefficient corresponding to the jth candidate image can represent the expression state predicted to be presented by the jth candidate image. It should be noted that the process of determining the predicted expression coefficient corresponding to the jth candidate image is similar to the process of determining the "predicted expression coefficient corresponding to the facial image" involved in this application. For the sake of brevity, it will not be repeated here. j is a positive integer, j≤the number of image frames in the at least one frame of the candidate image mentioned above.
[0100] In addition, the present application does not limit the implementation method of the above step 12. For example, when the predicted expression coefficients corresponding to each candidate image are multi-dimensional, such as 52 dimensions, the step 12 can be specifically as follows: first, all dimensions of the predicted expression coefficients corresponding to the jth candidate image are summed up to obtain the coefficient sum value corresponding to the jth candidate image, where j is a positive integer, and j≤the number of image frames in the at least one frame of candidate images above; then, based on the coefficients and values corresponding to each candidate image, the image with the smallest coefficient sum value is determined from the at least one frame of candidate images as the reference image corresponding to the above facial image, so that the reference image can represent the candidate image in the above image stream that is closest to the expressionless state, so that the reference image can represent the image that has been collected and is closest to the expressionless state as of the time of collection of the current frame image in the image stream.
[0101] Based on the above three paragraphs, it can be seen that for the image stream collected in real time, the current frame image in the image stream can be regarded as the facial image mentioned above; and if the frame number of the candidate image in the image stream is higher than or equal to the preset frame number threshold, it can be determined that the image stream is in the later acquisition stage, so that it can be determined that the information carried by the image stream is relatively sufficient, and then it can be determined that the expressionless state can be analyzed more accurately from the image stream. Therefore, in order to improve the accuracy of the expression coefficient, the candidate image in the image stream that is closest to the expressionless state can be determined as the reference image corresponding to the current frame image, so that in the later acquisition stage of the image stream, the expression coefficient can be determined based on multiple frame images. This can effectively overcome the adverse effects caused by the less information carried by a single frame image, thereby helping to improve the accuracy of the expression coefficient.
[0102] Based on the relevant contents of steps 11 to 12 above, it can be known that when the expression coefficient determination method provided in the present application is used to determine the predicted expression coefficient corresponding to the current frame image in the image stream, and the facial image above represents the current frame image in the image stream, if there is no to-be-selected image in the image stream, the reference image corresponding to the facial image can be directly implemented using the facial image; if the image stream includes at least one to-be-selected image, and the number of image frames in the at least one to-be-selected image is lower than the preset frame number threshold, the reference image corresponding to the facial image can be directly implemented using the facial image; if the image stream includes at least one to-be-selected image, and the number of image frames in the at least one to-be-selected image is equal to or higher than the preset frame number threshold, the reference image corresponding to the facial image can be implemented using the to-be-selected image with the smallest predicted expression coefficient that exists in the image stream. In this way, different reference image determination schemes can be adopted for different stages of the image stream, so that each stage of the image stream can achieve relatively accurate predicted expression coefficient determination processing, thereby helping to improve the determination effect of the expression coefficient. In addition, because the reference image can represent the candidate image in the image stream that is closest to the expressionless state, the reference image can be dynamically adjusted as the image stream is continuously updated, so that the reference image corresponding to each frame image in the image stream can provide as much supplementary information on facial features as possible, which is conducive to improving the accuracy of the expression coefficient.
[0103] In addition, the present application does not limit the implementation method of S1 above. For example, when the expression coefficient determination method provided in the present application is applied to the model update scenario of the expression coefficient determination model, the S1 can specifically be: randomly extracting an image from the sample video as a facial image, and selecting an image with the smallest expression coefficient label from the sample video as the reference image corresponding to the facial image, so that the expression coefficient determination model can subsequently determine the predicted expression coefficient corresponding to the facial image based on the facial image and the reference image.
[0104] For another example, when the expression coefficient determination method provided in the present application is used to determine the predicted expression coefficient corresponding to the current frame image in an image stream, the above S1 can specifically be: determining the current frame image in the image stream as a facial image, and if the number of frames of the to-be-selected image in the image stream is lower than the preset frame number threshold, then determining the current frame image as the reference image corresponding to the facial image; if the number of frames of the to-be-selected image in the image stream is equal to or higher than the preset frame number threshold, then determining the to-be-selected image in the image stream with the smallest predicted expression coefficient as the reference image corresponding to the facial image.
[0105] S2: Determine a predicted expression coefficient corresponding to the facial image based on the facial image and a reference image corresponding to the facial image.
[0106] The predicted expression coefficient corresponding to the facial image is used to describe the expression state predicted to be presented in the facial image. For example, the predicted expression coefficient corresponding to the facial image can be implemented using the predicted expression coefficient shown in Figure 2 or Figure 3.
[0107] In addition, the present application does not limit the implementation method of the above S2. For example, the S2 may specifically include the following steps 21 and 22.
[0108] Step 21: Perform feature fusion processing on the above facial image and the reference image corresponding to the facial image to obtain fused features.
[0109] The fusion feature refers to a feature fusion processing result between the facial image and the reference image corresponding to the facial image, so that the fusion feature can represent the image information carried by the facial image and the reference image.
[0110] In addition, the present application does not limit the implementation of the above step 21. For example, it may specifically include the following steps 211 to 213.
[0111] Step 211: performing image feature extraction processing on the facial image to obtain image features of the facial image, so that the image features can represent the image information carried by the facial image.
[0112] The image features of a facial image refer to the image feature extraction and processing results for the facial image, so that the image features can represent the image information carried by the facial image.
[0113] In addition, the present application does not limit the implementation of the above step 211. For example, it can be implemented using any existing or future image feature extraction method.
[0114] In addition, to further improve the feature extraction effect, this application also provides a possible implementation of step 211 above. Under this implementation, step 211 can specifically be: inputting the facial image above into a pre-constructed feature extraction network, so that the feature extraction network performs image feature extraction processing on the facial image, and obtains and outputs the image features of the facial image. The feature extraction network is used to perform feature extraction processing on the input data of the feature extraction network; and this application does not limit the implementation method of the feature extraction network. For example, the feature extraction network can be implemented using the feature extraction network shown in Figure 2 or Figure 3.
[0115] In addition, in order to better improve the feature extraction effect, the present application also provides a possible implementation method of the above feature extraction network. Under this implementation method, the feature extraction network can be implemented using a convolutional neural network, so that the feature extraction network includes multiple network layers, and different network layers in the feature extraction network are used to obtain different visual information, such as edge information, texture information, etc., so that the feature extraction network has better image feature extraction performance.
[0116] Based on the content of the above paragraph, it can be seen that in a possible implementation manner, when the above feature extraction network includes N network layers, the above step 211 can be specifically as follows: after the above facial image is input into the feature extraction network, the first network layer in the feature extraction network can be processed based on the facial image to obtain the output data of the first network layer; then the second network layer in the feature extraction network is processed based on the output data of the first network layer to obtain the output data of the second network layer; then the third network layer in the feature extraction network is processed based on the output data of the second network layer to obtain the output data of the third network layer; ... (and so on); then the Nth network layer in the feature extraction network is processed based on the output data of the N-1th network layer to obtain the output data of the Nth network layer; then, the output data of the Nth network layer is used as the image feature of the facial image. It can be seen that in a possible implementation, when the feature extraction network includes multiple network layers, the image features of the facial image may include output data of the last network layer in the feature extraction network.
[0117] In fact, in order to better improve the feature extraction effect, the present application also provides a possible implementation method of the image features of the above facial image. In this implementation method, when the above feature extraction network includes multiple network layers, the image features of the facial image include the output data of each network layer, such as the output data of the first network layer above, the output data of the second network layer above,... and the output data of the Nth network layer above, etc., so that the image features can describe the image information carried by the facial image as comprehensively as possible.
[0118] In addition, the present application does not limit the method for constructing the above feature extraction network. For example, it can be implemented by using any existing or future method that can construct a model with image feature extraction function.
[0119] Step 212: performing image feature extraction processing on the reference image corresponding to the above facial image to obtain image features of the reference image, so that the image features can represent the image information carried by the reference image.
[0120] The image features of the reference image refer to the image feature extraction and processing results for the reference image, so that the image features can represent the image information carried by the reference image.
[0121] In addition, this application does not limit the implementation of step 212 above. For example, the implementation of step 212 is similar to the implementation of step 211 above. Therefore, in one possible implementation, step 212 may specifically include: after obtaining a reference image corresponding to the facial image above, inputting the reference image into a pre-built feature extraction network, so that the feature extraction network performs image feature extraction processing on the reference image, and obtains and outputs image features of the reference image. For details about the feature extraction network, please refer to the above.
[0122] In addition, this application does not limit the implementation method of the image features of the reference image. For example, when the feature extraction network includes multiple network layers, the image features of the reference image may include the output data of the last network layer in the feature extraction network. For another example, when the feature extraction network includes multiple network layers, the image features of the reference image may include the output data of each network layer in the feature extraction network.
[0123] Furthermore, the present application does not limit the execution order of step 212 and step 211. For example, the former may be executed before the latter. For another example, the latter may be executed before the former. For another example, both may be executed simultaneously.
[0124] Step 213: performing splicing processing on the image features of the facial image and the image features of the reference image to obtain the fused features.
[0125] It should be noted that the present application does not limit the implementation method of the above step 213. For example, the step 213 can be implemented using any existing or future feature splicing method.
[0126] For another example, when the image features of the facial image and the image features of the reference image both include features of multiple channels, the above step 213 may specifically be: splicing the image features of the facial image and the image features of the reference image by channel to obtain the above fused features, so that the fused features include the feature splicing results under each channel.
[0127] For example, when the image features of the above facial image and the image features of the above reference image both include output data of multiple network layers, the above step 213 can specifically be: splicing the image features of the facial image and the image features of the reference image according to the network layer to obtain the above fusion feature, so that the fusion feature includes the data splicing results under each network layer.
[0128] Based on the relevant content of steps 211 to 213 above, it can be seen that after obtaining the facial image above and the reference image corresponding to the facial image, the image features of the facial image and the image features of the reference image can be determined first; then the image features of the facial image and the image features of the reference image are spliced together to obtain fused features, so that the fused features can represent the image information carried by the facial image and the reference image.
[0129] Based on the relevant content of step 21 above, it can be known that after obtaining the above facial image and the reference image corresponding to the facial image, the facial image and the reference image can be subjected to feature fusion processing to obtain a fusion feature, so that the fusion feature can represent the image information carried by the facial image and the reference image, thereby enabling the fusion feature to represent the information mixing result for the facial image and the reference image, and further enabling the fusion feature to provide more representative and accurate information for subsequent processing processes, thereby enabling the technical solution of the present application to have better performance and robustness.
[0130] Step 22: Perform expression coefficient regression processing based on the above fusion features to obtain the predicted expression coefficient corresponding to the above facial image.
[0131] It should be noted that the present application does not limit the implementation method of the above step 22. For example, it can be implemented by adopting any existing or future method that can perform expression coefficient regression processing based on features.
[0132] In addition, to further improve the accuracy of the expression coefficient, the present application also provides an implementation of step 22 above. In this implementation, step 22 can specifically be: inputting the above fusion feature into a pre-constructed expression coefficient regression network, so that the expression coefficient regression network can perform expression coefficient regression processing on the fusion feature, and obtain and output the predicted expression coefficient corresponding to the above facial image. The expression coefficient regression network is used to perform expression coefficient regression processing on the input data of the expression coefficient regression network; and the present application does not limit the implementation method of the expression coefficient regression network.
[0133] In addition, in order to better improve the accuracy of the expression coefficient, the present application also provides a possible implementation of the above expression coefficient regression network, under which the expression coefficient regression network may include a first fully connected layer, a batch normalization (BatchNorm) layer, a rectified linear unit (Rectified Linear Unit, ReLU) activation layer, and a second fully connected layer; and the input data of the batch normalization layer includes the output data of the first fully connected layer, the input data of the rectified linear unit activation layer includes the output data of the batch normalization layer, and the input data of the second fully connected layer includes the output data of the rectified linear unit activation layer. Among them, the first fully connected layer refers to a fully connected layer existing in the expression coefficient regression network, and the second fully connected layer refers to another fully connected layer existing in the expression coefficient regression network.
[0134] Based on the content of the previous paragraph, it can be seen that in one possible implementation method, in order to better ensure that the structure of the above expression coefficient regression network is concise and efficient, the present application adopts a fully connected neural network as the core component of the expression coefficient regression network, so that the expression coefficient regression network consists of two fully connected layers, and a batch normalization layer and a rectified linear unit activation layer are embedded between the two fully connected layers to enhance the nonlinear expression ability of the expression coefficient regression network.
[0135] It can be seen that for the expression coefficient regression network shown in the above two paragraphs, the expression coefficient regression network can establish a mapping relationship between the input features and the expression coefficients by using a fully connected layer. Among them, each neuron in the fully connected layer is connected to all neurons in the previous layer, thereby realizing the transmission and processing of global information. In addition, the expression coefficient regression network can learn the complex relationship between the input features and the expression coefficients through a multi-layer fully connected structure to accurately predict and regress the expression coefficients. In addition, in order to better improve the robustness and training effect of the expression coefficient regression network, the present application introduces a BatchNorm layer between the two fully connected layers of the expression coefficient regression network. Among them, because the BatchNorm layer can normalize the input of each batch, so that the expression coefficient regression network is more stable to the scale change of the input data, thereby accelerating the convergence speed of the network and improving the accuracy of the prediction. In addition, the introduction of the ReLU activation layer can enhance the nonlinear expression ability of the expression coefficient regression network, so that it can better adapt to complex expression features.
[0136] Based on the relevant content of steps 21 to 22 above, it can be seen that in some application scenarios, such as the scenario shown in Figure 2 or Figure 3, after obtaining the above facial image and the reference image corresponding to the facial image, the feature extraction network can be used to first determine the image features of the facial image and the image features of the reference image; then, the image features of the facial image and the image features of the reference image are spliced to obtain fused features; then, the expression coefficient regression network is used to perform expression coefficient regression on the fused features to obtain the predicted expression coefficient corresponding to the facial image, which is conducive to improving the accuracy of the expression coefficient.
[0137] In fact, in order to better improve the accuracy of the expression coefficient, the present application also provides a possible implementation method of the predicted expression coefficient corresponding to the facial image above. Under this implementation method, the predicted expression coefficient corresponding to the facial image is determined by using an expression coefficient determination model. It can be seen that, under a possible implementation method, the above S2 can specifically be: inputting the facial image and the reference image corresponding to the facial image into the expression coefficient determination model, so that the expression coefficient determination model can perform expression coefficient determination processing on the facial image based on the reference image, and obtain and output the predicted expression coefficient corresponding to the facial image. Among them, the expression coefficient determination model is used to perform expression coefficient determination processing on the input data of the expression coefficient determination model; and the present application does not limit the implementation method of the expression coefficient determination model. For example, the expression coefficient determination model can include the above feature extraction network and the above expression coefficient regression network.
[0138] Based on the relevant contents of S1 to S2 above, it can be seen that for the expression coefficient determination method provided in the embodiment of the present application, a facial image and a reference image corresponding to the facial image are first obtained; then, based on the facial image and the reference image, the predicted expression coefficient corresponding to the facial image is determined, so that the predicted expression coefficient can represent the expression state presented by the facial image. Among them, because the reference image and the facial image are both used to describe the same object, the facial features described by the reference image and the facial features described by the facial image are both facial features possessed by the object, so that the reference image can supplement some facial features for the facial image, and then when the expression coefficient of the facial image is determined based on the reference image, as many facial features as possible can be used, so that the predicted expression coefficient determined for the facial image can more accurately represent the expression state presented by the facial image, which is conducive to improving the accuracy of the expression coefficient. Also, because the facial image and the reference image are both determined from the same image sequence, the acquisition configuration of the facial image is consistent with the acquisition configuration of the reference image, so that the facial features described by the reference image are more adaptable to the facial features described by the facial image, and thus the reference image can provide as much positive influence as possible for the expression coefficient prediction process of the facial image, which is conducive to improving the accuracy of the expression coefficient.
[0139] In addition, for the reference image corresponding to the above facial image, the reference image can be used to represent the expressionless state of the object described by the facial image, so that the reference image can better express the facial features of the object, such as the facial features in the expressionless state, etc., so that when determining the predicted expression coefficient corresponding to the above facial image based on the reference image, the facial features can be referred to, and then the predicted expression coefficient finally determined for the facial image can more accurately express the expression state of the object presented in the facial image, thereby effectively avoiding defects caused by different objects having different facial features, such as ambiguity problems, etc., which is conducive to improving the accuracy of the expression coefficient.
[0140] In addition, this application does not limit the application scenarios of the expression coefficient determination method provided in this application. For ease of understanding, some scenarios are explained below.
[0141] Scenario 1: When the expression coefficient determination method provided in this application is applied to the model update scenario of the expression coefficient determination model, the expression coefficient determination method provided in this application may include some or all of the following steps 31-34.
[0142] Step 31: Acquire a facial image and a reference image corresponding to the facial image; the facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence.
[0143] It should be noted that for the relevant content of step 31 above, please refer to the relevant content of S1 above. It can be seen that under one possible implementation, step 31 can specifically be: after obtaining the above sample video, randomly extract a frame of image that has not been traversed from the sample video as the facial image, and use the image with the smallest expression coefficient label in the sample video as the reference image corresponding to the facial image.
[0144] Step 32: Input the above facial image and the reference image corresponding to the facial image into the expression coefficient determination model to obtain the predicted expression coefficient corresponding to the facial image output by the expression coefficient determination model.
[0145] The expression coefficient determination model refers to a model that needs to be updated; and this application does not limit the implementation method of the expression coefficient determination model.
[0146] In addition, the present application does not limit the implementation method of the above step 32. For example, when the above expression coefficient determination model includes a feature extraction network and an expression coefficient regression network, the step 32 can be specifically as follows: after the above facial image and the reference image corresponding to the facial image are input into the expression coefficient determination model, the feature extraction network in the expression coefficient determination model first determines the image features of the facial image and the image features of the reference image; then, the image features of the facial image and the image features of the reference image are spliced to obtain fused features; then, the expression coefficient regression network in the expression coefficient determination model performs expression coefficient regression on the fused features to obtain and output the predicted expression coefficient corresponding to the facial image.
[0147] Step 33: Determine the loss representation data of the above expression coefficient determination model based on the predicted expression coefficient corresponding to the above facial image and the expression coefficient label corresponding to the facial image.
[0148] The expression coefficient label corresponding to a facial image is used to represent the actual expression state presented in the facial image. Furthermore, this application does not limit the method for obtaining the expression coefficient label corresponding to the facial image; for example, it can be implemented through manual annotation. For another example, when the facial image is a frame extracted from the sample video above, the expression coefficient label corresponding to this frame can be determined as the expression coefficient label corresponding to the facial image.
[0149] The loss characterization data of the expression coefficient determination model is used to characterize the performance of the expression coefficient determination model; and this application does not limit the determination process of the loss characterization data of the expression coefficient determination model.
[0150] In fact, in some application scenarios, if the current round involves multiple facial images, and there are a large number of images in small expression states and a small number of images in other expression states among the multiple facial images, in order to better improve the learning effect of the above expression coefficient determination model for various expression states, the present application also provides a possible implementation method of the above step 33. Under this implementation method, when the number of the above facial images is multiple, such as when multiple facial images are required in each round, the step 33 can specifically include the following steps 331-333.
[0151] Step 331: For any facial image, determine the expression distribution range to which the facial image belongs from at least two expression distribution ranges based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image.
[0152] The at least two expression distribution ranges refer to expression state divisions that are pre-set based on application scenarios; and different expression distribution ranges represent different expression state intervals.
[0153] In addition, the present application does not limit the at least two expression distribution ranges mentioned above. For example, in some application scenarios, the at least two expression distribution ranges may include a small expression distribution range, a medium expression distribution range, and a large expression distribution range.
[0154] Regarding the small expression distribution range mentioned above, the small expression distribution range is used to represent the expression state interval to which small expressions belong, so that the small expression distribution range can represent the characteristics of various small expressions. In addition, this application does not limit the implementation method of the small expression distribution range. For example, the small expression distribution range can be: the maximum value of the predicted expression coefficient corresponding to an image is not higher than 0.3, and the maximum value of the expression coefficient label corresponding to the image is not higher than 0.3.
[0155] For the expression distribution range mentioned above, the medium expression distribution range is used to represent the expression state interval to which the medium-amplitude expression belongs, so that the medium expression distribution range can represent the characteristics of various medium-amplitude expressions; and the present application does not limit the implementation method of the medium expression distribution range. For example, the medium expression distribution range can be: the maximum value of the predicted expression coefficient corresponding to an image is not higher than 0.6, the maximum value of the predicted expression coefficient corresponding to the image is higher than 0.3, the maximum value of the expression coefficient label corresponding to the image is not higher than 0.6, and the maximum value of the expression coefficient label corresponding to the image is higher than 0.3.
[0156] For the large expression distribution range mentioned above, the large expression distribution range is used to represent the expression state interval to which large-scale expressions belong, so that the large expression distribution range can represent the characteristics of various large-scale expressions; and the present application does not limit the implementation method of the large expression distribution range. For example, the large expression distribution range can be: the maximum value of the predicted expression coefficient corresponding to an image is higher than 0.6, and the maximum value of the expression coefficient label corresponding to the image is higher than 0.6.
[0157] In addition, for the qth facial image involved in the current round, the qth facial image refers to the qth sample image involved in the current round; and the expression distribution range to which the qth facial image belongs refers to the expression state classification result for the qth facial image. For example, the expression distribution range to which the qth facial image belongs can be a small expression distribution range, a medium expression distribution range, or a large expression distribution range. Where q is a positive integer, and q≤the number of facial images involved in the current round.
[0158] In addition, the present application does not limit the process of determining the expression distribution range to which the qth facial image belongs. For example, when the at least two expression distribution ranges mentioned above include H expression distribution ranges, the process of determining the expression distribution range to which the qth facial image belongs may include: determining whether the maximum value of the predicted expression coefficient corresponding to the qth facial image belongs to the prediction coefficient interval in the hth expression distribution range, and determining whether the maximum value of the expression coefficient label corresponding to the qth facial image belongs to the label coefficient interval in the hth expression distribution range. If both belong, then the hth expression distribution range is determined as the expression distribution range to which the qth facial image belongs. The prediction coefficient interval refers to the interval constraint that exists in the hth expression distribution range and is pre-set for the predicted expression coefficient. The label coefficient interval refers to the interval constraint that exists in the hth expression distribution range and is pre-set for the expression coefficient label, where h is a positive integer, h≤H, and H is a positive integer.
[0159] Based on the relevant content of step 331 above, it can be known that in some application scenarios, for the qth facial image involved in the current round, based on the maximum value of the predicted expression coefficient corresponding to the qth facial image and the maximum value of the expression coefficient label corresponding to the qth facial image, the expression distribution range to which the facial image belongs is determined from the at least two expression distribution ranges above, so that both maximum values satisfy the range constraint described by the expression distribution range to which the facial image belongs.
[0160] Step 332 : Determine loss representation data for at least two expression distribution ranges based on the predicted expression coefficients corresponding to the plurality of facial images, the expression coefficient labels corresponding to the plurality of facial images, and the expression distribution ranges to which the plurality of facial images belong.
[0161] The loss representation data for the hth expression distribution range is used to represent the performance of the expression coefficient determination model described above under the expression state range described by the hth expression distribution range. h is a positive integer, h≤H.
[0162] In addition, the present application does not limit the implementation method of the above step 332. For example, when the above multiple facial images include at least one image to be used, and the expression distribution range of each image to be used is the target range in the above at least two expression distribution ranges, the step 332 can specifically include the following steps 3321-3322.
[0163] Step 3321: For any image to be used, determine the loss representation data of the image to be used based on the difference between the predicted expression coefficient corresponding to the image to be used and the expression coefficient label corresponding to the image to be used.
[0164] The target range is used to represent any one of the at least two expression distribution ranges mentioned above. For example, the target range is implemented using the hth expression distribution range mentioned above.
[0165] The image to be used refers to an image existing in the plurality of facial images and belonging to the target range, so that the at least one image to be used can represent all images existing in the plurality of facial images and belonging to the target range.
[0166] In addition, for the yth image to be used, the loss characterization data of the image to be used is used to represent the performance of the above expression coefficient determination model on the yth image to be used; and the present application does not limit the determination process of the loss characterization data of the image to be used. For example, it can be specifically: based on the preset loss function and the difference between the predicted expression coefficient corresponding to the yth image to be used and the expression coefficient label corresponding to the yth image to be used, the loss characterization data of the image to be used is determined. The preset loss function can be set in advance according to the application scenario. For example, the preset loss function can be implemented using smoothL1Loss. Wherein, y is a positive integer, and y≤the number of images in the above at least one image to be used.
[0167] Step 3322: Determine the loss characterization data of the target range based on the average value of the loss characterization data of at least one image to be used.
[0168] In the present application, for at least one image to be used that belongs to the target range, after obtaining the loss characterization data of each image to be used, the average value of the loss characterization data of these images to be used can be calculated as the loss characterization data of the target range, so that the loss characterization data can better represent the performance of the above expression coefficient determination model in the expression state range described by the target range.
[0169] Based on the relevant contents of steps 3321 to 3322 above, it can be known that in some application scenarios, for the multiple facial images above, after determining the expression distribution range to which each facial image belongs from at least two expression distribution ranges above, for any expression distribution range, the loss representation data of the expression distribution range can be determined based on the average value of the loss representation data of all images divided into the expression distribution range, so that the loss representation data can better represent the performance of the above expression coefficient determination model in the expression state range described by the expression distribution range, so that the performance of the expression coefficient determination model in different expression state ranges can be obtained.
[0170] Based on the relevant content of step 332 above, it can be known that in some application scenarios, if the current round involves multiple facial images, after determining the expression distribution range to which each facial image belongs from at least two expression distribution ranges above, the loss representation data of each expression distribution range can be determined separately based on the images under each expression distribution range, so that the loss representation data of these expression distribution ranges can better represent the performance of the expression coefficient determination model in different expression state intervals. This can effectively avoid defects caused by the large difference between the number of appearances of the multiple facial images in different expression state intervals, such as greatly weakening the impact of a smaller number of expression state intervals on the model update, so that the expression coefficient determination model can better learn the expression coefficient determination ability under various expression state intervals, which is beneficial to improving the model update effect.
[0171] Step 333: Determine the loss representation data of the expression coefficient determination model based on the loss representation data of the at least two expression distribution ranges.
[0172] It should be noted that the present application does not limit the implementation method of the above step 333. For example, it can be specifically as follows: adding the loss representation data of at least two expression distribution ranges above to obtain the loss sum value; and then determining the loss representation data of the expression coefficient determination model based on the loss sum value.
[0173] It should also be noted that the present application does not limit the implementation method of the step of "determining the loss representation data of the expression coefficient determination model based on the loss sum value" in the previous paragraph. For example, it can be specifically: directly determining the loss sum value as the loss representation data of the expression coefficient determination model. For example, in some application scenarios, in order to better improve the learning efficiency of the expression coefficient determination model for small expression state intervals, the present application also provides a possible implementation method of this step. Under this implementation method, the step can be specifically: determining the loss representation data of the expression coefficient determination model based on the loss sum value and the average value between the loss representation data of the multiple facial images above, so that the loss representation data of the expression coefficient determination model can further emphasize the performance of the expression coefficient determination model in the small expression state interval, thereby helping to improve the learning efficiency of the expression coefficient determination model for small expression states.
[0174] Based on the relevant contents of steps 331 to 333 above, it can be known that in some application scenarios, if the current round involves multiple facial images, the expression distribution range to which each facial image belongs can be determined from the at least two expression distribution ranges above based on the predicted expression coefficients and expression coefficient labels of the multiple facial images; then, based on the images under each expression distribution range, the loss representation data of each expression distribution range is determined respectively; then, based on the loss representation data of these expression distribution ranges, the loss representation data of the expression coefficient determination model above is determined, so that the loss representation data of the expression coefficient determination model can better represent the performance of the expression coefficient determination model in different expression state intervals, thereby facilitating better optimization of the performance of the expression coefficient determination model in different expression state intervals, and thus making the expression coefficient determination model finally obtained suitable for processing images in various expression states, thereby facilitating improving the performance of the expression coefficient determination model.
[0175] Step 34: Update the expression coefficient determination model based on the loss representation data of the expression coefficient determination model, and return to execute the above step 31 and subsequent steps until the preset stop condition is reached.
[0176] The preset stopping condition can be set in advance based on the actual application scenario; and this application does not limit the preset stopping condition. For example, the preset stopping condition can be: the loss representation data of the expression coefficient determination model is lower than a preset loss threshold. In another example, the preset stopping condition can be: the rate of change of the loss representation data of the expression coefficient determination model is lower than a preset rate of change threshold. In another example, the preset stopping condition can be: the number of updates of the expression coefficient determination model reaches a preset number threshold.
[0177] In addition, the present application does not limit the model updating process in the above step 34. For example, the model parameter updating process can be performed by a gradient descent algorithm.
[0178] Based on the relevant content of steps 31 to 34 above, it can be seen that the present application provides a model updating scheme, and the scheme can be specifically as follows: first obtain the facial image to be used in the current round and the reference image corresponding to the facial image; then, the expression coefficient determination model performs expression coefficient determination processing on the facial image based on the reference image to obtain the predicted expression coefficient corresponding to the facial image; then, based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image, the expression coefficient determination model is updated, and the next round of process is continued based on the updated expression coefficient determination model, and the iteration is repeated until the preset stop condition is reached, so that the expression coefficient determination model finally obtained has better expression coefficient determination performance.
[0179] In addition, in order to better improve the accuracy of the expression coefficient, the above expression coefficient determination model not only needs to learn how to perform expression coefficient determination processing based on two different images, but also learn how to perform expression coefficient determination processing based on two identical images, so that the expression coefficient determination model can learn how to perform expression coefficient determination processing based on multiple frame images and learn how to perform expression coefficient determination processing based on a single frame image, so that the final expression coefficient determination model has both the ability to perform expression coefficient determination processing based on multiple frame images and the ability to perform expression coefficient determination processing based on a single frame image, so that the expression coefficient determination model can better complete the expression coefficient determination task at different stages in the above image stream, which is beneficial to improving the accuracy of the expression coefficient.
[0180] In order to better meet the requirements described in the above paragraph, the present application also provides an update scheme for the above expression coefficient determination model, which may specifically include the following steps 41 to 44.
[0181] Step 41: Obtain a facial image and a reference image corresponding to the facial image; the facial image and the reference image are both used to describe the same object; the facial image and the reference image are both determined from the same image sequence; the image sequence includes an expressionless image, and the expressionless image is used to represent the expressionless state of the object; the reference image is the facial image or the expressionless image.
[0182] Among them, the expressionless image is used to represent the expressionless state of the object described by the above facial image; and this application does not limit the expressionless image. For example, when the facial image is extracted from the above sample video, the expressionless image may refer to an image with the smallest expression coefficient label existing in the sample video, so that the expressionless image can represent the expressionless state of the object.
[0183] In addition, the present application does not limit the process for determining the reference image in step 41 above. For example, the reference image may be determined by a process randomly selected from at least one candidate process; the randomly selected process is: determining a neutral image as the reference image, or determining a facial image as the reference image. The candidate process refers to a pre-set process that can be selected when determining the reference image; and the present application does not limit the implementation method of the at least one candidate process. For example, the at least one candidate process may include two determination processes: determining a neutral image as the reference image, and determining a facial image as the reference image.
[0184] Based on the content of the previous paragraph, it can be seen that under a possible implementation method, the above step 41 can be specifically as follows: after obtaining the above sample video, first randomly extract a frame of image that has not been traversed from the sample video as a facial image; then, use a process of randomly selecting from at least one candidate process to determine the reference image corresponding to the facial image, so that the reference image may be the facial image or the expressionless image, which is conducive to realizing the random selection of a combination of facial image + facial image or a combination of facial image + expressionless image for iterative update processing, thereby helping to improve model performance.
[0185] Step 42: Determine an image distinction prediction result and a predicted expression coefficient corresponding to the facial image based on the facial image, a reference image corresponding to the facial image, and an expression coefficient determination model; the image distinction prediction result is used to indicate whether the facial image and the reference image are the same image.
[0186] For details about the predicted expression coefficients corresponding to facial images, please refer to the above text.
[0187] The image distinction prediction result is used to indicate whether the above facial image and the reference image corresponding to the facial image are predicted to be the same image, so that the image distinction prediction result can indicate whether the two input images of the expression coefficient determination model are the same image.
[0188] In addition, the present application does not limit the process of determining the above image distinction prediction result. For example, it can specifically be: determining the image distinction prediction result based on the image features of the facial image and the image features of the reference image.
[0189] In addition, the present application does not limit the implementation method of the step of "determining the image distinction prediction result based on the image features of the facial image and the image features of the reference image" in the previous paragraph. For example, this step can be implemented with the help of an image distinction network. The image distinction network is used to determine whether the two input images of the expression coefficient determination model are the same image; and the present application does not limit the implementation method of the image distinction network. For example, it can be implemented using a fully connected layer. In addition, the present application does not limit the working principle of the image distinction network. For example, after the image features of the facial image and the image features of the reference image are determined by the feature extraction network in the expression coefficient determination model, the image distinction network can be used to determine the image distinction prediction result based on the image features of the facial image and the image features of the reference image, so that the image distinction prediction result is used to indicate whether the facial image and the reference image are the same image. It can be seen that, under one possible implementation, the process of determining the image distinction prediction result can be: the image distinction network processes the image features of the above facial image and the image features of the above reference image to obtain and output the image distinction prediction result, so that the image distinction prediction result can indicate whether the two input images of the expression coefficient determination model are the same image.
[0190] In addition, this application does not limit the implementation method of the image differentiation network in the above paragraph. For example, the image differentiation network can be an independent network; and the image differentiation network can use the feature extraction network in the above expression coefficient determination model to perform some data processing.
[0191] For another example, in some application scenarios, the above-mentioned image discrimination network can be used as a partial network in the expression coefficient determination model. It can be seen that, in one possible implementation, the above-mentioned expression coefficient determination model can include a feature extraction network, an image discrimination network, and an expression coefficient regression network; the input data of the image discrimination network includes the output data of the feature extraction network; and the input data of the expression coefficient regression network includes the output data of the feature extraction network, so that the working principle of the expression coefficient determination model can at least include: after the above-mentioned facial image and the reference image corresponding to the facial image are input into the expression coefficient determination model, the expression coefficient determination model can also be used to determine the image discrimination prediction result based on the image features of the facial image and the image features of the reference image, so that the image discrimination prediction result is used to indicate whether the facial image and the reference image are the same image, so that the image discrimination prediction result can represent the image feature extraction performance of the expression coefficient determination model to a certain extent, so that the update process of the feature extraction network in the expression coefficient determination model can be guided based on the image discrimination prediction result.
[0192] In addition, the present application does not limit the implementation of the above step 42. For ease of understanding, two examples are used for illustration below.
[0193] Example 1: When the image distinction network can be an independent network, the above step 42 can be specifically as follows: input the above facial image and the reference image corresponding to the facial image into the expression coefficient determination model, obtain the predicted expression coefficient corresponding to the facial image output by the expression coefficient determination model, the image features of the facial image and the image features of the reference image, and the image distinction network determines the image distinction prediction result based on the image features of the facial image and the image features of the reference image.
[0194] Example 2: When the expression coefficient determination model includes an image discrimination network, the above step 42 can be specifically: input the above facial image and the reference image corresponding to the facial image into the expression coefficient determination model, and obtain the image discrimination prediction result output by the expression coefficient determination model and the predicted expression coefficient corresponding to the facial image.
[0195] It can be seen that in one possible implementation mode, when the above expression coefficient determination model includes a feature extraction network, an image discrimination network and an expression coefficient regression network, the above step 42 can be specifically as follows: after the above facial image and the reference image corresponding to the facial image are input into the expression coefficient determination model, the feature extraction network in the expression coefficient determination model first determines the image features of the facial image and the image features of the reference image; then, the image features of the facial image and the image features of the reference image are spliced to obtain fused features; then, the image discrimination network processes the fused features to obtain and output the image discrimination prediction results, and the expression coefficient regression network in the expression coefficient determination model performs expression coefficient regression processing on the fused features to obtain and output the predicted expression coefficient corresponding to the facial image.
[0196] Based on the relevant content of step 42 above, it can be known that for the current round, after obtaining the facial image and the reference image corresponding to the facial image, the two images and the expression coefficient determination model can be used to determine the image distinction prediction result and the predicted expression coefficient corresponding to the facial image, so that the model performance of the expression coefficient determination model can be determined based on the image distinction prediction result and the predicted expression coefficient.
[0197] Step 43: Determine the loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image, the expression coefficient label corresponding to the facial image, the image distinction prediction result, and the image distinction label; the image distinction label is determined based on the reference image selection process.
[0198] Among them, the image distinction label refers to the true value corresponding to the above image distinction prediction result, so that the image distinction label can accurately indicate whether the two input images of the above expression coefficient determination model, such as the above facial image and the reference image corresponding to the facial image, are the same image.
[0199] In addition, for the above-mentioned image distinction label, the image distinction label can be determined based on the selection process of the reference image corresponding to the above-mentioned facial image; and the present application does not limit the determination process. For example, it can be specifically: if the facial image is selected as the reference image in the selection process, the first label value can be used as the image distinction label, so that the image distinction label can accurately indicate that the facial image and the reference image corresponding to the facial image are the same image; if the above-mentioned expressionless image is selected as the reference image in the selection process, the second label value can be used as the image distinction label, so that the image distinction label can accurately indicate that the facial image and the reference image corresponding to the facial image are not the same image. Among them, the first label value refers to a pre-set label value used to indicate that the facial image and the reference image corresponding to the facial image are the same image. The second label value refers to a pre-set label value used to indicate that the facial image and the reference image corresponding to the facial image are not the same image.
[0200] In addition, the present application does not limit the implementation of the above step 43. For example, it may specifically include the following steps 431 to 433.
[0201] Step 431: Determine a first loss based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image.
[0202] Among them, the first loss is the loss determined based on the predicted expression coefficient corresponding to the above facial image and the expression coefficient label corresponding to the facial image; and this application does not limit the determination process of the first loss. For example, the implementation method of the determination process of the first loss is similar to the implementation method of the determination process of the loss representation data shown in steps 31-33 above.
[0203] It can be seen that, in one possible implementation, the above first loss can be obtained by adding the loss representation data of the above at least two expression distribution ranges and the average value between the above loss representation data of the multiple facial images, so that the first loss can better represent the expression coefficient determination performance of the above expression coefficient determination model.
[0204] Step 432: Determine a second loss based on the above image distinction prediction result and the above image distinction label.
[0205] Among them, the second loss is determined based on the difference between the above image distinction prediction result and the above image distinction label, so that the second loss can represent the performance of the above expression coefficient determination model in image feature extraction.
[0206] In addition, the present application does not limit the implementation of the above step 432. For example, the step 432 can be implemented with the help of cross entropy loss.
[0207] Step 433: Determine the loss representation data of the expression coefficient determination model based on the first loss and the second loss.
[0208] It should be noted that the present application does not limit the implementation method of the above step 433. For example, it can be specifically: adding the above first loss and the above second loss to obtain the loss representation data of the above expression coefficient determination model.
[0209] Based on the relevant content of steps 431 to 433 above, it can be known that in some application scenarios, after obtaining the predicted expression coefficient corresponding to the above facial image and the above image distinction prediction result, the first loss can be determined based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image, and the second loss can be determined based on the image distinction prediction result and the above image distinction label; then the first loss and the second loss are added to obtain the loss representation data of the above expression coefficient determination model, so that the loss representation data can better represent the performance of the expression coefficient determination model.
[0210] Based on the above content, it can be seen that in one possible implementation, the above step 43 can be implemented with the help of the formulas shown in the following formulas (1)-(10). cur =Backbone(x cur ) (1) feat ref =Backbone(x ref ) (2) feat cat =Concat(feat cur ,feat ref ) (3) exp pred =Regressor(feat cat ) (4) cls pred =FC(feat cat ) (5) loss = SmoothL1Loss(exp pred ,exp truth )+CrossEntropy(cls pred ,cls truth )+loss bs (10)
[0211] Where x cur Represents a face image; x ref Indicates the reference image corresponding to the facial image; feat cur Represents the image features of the facial image; feat ref Represents the image features of the reference image; Backbone(·) represents the image feature extraction process, such as the representation function of the feature extraction network mentioned above; Concat(feat cur ,feat ref ) indicates that the feat cur With this feat ref Splicing process; feat cat Indicates the above fusion features; exp pred Represents the predicted expression coefficient corresponding to the facial image; Regressor(·) represents the expression coefficient regression processing, such as the representation function of the expression coefficient regression network mentioned above; cls pred represents the image distinction prediction result between the facial image and the reference image; FC(·) represents the frame judgment processing, such as the representation function of the image distinction network mentioned above; pred[0.6≤gt] represents the prediction coefficient interval in the large expression distribution range; truth[0.6≤gt] represents the label coefficient interval in the large expression distribution range; smoothL1Loss(pred[0.6≤gt],truth[0.6≤gt]) represents the average value between the loss representation data of all facial images belonging to the large expression distribution range; Represents the loss representation data of the large expression distribution range; pred[0.3≤gt≤0.6] represents the prediction coefficient interval in the medium expression distribution range; truth[0.3≤gt≤0.6] represents the label coefficient interval in the medium expression distribution range; smoothL1Loss(pred[0.3≤gt≤0.6],truth[0.3≤gt≤0.6]) represents the average value between the loss representation data of all facial images belonging to the medium expression distribution range; Represents the loss representation data of the expression distribution range; pred[gt≤0.3] represents the prediction coefficient interval in the small expression distribution range; truth[gt≤0.3] represents the label coefficient interval in the small expression distribution range; smoothL1Loss(pred[gt≤0.3], truth[gt≤0.3]) represents the average value between the loss representation data of all facial images belonging to the small expression distribution range; Indicates the loss representation data of the distribution range of the small expression; gt represents the maximum value of the multidimensional expression coefficient; loss bs The sum of the loss representation data representing the distribution range of all expressions; exp truth Indicates the expression coefficient label corresponding to the facial image; cls truth Indicates the image distinction label above; SmoothL1Loss(exp pred ,exp truth ) represents the average value of the loss representation data of all facial images involved in the current round; the loss representation data of each facial image is determined by smoothL1Loss; CrossEntropy(cls pred ,cls truth ) indicates that the cls pred With the cls truth The cross entropy loss between ; loss represents the loss representation data of the model determined by the expression coefficient above.
[0212] Based on the relevant content of step 43 above, it can be known that for the current round, if the current round involves multiple facial images, after obtaining the predicted expression coefficients corresponding to these facial images, the expression coefficient labels corresponding to these facial images, the image distinction prediction results corresponding to these facial images, and the image distinction labels corresponding to these facial images, the loss representation data of the above expression coefficient determination model is determined based on these data, so that the loss representation data can better represent the performance of the expression coefficient determination model.
[0213] Step 44: Update the expression coefficient determination model based on the loss representation data of the expression coefficient determination model, and return to execute the above step 41 and subsequent steps until the preset stop condition is reached.
[0214] It should be noted that the relevant contents of step 44 are similar to those of step 34 above, and for the sake of brevity, they will not be repeated here. For ease of understanding, two examples are used for illustration below.
[0215] Example 1: When the image distinction network can be an independent network, the above step 44 can be specifically as follows: based on the loss representation data of the above expression coefficient determination model, update the expression coefficient determination model, and return to execute the above step 41 and its subsequent steps until the preset stop condition is reached, save the current round of expression coefficient determination model, so that the saved expression coefficient determination model can be used to perform the expression coefficient prediction task later.
[0216] Example 2. When the expression coefficient determination model includes an image discrimination network, the above step 44 can be specifically as follows: based on the loss representation data of the above expression coefficient determination model, update the expression coefficient determination model, and return to execute the above step 41 and its subsequent steps until the preset stop condition is reached, and then delete the image discrimination network from the expression coefficient determination model, so that the deleted expression coefficient determination model can be used to perform the expression coefficient prediction task subsequently. This can effectively avoid the resource overhead generated by the image discrimination network when performing the expression coefficient prediction task.
[0217] Based on the relevant contents of steps 41 to 44 above, it can be known that for the above expression coefficient determination model, if the expression coefficient determination model includes a feature extraction network and an expression coefficient regression network, the present application can introduce a simple fully connected layer after the feature extraction network, such as the image distinction network shown in FIG3 , and the fully connected layer is used to determine whether the two input images of the expression coefficient determination model are the same, so that the expression coefficient determination model can distinguish whether the two input images of the expression coefficient determination model are the same with the help of the comparison analysis process in the fully connected layer, thereby providing targeted guidance for subsequent expression coefficient prediction. In addition, in the process of iterative updating of the model, the present application adopts two different input combinations: a combination of facial image + facial image, and a combination of facial image + expressionless image, so that the present application can iteratively update the model by randomly selecting these two combinations. In addition, the present application also optimizes the network parameters through the loss function, so that the expression coefficient determination model can have the ability of single-frame prediction and multi-frame prediction at the same time, so that when the expression coefficient determination model is used to analyze the real-time acquired image stream, whether in the early acquisition stage or in the later acquisition stage, the expression coefficient determination model can make accurate expression coefficient determination processing according to the actual situation, which is conducive to improving the accuracy of the expression coefficient. It can be seen that the present application uses the aforementioned model update strategy and the improvement method of the above-mentioned model network structure to make the expression coefficient determination model finally obtained show better performance in single-frame prediction and multi-frame prediction. In addition, the present application provides a flexible and effective method for the case where it is impossible to immediately obtain an image representing an expressionless state, so that the expression coefficient determination model can accurately predict the expression coefficient at different stages.
[0218] Scenario 2: When the expression coefficient determination method provided in this application is used to determine the predicted expression coefficient corresponding to the current frame image in the image stream, the expression coefficient determination method provided in this application may include some or all of the following steps 51-53.
[0219] Step 51: Obtain a facial image and a reference image corresponding to the facial image; the facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence; the facial image is the current frame image in the image stream.
[0220] It should be noted that for the relevant content of step 51 above, please refer to the relevant content of S1 above. It can be seen that for the facial image in step 51, the facial image can refer to the current frame image in the image stream, such as the image with the latest acquisition time in the image stream. In addition, for the reference image in step 51, if the number of frames of the to-be-selected image in the image stream is lower than the preset frame number threshold, the reference image can be directly implemented using the current frame image; however, if the number of frames of the to-be-selected image in the image stream is equal to or higher than the preset frame number threshold, the reference image can be implemented using the to-be-selected image with the minimum coefficient and value in the image stream, such as the image used to represent the expressionless state determined in the historical processing process of the image stream.
[0221] Step 52: Determine a predicted expression coefficient corresponding to the facial image based on the facial image and a reference image corresponding to the facial image.
[0222] It should be noted that for the relevant content of step 52, please refer to the relevant content of S2 above.
[0223] Step 53: When the acquisition time of the above reference image is earlier than the acquisition time of the above facial image, if the predicted expression coefficient corresponding to the facial image is lower than the predicted expression coefficient corresponding to the reference image, then the reference image is updated using the facial image.
[0224] In the present application, for the reference image corresponding to the facial image above, if the acquisition time of the reference image is earlier than the acquisition time of the facial image, it can be determined that the reference image refers to the candidate image with the minimum coefficient and value existing in the current image stream. Therefore, in order to better improve the accuracy of subsequent expression coefficient determination, the predicted expression coefficient corresponding to the facial image can be compared with the predicted expression coefficient corresponding to the reference image, so that when it is determined that the predicted expression coefficient corresponding to the facial image is lower than the predicted expression coefficient corresponding to the reference image, it can be determined that the expression state described by the facial image is closer to the expressionless state than the expression state described by the reference image. Therefore, the reference image can be directly updated using the facial image so that the updated reference image is the facial image, so that the updated reference image can better express the expressionless state, and then the updated reference image can better guide the expression coefficient determination process of the subsequently acquired images. In this way, the reference image can be dynamically adjusted, so that it can better cope with uncontrollable expression changes, which is beneficial to improving the adaptability and robustness of the expression coefficient determination scheme.
[0225] Based on the relevant contents of steps 51 to 53 above, it can be known that in some application scenarios, for a real-time captured image stream, if a reference image that can represent an expressionless state has been analyzed from the image stream, then after a new frame of image is captured, the expression coefficient of the new image can be determined directly based on the reference image to obtain a predicted expression coefficient of the new image, so that the predicted expression coefficient can represent the expression state carried by the new image; and, the predicted expression coefficient of the new image can be further compared with the predicted expression coefficient of the reference image, so that when it is determined that the predicted expression coefficient of the new image is lower than the predicted expression coefficient of the reference image, the reference image can be directly replaced with the new image, so that dynamic adjustment of the reference image can be achieved, which is conducive to improving the accuracy of the expression coefficient.
[0226] Scenario three: When the expression coefficient determination method provided in this application is applied to a VR device, the expression coefficient determination method may specifically include at least the following steps 61 and 62.
[0227] Step 61: The VR device obtains a facial image and a reference image corresponding to the facial image; the facial image and the reference image are both used to describe the same object; the facial image and the reference image are both determined from the same image sequence.
[0228] It should be noted that for the relevant content of step 61 above, please refer to the relevant content of S1 above. It can be seen that for the facial image in step 1, the facial image can refer to the current frame image in the image stream captured by the VR device, such as the image with the latest capture time in the image stream. In addition, for the reference image in step 61, if the number of frames of the to-be-selected image in the image stream is lower than the preset frame number threshold, the reference image can be directly implemented using the current frame image; however, if the number of frames of the to-be-selected image in the image stream is equal to or higher than the preset frame number threshold, the reference image can be implemented using the to-be-selected image with the minimum coefficient and value in the image stream, such as the image used to represent the expressionless state determined in the historical processing process of the image stream.
[0229] Step 62: The VR device determines a predicted expression coefficient corresponding to the facial image based on the facial image and a reference image corresponding to the facial image.
[0230] It should be noted that for the relevant content of step 62, please refer to the relevant content of S2 above.
[0231] Based on the relevant contents of steps 61 to 62 above, it can be known that for the above VR device, when the VR device is in the initial acquisition stage, after the VR device acquires the current frame image in real time, the current frame image is directly used as the reference image corresponding to the current frame image, and the predicted expression coefficient of the current frame image is determined based on the reference image; and after the VR device has determined the predicted expression coefficients of the K frames of candidate images, an image with the smallest predicted expression coefficient can be selected from the K frames of candidate images as the reference image, and based on the reference image, the expression coefficient determination processing is performed on the next frame image acquired by the VR device to obtain the predicted expression coefficient of the next frame image, so that when it is determined that the predicted expression coefficient of the next frame image is less than the predicted expression coefficient of the reference image, the reference image is updated to the next frame image, which is conducive to realizing dynamic adjustment processing of the reference image, thereby achieving high restoration of expressions when the number of cameras in the VR device is limited and the imaging conditions are not ideal, thereby improving the expression restoration effect.
[0232] In addition, in some application scenarios, in order to better improve the expression restoration effect, the VR device described above may have multiple cameras, such as three cameras, so that the VR device can capture multiple image streams in real time. Based on this, the present application also provides a possible implementation of the operating principle of the VR device. In this implementation, when the VR device is used to capture at least two image streams, the operating principle of the VR device may include the following steps 71 and 72.
[0233] Step 71: For any image stream, determine the current frame image in the image stream as a facial image, obtain a reference image corresponding to the facial image, and determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image corresponding to the facial image.
[0234] Step 72: Determine the current frame expression coefficient of the VR device according to the predicted expression coefficient corresponding to the current frame image in the at least two image streams.
[0235] The VR device's current frame expression coefficient refers to the expression coefficient determined by the VR device for the current capture moment, such that the current frame expression coefficient can represent the captured subject's facial expression state at the current capture moment. The current capture moment refers to the capture moment of the latest image in the at least two image streams.
[0236] In addition, the present application does not limit the implementation method of the above step 72. For example, it can be specifically: performing some analysis processing on the predicted expression coefficient corresponding to the current frame image in the at least two image streams above, such as average value analysis processing, etc., to obtain the current frame expression coefficient of the above VR device.
[0237] Based on the relevant contents of steps 71 to 72 above, it can be known that in some application scenarios, if multiple cameras are deployed in the above VR device, the predicted expression coefficient of the current frame image in the image stream captured by each camera can be determined using the expression coefficient determination method provided in this application, so that the current frame expression coefficient of the VR device can be determined based on the predicted expression coefficient of the current frame image captured by all cameras, so that the current frame expression coefficient can more accurately represent the expression state of the captured object, which is conducive to improving the accuracy of the expression coefficient.
[0238] Based on the expression coefficient determination method provided in the embodiments of this application, the embodiments of this application also provide an expression coefficient determination device, which is explained and illustrated below in conjunction with Figure 4. Figure 4 is a schematic structural diagram of the expression coefficient determination device provided in the embodiments of this application. It should be noted that for technical details of the expression coefficient determination device provided in the embodiments of this application, please refer to the relevant content of the expression coefficient determination method above.
[0239] As shown in FIG4 , the expression coefficient determination device 400 provided in an embodiment of the present application includes:
[0240] An image acquisition unit 401 is configured to acquire a facial image and a reference image corresponding to the facial image; the facial image and the reference image are both used to describe the same object; and the facial image and the reference image are both determined from the same image sequence;
[0241] The coefficient determination unit 402 is configured to determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image.
[0242] In a possible implementation, the reference image is used to represent the expressionless state of the subject.
[0243] In one possible implementation, the expression coefficient determination method is applied to a model update scenario;
[0244] The image sequence is determined based on a sample video;
[0245] The process of determining the reference image includes: determining the reference image from the image sequence based on the expression coefficient label corresponding to each frame image in the image sequence, and the expression coefficient label corresponding to the reference image is not greater than the expression coefficient label corresponding to any other image in the image sequence except the reference image.
[0246] In one possible implementation, the expression coefficient determination method is used to determine a predicted expression coefficient corresponding to a current frame image in an image stream;
[0247] The image stream includes at least one frame of image to be selected; the acquisition time of the image to be selected is earlier than the acquisition time of the current frame of image, and the predicted expression coefficient corresponding to the image to be selected is lower than a preset coefficient threshold;
[0248] The number of image frames in the at least one frame to be selected is not less than a preset frame number threshold;
[0249] The process of determining the reference image includes: determining the reference image from the at least one frame of the to-be-selected images based on the predicted expression coefficient corresponding to each of the to-be-selected images, and the predicted expression coefficient corresponding to the reference image is not greater than the predicted expression coefficient corresponding to any other to-be-selected image in the at least one frame of the to-be-selected images except the reference image.
[0250] In one possible implementation, the facial image is the current frame image in the image stream; the frame number of the to-be-selected image in the image stream is lower than a preset frame number threshold; the acquisition time of the to-be-selected image is earlier than the acquisition time of the current frame image, and the predicted expression coefficient corresponding to the to-be-selected image is lower than a preset coefficient threshold; the reference image is determined based on the current frame image.
[0251] In a possible implementation manner, the facial image is a current frame image in an image stream;
[0252] If the acquisition time of the reference image is earlier than the acquisition time of the current frame image, the expression coefficient determination device 400 further includes:
[0253] An image updating unit is configured to update the reference image using the facial image if the predicted expression coefficient corresponding to the facial image is lower than the predicted expression coefficient corresponding to the reference image.
[0254] In one possible implementation, the predicted expression coefficient is determined using an expression coefficient determination model;
[0255] The expression coefficient determination device 400 further includes:
[0256] a loss determination unit, configured to determine loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image;
[0257] A model updating unit is used to update the expression coefficient determination model based on the loss representation data of the expression coefficient determination model.
[0258] In a possible implementation manner, there are multiple facial images;
[0259] The loss determination unit is specifically used to: for any facial image, determine the expression distribution range to which the facial image belongs from at least two expression distribution ranges based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image; determine the loss representation data of the at least two expression distribution ranges based on the predicted expression coefficients corresponding to multiple facial images, the expression coefficient labels corresponding to multiple facial images, and the expression distribution ranges to which multiple facial images belong; and determine the loss representation data of the expression coefficient determination model based on the loss representation data of the at least two expression distribution ranges.
[0260] In a possible implementation manner, the plurality of facial images include at least one image to be used, and the expression distribution range to which each of the images to be used belongs is a target range among the at least two expression distribution ranges;
[0261] The process of determining the loss characterization data of the target range includes: for any of the images to be used, determining the loss characterization data of the image to be used based on the difference between the predicted expression coefficient corresponding to the image to be used and the expression coefficient label corresponding to the image to be used; and determining the loss characterization data of the target range based on the average value between the loss characterization data of at least one image to be used.
[0262] In one possible implementation, the image sequence includes an expressionless image, and the expressionless image is used to represent an expressionless state of the subject;
[0263] The expression coefficient determination device 400 further includes:
[0264] an image distinguishing unit, configured to determine an image distinguishing prediction result based on image features of the facial image and image features of the reference image, wherein the image distinguishing prediction result is used to indicate whether the facial image and the reference image are the same image;
[0265] The loss determination unit is specifically used to determine the loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image, the expression coefficient label corresponding to the facial image, the image distinction prediction result, and the image distinction label; the image distinction label is determined based on the selection process of the reference image.
[0266] In one possible implementation, the reference image is determined by a process randomly selected from at least one candidate process; the random selection process is: determining the expressionless image as the reference image, or determining the facial image as the reference image.
[0267] In one possible implementation, the coefficient determination unit 402 is specifically configured to: perform feature fusion processing on the facial image and the reference image to obtain fusion features; and perform expression coefficient regression processing based on the fusion features to obtain a predicted expression coefficient corresponding to the facial image.
[0268] In one possible implementation, the expression coefficient determination method is applied to a virtual reality (VR) device; the VR device is used to capture at least two image streams;
[0269] The image acquisition unit 401 is specifically configured to: for any of the image streams, determine a current frame image in the image stream as the facial image;
[0270] The expression coefficient determination device 400 further includes:
[0271] An expression determination unit is used to determine the current frame expression coefficient of the VR device based on the predicted expression coefficient corresponding to the current frame image in the at least two image streams.
[0272] Based on the relevant contents of the above-mentioned expression coefficient determination device 400, it can be known that for the expression coefficient determination device 400 provided in the embodiment of the present application, a facial image and a reference image corresponding to the facial image are first obtained; then, based on the facial image and the reference image, a predicted expression coefficient corresponding to the facial image is determined, so that the predicted expression coefficient can represent the expression state presented by the facial image. Among them, because the reference image and the facial image are both used to describe the same object, the facial features described by the reference image and the facial features described by the facial image are both facial features possessed by the object, so that the reference image can supplement some facial features for the facial image, and then when the expression coefficient of the facial image is determined based on the reference image, as many facial features as possible can be used. In this way, the predicted expression coefficient determined for the facial image can more accurately represent the expression state presented by the facial image, which is conducive to improving the accuracy of the expression coefficient. Also, because the facial image and the reference image are both determined from the same image sequence, the acquisition configuration of the facial image is consistent with the acquisition configuration of the reference image, so that the facial features described by the reference image are more adaptable to the facial features described by the facial image, and thus the reference image can provide as much positive influence as possible for the expression coefficient prediction process of the facial image, which is conducive to improving the accuracy of the expression coefficient.
[0273] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the expression coefficient determination method provided in the embodiment of the present application.
[0274] Referring to FIG5 , a schematic diagram of the structure of an electronic device 500 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG5 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0275] As shown in Figure 5, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0276] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0277] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0278] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0279] An embodiment of the present application also provides a computer-readable medium, which stores instructions or computer programs. When the instructions or computer programs are executed on a device, the device executes any implementation of the expression coefficient determination method provided in the embodiment of the present application.
[0280] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0281] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0282] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0283] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0284] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0285] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0286] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0287] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0288] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0289] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0290] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0291] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0292] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0293] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for determining an expression coefficient, comprising: Acquiring a facial image and a reference image corresponding to the facial image; The facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence; Determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image. The method according to claim 1 , wherein the reference image is used to represent a neutral state of the subject.
3. The method according to claim 2, wherein the expression coefficient determination method is applied to a model update scenario; The image sequence is determined based on a sample video; The process of determining the reference image includes: The reference image is determined from the image sequence based on the expression coefficient labels corresponding to each frame image in the image sequence, and the expression coefficient label corresponding to the reference image is not greater than the expression coefficient label corresponding to any other image in the image sequence except the reference image.
4. The method according to claim 2, wherein the expression coefficient determination method is used to determine the predicted expression coefficient corresponding to the current frame image in the image stream; The image stream includes at least one frame of to-be-selected image; the acquisition time of the to-be-selected image is earlier than the acquisition time of the current frame of image, and the predicted expression coefficient corresponding to the to-be-selected image is lower than a preset coefficient threshold; The number of image frames in the at least one frame to be selected is not less than a preset frame number threshold; The process of determining the reference image includes: Based on the predicted expression coefficient corresponding to each of the candidate images, the reference image is determined from the at least one frame of candidate images, and the predicted expression coefficient corresponding to the reference image is not greater than the predicted expression coefficient corresponding to any other candidate image in the at least one frame of candidate images except the reference image.
5. The method according to claim 1, wherein the facial image is a current frame image in an image stream; The number of frames of the image to be selected in the image stream is lower than a preset frame number threshold; the acquisition time of the image to be selected is earlier than the acquisition time of the current frame image, and the predicted expression coefficient corresponding to the image to be selected is lower than a preset coefficient threshold; The reference image is determined based on the current frame image.
6. The method according to claim 1, wherein the facial image is a current frame image in an image stream; If the acquisition time of the reference image is earlier than the acquisition time of the current frame image, then after determining the predicted expression coefficient corresponding to the facial image, the method further includes: If the predicted expression coefficient corresponding to the facial image is lower than the predicted expression coefficient corresponding to the reference image, the reference image is updated using the facial image.
7. The method according to claim 1, wherein the predicted expression coefficient is determined using an expression coefficient determination model; The method further comprises: Determining loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image; The expression coefficient determination model is updated according to the loss representation data of the expression coefficient determination model.
8. The method according to claim 7, wherein the number of the facial images is multiple; The process of determining the loss representation data of the expression coefficient determination model includes: For any of the facial images, determining an expression distribution range to which the facial image belongs from at least two expression distribution ranges based on a predicted expression coefficient corresponding to the facial image and an expression coefficient label corresponding to the facial image; Determining loss representation data for the at least two expression distribution ranges based on the predicted expression coefficients corresponding to the plurality of facial images, the expression coefficient labels corresponding to the plurality of facial images, and the expression distribution ranges to which the plurality of facial images belong; The loss characterization data of the expression coefficient determination model is determined based on the loss characterization data of the at least two expression distribution ranges.
9. The method according to claim 8, wherein the plurality of facial images include at least one to-be-used image, and the expression distribution range to which each of the to-be-used images belongs is a target range among the at least two expression distribution ranges; The process of determining the loss characterization data of the target range includes: For any of the images to be used, determining loss representation data of the image to be used according to a difference between a predicted expression coefficient corresponding to the image to be used and an expression coefficient label corresponding to the image to be used; The loss characterization data of the target range is determined according to an average value of the loss characterization data of the at least one image to be used.
10. The method according to claim 7, wherein the image sequence includes a neutral image, the neutral image being used to represent a neutral state of the subject; The method further comprises: determining an image distinction prediction result based on the image features of the facial image and the image features of the reference image, wherein the image distinction prediction result is used to indicate whether the facial image and the reference image are the same image; The step of determining loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image and the expression coefficient label corresponding to the facial image includes: Determining loss representation data of the expression coefficient determination model based on the predicted expression coefficient corresponding to the facial image, the expression coefficient label corresponding to the facial image, the image distinction prediction result, and the image distinction label; The image distinction label is determined according to the selection process of the reference image.
11. The method of claim 10, wherein the reference image is determined using a process randomly selected from at least one candidate process; The random selection process is: determining the expressionless image as the reference image, or determining the facial image as the reference image.
12. The method according to claim 1, wherein the process of determining the predicted expression coefficient comprises: Performing feature fusion processing on the facial image and the reference image to obtain fusion features; Perform expression coefficient regression processing based on the fusion features to obtain the predicted expression coefficient corresponding to the facial image.
13. The method according to claim 1, wherein The expression coefficient determination method is applied to virtual reality equipment; The virtual reality device is used to collect at least two image streams; The acquiring of the facial image comprises: For any of the image streams, determining a current frame image in the image stream as the facial image; The method further comprises: The current frame expression coefficient of the virtual reality device is determined according to the predicted expression coefficient corresponding to the current frame image in the at least two image streams.
14. An expression coefficient determination device, comprising: an image acquisition unit, configured to acquire a facial image and a reference image corresponding to the facial image; The facial image and the reference image are used to describe the same object; the facial image and the reference image are determined from the same image sequence; A coefficient determination unit is used to determine a predicted expression coefficient corresponding to the facial image based on the facial image and the reference image.
15. An electronic device comprising: processor and memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 13.
16. A computer-readable medium storing instructions or a computer program, wherein when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 13.
17. A computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Emotion recognition method and device for character in video, computer equipment and medium
CN112699774A
Information processing method and device, computer equipment and storage medium
CN114782864A
Expression information identification method, device and equipment, readable storage medium and product
CN115984944A
Method and apparatus for recognizing facial expression
US20210192192A1
Facial expression analysis method and system, and facial expression-based satisfaction analysis method and system
WO2021143667A1