Image transformation method, image processing model training method, device and equipment

The image processing model trained by the deep learning model uses the characteristic deviation value of the eye infrared image to generate the characteristic deviation value of the facial visible light image, solving the problem of large image transformation error in the prior art, and realizing high-precision image processing.

CN120047552APending Publication Date: 2025-05-27GRAVITYXR ELECTRONICS & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311592235.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing image transformation methods have large errors, resulting in low image processing accuracy and inability to provide accurate images.

Method used

The image processing model trained by deep learning model is used to obtain the characteristic deviation value of the infrared image of the eye, and generate the characteristic deviation value of the visible light image on the face, thereby realizing image transformation.

Benefits of technology

The error caused by mapping function calculation is reduced, and the error problem of the data itself is weakened by the feature deviation values, the details of the output image are increased, and a more accurate image is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047552A_ABST
    Figure CN120047552A_ABST
Patent Text Reader

Abstract

The invention provides an image transformation method, an image processing model training method, a device and equipment. The method comprises the following steps: acquiring a feature deviation value of an infrared image corresponding to a current to-be-processed eye; according to the feature deviation value of the infrared image corresponding to the current eye, a feature deviation value of a visible light image corresponding to the face containing the current eye is obtained through an image processing model, and the number of features of the visible light image is larger than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model. According to the method provided by the invention, the problem of low image processing precision caused by large errors in existing image transformation can be solved, and the accurate image is further provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to an image transformation method, a training method, an apparatus, and a device for an image processing model. Background Art

[0002] In traditional image display, infrared images can better reflect the thermal target characteristics of images and are commonly used in image fields with high precision requirements such as eye tracking. However, infrared images are not sensitive to the characteristics of scene brightness changes, and the image clarity is relatively low. On the other hand, visible light images can better reflect the scene detail information where the target is located and have higher clarity. Therefore, in order to ensure the clarity of image display, associating the key point information extracted from infrared images with visible light images has become an important issue in image display.

[0003] Taking the eye tracking scenario as an example, currently, the conventional method is to find the positions of corresponding key points in two sets of eye images, establish a projection transformation matrix between the two sets of eye images, and perform position correspondence through projection transformation. However, there are relatively large errors in the projection transformation of this method, and the errors will become larger and larger as the tracking time or state changes, resulting in relatively low accuracy of image processing.

[0004] Therefore, there are relatively large errors in the existing image transformation, resulting in relatively low accuracy of image processing, and thus it is impossible to provide accurate images. Summary of the Invention

[0005] The present application provides an image transformation method, a training method, an apparatus, and a device for an image processing model to solve the problem that there are relatively large errors in the existing image transformation, resulting in relatively low accuracy of image processing and inability to provide accurate images.

[0006] In a first aspect, an embodiment of the present application provides an image transformation method, the method including:

[0007] Obtaining a feature deviation value of an infrared image corresponding to an eye to be processed currently;

[0008] According to the feature deviation value of the infrared image corresponding to the eye currently, obtaining a feature deviation value of a visible light image corresponding to a face including the eye currently through an image processing model, where the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model.

[0009] In a possible design, the obtaining a feature deviation value of an infrared image corresponding to an eye currently includes:

[0010] Obtain the features of the infrared image corresponding to the current eye and the features of the reference infrared image corresponding to the eye; the features of the reference infrared image corresponding to the eye are the features of the infrared image in any eye movement state corresponding to the eye;

[0011] Take the difference between the features of the reference infrared image corresponding to the eye and the features of the infrared image corresponding to the current eye as the feature deviation value of the infrared image corresponding to the current eye.

[0012] In a possible design, the method further includes:

[0013] Obtain the features of the reference visible light image corresponding to the face including the eye; wherein, the reference infrared image corresponding to the eye and the reference visible light image corresponding to the face are collected at the same moment or when the eye corresponds to the same eye movement state;

[0014] Generate a facial expression according to the feature deviation value of the visible light image corresponding to the current face including the eye and the features of the reference visible light image corresponding to the face including the eye.

[0015] In a possible design, the obtaining the features of the reference visible light image corresponding to the face including the eye includes:

[0016] Extract the initial features of the reference visible light image corresponding to the face including the eye, and extract the initial features of the target visible light image corresponding to the face including the eye. The reference visible light image is the visible light image corresponding to the eye in any eye movement state, and the target visible light image is the visible light image corresponding to the eye at other moments or in other eye movement states;

[0017] Identify the first region and the second region in the target visible light image. The first region in the target visible light image contains feature points that change with time or eye movement state, and the second region in the target visible light image contains reference points that do not change with time or eye movement state or change within a preset condition. The reference points and the feature points are both used to represent pixel points, and one pixel point corresponds to two features;

[0018] Adjust the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face including the eye.

[0019] In a possible design, the adjusting the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face including the eye includes:

[0020] Determine a first region in the target visible light image and the boundary points of the first region according to the reference points of the target visible light image, and determine the boundary points of a second region in the target visible light image according to the feature points of the target visible light image;

[0021] Perform smoothing processing on the second region in the target visible light image by means of filtering to obtain a visible light smoothed image corresponding to the second region in the target visible light image;

[0022] Calculate the position change information of the reference visible light image relative to the target visible light image according to the initial features of the target reference points in the visible light smoothed image;

[0023] According to the position change information, obtain the features of the reference visible light image of the face including the eyes by adjusting the initial features of the respective feature points in the first region of the reference visible light image.

[0024] In a possible design, the obtaining the features of the infrared image corresponding to the current eyes includes:

[0025] Extract the initial features of the infrared image corresponding to the current eyes, and use the initial features of the infrared image corresponding to the current eyes as the features of the infrared image corresponding to the current eyes; or,

[0026] Extract the initial features of the infrared image corresponding to the current eyes and the initial features of the target infrared image corresponding to the eyes. The infrared image corresponding to the current eyes is the infrared image corresponding to the eyes in any eye movement state, and the target infrared image corresponding to the eyes is the infrared image corresponding to the eyes at other times or in other eye movement states;

[0027] Identify a first region and a second region in the target infrared image. The first region in the target infrared image includes feature points that change with time or eye movement state, and the second region in the target infrared image includes reference points that do not change with time or eye movement state or change within a preset condition. The reference points and the feature points are both used to represent pixel points, and one pixel point corresponds to two features;

[0028] Adjust the initial features of the feature points of the infrared image corresponding to the current eyes according to the initial features of the reference points of the target infrared image to obtain the features of the infrared image corresponding to the current eyes.

[0029] In a possible design, the learning rate of the image processing model is associated with the number and distribution of the feature deviation values of the input and output of the image processing model.

[0030] Second aspect, an embodiment of the present application provides a method for training an image processing model, the method comprising:

[0031] Determine a training sample set, where the training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face; wherein, the sample face includes the sample eye, the feature deviation value of the sample visible light image is used as the true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye;

[0032] Train a deep learning network model according to the training sample set to obtain an image processing model; wherein, the number of hidden layers in the deep learning network model is determined based on the number of features of one sample infrared image;

[0033] Wherein, the image processing model is used to determine the feature deviation value of the visible light image corresponding to the face including the eye according to the feature deviation value of the infrared image corresponding to the eye to be processed.

[0034] In a possible design, the sample eye is at least one; the determining the training sample set includes:

[0035] Obtain the features of the sample image, where the sample image is taken under different action states of each sample eye; for the same eye action state of each sample eye, the features of the sample image include the features of the sample infrared image corresponding to the sample eye and the features of the sample visible light image corresponding to the sample face;

[0036] According to the features of the sample infrared image corresponding to each sample eye and the corresponding features of the sample visible light image, respectively determine the feature deviation value of the sample infrared image corresponding to each sample eye and the feature deviation value of the sample visible light image corresponding to each sample face;

[0037] Determine the training sample set according to the feature deviation value of the sample infrared image corresponding to each sample eye and the feature deviation value of the sample visible light image corresponding to each sample face.

[0038] In a possible design, the respectively determining the feature deviation value of the sample infrared image corresponding to each sample eye and the feature deviation value of the sample visible light image corresponding to each sample face according to the features of the sample infrared image corresponding to each sample eye and the corresponding features of the sample visible light image includes:

[0039] For each sample eye, select the features of any sample infrared image as the features of the reference sample infrared image from the features of the sample infrared images in different action states corresponding to the sample eye, and select the features of the sample visible light image in the same action state as the reference sample infrared image from the features of the sample infrared images in different action states corresponding to the sample face as the features of the reference sample visible light image;

[0040] Take the difference between the features of the reference sample infrared image corresponding to the sample eye and the features of the other sample infrared images corresponding to the sample eye as the feature deviation value of the sample infrared image, and take the difference between the features of the reference sample visible light image corresponding to the sample face and the features of the other sample visible light images corresponding to the sample face as the feature deviation value of the sample infrared visible light image.

[0041] In a possible design, the number of hidden layers in the deep learning network model is determined based on the number of features of a sample infrared image; the learning rate of the deep learning network model is associated with the number and distribution of the feature deviation values of the input and output of the deep learning network model.

[0042] In a third aspect, an embodiment of the present application provides an image transformation device, and the device includes:

[0043] An acquisition module, configured to acquire the feature deviation value of the infrared image corresponding to the current eye to be processed;

[0044] A processing module, configured to obtain the feature deviation value of the visible light image corresponding to the face including the current eye through an image processing model according to the feature deviation value of the infrared image corresponding to the current eye, and the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model.

[0045] In a fourth aspect, an embodiment of the present application provides a training device for an image processing model, and the device includes:

[0046] A sample determination module, configured to determine a training sample set, where the training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face; wherein, the sample face includes the sample eye, and the feature deviation value of the sample visible light image is used as the true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye;

[0047] A training module, configured to train a deep learning network model according to the training sample set to obtain an image processing model;

[0048] Among them, the image processing model is used to determine the feature deviation value of the visible light image corresponding to the face including the eye according to the feature deviation value of the infrared image corresponding to the eye to be processed.

[0049] In a fifth aspect, an embodiment of the present application provides an image transformation device, including: an image transformation device and a data acquisition device; the image transformation device is used to execute the method according to any one of the first aspect; wherein, the data acquisition device includes an infrared camera, a visible light camera, and a transparent frame; the infrared camera is installed on the transparent frame;

[0050] The infrared camera is used to collect infrared images of the eye at any moment or in any eye movement state;

[0051] The visible light camera is used to collect visible light images of the eye located within the transparent frame at any moment or in any eye movement state.

[0052] In a sixth aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory;

[0053] The memory stores computer-executable instructions;

[0054] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the method according to any one of the first aspect and the second aspect.

[0055] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the method according to any one of the first aspect and the second aspect is implemented.

[0056] In an eighth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of the first aspect and the second aspect is implemented.

[0057] The image transformation method, the training method of the image processing model, the device and the equipment provided in this embodiment first obtain the feature deviation value of the infrared image corresponding to the current eye to be processed; then, according to the feature deviation value of the infrared image corresponding to the current eye, through the image processing model, obtain the feature deviation value of the visible light image corresponding to the face including the current eye, and the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model. By using the image processing model obtained by training the deep learning model to realize image transformation, it is possible to reduce the image processing error caused by calculation through the mapping function in the prior art. At the same time, since the input of the image processing model is the feature deviation value, the error problems brought by the data itself can be weakened mutually, and by corresponding the features of the infrared image with fewer quantities to the visible light image with more quantities, the details of the output image are increased. Therefore, a more accurate image can be provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 Schematic diagram of the image transformation device provided in the embodiment of the present application;

[0060] Figure 2 Schematic flow chart of the image transformation method provided in the embodiment of the present application;

[0061] Figure 3 Schematic flow chart of the image transformation method provided in another embodiment of the present application;

[0062] Figure 4 Schematic flow chart of the image transformation method provided in yet another embodiment of the present application;

[0063] Figure 5 Schematic flow chart of the image transformation method provided in still another embodiment of the present application;

[0064] Figure 6 Schematic diagram of the scenario of the image transformation method provided in the embodiment of the present application;

[0065] Figure 7 Schematic structural diagram of the image transformation device provided in the embodiment of the present application;

[0066] Figure 8 Schematic structural diagram of the electronic device provided in the embodiment of the present application. Detailed implementation manners

[0067] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0068] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0069] Currently, in traditional image display, infrared images are commonly used in image fields with high precision requirements such as eye movement tracking. Generally, it is necessary to establish a connection between the key point information extracted from the infrared images and the world coordinate system. Conventional methods include establishing a projection transformation matrix between two images by finding the positions of corresponding key points in two sets of images and performing coordinate correspondence through projection transformation. However, due to the factor of parallax in the depth direction, this method has low precision, large errors, and poor robustness. Therefore, there are large errors in the existing image transformation, resulting in low precision of image processing, and thus it is impossible to provide accurate images.

[0070] Therefore, in view of the above problems, the technical concept of the present application is to train an image processing model using a deep learning model, and realize image transformation by using the image processing model, which can reduce the image processing errors caused by calculation through a mapping function in the prior art. At the same time, in order to further improve the precision, the features of a smaller number of infrared images are mapped to a larger number of visible light images, increasing the details of the output image. And in order to reduce the errors brought by the data itself, the feature deviation value can be used as the input quantity of the image processing model, thereby being able to mutually weaken the error problems brought by the data itself, and further providing relatively accurate images.

[0071] In practical applications, as shown in Figure 1 shown Figure 1Schematic diagram of an image transformation device provided by an embodiment of the present application. The image transformation device can be a head-mounted device, including a data acquisition device and an image transformation device. Among them, the data acquisition device includes an infrared camera 101, a visible light camera 102 (which can be a front visible light camera, such as an RGB camera, a mobile phone, a single-lens reflex camera, etc., hereinafter taking the RGB camera as an example), and a transparent frame 103. The infrared camera is installed on the transparent frame. The infrared camera is used to collect infrared images of the eyes at any moment or in any eye movement state. The RGB camera is used to collect visible light images of the eyes located within the transparent frame at any moment or in any eye movement state. Specifically, as Figure 1 shown, by making the partial structure of the eyes facing the person being collected transparent (such as Figure 1 the transparent frame 103 in, which contains two transparent lens pieces), an infrared camera 101 (at the circled position) is arranged below the side of the eyes. After the person being collected wears the device (head-mounted device), the infrared camera of the device is started to collect data, and at the same time, a visible light camera is used to take pictures on the front of the device. Since the device part of the eyes is a transparent structure, the key points of the front visible light image structure of the eyes can be extracted.

[0072] Combined with Figure 2 shown, Figure 2 Schematic flow diagram of an image transformation method provided by an embodiment of the present application. The application scenario of this process includes a data acquisition device and an image transformation device (such as a computing unit, including a first computing unit and / or a second computing unit). The process of image processing in this scenario includes data acquisition, data preprocessing, (image processing) model training, and model inference application (here referring to the application of image transformation). Among them, data preprocessing, (image processing) model training, and model inference application can be performed by the first computing unit, or can be performed by the second computing unit, or can also be performed by combining the first computing unit and the second computing unit to perform the corresponding operations, which are not specifically limited here.

[0073] Exemplarily, taking the operation of the first computing unit to perform data preprocessing and model inference applications and the operation of the second processing unit to perform model training as an example, the specific process is as follows: First, an infrared image corresponding to the eye is obtained through the infrared camera in the data acquisition device, and a visible light image corresponding to the face including the eye is obtained through the visible light camera in the data acquisition device. Key points of the infrared image and the visible light image are respectively extracted. Then, after obtaining the corresponding key points, the first computing unit (for example, a CPU computing unit, hereinafter taking the first computing unit as a CPU computing unit as an example) can perform data preprocessing because there are errors in the technology of the initial features of the extracted visible light image (and infrared image) itself. In order to improve the image processing accuracy. Then, the CPU computing unit trains a deep learning network model (i.e., a deep learning model) based on the preprocessed data to obtain an image processing model.

[0074] In the actual application process, exemplarily, as shown in Figure 3 Model training is realized through a model training framework. Samples are collected by using a data acquisition device (including an infrared camera, a visible light camera (such as a front visible light camera), and a transparent glasses frame) to construct a training sample set, and model training is realized through a second computing unit (for example, a GPU computing unit, hereinafter taking the second computing unit as a GPU computing unit as an example). During use, infrared eye movement data (here referring to the infrared image corresponding to the eye and / or the features of the infrared image corresponding to the eye) is collected in real time through the infrared camera. The CPU computing unit performs image transformation on the infrared eye movement data through a key point conversion mapping model to obtain key points of the front posture, and then the GPU computing unit realizes expression generation.

[0075] Specifically, corresponding infrared images and visible light images can be obtained through the data acquisition device, and corresponding key points are extracted. Based on the extracted key points, data preprocessing is performed to obtain preprocessed data (here referring to the feature deviation value of the infrared image corresponding to the eye). The feature deviation value of the infrared image corresponding to the eye is input into the image processing model, and the feature deviation value of the visible light image corresponding to the face including the eye is output. The features of a smaller number of infrared images are corresponding to a larger number of visible light images, which increases the details of the output image. At the same time, it can mutually weaken the error problem brought by the data itself, thereby providing a more accurate image. Through the feature deviation value, the change in the shape of the current eye compared to the shape of the eye in the reference visible light image and the expression change of the face can be determined, thereby improving the accuracy of image processing.

[0076] In this application, by determining the reference feature values under different acquisition conditions and using the feature deviation values as the input of the model, the generalization applicability of the image processing model for different scenarios can be realized, the robustness of the image processing model itself can be improved, and the model can be applied to different scenarios. For example, for infrared light, the position of the head-mounted display (here refers to the image transformation device, such as: head-mounted device) may be different each time, and for visible light, the shooting angle can also be different. In addition, it can be applied to different people. That is, each time the scenario is different, the position of the current object and the reference data of the person are collected as a basis.

[0077] The technical solution of this application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0078] Figure 4 The following is a schematic flowchart of an image transformation method provided by another embodiment of this application. The image transformation method may include:

[0079] S401. Obtain the feature deviation value of the infrared image corresponding to the current eye to be processed.

[0080] In this embodiment, the execution subject may be an image transformation method device, which is installed in an image transformation device. The image transformation device may be an electronic device or a head-mounted device. The image transformation method device may also be configured in a server, and the server may be communicatively connected to a data acquisition device. The data acquisition device may include an infrared camera 101, a visible light camera 102, and a transparent frame 103. The structure of the data acquisition device will not be described in detail below.

[0081] Specifically, the feature deviation value of the infrared image corresponding to the current eye can be obtained through at least the following two methods:

[0082] Method 1: Obtain the infrared image corresponding to the current eye collected by the infrared camera, and extract the features of the infrared image.

[0083] Method 2: Directly obtain the features of the infrared image provided by the infrared camera, and the features of the infrared image are extracted after the infrared camera collects the infrared image corresponding to the current eye.

[0084] In a possible design, the obtaining of the feature deviation value of the infrared image corresponding to the current eye may include:

[0085] Obtain the features of the infrared image corresponding to the current eye and the features of the reference infrared image corresponding to the eye; the features of the reference infrared image corresponding to the eye are the features of the infrared image in any eye movement state corresponding to the eye.

[0086] The difference between the features of the reference infrared image corresponding to the eye and the features of the current infrared image corresponding to the eye is used as the feature deviation value of the current infrared image corresponding to the eye.

[0087] In this embodiment, first, the features of the infrared image corresponding to the current eye are obtained, and the features of the infrared image in any eye movement state corresponding to the eye are obtained as the features of the reference infrared image corresponding to the eye. Then, the difference between the features of the infrared image corresponding to the current eye and the features of the reference infrared image corresponding to the eye is calculated to obtain the feature deviation value of the infrared image corresponding to the current eye. By using the feature deviation value as the input of the model, the error problems brought by the data itself (such as the features of infrared images and visible light images) can be mutually weakened, thereby improving the accuracy of the model output result.

[0088] S402. According to the feature deviation value of the infrared image corresponding to the current eye, through an image processing model, obtain the feature deviation value of the visible light image corresponding to the face including the current eye; the image processing model is obtained by training a deep learning model.

[0089] In this embodiment, the feature deviation value of the infrared image corresponding to the current eye is input into the trained image processing model, and the feature deviation value of the visible light image corresponding to the face including the eye is output. The features of the infrared image with a smaller quantity are corresponding to the visible light images with a larger quantity, increasing the details of the output image. At the same time, the error problems brought by the data itself can be mutually weakened, thereby providing a precise image. Through the feature deviation value, the change in the shape of the current eye compared to the shape of the eye in the reference visible light image and the facial expression change can be determined, thereby improving the accuracy of image processing.

[0090] Among them, the training process of the image processing model is as follows: First, a training sample set is determined. The training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face. Among them, the sample face includes the sample eye, and the feature deviation value of the sample visible light image is used as the true value for training. The number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye. Then, according to the training samples, a deep learning network model is trained to obtain an image processing model. Among them, the number of hidden layers in the deep learning network model is determined based on the number of features of a sample infrared image, that is, an appropriate hidden layer is selected according to the number of features of the infrared image.

[0091] In a possible design, the learning rate of the image processing model is associated with the number and distribution of the feature deviation values of the input and output of the image processing model.

[0092] Among them, promote model convergence; where w represents the weight parameter, η represents the learning rate, represents the gradient of the loss function L with respect to w. Therefore, the update of the weight parameter is related to the learning rate, and the setting of the learning rate is related to the number and distribution of feature points of the input and output. A reasonable learning rate is set to make the model converge faster.

[0093] The image transformation method provided by the embodiments of the present application first obtains the feature deviation value of the infrared image corresponding to the current eye to be processed; then, according to the feature deviation value of the infrared image corresponding to the current eye, through an image processing model, obtains the feature deviation value of the visible light image corresponding to the face including the current eye, and the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model. By using the image processing model trained by the deep learning model to implement image transformation, it is possible to reduce the image processing error caused by calculating through the mapping function in the prior art. At the same time, since the input of the image processing model is the feature deviation value, the error problems brought by the data itself can be weakened with each other, and the features of the infrared image with a smaller number are corresponding to the visible light image with a larger number, increasing the details of the output image. Therefore, a more accurate image can be provided.

[0094] In a possible design, the method further includes:

[0095] obtain the features of the reference visible light image corresponding to the face including the eye; where the reference infrared image corresponding to the eye and the reference visible light image corresponding to the face are collected at the same time or when the eye corresponds to the same eye movement state;

[0096] generate a facial expression according to the feature deviation value of the visible light image corresponding to the current face including the eye and the features of the reference visible light image corresponding to the face including the eye.

[0097] Generally, the image in which the eyes face the camera directly is selected as the reference image (including the reference infrared image and the reference visible light image). Specifically, first obtain the visible light image corresponding to the face collected at the same time as the reference infrared image corresponding to the eye or when the eye corresponds to the same eye movement state, as the features of the reference visible light image corresponding to the face including the eye. According to the feature deviation value of the visible light image corresponding to the current face including the eye and the features of the reference visible light image corresponding to the face including the eye, obtain the state value of the visible light image corresponding to the current face including the eye, which can intuitively reflect the specific posture of the current eye feature, and can determine the transformation situation of the face including the eye, and then generate a facial expression.

[0098] For the image processing model, since the input is the feature deviation value of the infrared image corresponding to the current eye, and the output is the feature deviation value of the visible light image corresponding to the face containing the eye, therefore, by means of network learning, a smaller number of infrared key points are mapped to a larger number of frontal visible light images, increasing the details of the output image, thereby improving the accuracy of the image and making the generated facial expression more accurate.

[0099] In a possible design, obtaining the features of the reference visible light image corresponding to the face containing the eye includes:

[0100] Extracting the initial features of the reference visible light image corresponding to the face containing the eye, and extracting the initial features of the target visible light image corresponding to the face containing the eye, where the reference visible light image is the visible light image corresponding to the eye in any eye movement state, and the target visible light image is the visible light image corresponding to the eye at other times or in other eye movement states;

[0101] Identifying a first region and a second region in the target visible light image, where the first region in the target visible light image contains feature points that change with time or eye movement state, and the second region in the target visible light image contains reference points that do not change with time or eye movement state or change within a preset condition, and both the reference points and the feature points are used to represent pixel points, and one pixel point corresponds to two features;

[0102] Adjusting the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face containing the eye.

[0103] In this embodiment, since there are errors in the technology of extracting the initial features of the visible light image itself, in order to improve the image processing accuracy, data preprocessing can be performed on the obtained visible light image to remove jitter through data preprocessing. The preprocessing process for the reference visible light image corresponding to the face containing the eye is as follows:

[0104] First, extract the initial features of the visible light image corresponding to the eyes of the face containing the eyes in any eye movement state (i.e., the reference visible light image corresponding to the face containing the eyes), and extract the initial features of the visible light image corresponding to the eyes at other times or in other eye movement states, that is, the initial features of the target visible light image. Then, by identifying the first region containing feature points that change with time or eye movement state and the second region containing reference points that do not change with time or eye movement state or change within a preset condition in the target visible light image, according to the initial features of the reference points of the target visible light image in the second region, adjust the initial features of the feature points of the reference visible light image (here, it may include the feature points in the first region and other feature points outside the corresponding first region), so as to obtain the features of the reference visible light image corresponding to the face containing the eyes. Thereby, the obvious jitter phenomenon of the feature points is eliminated, the systematic error of the transformation is reduced, and further the accuracy of the data for image processing (here refers to image transformation or image processing model training, etc.) is improved.

[0105] It should be noted that in the model training, the preprocessing process of the sample visible light image corresponding to the face containing the eyes and the preprocessing process of the sample reference visible light image corresponding to the face containing the eyes are the same as or similar to the above-mentioned preprocessing process of the reference visible light image corresponding to the face containing the eyes, and will not be elaborated here.

[0106] In a possible design, the adjusting the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face containing the eyes includes:

[0107] Determine the first region and the boundary points of the first region in the target visible light image according to the reference points of the target visible light image, and determine the boundary points of the second region in the target visible light image according to the feature points of the target visible light image;

[0108] Perform smoothing processing on the second region in the target visible light image through a filtering method to obtain the visible light smoothed image corresponding to the second region in the target visible light image;

[0109] Calculate the position change information of the reference visible light image relative to the target visible light image according to the initial features of the target reference points in the visible light smoothed image;

[0110] According to the position change information, adjust the initial features of each feature point in the first region of the reference visible light image to obtain the features of the reference visible light image of the face containing the eyes.

[0111] In this embodiment, jitter is removed through data preprocessing to reduce the systematic error of image transformation. Combining with Figure 5 as shown in Figure 5 FIG. 4 is a schematic flowchart of an image transformation method provided by another embodiment of the present application. First, for the original image A (here referring to an infrared image or a visible light image at any moment or eye movement state, hereinafter taking the target visible light image as an example, the process of data preprocessing to remove jitter will be described in detail), a predefined reference point region C (here referring to the second region), a region of interest B (here referring to the first region), and feature points of interest (for example, feature points that change with time or eye movement state) are determined, and all points of interest are included in the region of interest B. For the region C, a mean filtering operation is performed to obtain the reference point coordinates (the pixel values meet certain conditions). For the region of interest B, the coordinates of the new region of interest relative to the reference point are given in combination with the reference point position to obtain a new region of interest B1. Based on the extracted feature points of interest and the new region of interest B1, the new coordinates of the original feature points (here referring to the extracted feature points of interest) relative to the new region of interest are obtained as the correctly labeled data GroundTruth.

[0112] Exemplarily, when extracting the feature point L as GroundTruth in the original image A by a third-party library, due to errors, the position of the feature point L deviates, resulting in an obvious jitter phenomenon. To solve the problems brought by the jitter phenomenon, first, the region of interest B in the original image A is determined, and points that change relatively little (points with obvious features, such as special color marks, or characteristic spots or moles of the person himself) are preselected as reference points when the eyes and eyebrows change to eliminate part of the jitter phenomenon. To accurately obtain the reference points, first, a larger rectangular region C (here referring to the region C) (C_left, C_right, C_upper, C_lower) where the reference points are located is obtained. Refer to Figure 6 as shown in Figure 6 FIG. 5 shows a schematic diagram of the region C in the original image A. Among them, the C region (i.e., the region C) can be independent of the B region (i.e., the region of interest B), and the reference points cannot fall within the B region because the points in the B region are constantly changing. Generally, the C region is much smaller than the B region. Then, a filtering method (in this embodiment, a 7*7 mean filter can be used) is adopted to smooth the image to obtain an image D (here referring to a visible light smoothed image), and then the specific coordinates (x_r, y_r) of a relatively special (obvious features: the minimum, maximum, or average value of the pixels can be used) feature point r in the D region (i.e., the region where the image D is located) are calculated. The coordinates of the reference point r in the image A are:

[0113] x_r_A = C_left + x_r

[0114] y_r_A = C_upper + y_r

[0115] Then, adjust the size of the region of interest B according to the coordinates of the reference point to obtain a new region of interest B1, and adjust the position L of the feature point to the relative coordinates with respect to the B1 region (i.e., the new region of interest B1), thereby eliminating the obvious jitter phenomenon of the feature point.

[0116] Therefore, by using the data acquisition device to obtain the true value of the training data (here referring to the preprocessed sample data) or the data to be processed (here referring to the image data to be processed, such as the features of the infrared image corresponding to the current eye to be processed), the data corresponding accuracy is high, and the image transformation result is relatively accurate.

[0117] In a possible design, obtaining the features of the infrared image corresponding to the current eye includes:

[0118] Extracting the initial features of the infrared image corresponding to the current eye, and using the initial features of the infrared image corresponding to the current eye as the features of the infrared image corresponding to the current eye; or,

[0119] Extracting the initial features of the infrared image corresponding to the current eye and the initial features of the target infrared image corresponding to the eye, where the infrared image corresponding to the current eye is the infrared image corresponding to the eye in any eye movement state, and the target infrared image corresponding to the eye is the infrared image corresponding to the eye at other times or in other eye movement states;

[0120] Identifying a first region and a second region in the target infrared image, where the first region in the target infrared image contains feature points that change with time or eye movement state, and the second region in the target infrared image contains reference points that do not change with time or eye movement state or change within a preset condition, and both the reference points and the feature points are used to represent pixel points, and one pixel point corresponds to two features;

[0121] Adjusting the initial features of the feature points of the infrared image corresponding to the current eye according to the initial features of the reference points of the target infrared image to obtain the features of the infrared image corresponding to the current eye.

[0122] In this embodiment, since the error of the feature extraction of the infrared image is usually small, it is also possible not to perform correction (here referring to the operation of data preprocessing to remove jitter). Therefore, the initial features of the infrared image corresponding to the current eye can be directly used as the features of the infrared image corresponding to the current eye for subsequent use.

[0123] To further improve the accuracy, the initial features of the infrared image can be corrected, and this correction process can refer to the embodiment shown above in combination with Figure 5 which is not described in detail here.

[0124] The present application also provides a method for training an image processing model. The execution subject of the method for training the image processing model may be a server or an image transformation device, and the image transformation device may be an electronic device or a head-mounted device. The method may include:

[0125] Determine a training sample set, where the training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face. Wherein, the sample face includes the sample eye, the feature deviation value of the sample visible light image is used as the true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye;

[0126] Train a deep learning network model according to the training sample set to obtain an image processing model. Wherein, the number of hidden layers in the deep learning network model is determined based on the number of features of one sample infrared image;

[0127] Wherein, the image processing model is used to determine the feature deviation value of the visible light image corresponding to the face including the eye according to the feature deviation value of the infrared image corresponding to the eye to be processed.

[0128] In this embodiment, first, a training sample set for training a deep learning network model is determined. The training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face. Wherein, the sample face includes the sample eye, the feature deviation value of the sample visible light image is used as the true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye. Then, according to the training samples, a deep learning network model is trained to obtain an image processing model. Wherein, the number of hidden layers in the deep learning network model is determined based on the number of features of one sample infrared image, that is, a suitable hidden layer is selected according to the number of features of the infrared image.

[0129] In a possible design, the number of hidden layers in the deep learning network model is determined based on the number of features of one sample infrared image; the learning rate of the deep learning network model is associated with the number and distribution of the feature deviation values of the input and output of the deep learning network model.

[0130] In this embodiment, the learning rate of the image processing model is associated with the number and distribution of the feature deviation values of the input and output of the image processing model. Promote model convergence. Wherein, w represents the weight parameter, and η represents the learning rate. It represents the gradient of the loss function L with respect to w. Therefore, the update of the weight parameter is related to the learning rate, and the setting of the learning rate is related to the number and distribution of feature points of the input and output. A reasonable learning rate is set to make the model converge faster.

[0131] Exemplarily, two sets of image coordinate correspondence relationships are established through a deep learning network (for example, using a Multilayer Perceptron (MLP), also called an artificial neural network, to learn the relationship of image transformation):

[0132] Input: 96 features (one key point or pixel point corresponds to two coordinates, and one coordinate is one feature); Hidden layer: 4 layers, with sizes of 400, 200, 200, 200; Output: 204 features;

[0133] Optimizer: sgd optimization; Error: MSE mean square error;

[0134] Input: 48 key points of both eyes (or the deviation values of 48 key points of both eyes), and each point is the relative coordinate with respect to the inner canthus;

[0135] Output: The deviation coordinates of 102 key points of both eyes. The output is the deviation value with respect to the normal eye-opening state, and it can adapt to multiple output objects according to the initial shapes of different eyes of the input.

[0136] Through a deep learning network that establishes the coordinate transformation relationship of different numbers of key points in two sets of images, training the deep learning network model, the obtained image processing model can achieve image transformation and can reduce the image processing error caused by calculation through a mapping function in the prior art.

[0137] In the training method of the image processing model provided in the embodiments of the present application, the feature deviation value of the infrared image and the feature deviation value of the corresponding sample visible light image of the face are used as training samples to train the deep learning network model to obtain an image processing model, which can reduce the image processing error caused by calculation of the mapping function; in addition, since the training samples are calculated based on the deviation values, the error problems brought by the data itself can be mutually weakened. At the same time, training the model based on the feature deviation values obtained under any eye state can be applicable to face detection under any eye state, and by corresponding the features of a smaller number of infrared images to a larger number of visible light images, the details of the output image are increased. Therefore, a more accurate image can be provided.

[0138] In a possible design, the sample eye part is at least one; the determining of the training sample set includes:

[0139] Obtain the features of the sample images, where the sample images are captured under different action states of each sample eye; for the same eye action state of each sample eye, the features of the sample images include the features of the sample infrared image corresponding to the sample eye and the features of the sample visible light image corresponding to the sample face;

[0140] According to the features of the sample infrared images corresponding to each sample eye and the features of the corresponding sample visible light images, respectively determine the feature deviation values of the sample infrared images corresponding to each sample eye and the feature deviation values of the sample visible light images corresponding to each sample face;

[0141] Determine the training sample set according to the feature deviation values of the sample infrared images corresponding to each sample eye and the feature deviation values of the sample visible light images corresponding to each sample face.

[0142] In this embodiment, in order to be applicable to the shapes of different eyes in the input, when constructing the training samples, the infrared images corresponding to the eyes under different action states of each sample eye and the visible light images corresponding to the faces containing the eyes can be collected, and the corresponding features can be extracted respectively, that is, the features of the sample infrared images corresponding to the sample eyes and the features of the sample visible light images corresponding to the sample faces containing the sample eyes. Then, according to the features of the sample infrared images corresponding to each sample eye and the features of the corresponding sample visible light images, respectively determine the feature deviation values of the sample infrared images corresponding to each sample eye and the feature deviation values of the sample visible light images corresponding to each sample face, and further generate the training sample set.

[0143] Therefore, the image processing model trained by the training sample set constructed in this way can be applicable to the image changes under different eye action states, and further ensure the accuracy of the output results corresponding to the input applicable to different eye action states.

[0144] In a possible design, the step of respectively determining the feature deviation values of the sample infrared images corresponding to each sample eye and the feature deviation values of the sample visible light images corresponding to each sample face according to the features of the sample infrared images corresponding to each sample eye and the features of the corresponding sample visible light images includes:

[0145] For each sample eye, select the features of any sample infrared image from the features of the sample infrared images corresponding to different action states of the sample eye as the features of the reference sample infrared image, and select the features of the sample visible light image in the same action state as the reference sample infrared image from the features of the sample infrared images corresponding to different action states of the sample face as the features of the reference sample visible light image;

[0146] The difference between the features of the reference sample infrared image corresponding to the sample eye and the features of other sample infrared images corresponding to the sample eye is used as the feature deviation value of the sample infrared image, and the difference between the features of the reference sample visible light image corresponding to the sample face and the features of other sample visible light images corresponding to the sample face is used as the feature deviation value of the sample infrared visible light image.

[0147] In this embodiment, in order to mutually weaken the error problem brought by the data itself, the feature deviation value can be used as a training sample. First, from the features of the sample infrared images in different action states corresponding to the sample eye, select the features of any sample infrared image as the features of the reference sample infrared image, and from the features of the sample infrared images in different action states corresponding to the sample face, select the features of the sample visible light image in the same action state as the reference sample infrared image as the features of the reference sample visible light image. Usually, an image with the eyes facing directly at the camera can be selected as the reference sample image (including the reference sample infrared image and the reference sample visible light image). In this way, after the output result, a facial expression can be directly generated according to the output visible light feature deviation value corresponding to the face, which is convenient for calculation and simplifies the processing logic.

[0148] Then, calculate the difference between the features of the reference sample infrared image corresponding to the sample eye and the features of other sample infrared images corresponding to the sample eye as the feature deviation value of the sample infrared image, and calculate the difference between the features of the reference sample visible light image corresponding to the sample face and the features of other sample visible light images corresponding to the sample face as the feature deviation value of the sample infrared visible light image, thereby generating a sample training set.

[0149] Exemplarily, establish two groups of image coordinate correspondence relationships through a deep learning network:

[0150] Input: 48 key points for both eyes, each point is the relative value of the inner corner of the eye position, with two dimensions x and y, a total of 96 values as the input to the network input layer;

[0151] After entering the network input layer, use 4 intermediate hidden layers to replace the traditional interpolation mapping function to implement the coordinate system conversion and the interpolation process of mapping from fewer points to more points. Among them, the role of the hidden layer is to learn the feature representation in the form of a neural network. Too few hidden layers are not conducive to describing relatively complex input features, and too many hidden layers will make the network too complex and large. Therefore, it is necessary to select an appropriate hidden layer according to the number of features for feature representation.

[0152] Output: 204 feature values for 102 key points of both eyes, each key point contains two coordinates x and y.

[0153] The SGD optimization method can be adopted: (Stochastic Gradient Descent)

[0154] 1. Set the loss function of the neural network to be minimized: min L(w),

[0155] where x i is the corresponding value of the output feature vector, and x j is the actual true value corresponding to x i and can be obtained through the above data acquisition device;

[0156] 2. Parameter update (The learning rate is related to the number and distribution of eigenvalues of the input and output.)

[0157]

[0158] where w represents the weight parameter, η represents the learning rate, represents the gradient of the loss function L with respect to w.

[0159] 3. Traversal:

[0160] (1) Calculate the gradient of the loss function with respect to the weight w

[0161] (2) Update the weight parameter w using the above formula;

[0162] (3) Repeat the above steps until the loss function converges.

[0163] Therefore, the update of the weight parameter is related to the learning rate, and the setting of the learning rate is related to the number and distribution of feature points of the input and output. A reasonable learning rate is set to make the model converge faster.

[0164] In a possible design, the obtaining of the features of the sample image includes:

[0165] For each sample eye part, respectively extract the initial features of the sample infrared image corresponding to different action states of the sample eye part and the initial features of the sample visible light image corresponding to the sample face;

[0166] Determine the target sample visible light image from the sample visible light images corresponding to different eye action states, and identify the first region and the second region in the target sample visible light image. The target sample visible light image is the sample visible light image corresponding to the sample eye part in any eye action state. The first region contains feature points that change with time or eye action state, and the second region contains reference points that do not change with time or eye action state or change within a preset condition. The reference points and the feature points are both used to represent pixel points, and one pixel point corresponds to two features;

[0167] Adjust the initial features of the feature points of each sample visible light image according to the initial features of the reference points of the target sample visible light image to obtain the features of the sample visible light image corresponding to the sample face;

[0168] Determine the features of the sample image according to the initial features of the sample infrared image and the features of the sample visible light image.

[0169] In this embodiment, since there are errors in the technology of the initial features of the extracted visible light images, data preprocessing can be performed on the obtained sample visible light images and / or sample infrared images, and the data preprocessing can be used to remove jitter and improve the image processing accuracy.

[0170] Specifically, taking a certain sample eye and the corresponding sample face as an example, extract the initial features of the sample infrared image corresponding to different action states of the sample eye and the initial features of the sample visible light image corresponding to the sample face. Select the target sample visible light image from the sample visible light images corresponding to different eye action states; then identify the first region including the feature points that change with time or eye action state and the second region including the reference points that do not change with time or eye action state or change within a preset condition in the target sample visible light image. According to the initial features of the reference points of the target sample visible light image in the second region, adjust the initial features of the feature points of each sample visible light image to obtain the features of the sample visible light image corresponding to the sample face, and determine the features of the sample image. Thereby, the obvious jitter phenomenon of the feature points is eliminated, the systematic error of the transformation is reduced, and the accuracy of the data for image processing (here referring to the training of the image processing model) is improved.

[0171] It should be noted that for the preprocessing process of the sample visible light image corresponding to the face including the eye and the preprocessing process of the sample reference visible light image corresponding to the face including the eye during model training are both the same as or similar to the above-mentioned preprocessing process of the reference visible light image corresponding to the face including the eye, and will not be elaborated here.

[0172] In a possible design, the adjusting the initial features of the feature points of each sample visible light image according to the initial features of the reference points of the target sample visible light image to obtain the features of the sample visible light image corresponding to the sample face includes:

[0173] Determine the first region and the boundary points of the first region in the target sample visible light image according to the reference points of the target sample visible light image, and determine the boundary points of the second region in the target sample visible light image according to the feature points of the target sample visible light image;

[0174] By means of filtering, smooth processing is performed on the second region in the visible light image of the target sample to obtain a sample visible light smoothed image corresponding to the second region in the visible light image of the target sample;

[0175] According to the initial features of the target reference points in the sample visible light smoothed image, traverse the sample visible light images corresponding to other action states of the sample face:

[0176] Calculate the position change information of the sample visible light image corresponding to the other action state relative to the target sample visible light image;

[0177] According to the position change information, by adjusting the initial features of the feature points in the first region of the sample visible light image corresponding to the other action state, obtain the features of the sample visible light image corresponding to the other action state.

[0178] In this embodiment, combined with Figure 5 As shown, the training data true value (here referring to the preprocessed sample data) is obtained through the data acquisition device, the data has high corresponding accuracy, and the image transformation result is relatively accurate. For the specific implementation process, reference can be made to the embodiments of the image transformation method, which will not be elaborated here.

[0179] In a possible design, the determining the features of the sample image according to the initial features of the sample infrared image and the features of the sample visible light image includes:

[0180] Taking the initial features of the sample infrared image as the features of the sample infrared image, and forming the features of the sample image by the features of the sample infrared image and the features of the sample visible light image; or,

[0181] Determine a target sample infrared image from the sample infrared images corresponding to different eye action states, and identify the first region and the second region in the target sample infrared image. The first region in the target sample infrared image contains the feature points, and the second region in the target sample infrared image contains the reference points;

[0182] According to the initial features of the reference points in the target sample infrared image, adjust the initial features of the feature points in other sample infrared images to obtain the features of the sample visible infrared image corresponding to the sample eyes;

[0183] Wherein, the features of the sample image are formed by the features of the sample infrared image and the features of the sample visible light image.

[0184] In this embodiment, since the error in feature extraction of infrared images is usually small, it is also possible not to perform correction (here, the operation of data preprocessing to remove jitter). Therefore, the initial features of the sample infrared image corresponding to the sample eye can be directly used as the features of the sample infrared image corresponding to the sample eye for subsequent use.

[0185] To further improve the accuracy, the initial features of the sample infrared image can be corrected. The correction process can refer to the embodiment described above in combination with Figure 5 shown, and will not be elaborated here.

[0186] To implement the image transformation method, this embodiment provides an image transformation device. Refer to Figure 7 , Figure 7 which is the structural schematic diagram of the image transformation device provided by the embodiment of the present application; the image transformation device includes: an acquisition module 701 and a processing module 702.

[0187] Among them, the acquisition module 701 is used to acquire the feature deviation value of the infrared image corresponding to the current eye to be processed;

[0188] The processing module 702 is used to obtain the feature deviation value of the visible light image corresponding to the face including the current eye through an image processing model according to the feature deviation value of the infrared image corresponding to the current eye. The number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model.

[0189] In this embodiment, through the acquisition module 701 and the processing module 703, the feature deviation value of the infrared image corresponding to the current eye to be processed is obtained; then, according to the feature deviation value of the infrared image corresponding to the current eye, the feature deviation value of the visible light image corresponding to the face including the current eye is obtained through an image processing model. The number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model. By using the image processing model trained by the deep learning model to implement image transformation, it is possible to reduce the image processing error caused by calculation through a mapping function in the prior art. At the same time, since the input of the image processing model is the feature deviation value, it is possible to mutually weaken the error problem brought by the data itself. And by mapping the features of the infrared image with a smaller number to the visible light image with a larger number, the details of the output image are increased. Therefore, a more accurate image can be provided.

[0190] The image transformation device provided in this embodiment can be used to execute the technical solutions of the above image transformation method embodiment. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0191] In a possible design, the obtaining module is specifically configured to:

[0192] Obtain the features of the infrared image corresponding to the eye currently and the features of the reference infrared image corresponding to the eye; the features of the reference infrared image corresponding to the eye are the features of the infrared image in any eye movement state corresponding to the eye;

[0193] Take the difference between the features of the reference infrared image corresponding to the eye and the features of the infrared image corresponding to the eye currently as the feature deviation value of the infrared image corresponding to the eye currently.

[0194] In a possible design, the device further includes: a generating module; the generating module includes a first processing unit and a second processing unit;

[0195] The first processing unit is configured to obtain the features of the reference visible light image corresponding to the face including the eye; wherein, the reference infrared image corresponding to the eye and the reference visible light image corresponding to the face are collected at the same moment or in the same eye movement state corresponding to the eye;

[0196] The second processing unit is configured to generate a facial expression according to the feature deviation value of the visible light image corresponding to the face including the eye currently and the features of the reference visible light image corresponding to the face including the eye.

[0197] In a possible design, the first processing unit is specifically configured to:

[0198] Extract the initial features of the reference visible light image corresponding to the face including the eye, and extract the initial features of the target visible light image corresponding to the face including the eye, the reference visible light image is the visible light image corresponding to the eye in any eye movement state, and the target visible light image is the visible light image corresponding to the eye at other moments or in other eye movement states;

[0199] Identify a first region and a second region in the target visible light image, the first region in the target visible light image includes feature points that change with time or eye movement state, the second region in the target visible light image includes reference points that do not change with time or eye movement state or change within a preset condition, the reference points and the feature points are both used to represent pixel points, and one pixel point corresponds to two features;

[0200] Adjust the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face including the eye.

[0201] In a possible design, the first processing unit is specifically configured to:

[0202] Determine a first region and boundary points of the first region in the target visible light image according to reference points of the target visible light image, and determine boundary points of a second region in the target visible light image according to feature points of the target visible light image;

[0203] Perform smoothing processing on the second region in the target visible light image by a filtering method to obtain a visible light smoothed image corresponding to the second region in the target visible light image;

[0204] Calculate position change information of the reference visible light image relative to the target visible light image according to initial features of a target reference point in the visible light smoothed image;

[0205] According to the position change information, obtain features of the reference visible light image of the face including the eyes by adjusting initial features of each feature point in the first region of the reference visible light image.

[0206] In a possible design, the acquisition module is specifically configured to:

[0207] Extract initial features of the infrared image corresponding to the current eye, and use the initial features of the infrared image corresponding to the current eye as features of the infrared image corresponding to the current eye; or,

[0208] Extract initial features of the infrared image corresponding to the current eye and initial features of the target infrared image corresponding to the eye, where the infrared image corresponding to the current eye is the infrared image corresponding to the eye in any eye movement state, and the target infrared image corresponding to the eye is the infrared image corresponding to the eye at other times or in other eye movement states;

[0209] Identify a first region and a second region in the target infrared image, where the first region in the target infrared image includes feature points that change with time or eye movement state, and the second region in the target infrared image includes reference points that do not change with time or eye movement state or change within a preset condition, and both the reference points and the feature points are used to represent pixel points, and one pixel point corresponds to two features;

[0210] Adjust initial features of feature points of the infrared image corresponding to the current eye according to initial features of the reference points of the target infrared image to obtain features of the infrared image corresponding to the current eye.

[0211] In a possible design, the learning rate of the image processing model is associated with the number and distribution of feature deviation values of the input and output of the image processing model.

[0212] To implement the training method of the image processing model, this embodiment provides a training device for the image processing model. The training device for the image processing model includes:

[0213] A sample determination module, configured to determine a training sample set, where the training sample set includes the feature deviation value of the sample infrared image corresponding to the sample eye and the feature deviation value of the sample visible light image corresponding to the sample face; wherein, the sample face includes the sample eye, the feature deviation value of the sample visible light image is used as the true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye;

[0214] A training module, configured to train a deep learning network model according to the training sample set to obtain an image processing model;

[0215] Wherein, the image processing model is configured to determine the feature deviation value of the visible light image corresponding to the face including the eye according to the feature deviation value of the infrared image corresponding to the eye to be processed.

[0216] The training device for the image processing model provided in this embodiment can be used to execute the technical solutions of the above-mentioned embodiments of the training method of the image processing model. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0217] In a possible design, the number of hidden layers in the deep learning network model is determined based on the number of features of one sample infrared image; the learning rate of the deep learning network model is associated with the number and distribution of the feature deviation values of the input and output of the deep learning network model.

[0218] In a possible design, the sample eye is at least one; the sample determination module includes: a feature acquisition unit, a first processing unit, and a second processing unit;

[0219] The feature acquisition unit is configured to acquire the features of the sample image, where the sample image is taken under different action states of each sample eye; for the same eye action state of each sample eye, the features of the sample image include the features of the sample infrared image corresponding to the sample eye and the features of the sample visible light image corresponding to the sample face;

[0220] The first processing unit is configured to respectively determine the feature deviation value of the sample infrared image corresponding to each sample eye and the feature deviation value of the sample visible light image corresponding to each sample face according to the features of the sample infrared image corresponding to each sample eye and the corresponding features of the sample visible light image;

[0221] A second processing unit, configured to determine a training sample set according to the feature deviation values of the sample infrared images corresponding to the sample eyes and the feature deviation values of the sample visible light images corresponding to the sample faces.

[0222] In a possible design, the first processing unit is specifically configured to:

[0223] For each sample eye, select the feature of any sample infrared image as the feature of the reference sample infrared image from the features of the sample infrared images corresponding to different action states of the sample eye, and select the feature of the sample visible light image in the same action state as the reference sample infrared image from the features of the sample infrared images corresponding to different action states of the sample face as the feature of the reference sample visible light image;

[0224] Use the difference between the feature of the reference sample infrared image corresponding to the sample eye and the features of the other sample infrared images corresponding to the sample eye as the feature deviation value of the sample infrared image, and use the difference between the feature of the reference sample visible light image corresponding to the sample face and the features of the other sample visible light images corresponding to the sample face as the feature deviation value of the sample infrared visible light image.

[0225] In a possible design, the feature acquisition unit is specifically configured to:

[0226] For each sample eye, respectively extract the initial features of the sample infrared images corresponding to different action states of the sample eye and the initial features of the sample visible light images corresponding to the sample face;

[0227] Determine a target sample visible light image from the sample visible light images corresponding to different eye action states, and identify a first region and a second region in the target sample visible light image. The target sample visible light image is the sample visible light image corresponding to the sample eye in any eye action state. The first region contains feature points that change with time or eye action state, and the second region contains reference points that do not change with time or eye action state or change within a preset condition. The reference points and the feature points are both used to represent pixel points, and one pixel point corresponds to two features;

[0228] Adjust the initial features of the feature points of each sample visible light image according to the initial features of the reference points of the target sample visible light image to obtain the features of the sample visible light image corresponding to the sample face;

[0229] Determine the features of the sample image according to the initial features of the sample infrared image and the features of the sample visible light image.

[0230] In a possible design, the feature acquisition unit is specifically configured to:

[0231] Determine a first region and boundary points of the first region in the target sample visible light image according to the reference points of the target sample visible light image, and determine boundary points of a second region in the target sample visible light image according to the feature points of the target sample visible light image;

[0232] Smooth the second region in the target sample visible light image by a filtering method to obtain a sample visible light smoothed image corresponding to the second region in the target sample visible light image;

[0233] According to the initial features of the target reference points in the sample visible light smoothed image, traverse the sample visible light images corresponding to other action states of the sample face:

[0234] Calculate the position change information of the sample visible light image corresponding to the other action state relative to the target sample visible light image;

[0235] According to the position change information, obtain the features of the sample visible light image corresponding to the other action state by adjusting the initial features of the feature points in the first region of the sample visible light image corresponding to the other action state.

[0236] In a possible design, the feature acquisition unit is specifically configured to:

[0237] Use the initial features of the sample infrared image as the features of the sample infrared image, and form the features of the sample image by the features of the sample infrared image and the features of the sample visible light image; or,

[0238] Determine a target sample infrared image from the sample infrared images corresponding to different eye action states, and identify a first region and a second region in the target sample infrared image, where the first region in the target sample infrared image contains the feature points, and the second region in the target sample infrared image contains the reference points;

[0239] Adjust the initial features of the feature points of other sample infrared images according to the initial features of the reference points of the target sample infrared image to obtain the features of the sample visible infrared image corresponding to the sample eye;

[0240] Wherein, the features of the sample image are formed by the features of the sample infrared image and the features of the sample visible light image.

[0241] To implement the above image transformation method, this embodiment provides an image transformation device, including: an image transformation device and a data acquisition device; the image transformation device is configured to execute the method described in any item of the first aspect; wherein, the data acquisition device includes an infrared camera, a visible light camera, and a transparent frame; the infrared camera is mounted on the transparent frame;

[0242] The infrared camera is configured to acquire an infrared image of the eye at any moment or in any eye movement state;

[0243] The visible light camera is configured to acquire a visible light image of the eye located within the transparent frame at any moment or in any eye movement state.

[0244] The image transformation device provided in this embodiment can be used to execute the technical solutions of the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0245] To implement the method of the above embodiment, this embodiment provides an electronic device. Figure 8 It is a schematic structural diagram of the electronic device provided in the embodiment of the present application. As Figure 8 shown, the electronic device of this embodiment includes: a processor 801 and a memory 802; wherein, the memory 802 is used to store computer execution instructions; the processor 801 is used to execute the computer execution instructions stored in the memory to implement each step executed in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiment.

[0246] This application embodiment also provides a computer-readable storage medium, in which computer execution instructions are stored, and when the processor executes the computer execution instructions, the above method is implemented.

[0247] This application embodiment also provides a computer program product, including a computer program, and when the computer program is executed by the processor, the above method is implemented.

[0248] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices or modules, and can be in electrical, mechanical or other forms. In addition, in each embodiment of the present application, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0249] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods in various embodiments of the present application. It should be understood that the above processor can be a central processing unit (English: Central Processing Unit, abbreviated: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0250] The memory may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus. The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0251] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0252] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disks, or optical discs and other media that can store program codes.

[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An image transformation method, It is characterized in that The method comprises: Obtaining a characteristic deviation value of the infrared image corresponding to the eye to be processed; According to the feature deviation value of the infrared image corresponding to the current eye, a feature deviation value of the visible light image corresponding to the face containing the current eye is obtained through an image processing model, and the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model.

2. The method according to claim 1, It is characterized in that The step of obtaining the characteristic deviation value of the infrared image corresponding to the current eye includes: Acquire features of the infrared image corresponding to the current eye and features of a reference infrared image corresponding to the eye; the features of the reference infrared image corresponding to the eye are features of an infrared image corresponding to the eye in any eye action state; The difference between the feature of the reference infrared image corresponding to the eye and the feature of the current infrared image corresponding to the eye is used as the feature deviation value of the current infrared image corresponding to the eye.

3. The method according to claim 1, It is characterized in that The method further comprises: Acquire features of a reference visible light image corresponding to the face including the eye; wherein the reference infrared image corresponding to the eye and the reference visible light image corresponding to the face are acquired at the same time or when the eye corresponds to the same eye action state; A facial expression is generated based on a feature deviation value of a current visible light image corresponding to the face including the eyes and features of a reference visible light image corresponding to the face including the eyes.

4. The method according to claim 3, It is characterized in that The acquiring the features of the reference visible light image corresponding to the face including the eyes comprises: Extracting initial features of a reference visible light image corresponding to a face including the eye, and extracting initial features of a target visible light image corresponding to a face including the eye, wherein the reference visible light image is a visible light image corresponding to the eye in any eye movement state, and the target visible light image is a visible light image corresponding to the eye at other times or other eye movement states; Identify a first area and a second area in the target visible light image, wherein the first area in the target visible light image includes feature points that change with time or eye movement state, and the second area in the target visible light image includes reference points that do not change with time or eye movement state or change within a preset condition, wherein both the reference points and the feature points are used to represent pixel points, and one pixel point corresponds to two features; The initial features of the feature points of the reference visible light image are adjusted according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face including the eyes.

5. The method according to claim 4, It is characterized in that The adjusting the initial features of the feature points of the reference visible light image according to the initial features of the reference points of the target visible light image to obtain the features of the reference visible light image corresponding to the face including the eye includes: Determining a first region and a boundary point of the first region in the target visible light image according to a reference point of the target visible light image, and determining a boundary point of a second region in the target visible light image according to a feature point of the target visible light image; Smoothing the second region in the target visible light image by filtering to obtain a visible light smoothed image corresponding to the second region in the target visible light image; Calculating position change information of the reference visible light image relative to the target visible light image according to initial features of the target reference point in the visible light smoothed image; The features of the reference visible light image of the face including the eyes are obtained by adjusting the initial features of the feature points of the first area in the reference visible light image according to the position change information.

6. The method according to any one of claims 2 to 5, It is characterized in that The step of obtaining the feature of the infrared image corresponding to the current eye includes: Extracting the initial features of the infrared image currently corresponding to the eye, and using the initial features of the infrared image currently corresponding to the eye as the features of the infrared image currently corresponding to the eye; or, Extracting initial features of the infrared image currently corresponding to the eye and initial features of the target infrared image corresponding to the eye, wherein the infrared image currently corresponding to the eye is an infrared image corresponding to the eye in any eye movement state, and the target infrared image corresponding to the eye is an infrared image corresponding to the eye at other times or other eye movement states; Identify a first area and a second area in the target infrared image, wherein the first area in the target infrared image contains feature points that change with time or eye movement state, and the second area in the target infrared image contains reference points that do not change with time or eye movement state or change within a preset condition, wherein both the reference points and the feature points are used to represent pixel points, and one pixel point corresponds to two features; According to the initial features of the reference points of the target infrared image, the initial features of the feature points of the infrared image currently corresponding to the eye are adjusted to obtain the features of the infrared image currently corresponding to the eye.

7. The method according to any one of claims 2 to 5, It is characterized in that The learning rate of the image processing model is associated with the number and distribution of feature deviation values ​​of the input and output of the image processing model.

8. A training method for an image processing model, It is characterized in that The method comprises: Determine a training sample set, wherein the training sample set includes feature deviation values ​​of a sample infrared image corresponding to a sample eye and feature deviation values ​​of a sample visible light image corresponding to a sample face; wherein the sample face includes the sample eye, the feature deviation values ​​of the sample visible light image are used as true values ​​for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye; According to the training sample set, a deep learning network model is trained to obtain an image processing model; wherein the number of hidden layers in the deep learning network model is determined based on the number of features of one of the sample infrared images; The image processing model is used to determine the characteristic deviation value of the visible light image corresponding to the face including the eye according to the characteristic deviation value of the infrared image corresponding to the eye to be processed.

9. The method according to claim 8, It is characterized in that The sample eye is at least one; and determining the training sample set comprises: Acquire features of sample images, where the sample images are taken for each sample eye in different motion states; for the same eye motion state of each sample eye, the features of the sample images include features of a sample infrared image corresponding to the sample eye and features of a sample visible light image corresponding to the sample face; Determine, according to the features of the sample infrared images corresponding to each sample eye and the features of the sample visible light images corresponding to each sample face, the feature deviation values ​​of the sample infrared images corresponding to each sample eye and the feature deviation values ​​of the sample visible light images corresponding to each sample face; A training sample set is determined according to the characteristic deviation value of the sample infrared image corresponding to each sample eye and the characteristic deviation value of the sample visible light image corresponding to each sample face.

10. The method according to claim 9, It is characterized in that The method of determining the characteristic deviation value of the sample infrared image corresponding to each sample eye and the characteristic deviation value of the sample visible light image corresponding to each sample face respectively according to the characteristic of the sample infrared image corresponding to each sample eye and the characteristic deviation value of the sample visible light image corresponding to each sample face includes: For each sample eye, from the features of the sample infrared images in different motion states corresponding to the sample eye, select the features of any sample infrared image as the features of the reference sample infrared image, and from the features of the sample infrared images in different motion states corresponding to the sample face, select the features of the sample visible light image in the same motion state as the reference sample infrared image as the features of the reference sample visible light image; The difference between the features of the baseline sample infrared image corresponding to the sample eyes and the features of other sample infrared images corresponding to the sample eyes is used as the feature deviation value of the sample infrared image, and the difference between the features of the baseline sample visible light image corresponding to the sample face and the features of other sample visible light images corresponding to the sample face is used as the feature deviation value of the sample infrared visible light image.

11. The method according to any one of claims 8 to 10, It is characterized in that The number of hidden layers in the deep learning network model is determined based on the number of features of a sample infrared image; the learning rate of the deep learning network model is associated with the number and distribution of feature deviation values ​​of the input and output of the deep learning network model.

12. An image conversion device, It is characterized in that The device comprises: An acquisition module, used for acquiring a characteristic deviation value of an infrared image corresponding to an eye to be processed; A processing module is used to obtain a feature deviation value of a visible light image corresponding to a face containing the current eye according to a feature deviation value of an infrared image corresponding to the current eye through an image processing model, wherein the number of features of the visible light image is greater than the number of features of the infrared image; the image processing model is obtained by training a deep learning network model.

13. A training device for an image processing model, It is characterized in that The device comprises: A sample determination module is used to determine a training sample set, wherein the training sample set includes a feature deviation value of a sample infrared image corresponding to a sample eye and a feature deviation value of a sample visible light image corresponding to a sample face; wherein the sample face includes the sample eye, the feature deviation value of the sample visible light image is used as a true value for training, and the number of features of the sample visible light image is greater than the number of features of the infrared image of the sample eye; A training module, used to train a deep learning network model according to the training sample set to obtain an image processing model; The image processing model is used to determine the characteristic deviation value of the visible light image corresponding to the face including the eye according to the characteristic deviation value of the infrared image corresponding to the eye to be processed.

14. An image transformation device, It is characterized in that include: An image conversion device and a data acquisition device; the image conversion device is used to execute the method according to any one of claims 1 to 7; wherein the data acquisition device comprises an infrared camera, a visible light camera and a transparent frame; the infrared camera is mounted on the transparent frame; The infrared camera is used to collect infrared images of the eye at any time or in any eye movement state; The visible light camera is used to collect visible light images of the eye located in the transparent frame at any time or in any eye movement state.

15. An electronic device, It is characterized in that include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 11.

16. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.

17. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.