Image processing methods and apparatuses, electronic devices and computer-readable storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-08-11
Smart Images

Figure CN115620375B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to an image processing method, an image processing apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of computer technology, mobile phones, computers, and other electronic devices are increasingly integrated into people's lives and work, and their functions are becoming more and more diverse. In some scenarios, electronic devices can detect user attention levels. For example, in online learning scenarios, with user authorization, electronic devices can collect user facial images, analyze these images to determine the user's real-time attention level, and provide reminders based on this attention level to help the user better complete learning tasks. Summary of the Invention
[0003] At least one embodiment of this disclosure provides an image processing method, comprising: acquiring a detected image, wherein the detected image includes a target part of a detected object; determining K driving parameters of the target part based on the detected image, wherein the K driving parameters reflect the overall movement of the target part; inputting the K driving parameters into a first detection model to obtain an output result of the first detection model, wherein the output result of the first detection model includes the focus data of the detected object, wherein K is a positive integer.
[0004] For example, in an image processing method provided in one embodiment of this disclosure, the target region is the face.
[0005] For example, in an image processing method provided in one embodiment of this disclosure, the K driving parameters are parameters relating to the driving state of the target region.
[0006] For example, in an image processing method provided in an embodiment of this disclosure, the target area includes multiple drivable regions, each of the multiple drivable regions includes multiple muscle points, and the K driving parameters include at least one parameter reflecting the driving state of each drivable region, wherein the driving state of the drivable region is controlled based on the multiple muscle points of the drivable region.
[0007] For example, in an image processing method provided in one embodiment of this disclosure, determining K driving parameters of the target region based on the detected image includes: inputting the detected image into a second detection model to obtain the output result of the second detection model, wherein the output result of the second detection model includes an S-dimensional vector about the target region, the S-dimensional vector including S driving parameters; performing dimensionality reduction processing on the S-dimensional vector to obtain a K-dimensional vector, wherein the K-dimensional vector includes the K driving parameters, and S is a positive integer greater than K.
[0008] For example, an embodiment of the image processing method provided in this disclosure further includes: acquiring a plurality of sample images, wherein each sample image includes a sample part of a sample object; acquiring focus label data of the sample object in each of the plurality of sample images; determining a sample driving parameter set for the sample part of each sample image, wherein the sample driving parameter set includes K sample driving parameters, the K sample driving parameters reflecting the overall action of the sample part; and training the first detection model based on the focus label data corresponding to the plurality of sample images and the plurality of sample driving parameter sets corresponding to the plurality of sample images.
[0009] For example, in an image processing method provided in one embodiment of this disclosure, acquiring multiple sample images includes: acquiring images of each sample object in N sample objects over P time periods to obtain Q sample images for each sample object, wherein the multiple sample images include Q sample images corresponding to each of the N sample objects; acquiring focus label data of the sample object for each sample image in the multiple sample images includes: for each sample image, determining focus label data corresponding to each sample image based on the time period in which the acquisition time of each sample image is located, wherein N, P, and Q are all positive integers.
[0010] For example, in an image processing method provided in one embodiment of this disclosure, acquiring images of each of N sample objects over P time periods includes: acquiring images of each sample object having P behaviors over the P time periods, wherein the P behaviors correspond to P attention tag data respectively.
[0011] For example, in an image processing method provided in an embodiment of this disclosure, the P types of behaviors include a first behavior and a second behavior, the P attention label data include a first attention label value and a second attention label value, the attention label data of the sample object when it has the first behavior is the first attention label value, the attention label data of the sample object when it has the second behavior is the second attention label value, and the first attention label value is greater than the second attention label value.
[0012] For example, in an image processing method provided in one embodiment of this disclosure, the first action includes watching a first type of video, and the second action includes watching a second type of video.
[0013] For example, in an image processing method provided in an embodiment of this disclosure, determining K sample driving parameters for the sample region in each sample image includes: inputting each sample image into a second detection model to obtain a sample output result of the second detection model, wherein the sample output result includes an S-dimensional sample vector for the sample region, wherein the S-dimensional sample vector includes S sample driving parameters; performing dimensionality reduction processing on the S-dimensional sample vector to obtain a K-dimensional sample vector, wherein the K-dimensional sample vector includes the K sample driving parameters, wherein K is a positive integer and S is a positive integer greater than K.
[0014] For example, in an image processing method provided in an embodiment of this disclosure, dimensionality reduction processing is performed on the S-dimensional sample vector to obtain the K-dimensional sample vector, including: forming a first matrix based on the S-dimensional sample vectors corresponding to N sample images in the plurality of sample images, wherein the N sample images correspond to the N sample objects respectively; determining a second matrix and a third matrix based on the first matrix, wherein the second matrix is a matrix composed of multiple feature vectors of the first matrix, and the third matrix is a matrix composed of multiple feature values of the first matrix, wherein the multiple feature vectors correspond one-to-one with the multiple feature values, and the multiple feature values are sorted in descending order in the third matrix; selecting the top K feature values in the third matrix with a proportion not less than a first ratio, and deleting the remaining feature vectors in the second matrix except for the K feature vectors corresponding to the top K feature values, to obtain a dimensionality-reduced second matrix; obtaining a dimensionality-reduced first matrix based on the first matrix and the dimensionality-reduced second matrix; and obtaining the K-dimensional sample vector based on the dimensionality-reduced first matrix.
[0015] For example, in an image processing method provided in one embodiment of this disclosure, a first detection model is trained based on the attention label data corresponding to the plurality of sample images and the K sample driving parameters corresponding to the plurality of sample images. This includes: updating and iterating the parameters of an initial model using the attention label data corresponding to the plurality of sample images and the K sample driving parameters corresponding to the plurality of sample images until the training completion condition is met, and using the trained initial model as the first detection model. The updating and iterating of the parameters of the initial model includes: performing the following operations for each sample image: inputting the K sample driving parameters of the sample image into the initial model to obtain the initial output result of the initial model; calculating loss information based on the initial output result of the initial model and the attention label data corresponding to the sample image; and updating the parameters of the initial model based on the loss information.
[0016] At least one embodiment of this disclosure provides an image processing apparatus, including an acquisition module, a determination module, and a result module. The acquisition module is configured to acquire a detected image, wherein the detected image includes a target part of a detected object. The determination module is configured to determine K driving parameters of the target part based on the detected image, wherein the K driving parameters reflect the overall movement of the target part. The result module is configured to input the K driving parameters into a first detection model to obtain the output result of the first detection model, wherein the output result of the first detection model includes the focus data of the detected object, and K is a positive integer.
[0017] At least one embodiment of this disclosure provides an electronic device, including an imaging device and an image processing device. The imaging device is configured to capture a detected image; the image processing device is configured to receive the detected image and perform an image processing method provided in any embodiment of this disclosure based on the detected image.
[0018] At least one embodiment of this disclosure provides an electronic device, including a memory and a processor, wherein the memory stores computer-executable instructions non-transitoryly; the processor is configured to execute the computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the image processing method provided in any embodiment of this disclosure.
[0019] At least one embodiment of this disclosure provides a non-transitory computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image processing method provided in any embodiment of this disclosure. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0021] Figure 1 A flowchart of an image processing method provided by at least one embodiment of the present disclosure is shown;
[0022] Figure 2 A flowchart illustrating a training method for a first detection model provided in at least one embodiment of this disclosure is shown;
[0023] Figure 3 A flowchart illustrating the acquisition of K sample driving parameters is shown, according to at least one embodiment of this disclosure.
[0024] Figure 4 A schematic block diagram of an image processing apparatus provided in at least one embodiment of the present disclosure;
[0025] Figure 5 A schematic block diagram of an electronic device provided for at least one embodiment of this disclosure;
[0026] Figure 6 A schematic block diagram of another electronic device provided for at least one embodiment of this disclosure;
[0027] Figure 7 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of this disclosure; and
[0028] Figure 8 This is a schematic diagram of a hardware environment provided for at least one embodiment of the present disclosure. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0030] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0031] The inventors discovered that when detecting user attention, analysis is usually performed only on local facial features such as the eyes or mouth, and typically based on the position or distance information of these features, such as the position of the eyeballs, the distance between the upper and lower eyelids, and the distance between the upper and lower lips. However, attention levels obtained from local features are not accurate enough, and the position or distance information of these features cannot accurately reflect the overall state of the face.
[0032] At least one embodiment of this disclosure provides an image processing method, an image processing apparatus, an electronic device, and a computer-readable storage medium. The image processing method includes: acquiring a detection image, wherein the detection image includes a target part of a detected object; determining K driving parameters of the target part based on the detection image, wherein the K driving parameters reflect the overall movement of the target part; inputting the K driving parameters into a first detection model to obtain an output result of the first detection model, wherein the output result of the first detection model includes focus data of the detected object, and K is a positive integer.
[0033] According to the image processing method of this disclosure, K driving parameters can reflect the overall movement of the target part, covering the entire area of the target part rather than a local area, and can also reflect the driving state of each area of the target part rather than simply position or distance information. Therefore, the attention detection result obtained based on the K driving parameters can accurately reflect the attention level of the detected object, thus improving the accuracy of attention detection.
[0034] Figure 1 A flowchart of an image processing method provided by at least one embodiment of the present disclosure is shown.
[0035] like Figure 1As shown, the image processing method may include steps S110 to S130.
[0036] Step S110: Obtain the image to be detected, which includes the target part of the object to be detected.
[0037] Step S120: Based on the detected image, determine K driving parameters of the target part. The K driving parameters reflect the overall movement of the target part, and K is a positive integer.
[0038] Step S130: Input the K driving parameters into the first detection model to obtain the output of the first detection model, which includes the focus data of the detected object.
[0039] For example, in step S110, the image to be detected can be acquired through image acquisition. The object to be detected in the image can be, for example, a person, and the target part can be a face. In some embodiments of this disclosure, the example of the object to be detected being a person and the target part being a face is used for illustration. It should be noted that in practical applications, the types of the object to be detected and the target part can be set according to actual needs. For example, the object to be detected can be an animal in addition to a person. It should be noted that in the embodiments of this disclosure, image acquisition and other operations are performed with the user's authorization.
[0040] For example, the K driving parameters are parameters relating to the driving state of the target area. Driving parameters can be understood as parameters driven by muscle points. Taking the human face as an example, the face has multiple muscle points, which can drive multiple areas of the face to present different states. The driving state can be the state presented by the face after being driven by multiple muscle points of the face.
[0041] For example, the target site includes multiple drivable regions, each of which includes multiple muscle points, and the K driving parameters include at least one parameter reflecting the driving state of each drivable region, wherein the driving state of the drivable region is controlled based on the multiple muscle points of the drivable region.
[0042] For example, taking the human face as an example, multiple drivable regions can include the mouth region, eye region, cheek region, eyebrow region, and other movable areas of the face. Each drivable region can correspond to at least one driving parameter, and the driving state corresponding to the drivable region can be the state presented after the drivable region is driven by the corresponding muscle point. For example, the mouth, when driven by the mouth muscles, presents a pouting state, and the eyebrows, when driven by the eyebrow muscles, present a raised eyebrow state, and so on. These K driving parameters can represent the state presented after each area of the face is driven; therefore, the combination of these K driving parameters can represent the overall driving state of the face.
[0043] For example, the driving parameters can be blendshape parameters, which may include 52 facial parameters and 9 pose angle parameters. For instance, the pose angle parameters can be represented using Euler angles. The 9 pose angle parameters may include three angle parameters corresponding to head rotation (the three angle parameters are angles in the three dimensions (XYZ) of virtual 3D space), three angle parameters corresponding to left eye rotation, and three angle parameters corresponding to right eye rotation. These 9 pose angle parameters can drive the rotation of the virtual model's head and eyes, thus simulating head and eye rotation while driving facial expressions based on multiple blendshape parameters.
[0044] For example, the 52 facial parameters are: eyeBlinkLeft (left eye blink), eyeLookDownLeft (left eye looking down), eyeLookInLeft (left eye looking at the tip of the nose), eyeLookOutLeft (left eye looking to the left), eyeLookUpLeft (left eye looking up), eyeSquintLeft (left eye squinting), eyeWideLeft (left eye widening), eyeBlinkRight (right eye blink), eyeLookDownRight (right eye looking down), eyeLookInRight (right eye looking at the tip of the nose), and eyeLookOutRight (right eye blinking). (Look left), eyeLookUpRight (right eye looking upward), eyeSquintRight (right eye squinting), eyeWideRight (right eye wide open), jawForward (chin forward when pursing), jawLeft (chin to the left when pouting), jawRight (chin to the right when pouting), jawOpen (chin down when opening mouth), mouthClose (close mouth), mouthFunnel (mouth slightly open with lips parted), mouthPucker (pursed lips), mouthLeft (pouting to the left), mouthRight (pouting to the right), mouthSmileLeft (smile with a left pout) Mouth Smile Right (smile with the right lip turned down), Mouth Frown Left (press the left lip down), Mouth Frown Right (press the right lip down), Mouth Dimple Left (left lip turned back), Mouth Dimple Right (right lip turned back), Mouth Stretch Left (left corner of the mouth turned to the left), Mouth Stretch Right (right corner of the mouth turned to the right), Mouth Roll Lower (lower lip curled inward), Mouth Roll Upper (lower lip curled upward), Mouth Shrug Lower (lower lip turned downward), Mouth Shrug Up per (upper lip up), mouthPressLeft (lower lip down to the left), mouthPressRight (lower lip down to the right), mouthLowerDownLeft (lower lip down to the lower left), mouthLowerDownRight (lower lip down to the lower right), mouthUpperUpLeft (upper lip up to the left), mouthUpperUpRight (upper lip up to the right), browDownLeft (left eyebrow outward), browDownRight (right eyebrow outward), browInnerUp (frowning), browOuterUpLeft (left eyebrow up to the left)The parameters are: browOuterUpRight (right eyebrow pointing upwards to the right), cheekPuff (cheek pointing outwards), cheekSquintLeft (left cheek pointing upwards and in a circular motion), cheekSquintRight (right cheek pointing upwards and in a circular motion), noseSneerLeft (nose pointing left), noseSneerRight (nose pointing right), and tongueOut (tongue sticking out). Of these 52 facial parameters, 14 represent the specific state of the eyes, 4 represent the state of the chin, 23 represent the state of the mouth, 5 represent the state of the eyebrows, 3 represent the state of the cheeks, 2 represent the state of the nose, and 1 represents the state of the tongue.
[0045] For example, each blendshape parameter can be represented by a floating-point value, and the value range of each blendshape parameter can be [0,1]. For example, for an open mouth, a value of 0 indicates that the mouth is not open, and a value of 1 indicates that the mouth is fully open. The specific meaning of the value of each parameter can be set according to the actual situation.
[0046] For example, in some embodiments, the K driving parameters can be the aforementioned 61 blendshape parameters (K, for example, equal to 61). In other embodiments, the K driving parameters can also be some of the aforementioned 61 blendshape parameters (K, for example, less than 61), depending on the actual needs.
[0047] For example, the first detection model can be an Artificial Neural Network (ANN), and its activation function can be, for example, the sigmoid function. For instance, the first detection model can be trained using K driving parameters and attention label data from multiple sample images to obtain a trained first detection model. In step S130, the K driving parameters of the target area are input into the trained first detection model, and the first detection model outputs the attention data of the detected object.
[0048] For example, focus data can be a numerical value, such as a value between 0 and 1. The closer the focus data is to 1, the higher the focus level.
[0049] According to the image processing method of this disclosure, K driving parameters can reflect the overall movement of the target part, covering the entire area of the target part rather than a local area, and can also reflect the driving state of each area of the target part rather than simply position or distance information. Therefore, the attention detection result obtained based on the K driving parameters can accurately reflect the attention level of the detected object, thus improving the accuracy of attention detection.
[0050] The training process of the first detection model is explained below.
[0051] Figure 2 A flowchart is shown of a training method for a first detection model provided in at least one embodiment of the present disclosure.
[0052] like Figure 2 As shown, for example, the image processing method of this disclosure embodiment may further include steps S210 to S240.
[0053] Step S210: Acquire multiple sample images, each sample image including a sample part of the sample object.
[0054] Step S220: Obtain the focus label data of the sample object in each of the multiple sample images.
[0055] Step S230: Determine the sample driving parameter group for the sample region of each sample image. The sample driving parameter group includes K sample driving parameters, which reflect the overall movement of the sample region.
[0056] Step S240: Based on the focus label data corresponding to the multiple sample images and the multiple sample driving parameter groups corresponding to the multiple sample images, the first detection model is trained.
[0057] For example, the type of the sample object in the sample image is consistent with the type of the detected object in the input image (e.g., both are people), and the type of the sample part in the sample image is consistent with the type of the target part in the input image (e.g., both are faces). This embodiment of the disclosure uses a human face as an example for illustration.
[0058] For example, each sample image has corresponding label data, which can be calculated using a predetermined algorithm or obtained through manual annotation. The label data includes at least the focus label data of the sample object (e.g., a person). For example, the focus label data can be a value between 0 and 1.
[0059] For example, for each sample image, K corresponding sample driving parameters are obtained to form a sample driving parameter group. For example, the K sample driving parameters can be K blendshape parameters, which are used to reflect the overall driving state of the sample parts of the sample object. The relevant characteristics of these K sample driving parameters can be referred to the K driving parameters mentioned above, and will not be repeated here.
[0060] For example, after obtaining the sample-driven parameter set and attention label data corresponding to each sample image, a first detection model can be trained using the sample-driven parameter set and attention label data. Step S240 may include: updating and iterating the parameters of the initial model using the attention label data corresponding to the multiple sample images and the K sample-driven parameters corresponding to the multiple sample images until the training completion condition is met, and using the trained initial model as the first detection model.
[0061] For example, updating and iterating the parameters of the initial model includes the following steps for each sample image: inputting the K sample-driving parameters of the sample image into the initial model to obtain the initial output of the initial model; calculating the loss information based on the initial output of the initial model and the attention label data corresponding to the sample image; and updating the parameters of the initial model based on the loss information.
[0062] For example, the sample-driven parameter set of the sample image is input into the initial model to be trained. The output of the initial model is compared with the attention label data corresponding to the sample image to obtain the loss information of the initial model. The parameters of the initial model are adjusted using the loss information to complete one iteration. After multiple iterations, when the loss information of the adjusted model meets the predetermined conditions, the first detection model that has been trained can be obtained. For example, the adjusted model whose loss information meets the predetermined conditions is the first detection model that has been trained.
[0063] For example, in other embodiments, each sample image may also have other label data besides focus label data, and the focus label data and other label data may be combined to train a first detection model.
[0064] For example, in step S210, images of each of the N sample objects can be acquired over P time periods to obtain Q sample images for each sample object. These multiple sample images include the Q sample images corresponding to each of the N sample objects. In step S220, for each sample image, based on the time period in which the acquisition time of each sample image falls, the attention label data corresponding to each sample image is determined, where N, P, and Q are all positive integers.
[0065] For example, the focus level of an object, such as a person, may change over time. For instance, if a sample object (e.g., a person) participates in different activities (i.e., does different things, exhibits different behaviors) over P time periods, the person's focus level may differ across these P time periods because the focus level may vary depending on the activity. Similarly, even when participating in the same activity over the P time periods, a person's focus level may change over time; for example, it might be higher at the beginning and gradually decrease over time. Therefore, images of each sample object can be collected over P different time periods, with D images collected in each time period (where D is a positive integer). This would result in P*D images for each sample object, and Q = P*D.
[0066] For example, attention level labels can be pre-set for each time period, such as pre-setting attention levels for different activities or the same activity at different time periods. After collecting sample images from different time periods, the attention level corresponding to each sample image can be obtained based on the collection time of each sample image.
[0067] For example, each sample object exhibits P behaviors over P time periods, and these P behaviors correspond to P attention level labels. When collecting sample images, images of each sample object exhibiting each of the P behaviors over the P time periods can be captured.
[0068] It should be noted that the P time periods here are only used to distinguish different time periods, and are not used to restrict which specific time periods they are. Each sample object can perform P different behaviors within any P different time periods.
[0069] For example, the P time periods corresponding to different sample objects can be the same or different. For instance, taking each time period as 10 minutes, the first sample object performs the first action from 10:00 to 10:10 and the second action from 10:10 to 10:20. The second sample object can remain the same as the first sample object, or the second sample object can also perform the first action and the second action in two other time periods (e.g., 12:00 to 12:10 and 12:10 to 12:20).
[0070] For example, the order in which different sample objects perform P behaviors can be the same or different. For instance, in some examples, the first sample object may perform the first behavior in a first time period and the second behavior in a subsequent second time period. The second sample object may remain the same as the first sample object, or the second sample object may also perform the second behavior in a first time period and the first behavior in a subsequent second time period.
[0071] For example, different sample objects may perform the same or different P behaviors. For instance, in some examples, the first sample object may perform the first behavior in the first time period and the second behavior in the second time period. The second sample object may remain the same as the first sample object, or the second sample object may perform the third behavior in the first time period and the fourth behavior in the second time period.
[0072] For example, the durations of the P time periods corresponding to different sample objects can be the same or different. For instance, in some examples, the duration of each time period corresponding to the first sample object is 10 minutes, and the duration of each time period corresponding to the second sample object can be the same as that of the first sample object, or the duration of each time period corresponding to the second sample object can be 5 minutes.
[0073] For example, for the same sample object, the duration of P time periods can be the same or different, and the P time periods can be continuous or discontinuous.
[0074] For example, the P types of behavior include a first behavior and a second behavior, and the P attention label data may include a first attention label value and a second attention label value. The attention label data of the sample object when it has the first behavior is the first attention label value, and the attention label data of the sample object when it has the second behavior is the second attention label value. The first attention label value is greater than the second attention label value.
[0075] For example, the first action includes watching a first type of video, and the second action includes watching a second type of video. For example, the first type of video could be an interesting, highly engaging video, while the second type of video could be a boring, less engaging video. The first focus tag value corresponding to the first type of video is, for example, 1, and the second focus tag value corresponding to the second type of video is, for example, 0.
[0076] For example, in other embodiments, the P-type behavior may also include a third behavior, such as watching a third type of video, the third type of video corresponding to a third focus tag value, such as between medium 0 and 1.
[0077] For example, in other embodiments, the P-type behavior may also include other behaviors besides watching videos, such as doing exercises, taking online classes, playing games, etc.
[0078] Based on the above-mentioned methods of collecting sample images and obtaining attention label data, the training samples can reflect the differences in attention, achieve the diversity of training samples, and contain multiple levels of attention for the same sample object, which can effectively form a comparison and make the first detection model trained more accurate.
[0079] For example, step S230 may include: inputting each sample image into the second detection model to obtain the sample output result of the second detection model, wherein the sample output result includes an S-dimensional sample vector about the sample part, wherein the S-dimensional sample vector includes S sample driving parameters; performing dimensionality reduction processing on the S-dimensional sample vector to obtain the K-dimensional sample vector, wherein the K-dimensional sample vector includes the K sample driving parameters, wherein K is a positive integer and S is a positive integer greater than K.
[0080] Figure 3 A flowchart for obtaining K sample driving parameters is shown, according to at least one embodiment of the present disclosure.
[0081] like Figure 3 As shown, Q sample images corresponding to each of the N sample objects are input into the second detection model to obtain an S-dimensional sample vector for each sample image. For example, for each sample object, two minutes of video of type 1 and two minutes of video of type 2 are watched, for a total of four minutes. Within these four minutes, sample images are acquired at a frequency of m frames / second, resulting in 240*m sample images for each sample object.
[0082] For example, the second detection model can be a blendshape model, which is a model capable of outputting the aforementioned 61 blendshape parameters (52 facial parameters and 9 pose angle parameters) corresponding to the input image. This second detection model can be a neural network model, which can be pre-trained using multiple sample images and the corresponding blendshape parameter labels for each sample image to obtain a trained second detection model. In step S230, the sample image can be input into the trained second detection model, and the second detection model outputs the 61 blendshape parameters corresponding to the sample image, forming a 61-dimensional vector (S, for example, equals 61).
[0083] For example, after the second detection model, a 61-dimensional vector can be obtained for each sample image. There are a total of 240*m*N sample images for N sample objects. If the 61-dimensional vectors of each sample image are combined together, a (240*m*N)*61 matrix can be obtained, where (240*m*N) and 61 are, for example, the number of rows and columns of the matrix.
[0084] For example, if the first detection model is trained based on the 61-dimensional vector of each sample image, the large amount of data may result in a long training time. Therefore, the 61-dimensional vector can be reduced to a K-dimensional vector, and then the first detection model can be trained based on the reduced K-dimensional vector. For example, the matrix (240*m*N)*61 mentioned above can be reduced to K dimensions. For example, K is less than S, and K can be a value between 2 and 5 (inclusive).
[0085] For example, dimensionality reduction can be performed as follows: A first matrix is formed based on the S-dimensional sample vectors corresponding to the N sample images, where each of the N sample images corresponds to one of the N sample objects; a second and a third matrix are determined based on the first matrix, where the second matrix is composed of multiple eigenvectors of the first matrix, and the third matrix is composed of multiple eigenvalues of the first matrix, where each eigenvector corresponds one-to-one with each eigenvalue, and the eigenvalues are sorted in descending order in the third matrix; the top K eigenvalues with a proportion not less than a first ratio are selected from the third matrix, and all other eigenvectors in the second matrix except those corresponding to the top K eigenvalues are deleted to obtain the dimensionality-reduced second matrix; the dimensionality-reduced first matrix is obtained based on the first matrix and the dimensionality-reduced second matrix; and the K-dimensional sample vector is obtained based on the dimensionality-reduced first matrix.
[0086] For example, taking S as 61, first let A = N*61, then calculate the covariance matrix C = A T A. Matrix C has dimensions 61*61. Performing matrix diagonalization on matrix C yields C = QΛQ. -1 Here, Q is a matrix composed of eigenvectors of matrix A, and Λ is an N*N diagonal matrix. Each diagonal element of matrix Λ represents an eigenvalue (an eigenvalue of matrix A). In other words, the information of matrix A can be represented by eigenvalues and eigenvectors. Matrix Λ is a matrix composed of the eigenvalues of matrix A, and matrix Q is a matrix composed of the eigenvectors of matrix A. Matrix A can be considered the first matrix, matrix Q can be considered the second matrix, and matrix Λ can be considered the third matrix. In matrix Λ, the eigenvalues can be arranged from largest to smallest. There can be a one-to-one correspondence between the eigenvectors in matrix Q and the eigenvalues in matrix Λ. The eigenvectors corresponding to these eigenvalues in matrix Λ can describe the direction of change of matrix Q (arranged from primary to secondary changes).
[0087] For example, matrix Λ can be represented as:
[0088]
[0089] Wherein, λ1~λ all These are elements of matrix Λ, for example, representing eigenvalues.
[0090] For example, λ1≥…≥λ all We can select the first K characteristic values whose numerical percentages are equal to or greater than the first ratio (the first ratio is a value between 0 and 1, such as 90%). K is the target dimension for dimensionality reduction. For example, if the dimension of matrix Q is 61*N, since we take the first k eigenvalues, we take the first k rows of matrix Q (matrix Q and matrix Λ have a one-to-one correspondence, i.e., λ1 corresponds to the first row of matrix Q), resulting in the dimensionality-reduced matrix Q. k Given a dimension of K*N, the final matrix A after dimensionality reduction is... The dimension is N*K. Based on this, we can obtain the dimensionality-reduced K-dimensional vector corresponding to each sample image.
[0091] Based on the dimensionality reduction method described above, the data corresponding to N sample objects are aggregated together for dimensionality reduction processing, which helps to extract the K-dimensional data with greater influence from the 61-dimensional data, thereby making the trained first detection model more accurate.
[0092] For example, after dimensionality reduction, a matrix of (240*m*N)*K dimensions is obtained. This matrix includes a time dimension of 240*m, a sample object quantity dimension of N, and a feature dimension of K. Attention label data for sample objects in the time dimension can be statistically analyzed, as shown below.
[0093] k i (t), t∈[1, 240m], i∈[1, N]
[0094] Where i represents the i-th sample object out of N sample objects, t represents the time of the sample image acquisition, the acquisition frequency is m frames / second, and 240m (i.e., 240*m) sample images can be acquired within 4 minutes for each sample object, and [1, 240m] represents the time from acquiring the first sample image to acquiring the 240m sample image. k i (t) represents the focus label data corresponding to the sample image collected at time t.
[0095] For example, in embodiments of this disclosure, k i(t) is regressed to a numerical value. As mentioned above, for example, the first two minutes are an interesting video, and the last two minutes are a boring video. Then, the focus label data of the first 120*m frames is higher than that of the last 120*m frames. Because focus itself is a relatively subjective value and cannot be imagined out of thin air, this method can distinguish the level of focus. For example, the level of focus can be represented by quantifying focus into numbers. For example, the value of the highest focus can be 1, and the value of the lowest focus can be 0. Therefore, the ideal output value (i.e., focus label data) corresponding to the first 120*m frames can be 1, and the ideal output value corresponding to the last 120*m frames can be 1. Then the output of the first detection model trained can be a value between 0 and 1 (inclusive), for example, an output of 0.56, which can obtain a quantified value of focus.
[0096] For example, the loss function during training can be expressed as: Loss = (y - y0) 2 y represents the model's output value, and y0 represents the focus label data. The process of updating the weights of the first detection model based on the loss value can include forward computation, backward layer-by-layer gradient computation, and layer-by-layer update of network parameters.
[0097] For example, after training the first detection model, the trained first detection model can be used to perform the above step S130 to obtain the focus data of the detected object.
[0098] For example, in step S120, the image to be detected can be input into the second detection model to obtain the output result of the second detection model. The output result of the second detection model includes an S-dimensional vector about the target part, which includes S driving parameters. The S-dimensional vector is then reduced in dimensionality to obtain a K-dimensional vector, which includes the K driving parameters, where S is a positive integer greater than K.
[0099] For example, as mentioned above, the second detection model can be a blendshape model. In step S120, the image to be detected can be input into the second detection model, and the second detection model outputs 61 blendshape parameters corresponding to the image to be detected, forming a 61-dimensional vector (S is, for example, equal to 61). If the first detection model is trained based on a K-dimensional vector during training, then during use, the 61-dimensional vector also needs to be reduced to a K-dimensional vector, and then the reduced K-dimensional vector is input into the first detection model to obtain the focus data of the detected object.
[0100] For example, in the training process described above, the S-dimensional sample vector has been reduced to a K-dimensional sample vector. This reduced K-dimensional sample vector includes K selected blendshape parameters. The types of these K selected blendshape parameters can be recorded so that the same K types of parameters can be selected during application. In other words, in step S120, data corresponding to the types of the K blendshape parameters can be selected from the S-dimensional vector to form a K-dimensional vector, thus achieving dimensionality reduction of the S-dimensional vector.
[0101] At least one embodiment of this disclosure also provides an image processing apparatus. Figure 4 This is a schematic block diagram of an image processing apparatus provided for at least one embodiment of the present disclosure.
[0102] like Figure 4 As shown, the image processing apparatus 200 may include an acquisition module 201, a determination module 202, and a result module 203. These components are interconnected via a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented as hardware (e.g., circuit) modules, software modules, or any combination of both, as is the case in the following embodiments, and will not be repeated here. For example, these units can be implemented using a central processing unit (CPU), a graphics processing unit (GPU), a tensor processor (TPU), a field-programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions. It should be noted that... Figure 4 The components and structure of the image processing apparatus 200 shown are merely exemplary and not limiting. The image processing apparatus 200 may also have other components and structures as needed.
[0103] For example, the acquisition module 201 is configured to acquire a detected image, wherein the detected image includes the target part of the detected object.
[0104] For example, the determining module 202 is configured to determine K driving parameters of the target region based on the detected image, wherein the K driving parameters reflect the overall movement of the target region.
[0105] For example, the result module 203 is configured to input the K driving parameters into the first detection model to obtain the output result of the first detection model, wherein the output result of the first detection model includes the focus data of the detected object, and K is a positive integer.
[0106] For example, the acquisition module 201, determination module 202, and result module 203 may include code and programs stored in memory; the processor may execute the code and programs to implement some or all of the functions of the acquisition module 201, determination module 202, and result module 203 as described above. For example, the acquisition module 201, determination module 202, and result module 203 may be dedicated hardware devices used to implement some or all of the functions of the acquisition module 201, determination module 202, and result module 203 as described above. For example, the acquisition module 201, determination module 202, and result module 203 may be a circuit board or a combination of multiple circuit boards used to implement the functions described above. In the embodiments of this application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processor; and (3) processor-executable firmware stored in memory.
[0107] It should be noted that the acquisition module 201 can be used to implement Figure 1 As shown in step S110, module 202 is determined to be able to implement Figure 1 The result module 203, as shown in step S120, can be used to implement... Figure 1 The step S130 is shown. Therefore, for a detailed description of the functions that the acquisition module 201, the determination module 202, and the result module 203 can achieve, please refer to the relevant descriptions of steps S110 to S130 in the embodiments of the above image processing method; repeated details will not be repeated here. Furthermore, the image processing apparatus 200 can achieve similar technical effects to the aforementioned image processing method, which will not be described further here.
[0108] It should be noted that, in the embodiments of this disclosure, the image processing apparatus 200 may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited; it can be constructed from analog devices, digital chips, or other suitable methods according to circuit principles.
[0109] For example, in some embodiments, the image processing apparatus 200 may further include a training module configured to: acquire a plurality of sample images, wherein each sample image includes a sample part of a sample object; acquire focus label data of the sample object in each of the plurality of sample images; determine a set of sample driving parameters for the sample part of each sample image, wherein the set of sample driving parameters includes K sample driving parameters that reflect the overall action of the sample part; and train the first detection model based on the focus label data corresponding to the plurality of sample images and the set of sample driving parameters corresponding to the plurality of sample images.
[0110] For example, the acquisition module 201, the determination module 202, the result module 203, and the training module can also implement more or further functions.
[0111] For example, in some embodiments, the determining module 202 is further configured to: input the image to be detected into a second detection model to obtain the output result of the second detection model, wherein the output result of the second detection model includes an S-dimensional vector about the target part, the S-dimensional vector including S driving parameters; and perform dimensionality reduction processing on the S-dimensional vector to obtain a K-dimensional vector, wherein the K-dimensional vector includes the K driving parameters, and S is a positive integer greater than K.
[0112] For example, in some embodiments, the training module is further configured to: acquire images of each of the N sample objects over P time periods to obtain Q sample images for each sample object, wherein the plurality of sample images includes the Q sample images corresponding to each of the N sample objects, and N, P and Q are all positive integers.
[0113] For example, in some embodiments, the training module is further configured to: for each sample image, determine the attention label data corresponding to each sample image based on the time period in which the acquisition time of each sample image is located.
[0114] For example, in some embodiments, the training module is further configured to: collect images of each sample object when it has P behaviors in P time periods, wherein the P behaviors correspond to P attention label data respectively.
[0115] For example, in some embodiments, the training module is further configured to: input each sample image into a second detection model to obtain a sample output result of the second detection model, wherein the sample output result includes an S-dimensional sample vector about the sample region, wherein the S-dimensional sample vector includes S sample driving parameters; and perform dimensionality reduction processing on the S-dimensional sample vector to obtain a K-dimensional sample vector, wherein the K-dimensional sample vector includes the K sample driving parameters, wherein K is a positive integer and S is a positive integer greater than K.
[0116] For example, in some embodiments, the training module is further configured to: form a first matrix based on the S-dimensional sample vectors corresponding to the N sample images in the plurality of sample images, wherein the N sample images correspond to the N sample objects respectively; determine a second matrix and a third matrix based on the first matrix, wherein the second matrix is a matrix composed of multiple feature vectors of the first matrix, and the third matrix is a matrix composed of multiple feature values of the first matrix, wherein the multiple feature vectors correspond one-to-one with the multiple feature values, and the multiple feature values are sorted in descending order in the third matrix; select the top K feature values in the third matrix with a proportion not less than a first ratio, and delete the remaining feature vectors in the second matrix except for the K feature vectors corresponding to the top K feature values, to obtain a dimension-reduced second matrix; obtain a dimension-reduced first matrix based on the first matrix and the dimension-reduced second matrix; and obtain the K-dimensional sample vector based on the dimension-reduced first matrix.
[0117] For example, in some embodiments, the training module is further configured to: update and iterate the parameters of the initial model using the attention label data corresponding to the plurality of sample images and the K sample driving parameters corresponding to the plurality of sample images, until the training completion condition is met, and use the trained initial model as the first detection model; wherein, updating and iterating the parameters of the initial model includes: performing the following operations for each sample image: inputting the K sample driving parameters of the sample image into the initial model to obtain the initial output result of the initial model; calculating loss information based on the initial output result of the initial model and the attention label data corresponding to the sample image; and updating the parameters of the initial model based on the loss information.
[0118] This disclosure also provides an electronic device in some embodiments. Figure 5 This is a schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.
[0119] For example, such as Figure 5 As shown, the electronic device 300 includes a shooting device 301 and an image processing device 302.
[0120] For example, the imaging device 301 is configured to capture an image to be detected, the image including the target part of the object to be detected.
[0121] For example, the image processing apparatus 302 is configured to receive a detected image and perform the image processing method as described in any of the above embodiments based on the detected image.
[0122] For example, the shooting device 301 can be a rear camera of an electronic device, or a front camera and reflective device of an electronic device.
[0123] For example, the image processing device 302 can be implemented as a central processing unit, a dedicated processing chip, a digital signal processor, etc., and this disclosure does not impose any specific limitations on it.
[0124] For example, electronic device 300 can be a terminal device such as a learning machine, and can also provide a display unit (such as a touch screen). For example, the display unit can provide a corresponding human-computer interaction interface for displaying the response of interactive operation, interactive action prompt information, etc. This disclosure does not impose specific limitations in this regard.
[0125] For example, a detailed description of the process by which the electronic device 300 performs the image processing method can be found in the relevant description in the embodiments of the above-described image processing method, and repeated descriptions will not be repeated here.
[0126] Some embodiments of this disclosure also provide another electronic device. Figure 6 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0127] For example, such as Figure 6 As shown, the electronic device 400 includes a processor 401 and a memory 402. It should be noted that... Figure 6 The components of the electronic device 400 shown are merely exemplary and not limiting. The electronic device 400 may have other components as needed for the actual application.
[0128] For example, processor 401 and memory 402 can communicate with each other directly or indirectly.
[0129] For example, processor 401 and memory 402 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. Processor 401 and memory 402 can also communicate with each other via a system bus; this disclosure does not limit this.
[0130] For example, in some embodiments, memory 402 is used to store computer-readable instructions non-transitory. When processor 401 executes the computer-readable instructions, the computer-readable instructions are executed by processor 401 to implement the image processing method according to any of the above embodiments. For specific implementations and related explanations of the various steps of this image processing method, please refer to the embodiments of the image processing method described above; repeated details will not be elaborated here.
[0131] For example, processor 401 and memory 402 can be located on the server side (or in the cloud).
[0132] For example, processor 401 can control other components in electronic device 400 to perform desired functions. Processor 401 can be a central processing unit (CPU), graphics processing unit (GPU), network processor (NP), etc.; it can also be a digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.
[0133] For example, memory 402 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and processor 401 may execute the computer-readable instructions to implement various functions of electronic device 400. Various application programs and various data may also be stored in the storage medium.
[0134] For example, in some embodiments, the electronic device 400 can be a mobile phone, tablet computer, electronic paper, television, monitor, laptop computer, digital photo frame, navigator, wearable electronic device, smart home device, etc.
[0135] For example, electronic device 400 may include a display panel, which can be used for image segmentation, etc. For example, the display panel can be a rectangular panel, a circular panel, an elliptical panel, or a polygonal panel. Furthermore, the display panel can be not only a flat panel, but also a curved panel, or even a spherical panel.
[0136] For example, electronic device 400 can have touch functionality, that is, electronic device 400 can be a touch device.
[0137] For example, a detailed description of the process by which the electronic device 400 performs the image processing method can be found in the relevant description in the embodiments of the above-described image processing method, and repeated descriptions will not be repeated here.
[0138] Figure 7 This is a schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. For example, such as Figure 7As shown, computer-readable instructions 501 are non-transitory stored on the computer-readable storage medium 500. For example, when the computer-readable instructions 501 are executed by a processor, one or more steps in the image processing method described above can be performed.
[0139] For example, the storage medium 500 can be used in the aforementioned electronic device 400. For example, the storage medium 500 may include the memory 402 in the electronic device 400.
[0140] For example, the description of storage medium 500 can be found in the description of memory 402 in the embodiment of electronic device 400, and the repeated parts will not be repeated.
[0141] Figure 8 This is a schematic diagram of a hardware environment provided for at least one embodiment of the present disclosure. The electronic device provided in this disclosure can be applied to an Internet system.
[0142] use Figure 8 The computer system provided herein can implement the functions of the image processing apparatus and / or electronic device involved in this disclosure. Such computer systems may include personal computers, laptops, tablets, mobile phones, personal digital assistants, smart glasses, smartwatches, smart rings, smart helmets, and any smart portable or wearable device. A specific system in this embodiment uses a functional block diagram to explain a hardware platform including a user interface. This computer device can be a general-purpose computer device or a purpose-specific computer device. Both types of computer devices can be used to implement the image processing apparatus and / or electronic device of this embodiment. The computer system may include any components necessary to implement the image processing described herein. For example, the computer system can be implemented by a computer device through its hardware, software programs, firmware, and combinations thereof. For convenience, Figure 8 Although only one computer device is shown in the figure, the computer functions related to the information required for image processing described in this embodiment can be implemented in a distributed manner by a set of similar platforms, thus distributing the processing load of the computer system.
[0143] like Figure 8As shown, the computer system may include a communication port 650, connected to a network for data communication. For example, the computer system can send and receive information and data through the communication port 650, enabling wireless or wired communication between the computer system and other electronic devices to exchange data. The computer system may also include a processor group 620 (i.e., the processor described above) for executing program instructions. The processor group 620 may consist of at least one processor (e.g., a CPU). The computer system may include an internal communication bus 610. The computer system may include different forms of program storage units and data storage units (i.e., the memory or storage media described above), such as a hard disk 670, read-only memory (ROM) 630, and random access memory (RAM) 640, capable of storing various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor group 620. The computer system may also include an input / output component 660 for implementing input / output data flow between the computer system and other components (e.g., user interface 680, etc.).
[0144] Typically, the following devices can be connected to the input / output component 660: input devices such as touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices such as displays (e.g., LCD, OLED displays, etc.), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication interfaces.
[0145] Although Figure 8 A computer system with various devices is shown, but it should be understood that the computer system is not required to have all the devices shown, and alternatively, the computer system may have more or fewer devices.
[0146] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0147] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0148] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
[0149] The following points should be noted regarding this disclosure:
[0150] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.
[0151] (2) For clarity, the thickness and dimensions of layers or structures are enlarged in the accompanying drawings used to describe embodiments of the invention. It will be understood that when an element such as a layer, film, region, or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element, or there may be intermediate elements present.
[0152] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0153] The above description is only a specific embodiment of this disclosure, but the protection scope of this disclosure is not limited thereto. The protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. An image processing method, comprising: Acquire a detection image, wherein the detection image includes the target part of the object being detected; Based on the detected image, K driving parameters of the target part are determined, wherein the K driving parameters reflect the overall movement of the target part; The K driving parameters are input into the first detection model to obtain the output result of the first detection model, wherein the output result of the first detection model includes the attention data of the detected object. Specifically, based on the detected image, K driving parameters of the target region are determined, including: The image to be detected is input into a second detection model to obtain the output of the second detection model. The output of the second detection model includes an S-dimensional vector relating to the target region, wherein the S-dimensional vector includes S driving parameters; and The S-dimensional vector is reduced in dimension to obtain a K-dimensional vector, wherein the K-dimensional vector includes the K driving parameters. Where K is a positive integer and S is a positive integer greater than K.
2. The method according to claim 1, wherein, The target area is the face.
3. The method according to claim 1, wherein, The K driving parameters are parameters relating to the driving state of the target part.
4. The method according to claim 3, wherein, The target area includes multiple drivable regions, each of which includes multiple muscle points. The K driving parameters include at least one parameter reflecting the driving state of each drivable region, wherein the driving state of the drivable region is controlled based on multiple muscle points of the drivable region.
5. The method according to any one of claims 1 to 4, wherein, Also includes: Acquire multiple sample images, where each sample image includes a sample part of the sample object; Obtain the focus label data of the sample object in each of the multiple sample images; Determine a sample driving parameter set for the sample region of each sample image, wherein the sample driving parameter set includes K sample driving parameters, and the K sample driving parameters reflect the overall movement of the sample region; The first detection model is trained based on the focus label data corresponding to the multiple sample images and the multiple sample driving parameter sets corresponding to the multiple sample images.
6. The method according to claim 5, wherein, Acquire multiple sample images, including: Images of each of the N sample objects are collected over P time periods to obtain Q sample images for each sample object, wherein the multiple sample images include the Q sample images corresponding to each of the N sample objects; Obtain the focus label data of the sample object in each of the plurality of sample images, including: For each sample image, based on the time period in which the corresponding acquisition time of each sample image falls, the attention label data corresponding to each sample image is determined. Where N, P and Q are all positive integers.
7. The method according to claim 6, wherein, Collect images of each of N sample objects over P time periods, including: Images of each sample object are collected when it exhibits P behaviors over P time periods, where each of the P behaviors corresponds to a different focus label.
8. The method according to claim 7, wherein, The P types of behavior include the first behavior and the second behavior. The P focus label data include a first focus label value and a second focus label value. The focus label data of the sample object when it has the first behavior is the first focus label value, and the focus label data of the sample object when it has the second behavior is the second focus label value. The first focus label value is greater than the second focus label value.
9. The method according to claim 8, wherein, The first action includes watching a first type of video, and the second action includes watching a second type of video.
10. The method according to claim 5, wherein, Determining the sample driving parameter set for the sample region of each sample image includes: Each sample image is input into the second detection model to obtain the sample output result of the second detection model, wherein the sample output result includes an S-dimensional sample vector about the sample region, wherein the S-dimensional sample vector includes S sample driving parameters; The S-dimensional sample vector is reduced in dimensionality to obtain a K-dimensional sample vector, wherein the K-dimensional sample vector includes the K sample driving parameters. Where K is a positive integer and S is a positive integer greater than K.
11. The method according to claim 10, wherein, The S-dimensional sample vector is reduced in dimensionality to obtain the K-dimensional sample vector, including: A first matrix is formed based on the S-dimensional sample vectors corresponding to the N sample images in the plurality of sample images, wherein the N sample images correspond to the N sample objects respectively; Based on the first matrix, a second matrix and a third matrix are determined, wherein the second matrix is a matrix composed of multiple eigenvectors of the first matrix, and the third matrix is a matrix composed of multiple eigenvalues of the first matrix, wherein the multiple eigenvectors correspond one-to-one with the multiple eigenvalues, and the multiple eigenvalues are sorted in descending order in the third matrix; Select the top K eigenvalues in the third matrix whose proportion is not less than the first ratio, and delete the remaining eigenvectors in the second matrix except for the K eigenvectors corresponding to the top K eigenvalues, to obtain the second matrix after dimensionality reduction; Based on the first matrix and the dimension-reduced second matrix, the dimension-reduced first matrix is obtained; The K-dimensional sample vector is obtained based on the first matrix after dimensionality reduction.
12. The method according to claim 5, wherein, Based on the focus label data corresponding to the multiple sample images and the K sample driving parameters corresponding to the multiple sample images, the first detection model is trained, including: Using the attention label data corresponding to the multiple sample images and the K sample driving parameters corresponding to the multiple sample images, the parameters of the initial model are updated and iterated until the training completion condition is met, and the trained initial model is used as the first detection model. The process of updating and iterating the parameters of the initial model includes performing the following operations for each sample image: The initial output result of the initial model is obtained by inputting the K sample driving parameters of the sample image into the initial model. Based on the initial output of the initial model and the attention label data corresponding to the sample images, the loss information is calculated. Based on the loss information, the parameters of the initial model are updated.
13. An image processing apparatus, comprising: The acquisition module is configured to acquire the image to be detected, wherein the image to be detected includes the target part of the object to be detected; The determination module is configured to determine K driving parameters of the target region based on the detected image, wherein the K driving parameters reflect the overall movement of the target region; The results module is configured to input the K driving parameters into a first detection model to obtain the output results of the first detection model, wherein the output results of the first detection model include the attention data of the detected object. The determining module is further configured as follows: The image to be detected is input into a second detection model to obtain the output of the second detection model. The output of the second detection model includes an S-dimensional vector relating to the target region, wherein the S-dimensional vector includes S driving parameters; and The S-dimensional vector is reduced in dimension to obtain a K-dimensional vector, wherein the K-dimensional vector includes the K driving parameters. Where K is a positive integer and S is a positive integer greater than K.
14. An electronic device, comprising: An imaging device configured to capture an image of the object being inspected; An image processing apparatus configured to receive the detected image and perform the image processing method according to any one of claims 1-12 based on the detected image.
15. An electronic device comprising: Memory stores computer-executable instructions non-transiently; The processor is configured to run computer-executable instructions. The computer-executable instructions are executed by the processor to implement the image processing method according to any one of claims 1-12.
16. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the image processing method according to any one of claims 1-12.
Citation Information
Patent Citations
Facial concentration detection system and detection method
CN107392159A