An electronic device adjustment method, device, electronic device and storage medium

The camera collects images and detects key points of the human body, calculates the image distance across the shoulder to adjust the fan wind speed, solving the problem of inconvenient control methods of existing smart home devices and achieving more intelligent equipment adjustment.

CN119625839BActive Publication Date: 2025-06-27INTELLINDUST INFORMATION TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510154076.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-27
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The control methods of existing smart home devices have limitations, requiring users to operate buttons or hold remote controls, which lack convenience.

Method used

Images are collected by the camera, key points of the human body are detected, the shoulder-span image distance is calculated, and it is converted into the first distance indicating the distance between the target person and the electronic device, and the fan wind speed is adjusted according to the preset wind speed correspondence.

Benefits of technology

It realizes the convenient adjustment of electronic devices without user button operation, improving the intelligence of user experience and device control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625839B_ABST
    Figure CN119625839B_ABST
Patent Text Reader

Abstract

An electronic device adjustment method, device, electronic device, and storage medium provided by an embodiment of the present application are applied to the field of smart home technology. By applying the method of the embodiment of the present application, key point detection can be performed on the first image to be processed collected, the shoulder key points and hip key points corresponding to the target person are determined, the shoulder-hip image distance is calculated according to the coordinates of the shoulder key points and hip key points in the image, and the shoulder-hip image distance is converted into a first distance representing the actual distance between the target person and the electronic device, so that the target wind speed can be determined according to the first distance, and the electronic device can adjust the wind speed according to the distance between the user and the electronic device. Compared with the operation method in the related art, the device adjustment method of the embodiment of the present application can eliminate the need for the user to perform key operations and is more convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart home, and particularly to a method and device for adjusting an electronic device, an electronic device, and a storage medium. Background Art

[0002] Today, with the increasing popularity of smart home, people's demand for the intelligence and convenience of home appliances is growing. In the related art, the common control methods for devices are buttons, knobs, remote controls, etc. However, such control methods usually have certain limitations. For example, when adjusting a device through button or knob operations, the user needs to reach the device to perform the operation; when adjusting a device through remote control operation, the user needs to hold the remote control and aim it at the device for adjustment. Such methods are obviously not intelligent enough. Therefore, how to enable users to adjust devices more conveniently is an urgent problem to be solved. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a method and device for adjusting an electronic device, an electronic device, and a storage medium, so as to achieve more convenient adjustment of the device. The specific technical solutions are as follows:

[0004] In the first aspect of the embodiments of this application, a method for adjusting an electronic device is provided. The electronic device includes a fan and a camera. The method includes:

[0005] Obtain a first image to be processed collected by the camera; wherein, the first image to be processed includes a target person;

[0006] Perform human key point detection on the first image to be processed to determine the shoulder key point and hip key point of the target person in the first image to be processed;

[0007] According to the image coordinates of the shoulder key point and hip key point in the first image to be processed, calculate the shoulder-hip image distance, and convert the shoulder-hip image distance into a first distance, where the first distance represents the distance between the target person and the electronic device;

[0008] Query a preset wind speed correspondence relationship to determine the target wind speed corresponding to the first distance, and adjust the fan to the target wind speed, where the wind speed correspondence relationship is a pre-set correspondence relationship between wind speed and distance.

[0009] In a possible implementation manner, the calculating the shoulder-hip image distance according to the image coordinates of the shoulder key point and hip key point in the first image to be processed and converting the shoulder-hip image distance into a first distance includes:

[0010] Calculate the coordinates of the shoulder key point and / or hip key point through the following formula:

[0011] ;

[0012] ;

[0013] Among them, the k x represents the abscissa of the shoulder key point or the hip key point, and the k y represents the ordinate of the shoulder key point or the hip key point. The abscissa and ordinate of the offset coordinates of the key points predicted by the first neural network model relative to the anchor point are t kx and t ky , respectively. The abscissa and ordinate of the anchor point coordinates are c x and c y respectively; the strides represent the downsampling scaling ratio;

[0014] Calculate the coordinates of the center point of the two shoulder key points and the coordinates of the center point of the two hip key points through the following formula:

[0015] ;

[0016] Among them, the k csx represents the abscissa of the center point of the two shoulder key points, the k lsx represents the abscissa of the left shoulder key point of the target person, and the k rsx represents the abscissa of the right shoulder key point of the target person;

[0017] ;

[0018] Among them, the k csy represents the ordinate of the center point of the two shoulder key points, the k lsy represents the ordinate of the left shoulder key point of the target person, and the k rsy represents the ordinate of the right shoulder key point of the target person;

[0019] ;

[0020] Among them, the k chx represents the abscissa of the center point of the two hip key points, the k lhx represents the abscissa of the left hip key point of the target person, and the k rhx represents the abscissa of the right hip key point of the target person;

[0021] ;

[0022] Among them, the k chy represents the ordinate of the center point of the two hip key points, and the k lhyrepresents the vertical coordinate of the left hip key point of the target person, where k rhy represents the vertical coordinate of the right hip key point corresponding to the target;

[0023] According to the coordinates of the center points of the two shoulder key points and the coordinates of the center points of the two hip key points, calculate the shoulder-hip image distance through the following formula;

[0024] ;

[0025] where, the d pix represents the shoulder-hip image distance;

[0026] Convert the shoulder-hip image distance into the first distance through the following formula;

[0027] ;

[0028] where, the D represents the first distance, the f represents the camera focal length, and the d body represents the preset average shoulder-hip distance.

[0029] In a possible implementation manner, the performing human key point detection on the first to-be-processed image to determine the shoulder key points and hip key points of the target person in the first to-be-processed image includes:

[0030] Input the first to-be-processed image into a pre-trained first neural network model to obtain the shoulder key points and hip key points of the target person in the first to-be-processed image;

[0031] where, the first neural network model also outputs a gesture box of the target person, and the method further includes:

[0032] Cut the first to-be-processed image according to the gesture box to obtain a gesture image;

[0033] Input the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, where the gesture classification result represents the service gesture category corresponding to the gesture in the gesture image;

[0034] Determine a device control instruction according to the gesture classification result to enable the fan to perform an adjustment operation corresponding to the device control instruction.

[0035] In a possible implementation manner, before the inputting the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, the method further includes:

[0036] Input the gesture image into a pre-trained third neural network model to obtain the clarity score and blurriness score of the gesture image; wherein, the third neural network model is used to detect the clarity of the gesture image.

[0037] In the case where the blurriness score is greater than or equal to the clarity score, discard the gesture image.

[0038] In the case where the clarity score is greater than the blurriness score, perform the steps: Input the gesture image into a pre-trained second neural network model to obtain the gesture classification result corresponding to the gesture image.

[0039] In a possible implementation manner, after obtaining the first image to be processed collected by the fan, the method further includes:

[0040] Obtain a second image to be processed collected by the fan, where the second image to be processed is an image collected after the first image to be processed.

[0041] Input the first image to be processed and the second image to be processed into a first neural network model respectively to obtain a first target box and a second target box, where the first target box represents the position of the target person in the first image to be processed, and the second target box represents the position of the target person in the second image to be processed.

[0042] Calculate a third target box of the target person according to the first target box, where the third target box represents the predicted position of the target person in the second image to be processed.

[0043] Match the positional relationship between the second target box and the third target box to obtain a matching result.

[0044] When the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person, calculate the difference between the first target box and the second target box to obtain position transformation information.

[0045] Determine a device adjustment instruction according to the position transformation information so that the fan adjusts according to the device adjustment instruction.

[0046] In a possible implementation manner, the calculating the third target box of the target person according to the first target box includes:

[0047] Calculate an initial state vector of the target person in the first image to be processed according to the first target box.

[0048] Calculate the error between the first predicted state vector and the first predicted state vector according to the initial state vector;

[0049] Calculate the error between the measurement value of the target person in the second image to be detected and the measurement value according to the measurement matrix of the measuring device;

[0050] Combine the error of the measurement value and the error of the first predicted state vector to calculate the Kalman gain;

[0051] Correct the first predicted state vector through the Kalman gain to obtain the first corrected state vector;

[0052] Calculate the third target box according to the first corrected state vector.

[0053] In a possible implementation manner, the step of matching the positional relationship between the second target box and the third target box to obtain a matching result includes:

[0054] Calculate the intersection over union (IoU) of the second target box and the third target box;

[0055] When the intersection over union is greater than or equal to a preset threshold, determine that the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person.

[0056] In a possible implementation manner, the step of inputting the first image to be processed into a pre-trained first neural network model to obtain the shoulder key point and hip key point of the target person in the first image to be processed includes:

[0057] Input the first image to be processed into the first neural network model. The first network structure in the first neural network model performs feature extraction and downsampling processing on the first image to be processed to obtain a first image feature;

[0058] Input the first image feature into the second network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a second image feature; wherein, the network structure of the second network structure is the same as that of the first network structure;

[0059] Input the second image feature into the third network structure for feature enhancement processing to obtain a third image feature;

[0060] Input the third image feature into the fourth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a fourth image feature; wherein, the network structure of the fourth network structure is the same as that of the second network structure;

[0061] Input the fourth image feature into the fifth network structure for feature enhancement processing to obtain the fifth image feature; wherein, the network structure of the fifth network structure is the same as that of the third network structure;

[0062] Input the fifth image feature into the sixth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the sixth image feature, wherein the network structure of the sixth network structure is the same as that of the fourth network structure;

[0063] Input the sixth image feature into the seventh network structure for feature enhancement processing to obtain the seventh image feature, wherein the network structure of the seventh network structure is the same as that of the fifth network structure;

[0064] Input the seventh image feature into the eighth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the eighth image feature; wherein the network structure of the eighth network structure is the same as that of the sixth network structure;

[0065] Input the eighth image feature into the ninth network structure for multi-scale pooling operation to fuse image features of different scales to obtain the ninth image feature;

[0066] Input the ninth image feature into the tenth network structure for convolution processing and feature enhancement processing to obtain the tenth image feature, wherein the sizes of the tenth image feature, the seventh image feature, and the fifth image feature are different;

[0067] Perform convolution processing and feature enhancement processing on the tenth image feature through the eleventh network structure to obtain the eleventh image feature, wherein the network structure of the eleventh network structure is the same as that of the eighth network structure;

[0068] Perform upsampling processing on the eleventh image feature through the first upsampling module to obtain the twelfth image feature, wherein the size of the twelfth image feature is the same as that of the seventh image feature;

[0069] Perform feature fusion on the twelfth image feature and the seventh image feature through the first feature fusion layer to obtain the thirteenth image feature;

[0070] Input the thirteenth image feature into the twelfth network structure for convolution processing and feature enhancement processing to obtain the fourteenth image feature, wherein the network structure of the twelfth network structure is the same as that of the tenth network structure;

[0071] Input the fourteenth image feature into the thirteenth network structure for convolution processing and feature enhancement processing to obtain the fifteenth image feature, wherein the network structure of the thirteenth network structure is the same as that of the eleventh network structure;

[0072] The sixteenth image feature is obtained by performing upsampling processing on the fifteenth image feature through a second upsampling module, where the size of the sixteenth image feature is the same as that of the fifth image feature;

[0073] The seventeenth image feature is obtained by performing feature fusion on the fifth image feature and the sixteenth image feature through a second feature fusion layer;

[0074] The eighteenth image feature is obtained by inputting the seventeenth image feature into a fourteenth network structure for convolution processing and feature enhancement processing, where the network structure of the fourteenth network structure is the same as that of the twelfth network structure;

[0075] The nineteenth image feature is obtained by inputting the eighteenth image feature into a fifteenth network structure for convolution processing, feature enhancement processing, and downsampling processing, where the network structure of the fifteenth network structure is the same as that of the thirteenth network structure;

[0076] The twentieth image feature is obtained by performing feature fusion on the nineteenth image feature and the fifteenth image feature through a third feature fusion layer;

[0077] The twenty - first image feature is obtained by inputting the twentieth image feature into a sixteenth network structure for convolution processing and feature enhancement processing, where the network structure of the sixteenth network structure is the same as that of the fourteenth network structure;

[0078] The twenty - second image feature is obtained by inputting the twenty - first image feature into a seventeenth network structure for convolution processing, feature enhancement processing, and downsampling processing, where the network structure of the seventeenth network structure is the same as that of the fifteenth network structure;

[0079] The twenty - third image feature is obtained by performing feature fusion on the twenty - second image feature and the eleventh image feature through a fourth feature fusion layer;

[0080] The twenty - fourth image feature is obtained by inputting the twenty - third image feature into an eighteenth network structure for convolution processing and feature enhancement processing, where the network structure of the eighteenth network structure is the same as that of the sixteenth network structure;

[0081] The twenty - fourth image feature, the twenty - first image feature, and the eighteenth image feature are respectively input into network head modules corresponding to their respective sizes for classification prediction to obtain the shoulder key points and hip key points of the target person.

[0082] In a second aspect of the embodiments of the present application, an electronic device adjustment device is provided. The electronic device includes a fan and a camera, and the device includes:

[0083] The first image to be processed acquisition module is used to acquire the first image to be processed collected by the camera; wherein, the first image to be processed includes a target person;

[0084] The key point detection module is used to perform human key point detection on the first image to be processed, and determine the shoulder key points and hip key points of the target person in the first image to be processed;

[0085] The first distance calculation module is used to calculate the shoulder-hip image distance according to the image coordinates of the shoulder key points and hip key points in the first image to be processed, and convert the shoulder-hip image distance into a first distance, where the first distance represents the distance between the target person and the electronic device;

[0086] The target wind speed determination module is used to query the preset wind speed correspondence relationship, determine the target wind speed corresponding to the first distance, and adjust the fan to the target wind speed, where the wind speed correspondence relationship is a pre-set correspondence relationship between wind speed and distance.

[0087] In a possible implementation manner, the first distance calculation module includes:

[0088] The key point coordinate calculation sub-module is specifically used to calculate the coordinates of the shoulder key point and / or hip key point through the following formula:

[0089] ;

[0090] ;

[0091] Wherein, the k x represents the abscissa of the shoulder key point or hip key point, the k y represents the ordinate of the shoulder key point or hip key point, the abscissa and ordinate of the offset coordinates of the key point predicted by the first neural network model relative to the anchor point are t kx and t ky , the abscissa and ordinate of the anchor point coordinates are c x and c y ; the strides represents the downsampling scaling ratio;

[0092] The center point coordinate calculation sub-module is specifically used to calculate the coordinates of the center point of the two shoulder key points and the coordinates of the center point of the two hip key points through the following formula:

[0093] ;

[0094] Wherein, the k csx represents the abscissa of the center point of the two shoulder key points, the k lsxrepresents the abscissa of the left shoulder key point of the target person, and the k rsx represents the abscissa of the right shoulder key point of the target person;

[0095] ;

[0096] wherein, the k csy represents the ordinate of the center point of the two shoulder key points, and the k lsy represents the ordinate of the left shoulder key point of the target person, and the k rsy represents the ordinate of the right shoulder key point of the target person;

[0097] ;

[0098] wherein, the k chx represents the abscissa of the center point of the two hip key points, and the k lhx represents the abscissa of the left hip key point of the target person, and the k rhx represents the abscissa of the right hip key point of the target person;

[0099] ;

[0100] wherein, the k chy represents the ordinate of the center point of the two hip key points, and the k lhy represents the ordinate of the left hip key point of the target person, and the k rhy represents the ordinate of the corresponding right hip key point of the target;

[0101] The shoulder-hip image distance calculation sub-module is specifically configured to calculate the shoulder-hip image distance according to the coordinates of the center points of the two shoulder key points and the coordinates of the center points of the two hip key points through the following formula;

[0102] ;

[0103] wherein, the d pix represents the shoulder-hip image distance;

[0104] The first distance calculation sub-module is specifically configured to convert the shoulder-hip image distance into the first distance through the following formula;

[0105] ;

[0106] wherein, the D represents the first distance, the f represents the camera focal length, and the d body represents the preset average shoulder-hip distance.

[0107] In a possible implementation manner, the key point detection module includes:

[0108] The first neural network model prediction sub-module is specifically configured to input the first image to be processed into a pre-trained first neural network model to obtain the shoulder key points and hip key points of the target person in the first image to be processed;

[0109] Wherein, the first neural network model also outputs a gesture box of the target person, and the device further includes:

[0110] An image cutting module, configured to cut the first image to be processed according to the gesture box to obtain a gesture image;

[0111] A gesture detection module, configured to input the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, where the gesture classification result represents a service gesture category corresponding to the gesture in the gesture image;

[0112] A device control instruction determination module, configured to determine a device control instruction according to the gesture classification result so that the fan executes an adjustment operation corresponding to the device control instruction.

[0113] In a possible implementation manner, the device further includes:

[0114] An image clarity detection module, configured to input the gesture image into a pre-trained third neural network model to obtain a clarity score and a blurriness score of the gesture image; wherein, the third neural network model is used to perform clarity detection on the gesture image;

[0115] A blurred image processing module, configured to discard the gesture image when the blurriness score is greater than or equal to the clarity score;

[0116] A clear image processing module, configured to, when the clarity score is greater than the blurriness score, perform the following steps: input the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image.

[0117] In a possible implementation manner, the device further includes:

[0118] A second image to be processed acquisition module, configured to acquire a second image to be processed collected by the fan, where the second image to be processed is an image collected after the first image to be processed;

[0119] The target box acquisition module is configured to input the first image to be processed and the second image to be processed into a first neural network model respectively, to obtain a first target box and a second target box, where the first target box represents the position of the target person in the first image to be processed, and the second target box represents the position of the target person in the second image to be processed;

[0120] The third target box prediction module is configured to calculate a third target box of the target person according to the first target box, where the third target box represents the predicted position of the target person in the second image to be processed;

[0121] The position relationship matching module is configured to match the position relationship between the second target box and the third target box to obtain a matching result;

[0122] The position transformation information calculation module is configured to calculate the difference between the first target box and the second target box to obtain position transformation information when the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person;

[0123] The device adjustment instruction determination module is configured to determine a device adjustment instruction according to the position transformation information, so that the fan is adjusted according to the device adjustment instruction.

[0124] In a possible implementation manner, the third target box prediction module includes:

[0125] The initial state vector calculation sub-module is specifically configured to calculate an initial state vector of the target person in the first image to be processed according to the first target box;

[0126] The prediction vector calculation sub-module is specifically configured to calculate the error between a first predicted state vector and the first predicted state vector according to the initial state vector;

[0127] The measurement value calculation sub-module is specifically configured to calculate the error between the measurement value of the target person in the second image to be detected and the measurement value according to the measurement matrix of the measurement device;

[0128] The Kalman gain calculation sub-module is specifically configured to calculate a Kalman gain by combining the error of the measurement value and the error of the first predicted state vector;

[0129] The prediction vector correction sub-module is specifically configured to correct the first predicted state vector through the Kalman gain to obtain a first corrected state vector;

[0130] The third target box calculation sub-module is specifically configured to calculate the third target box according to the first corrected state vector.

[0131] In a possible implementation, the position relationship matching module includes:

[0132] An intersection over union (IoU) calculation sub-module, specifically configured to calculate the IoU of the second target box and the third target box;

[0133] An IoU judgment sub-module, specifically configured to determine that the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person when the IoU is greater than or equal to a preset threshold.

[0134] In a possible implementation, the key point detection module includes:

[0135] An image processing sub-module, specifically configured to input the first image to be processed into the first neural network model. The first network structure in the first neural network model performs feature extraction and downsampling processing on the first image to be processed to obtain a first image feature;

[0136] Input the first image feature into a second network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a second image feature; wherein, the second network structure has the same network structure as the first network structure;

[0137] Input the second image feature into a third network structure for feature enhancement processing to obtain a third image feature;

[0138] Input the third image feature into a fourth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a fourth image feature; wherein, the fourth network structure has the same network structure as the second network structure;

[0139] Input the fourth image feature into a fifth network structure for feature enhancement processing to obtain a fifth image feature; wherein, the fifth network structure has the same network structure as the third network structure;

[0140] Input the fifth image feature into a sixth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a sixth image feature, wherein, the sixth network structure has the same network structure as the fourth network structure;

[0141] Input the sixth image feature into a seventh network structure for feature enhancement processing to obtain a seventh image feature, wherein, the seventh network structure has the same network structure as the fifth network structure;

[0142] Input the seventh image feature into the eighth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain an eighth image feature; wherein, the eighth network structure has the same network structure as the sixth network structure;

[0143] Input the eighth image feature into the ninth network structure for multi-scale pooling operation to fuse image features of different scales and obtain a ninth image feature;

[0144] Input the ninth image feature into the tenth network structure for convolution processing and feature enhancement processing to obtain a tenth image feature, wherein the tenth image feature, the seventh image feature, and the fifth image feature have different sizes;

[0145] Perform convolution processing and feature enhancement processing on the tenth image feature through the eleventh network structure to obtain an eleventh image feature, wherein the eleventh network structure has the same network structure as the eighth network structure;

[0146] Perform upsampling processing on the eleventh image feature through the first upsampling module to obtain a twelfth image feature, wherein the twelfth image feature has the same size as the seventh image feature;

[0147] Perform feature fusion on the twelfth image feature and the seventh image feature through the first feature fusion layer to obtain a thirteenth image feature;

[0148] Input the thirteenth image feature into the twelfth network structure for convolution processing and feature enhancement processing to obtain a fourteenth image feature, wherein the twelfth network structure has the same network structure as the tenth network structure;

[0149] Input the fourteenth image feature into the thirteenth network structure for convolution processing and feature enhancement processing to obtain a fifteenth image feature, wherein the thirteenth network structure has the same network structure as the eleventh network structure;

[0150] Perform upsampling processing on the fifteenth image feature through the second upsampling module to obtain a sixteenth image feature, wherein the sixteenth image feature has the same size as the fifth image feature;

[0151] Perform feature fusion on the fifth image feature and the sixteenth image feature through the second feature fusion layer to obtain a seventeenth image feature;

[0152] Input the seventeenth image feature into the fourteenth network structure for convolution processing and feature enhancement processing to obtain an eighteenth image feature; wherein, the fourteenth network structure has the same network structure as the twelfth network structure;

[0153] Input the eighteenth image feature into the fifteenth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the nineteenth image feature, where the fifteenth network structure has the same network structure as the thirteenth network structure;

[0154] Perform feature fusion on the nineteenth image feature and the fifteenth image feature through the third feature fusion layer to obtain the twentieth image feature;

[0155] Input the twentieth image feature into the sixteenth network structure for convolution processing and feature enhancement processing to obtain the twenty-first image feature, where the sixteenth network structure has the same network structure as the fourteenth network structure;

[0156] Input the twenty-first image feature into the seventeenth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the twenty-second image feature, where the seventeenth network structure has the same network structure as the fifteenth network structure;

[0157] Perform feature fusion on the twenty-second image feature and the eleventh image feature through the fourth feature fusion layer to obtain the twenty-third image feature;

[0158] Input the twenty-third image feature into the eighteenth network structure for convolution processing and feature enhancement processing to obtain the twenty-fourth image feature, where the eighteenth network structure has the same network structure as the sixteenth network structure;

[0159] Input the twenty-fourth image feature, the twenty-first image feature, and the eighteenth image feature into the network head modules corresponding to their respective sizes for classification prediction to obtain the shoulder key points and hip key points of the target person.

[0160] In another aspect of the embodiments of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0161] The memory is used to store a computer program;

[0162] When the processor is used to execute the program stored on the memory, it implements the method steps of any one of the first aspects of the embodiments of the present application.

[0163] In another aspect of the embodiments of the present application, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps of any one of the first aspects of the embodiments of the present application.

[0164] An embodiment of the present application also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the electronic device adjustment method described in any one of the above.

[0165] Advantages of the embodiment of the present application:

[0166] An electronic device adjustment method, device, electronic device, and storage medium provided by an embodiment of the present application can perform key point detection on a first image to be processed collected, determine the shoulder key points and hip key points corresponding to the target person, calculate the shoulder-hip image distance according to the coordinates of the shoulder key points and hip key points in the image, and convert the shoulder-hip image distance into a first distance representing the actual distance between the target person and the electronic device, so that the target wind speed can be determined according to the first distance, and the electronic device can adjust the wind speed according to the distance between the user and the electronic device. Compared with the operation method in the related art, the device adjustment method of the embodiment of the present application can eliminate the need for the user to perform key operations, which is more convenient.

[0167] Of course, implementing any product or method of the present application does not necessarily require achieving all the above advantages at the same time. Description of the Drawings

[0168] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.

[0169] Figure 1 It is a flowchart of an electronic device adjustment method provided by an embodiment of the present application;

[0170] Figure 2 It is a schematic diagram of the network structure of the second neural network model or the third neural network model provided by an embodiment of the present application;

[0171] Figure 3 It is a schematic diagram of the network structure of the second network unit provided by an embodiment of the present application;

[0172] Figure 4 It is a schematic diagram of the network structure of the third network unit provided by an embodiment of the present application;

[0173] Figure 5 It is a schematic diagram of the network structure of the first neural network model provided by an embodiment of the present application;

[0174] Figure 6 It is a schematic diagram of the network structure of the first network structure provided by an embodiment of the present application;

[0175] Figure 7 Schematic diagram of the third network structure provided by an embodiment of the present application;

[0176] Figure 8 Schematic diagram of the residual structure network provided by an embodiment of the present application;

[0177] Figure 9 Schematic diagram of the ninth network structure provided by an embodiment of the present application;

[0178] Figure 10 Schematic diagram of the tenth network structure provided by an embodiment of the present application;

[0179] Figure 11 Another flowchart of the electronic device adjustment method provided by an embodiment of the present application;

[0180] Figure 12 Schematic diagram of a structure of an electronic device adjustment device provided by an embodiment of the present application;

[0181] Figure 13 Schematic diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0182] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the protection scope of the present application.

[0183] Today, with the increasing popularity of smart homes, people's demand for the intelligence and convenience of home appliances is growing. In the related art, the control methods for devices are usually buttons, knobs, remote controls, etc. However, such control methods usually have certain limitations. For example, when adjusting a device by operating buttons or knobs, the user needs to reach the device to operate; when adjusting a device by operating a remote control, the user needs to hold the remote control and aim it at the device for adjustment. Such a method is obviously not intelligent enough. Therefore, how to enable users to adjust devices more conveniently is an urgent problem to be solved.

[0184] To solve at least one of the above problems, in the first aspect of the embodiments of the present application, an electronic device adjustment method is provided, and the method includes steps as Figure 1 shown below:

[0185] Step S101: Obtain a first image to be processed collected by a camera.

[0186] Among them, the first image to be processed includes a target person. The electronic device according to the embodiments of the present application includes a fan and a camera. In one example, the electronic device may be an electric fan, and the camera is disposed on the electric fan for collecting an image of the environment where the electric fan is located. In another example, the electronic device may be an air conditioner. Among them, the air conditioner includes a fan structure, and the camera may be disposed at the air conditioner panel for collecting an image of the environment where the air conditioner is located.

[0187] In practical applications, the camera can perform image acquisition regularly. For example, image acquisition is performed every 0.1 s, 0.5 s, or 1 s, and the acquired images are subjected to image recognition. If a person is recognized in the image, it is determined that the image is the first image to be processed; if no person is recognized, the image is discarded or not processed.

[0188] Step S102: Perform human key point detection on the first image to be processed, and determine the shoulder key points and hip key points of the target person in the first image to be processed.

[0189] Among them, when there are multiple human objects in the first image to be processed, the target person can be determined by the pixel area of each human object in the first image to be processed. In one example, contour detection is performed on the first image to be processed to obtain the image regions corresponding to each human object, the areas of the image regions corresponding to each human object are calculated, and the human object corresponding to the image region with the largest area is determined as the target person.

[0190] When performing human key point detection, key point detection can be directly performed through a pre-trained first neural network model. Among them, the first neural network model is a deep learning model for target detection. In one example, the target detection network model is a network model with a CSPNet (Cross Stage Partial Network, lightweight neural network structure) structure, OpenPose (a human pose estimation library based on deep learning), or a BlazePose (lightweight convolutional neural network structure) neural network, an HRNet (High-Resolution Network, high-resolution network) structure neural network model, or other network models.

[0191] When training the first neural network model, it can be trained using the collected dataset or the dataset in related technologies. In one example, training the first neural network using the COCO (Common Objects in Context, a dataset for image recognition) dataset enables the first neural network to detect 17 key points of a human body (nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle). In another example, when the first neural network model performs human key point detection on the first image to be processed, it can classify the features of each pixel point to obtain the key point classification result of that pixel point. For example, through the first neural network model, the scores of each key point of a certain pixel point are obtained, and based on these scores of each key point, the key points of the target person are determined, so as to select the shoulder key points and hip key points of the target person from these key points.

[0192] Among them, the key point score is the confidence score of the pixel point for each key point. In one example, taking the key point detection corresponding to the above 17 key points as an example, the scores of the above 17 key points corresponding to a pixel point are (0.3, 0.1, 0.45, 0.14, 0.65, 0.9, 0.1, 0.26, 0.56, 0.14, 0.05, 0.02, 0.04, 0.03, 0.74, 0.03, 0.35). Among them, the scores of each key point are independent of each other. When screening the key point scores, they can be screened according to a preset threshold. In one example, 0.5 is selected as the threshold for screening. Among them, the score of the nose key point is 0.9, and since this score is greater than the threshold, it is considered that this key point can be used as the nose key point. In practical applications, the selection and rejection of key points can also be based on the target box. Among them, the first neural network model can predict the target box of each pixel point, calculate the intersection over union between the target box and the ground truth box, screen out the target box with the largest intersection over union, and use this target box as the target box corresponding to the final target person. For the key points in this target box, another screening of key points can be performed according to the key point scores in the above example to determine the final shoulder key points and hip key points.

[0193] Step S103: Calculate the shoulder-hip image distance based on the image coordinates of the shoulder key points and hip key points in the first image to be processed, and convert the shoulder-hip image distance into a first distance.

[0194] Among them, the first distance represents the distance between the target person and the electronic device. When converting the shoulder-hip image distance into the actual distance between the target person and the electronic device, a mathematical model for converting the two-dimensional image distance into the three-dimensional actual distance can be determined according to the pre-calibrated parameters such as the camera's internal and external parameters, so as to convert the shoulder-hip image distance into the first distance.

[0195] In practical applications, when determining the shoulder key point and hip key point of the target person in step S102, the coordinates of the shoulder key point and hip key point in the first image to be processed can be directly obtained. Thus, the shoulder-hip image distance can be calculated based on the coordinates of the shoulder key point and hip key point.

[0196] Step S104: Query the preset wind speed correspondence relationship, determine the target wind speed corresponding to the first distance, and adjust the fan to the target wind speed.

[0197] Among them, the wind speed correspondence relationship is the pre-set correspondence relationship between wind speed and distance. This wind speed correspondence relationship can be set according to actual business requirements. For example, different wind speed levels can be set according to different distance ranges. In one example, when the first distance is within the range of 8m - 10m, the corresponding wind speed level is level 5; when the first distance is within the range of 6m - 8m, the corresponding wind speed level is level 4; when the first distance is within the range of 4m - 6m, the corresponding wind speed level is level 3; when the first distance is within the range of 2m - 4m, the corresponding wind speed level is level 2; when the first distance is within the range of 0m - 2m, the corresponding wind speed level is level 1.

[0198] By applying the method of the embodiment of the present application, key point detection can be performed on the collected first image to be processed to determine the shoulder key point and hip key point corresponding to the target person. According to the coordinates of the shoulder key point and hip key point in the image, the shoulder-hip image distance is calculated, and the shoulder-hip image distance is converted into the first distance representing the actual distance between the target person and the electronic device. Thus, the target wind speed can be determined according to the first distance, enabling the electronic device to adjust the wind speed according to the distance between the user and the electronic device. Compared with the operation method in the related art, the device adjustment method of the embodiment of the present application can eliminate the need for the user to perform key operations, making it more convenient.

[0199] In a possible implementation manner, step S103 can calculate the first distance through the following steps:

[0200] Step 1: Calculate the coordinates of the shoulder key point and / or hip key point through the following formula:

[0201] ;

[0202] ;

[0203] Among them, k x represents the abscissa of the shoulder key point or hip key point, k y represents the ordinate of the shoulder key point or hip key point, and the abscissa and ordinate of the offset coordinates of the key point predicted by the first neural network model relative to the anchor point are t kx and tky The abscissa and ordinate of the anchor point coordinates are c x and c y .

[0204] Among them, the first neural network model performs downsampling on the first image to be processed to obtain downsampled image features. Take the width and height of the downsampled image features to form a two-dimensional plane, and each coordinate point on this two-dimensional plane is an anchor point. Its anchor point coordinates can be determined according to the size of the feature map. In one example, if the size of the downsampled image features is 20x20, the anchor point coordinates in the downsampled image features are (0, 0), (0, 1), (0, 2)... (19, 19) in sequence.

[0205] Step 2: Calculate the coordinates of the center point of the two shoulder key points and the coordinates of the center point of the two hip key points through the following formula:

[0206] ;

[0207] where k csx represents the abscissa of the center point of the two shoulder key points, k lsx represents the abscissa of the left shoulder key point of the target person, k rsx represents the abscissa of the right shoulder key point of the target person;

[0208] ;

[0209] where k csy represents the ordinate of the center point of the two shoulder key points, k lsy represents the ordinate of the left shoulder key point of the target person, k rsy represents the ordinate of the right shoulder key point of the target person;

[0210] ;

[0211] where k chx represents the abscissa of the center point of the two hip key points, k lhx represents the abscissa of the left hip key point of the target person, k rhx represents the abscissa of the right hip key point of the target person;

[0212] ;

[0213] where k chy represents the ordinate of the center point of the two hip key points, k lhy represents the ordinate of the left hip key point of the target person, k rhy represents the ordinate of the corresponding right hip key point of the target;

[0214] Step 3: According to the coordinates of the center points of the two shoulder key points and the coordinates of the center points of the two hip key points, calculate the shoulder-hip image distance through the following formula:

[0215] ;

[0216] where d pix represents the shoulder-hip image distance;

[0217] Step 4: Convert the shoulder-hip image distance into the first distance through the following formula:

[0218] ;

[0219] where D represents the first distance, f represents the camera focal length, and d body represents the preset average shoulder-hip distance.

[0220] In one example, the electronic device can set different modes to distinguish users of different heights. For example, the electronic device has an adult mode and a child mode. Then, the average shoulder-hip distance of the adult population can be statistically calculated as the preset average shoulder-hip distance when the electronic device works in the adult mode, and the average shoulder-hip distance of the child population can be statistically calculated as the preset average shoulder-hip distance when the electronic device works in the child mode.

[0221] By applying the method of the embodiment of the present application, the conversion of the shoulder-hip image distance of the target person to the first distance can be realized through the above formula, so that the target wind speed can be determined according to the first distance, and further, the wind speed can be adjusted according to the distance between the user and the electronic device, so that the user does not need to reach the device to press a button, and the adjustment of the electronic device can be carried out more conveniently.

[0222] In a possible implementation manner, the electronic device can also recognize the gestures of the target person, so that the user can control the electronic device through gestures. Among them, step S102 can be realized through the following steps:

[0223] Input the first image to be processed into a pre-trained first neural network model to obtain the shoulder key points and hip key points of the target person in the first image to be processed.

[0224] In practical applications, the first neural network model can also output the gesture box of the target person. Then, the method of the embodiment of the present application can further include the following steps:

[0225] Step 1: Cut the first image to be processed according to the gesture box to obtain the gesture image.

[0226] In one example, when the first neural network model outputs the gesture box of the target person, it can first predict the gesture box for each pixel point to obtain the gesture box and the gesture box score corresponding to each pixel point, and select the gesture box with the highest gesture box score as the gesture box of the target person. Then, according to the gesture box, the gesture image is cropped from the first image to be processed.

[0227] Step 2: Input the gesture image into a pre-trained second neural network model to obtain the gesture classification result corresponding to the gesture image.

[0228] Among them, the gesture classification result represents the business gesture category corresponding to the gesture in the gesture image. The network structure of the second neural network model can be the neural network structure used for classification prediction in the related art. In one example, the second neural network model can be a neural network with a CNN (Convolutional Neural Networks) structure, such as a LeNet (a convolutional neural network), an AlexNet (a neural network), a VGG (Visual Geometry Group) network structure, etc.

[0229] Among them, the output value of the second neural network model can be the same as the number of gesture categories. In one example, there are 5 types of business gestures, so the second neural network model will output 5 values for the gesture image. These 5 values respectively correspond to the classification scores of the gesture image for each type of business gesture, and the sum of these 5 values is 1. For example, the output of the second neural network model is [0.1, 0.5, 0.2, 0.05, 0.15]. Among them, the score of 0.5 is the largest, and the subscript position of this value in the array is 1, so it is determined that the business gesture category corresponding to the gesture image is category 1.

[0230] Among them, the business gesture is a pre-set gesture pattern. In another example, when recognizing the gesture image, the gesture in the gesture image can also be directly matched with the pre-set gesture pattern in terms of shape, and the business gesture pattern with the highest matching degree is selected as the gesture category corresponding to the gesture image.

[0231] Step 3: Determine the device control instruction according to the gesture classification result so that the fan executes the adjustment operation corresponding to the device control instruction.

[0232] Similarly, in practical applications, the corresponding relationship between each business gesture category and the fan adjustment operation is preset. For example, business gesture category 0 corresponds to increasing the wind speed level, business gesture category 1 corresponds to decreasing the wind speed level, business gesture category 2 corresponds to increasing the outlet air temperature, business gesture category 3 corresponds to decreasing the outlet air temperature, and so on. In one example, the preset business gestures include aiming gestures, digital gestures, OK (good) gestures, like gestures, etc. The business functions corresponding to each business gesture can be preset. For example, the aiming gesture corresponds to the on / off of the fan. After the electronic device recognizes that the business gesture is an aiming gesture, it can execute the opposite state according to the current on / off state of the fan. If the current state of the fan is the on state, the electronic device sends a turn-off instruction to the fan after detecting the aiming gesture; if the current state of the fan is the off state, the electronic device sends a turn-on instruction to the fan after detecting the aiming gesture.

[0233] In another example, the digital (1, 2, 3, 4, 5) gestures can be set as the fan wind speed control gestures, corresponding to different fan wind speed level control instructions respectively. In another example, the OK gesture corresponds to the target tracking function of the fan. After the electronic device recognizes that the business gesture is an OK gesture, it sends a target tracking instruction to the fan, so that the fan can automatically adjust the wind speed and direction according to the user's distance. In practical applications, the business gestures and their corresponding functions can include but are not limited to the above business gestures and their corresponding functions. For example, the business gestures and their corresponding functions can also include that the vertical eight gesture corresponds to the function of the fan blowing in a fixed direction, the like gesture corresponds to the natural wind function of the fan, and the six gesture corresponds to the sleep wind function of the fan.

[0234] By applying the method of the embodiment of the present application, a gesture image can be obtained from the to-be-processed image through a gesture frame, so as to perform gesture detection on the gesture image, obtain the classification result of the gesture, and determine the relevant adjustment operation of the fan according to the classification result of the gesture. Furthermore, the user can directly control the electronic device through gestures. Compared with the prior art, the method of the embodiment of the present application can eliminate the need for the user to go to the device to perform the control, nor to find a control device such as a remote control to control the electronic device, realizing a more convenient way of controlling the electronic device.

[0235] In a possible implementation manner, since the gesture image is intercepted from the first to-be-processed image, its clarity may not necessarily meet the requirement for the second neural network model to clearly identify the gesture category. Therefore, before inputting the gesture image into the second neural network model in step 2, the embodiment of the present application can also perform quality detection on the gesture image first. In one example, the quality detection of the gesture image can be performed through the following steps:

[0236] Step A: Input the gesture image into a pre-trained third neural network model to obtain the clarity score and blurriness score of the gesture image.

[0237] Step B: If the blur score is greater than or equal to the clarity score, discard the gesture image.

[0238] Step C: If the clarity score is greater than the blur score, execute Step 2.

[0239] Among them, the third neural network model is used to detect the clarity of the gesture image. The network structure of the third neural network model can be the same as or different from that of the second neural network model. In one example, both the third neural network model and the second neural network model are neural networks with a CNN structure. In practical applications, the third neural network model has been pre-trained through a training set and a test set, where both the training set and the test set include clear photos and blurred photos.

[0240] When the third neural network model classifies and predicts the gesture image, it will respectively obtain the blur score and clarity score of the gesture image. When the blur score is greater than or equal to the clarity score, it is considered that the clarity of the gesture image is not sufficient for the second neural network model to clearly identify the business gesture category corresponding to the gesture in the gesture image. Then, the gesture image can be discarded or not processed. When the clarity score of the gesture image is greater than the blur score, it is considered that the gesture image is relatively clear, and thus Step 2 is executed to input it into the second neural network model.

[0241] In one example, the network structures of the second neural network model and the third neural network model are both Figure 2 the network structures shown. The gesture image (a gesture image with a size of 128x128x3) is input into the second neural network model or the third neural network model. After the first network unit, the second network unit, and multiple third network units perform feature extraction, convolution, feature enhancement, etc. on the gesture image to obtain the image features corresponding to the gesture image, and then the pooling layer and the fully connected layer are used to classify and predict the image features to obtain the category of the gesture image for output.

[0242] Among them, the first network unit includes a convolutional layer, a batch normalization layer, and an activation function layer. The network structure of the second network unit is as Figure 3 shown, including a network module with the same structure as the first network unit, a fourth network unit, a feature fusion layer, and an activation function layer. Among them, the fourth network unit includes a convolutional layer and a batch normalization layer. The network structure of the third network unit is as Figure 4 shown. Compared with the second network unit, the third network unit also includes a network module with the same structure as the fourth network unit, so that more image features of the gesture image can be obtained.

[0243] Applying the method of the embodiment of the present application, before inputting the gesture image into the second neural network model for gesture detection, the gesture image is input into the third neural network model for clarity detection, so as to ensure that the gesture images input into the second neural network model are all clear images, which is convenient for the second neural network model to perform gesture detection and also avoids the waste of computing power caused by performing gesture detection on blurred images.

[0244] In a possible implementation manner, the embodiment of the present application can also control the direction of the fan according to the position of the target person. In one example, after step S101, the embodiment of the present application may further include the following steps:

[0245] Step (1): Obtain the second image to be processed collected by the fan.

[0246] Wherein, the second image to be processed is the image collected after the first image to be processed. In practical applications, the camera collects images at regular intervals. In one example, the camera collects images every 5s. Then at 10s, the camera will collect two images, and the first collected image is determined as the first image to be processed, and the second collected image is determined as the second image to be processed.

[0247] Step (2): Input the first image to be processed and the second image to be processed into the first neural network model respectively to obtain the first target box and the second target box.

[0248] Wherein, the first target box represents the position of the target person in the first image to be processed, and the second target box represents the position of the target person in the second image to be processed. In practical applications, when the first neural network model outputs key points and gesture boxes, it will also output target boxes. The target box includes the coordinates, width, and height of the target box. The position of the target box can be calculated according to the width and height of the prediction box predicted by the first neural network model, the offset coordinates of the prediction box relative to the anchor point, the anchor point coordinates, and the downsampling ratio of the first image to be processed and the second image to be processed.

[0249] In one example, the first neural network model performs downsampling processing on the first image to be processed or the second image to be processed to obtain a two-dimensional image feature, and each coordinate point on the two-dimensional image feature is an anchor point. The first neural network model will output a set of prediction values for each anchor point, and the prediction values include the offset coordinates of the prediction box relative to the anchor point, the width of the prediction box, the height of the prediction box, and the score of the target box. Among them, the score of the target box is the confidence score of the anchor point being the pixel point corresponding to the target person. Select a set of prediction values with the highest target box score for target box calculation.

[0250] In one example, the target box can be calculated by the following formula:

[0251] ;

[0252] ;

[0253] Among them, the coordinates of the target box are (b x , b y ), the offset coordinates of the predicted box relative to the anchor point are (t x , t y ), the coordinates of the anchor point are (c x , c y ), and strides represents the downsampling scale ratio.

[0254] ;

[0255] ;

[0256] Among them, the width and height of the target box are (b w , b h ), and the width and height of the predicted target box are (t w , t h ).

[0257] Step (III): Calculate the third target box of the target person according to the first target box.

[0258] Among them, the third target box represents the predicted position of the target person in the second image to be processed. When predicting the third target box, the prediction and tracking of the target person can be realized by means of a Kalman filter. In another example, the prediction and tracking of the target person can also be realized by means of particle filtering, set membership filtering, etc.

[0259] Step (IV): Match the positional relationship between the second target box and the third target box to obtain a matching result.

[0260] In the embodiment of the present application, the second target box represents the position of the target person in the second image to be processed in the actually acquired image, and the third target box represents the predicted position of the target person in the second image to be processed. To match the positional relationship between the second target box and the third target box, it can be judged by calculating the intersection over union between the second target box and the third target box.

[0261] In one example, the following steps can be used to match the positional relationship between the second target box and the third target box:

[0262] Step a: Calculate the intersection over union between the second target box and the third target box.

[0263] Step b: When the intersection over union (IoU) is greater than or equal to a preset threshold, determine that the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person.

[0264] Specifically, the intersection and union of the second target box and the third target box can be calculated respectively based on the coordinates, width, and height of the second target box and the third target box, so as to calculate the ratio of the intersection to the union, and obtain the IoU of the second target box and the third target box.

[0265] The preset threshold can be determined according to actual business requirements. The larger the preset threshold is set, the higher the credibility of determining the same target person. However, correspondingly, if the moving speed of the target person is relatively fast, the target persons in the first image to be processed and the second image to be processed may be misjudged as different persons. In one example, the preset threshold is set to 0.5. If the IoU is greater than or equal to 0.5, it is considered that the second target box and the third target box can be associated, and thus it is determined that the target person in the first image to be processed and the target person in the second image to be processed are the same target person.

[0266] By applying the method of this embodiment of the present application, the position matching is performed by calculating the IoU of the second target box and the third target box, so that the position of the target person in the second image to be processed collected can be associated with the position of the target person in the second image to be processed predicted. If there is an association, it is determined that the first image to be processed and the second image to be processed are images of the same target person, realizing the tracking of the target person, and further, the device adjustment instruction can be determined according to the position transformation information of the target person.

[0267] Step (v): When the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person, calculate the difference between the first target box and the second target box to obtain the position transformation information.

[0268] Step (vi): Determine a device adjustment instruction according to the position transformation information, so that the fan is adjusted according to the device adjustment instruction.

[0269] After determining that the target person in the first image to be processed and the target person in the second image to be processed are the same target person, the position transformation information can be calculated according to the first target box and the second target box obtained in step (ii). The position transformation information can obtain the displacement transformation of the target person in the image by calculating the difference between the first target box and the second target box, and then convert the displacement transformation from two-dimensional information to three-dimensional position transformation information. Similarly, according to the parameters pre-calibrated by the camera, the mathematical model for two-dimensional and three-dimensional conversion can be determined, and through this mathematical model, the two-dimensional displacement transformation is converted into three-dimensional position transformation information.

[0270] In practical applications, in order to avoid the accumulation of errors, it is also possible to obtain the two-dimensional position information of the target person in the image in real time according to the image information collected by the camera, and directly convert the two-dimensional position information into three-dimensional position information through a pre-determined two-dimensional to three-dimensional conversion mathematical model, so as to determine the device adjustment instruction according to the three-dimensional position information. Among them, the device adjustment instruction includes but is not limited to instructions for controlling the fan to change the wind direction, instructions for controlling the fan to change the wind speed, and adjustment instructions such as instructions for controlling the fan to change the air outlet temperature.

[0271] By applying the method of the embodiment of the present application and matching the positional relationship between the second target box and the third target box, the tracking of the target person can be realized, so that the device adjustment instruction can be determined according to the position change information of the target person, and the fan of the electronic device can send air to the user's location more intelligently, reducing useless air supply and avoiding energy waste.

[0272] In a possible implementation manner, when predicting the third target box, the prediction can be performed by means of Kalman filtering, and the steps are as follows:

[0273] Step (a): Calculate the initial state vector of the target person in the first image to be processed according to the first target box.

[0274] Step (b): Calculate the error between the first predicted state vector and the first predicted state vector according to the initial state vector.

[0275] Step (c): Calculate the measurement value of the target person in the second image to be detected according to the measurement matrix of the measuring device.

[0276] Step (d): Combine the measurement value to correct the first predicted state vector to obtain the first corrected state vector.

[0277] Step (e): Calculate the third target box according to the first corrected state vector.

[0278] Among them, the state vector is used to represent the motion state and motion position of the target person, including but not limited to the coordinates of the target box corresponding to the target person in the image, the width and height of the target box, the motion data (speed, direction) of the target person, the probability distribution that the target person's motion conforms to, and other information. In one example, steps (a)-(e) of the embodiment of the present application can be directly implemented by the Kalman filter algorithm.

[0279] Taking the state vector corresponding to the first target box as the initial state vector, calculate the first predicted state vector through the prediction model of Kalman filtering. In one example, the first predicted state vector is calculated by the following formula:

[0280] ;

[0281] Among them, x k1 represents the first predicted state vector, and A k represents the state transition matrix, and x k-1 represents the initial state vector, and B k represents the control matrix, and u k represents the control vector, and w k represents the Gaussian noise in the prediction process.

[0282] In practical applications, there is a prediction error between the predicted value and the actual value, and this prediction error can be calculated by the following formula:

[0283] ;

[0284] Among them, P k1 represents the prediction error corresponding to the first predicted state vector, and P k-1 represents the prediction error corresponding to the initial state vector; Q k represents the noise covariance matrix, represents the transpose operation on the state transition matrix.

[0285] In Kalman filtering, it also includes the measured value obtained by the measurement device. Among them, the measurement device can be devices such as sensors and cameras, and the measured value can be calculated in the following way:

[0286] ;

[0287] Among them, z k represents the measured value, and H k represents the measurement matrix detected by the measurement device, and v k represents the Gaussian noise in the observation process.

[0288] Correspondingly, there is also a measurement error y k when the measurement device makes a measurement, and this measurement error can be represented by v k , then the measurement covariance can be calculated by the following formula:

[0289] ;

[0290] Among them, S k represents the measurement covariance, and R k represents the measurement noise covariance matrix, represents the transpose operation on the measurement matrix.

[0291] According to the prediction error and the measurement covariance, the Kalman gain can be calculated by the following formula:

[0292] ;

[0293] Among them, K k represents the Kalman gain.

[0294] The first predicted state vector is corrected through the following formula:

[0295] ;

[0296] Among them, x k2 represents the first corrected state vector.

[0297] Correspondingly, the prediction error is also corrected:

[0298] ;

[0299] Among them, P K2 represents the corrected prediction error.

[0300] After correcting the first predicted state vector to obtain the first corrected state vector, the Kalman filtering algorithm can also implement the conversion between the state vector and the target box to obtain the third target box.

[0301] Applying the method of the embodiment of the present application, the tracking of the target person can be realized through the Kalman filtering method, so that the device adjustment instruction can be determined according to the position change information of the target person, and the fan of the electronic device can send air more intelligently to the position where the user is located, reducing useless air supply and avoiding energy waste.

[0302] In a possible implementation manner, the network structure of the first neural network model is as Figure 5 shown. Taking the step of predicting the shoulder key point and hip key point of the target person by the first neural network model for the first image to be processed as an example, the processing process of the first neural network model for the first image to be processed is as follows:

[0303] Input the first image to be processed into the first neural network model. The first network structure in the first neural network model performs feature extraction and downsampling processing on the first image to be processed to obtain the first image feature;

[0304] Input the first image feature into the second network structure for convolution processing, feature enhancement processing and downsampling processing to obtain the second image feature; among them, the network structure of the second network structure is the same as that of the first network structure;

[0305] Input the second image feature into the third network structure for feature enhancement processing to obtain the third image feature;

[0306] Input the third image feature into the fourth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a fourth image feature; wherein, the fourth network structure has the same network structure as the second network structure;

[0307] Input the fourth image feature into the fifth network structure for feature enhancement processing to obtain a fifth image feature; wherein, the fifth network structure has the same network structure as the third network structure;

[0308] Input the fifth image feature into the sixth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a sixth image feature, wherein the sixth network structure has the same network structure as the fourth network structure;

[0309] Input the sixth image feature into the seventh network structure for feature enhancement processing to obtain a seventh image feature, wherein the seventh network structure has the same network structure as the fifth network structure;

[0310] Input the seventh image feature into the eighth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain an eighth image feature; wherein, the eighth network structure has the same network structure as the sixth network structure;

[0311] Input the eighth image feature into the ninth network structure for multi-scale pooling operation to fuse image features of different scales and obtain a ninth image feature;

[0312] Input the ninth image feature into the tenth network structure for convolution processing and feature enhancement processing to obtain a tenth image feature, wherein the tenth image feature, the seventh image feature, and the fifth image feature have different sizes;

[0313] Perform convolution processing and feature enhancement processing on the tenth image feature through the eleventh network structure to obtain an eleventh image feature, wherein the eleventh network structure has the same network structure as the eighth network structure;

[0314] Perform upsampling processing on the eleventh image feature through the first upsampling module to obtain a twelfth image feature, wherein the twelfth image feature has the same size as the seventh image feature;

[0315] Perform feature fusion on the twelfth image feature and the seventh image feature through the first feature fusion layer to obtain a thirteenth image feature;

[0316] Input the thirteenth image feature into the twelfth network structure for convolution processing and feature enhancement processing to obtain a fourteenth image feature, wherein the twelfth network structure has the same network structure as the tenth network structure;

[0317] Input the fourteenth image feature into the thirteenth network structure for convolution processing and feature enhancement processing to obtain the fifteenth image feature, where the network structure of the thirteenth network structure is the same as that of the eleventh network structure;

[0318] Perform upsampling processing on the fifteenth image feature through the second upsampling module to obtain the sixteenth image feature, where the size of the sixteenth image feature is the same as that of the fifth image feature;

[0319] Perform feature fusion on the fifth image feature and the sixteenth image feature through the second feature fusion layer to obtain the seventeenth image feature;

[0320] Input the seventeenth image feature into the fourteenth network structure for convolution processing and feature enhancement processing to obtain the eighteenth image feature; where the network structure of the fourteenth network structure is the same as that of the twelfth network structure;

[0321] Input the eighteenth image feature into the fifteenth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain the nineteenth image feature, where the network structure of the fifteenth network structure is the same as that of the thirteenth network structure;

[0322] Perform feature fusion on the nineteenth image feature and the fifteenth image feature through the third feature fusion layer to obtain the twentieth image feature;

[0323] Input the twentieth image feature into the sixteenth network structure for convolution processing and feature enhancement processing to obtain the twenty - first image feature, where the network structure of the sixteenth network structure is the same as that of the fourteenth network structure;

[0324] Input the twenty - first image feature into the seventeenth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain the twenty - second image feature, where the network structure of the seventeenth network structure is the same as that of the fifteenth network structure;

[0325] Perform feature fusion on the twenty - second image feature and the eleventh image feature through the fourth feature fusion layer to obtain the twenty - third image feature;

[0326] Input the twenty - third image feature into the eighteenth network structure for convolution processing and feature enhancement processing to obtain the twenty - fourth image feature, where the network structure of the eighteenth network structure is the same as that of the sixteenth network structure;

[0327] Input the twenty - fourth image feature, the twenty - first image feature and the eighteenth image feature into the network head modules corresponding to their respective sizes for classification prediction to obtain the shoulder key points and hip key points of the target person.

[0328] Among them, for each network module with the same network structure, although the structure is the same, its parameters can be set to different network parameters according to the corresponding image processing operations. In one example, although the network structures of the thirteenth network structure and the fifteenth network structure are the same, the downsampling parameter of the fifteenth network structure is 2, and the downsampling parameter of the thirteenth network structure is 1.

[0329] The network head module is the classification prediction module of the first neural network, and each network head module corresponds to a different image size. For example, Figure 5 as shown, after the first image to be processed (640x640x3) undergoes multiple downsamplings, three image features with different sizes (the twenty-fourth image feature 20x20x512, the twenty-first image feature 40x40x256, and the eighteenth image feature 80x80x128) are obtained. Then, there are 3 network head modules in this neural network, each corresponding to a different downsampling depth. The network structures of each network head module are the same, but due to different downsampling depths, their parameters may be different. Specifically, the parameters of the network structure are set or obtained through training according to actual business requirements.

[0330] The network head module obtains its corresponding prediction result by performing operations such as convolution and feature enhancement on the image feature. Taking the network head module corresponding to the size of "20x20x512" as an example, this network head module includes multiple CBR (Conv + Batch Norm + Relu, convolutional layer + batch normalization layer + activation function layer) modules with the same network structure as the first network structure. Through multiple convolutional layers, classification results of different channels (20x20x2, 20x20x1, 20x20x4, 20x20x51) can be output. Then, through the feature fusion layer, the classification results of each convolutional layer are fused to obtain the classification result (20x20x58) corresponding to the downsampled image feature of this size and output it (output 2). It should be noted that the first network structure in the network head module only represents the CBR module with the same network structure as the first network structure, and its parameters do not need to be the same as those of the first network structure, and the parameters of each convolutional layer and the feature fusion layer are not necessarily the same.

[0331] In practical applications, the output of the network head module includes multiple categories and the scores of each category. Still taking the network head module corresponding to the size of "20x20x512" as an example, its output (20x20x58) can include 58 values, which are the x (the x - coordinate offset of the predicted point relative to the anchor point), y (the y - coordinate offset of the predicted point relative to the anchor point), w (the width of the predicted target box), h (the height of the predicted target box), obj (the score of the target box), cls1 (the score of category 1), cls2 (the score of category 2), kx1 (the x - coordinate offset of key point 1), ky1 (the y - coordinate offset of key point 1), ks1 (the score of key point 1), kx2 (the x - coordinate offset of key point 2), ky2 (the y - coordinate offset of key point 1), ks2 (the score of key point 2) …, kx17 (the x - coordinate offset of key point 17), ky17 (the y - coordinate offset of key point 17), ks17 (the score of key point 17) for each point. In practical applications, the intersection - over - union ratio of the predicted target boxes can be used to screen the target boxes to determine the target box corresponding to the target person. Then, based on the scores of each category, each key point in the target box is screened, and finally, the target box, gesture box, and key points corresponding to the target person in the first image to be processed are obtained.

[0332] Among them, the first network structure is the CBR module, as Figure 6 shown. This network structure includes a convolutional layer, a batch normalization layer, and an activation function layer. The third network structure is as Figure 7 shown. Similar to the CSPNet (Cross Stage Partial Network) network structure, this network structure is the CSP - X (the network structure applied to the Backbone part of CSPNet) structure, which includes multiple CBR modules, residual structures, and feature fusion layers that are the same as the network structure of the first network structure. It can divide the original input into two branches, respectively perform convolutional operations to halve the number of channels, then one branch performs Bottleneck X N (residual connection) operations, and then the two branches are feature - fused so that the input and output of the third network structure are of the same size, enabling the network to be constructed deeper and obtain more image information. In one example, the residual structure is as Figure 8 shown. This network structure includes multiple CBR modules and feature fusion layers that are the same as the structure of the first network structure, so that the output of the network layer can be connected to the output across multiple network layers, avoiding model gradient disappearance or gradient explosion.

[0333] The ninth network structure is as Figure 9As shown, it is also similar to the CSPNet network structure. This network structure is an SPP structure (Spatial Pyramid Pooling), including multiple CBR modules, a max pooling layer, and a feature fusion layer that have the same structure as the first network structure. The tenth network structure is as Figure 10 shown. Similarly, it is similar to the CSPNet network structure. This structure is a CSP2-X structure (a network structure applied to the Neck (intermediate connection) part of CSPNet), including multiple CBR modules and a feature fusion layer that have the same structure as the first network structure.

[0334] By applying the method of the embodiments of the present application, the detection of the target box, gesture box, and key points corresponding to the target person can be achieved through the first neural network. Thus, the position information and / or gesture information of the target person can be calculated based on the target box, gesture box, and key points. Furthermore, the electronic device can be adjusted according to the position information and / or gesture information of the target person, providing a more convenient human-computer interaction method for the user and enhancing the user experience.

[0335] In a possible implementation manner, the present application can Figure 11 adjust the device according to the flowchart shown. First, the target detection network is used to perform target box detection, gesture box detection, and key point detection on the image to be processed, obtaining the target box, gesture box, and key points corresponding to the target person in the image to be processed. Among them, the target detection network is the first neural network model pre-trained in the above embodiments.

[0336] After obtaining the target box, gesture box, and key points, filtering can also be performed according to the confidence scores of each target box, each gesture box, and each key point, filtering out the target boxes, gesture boxes, and key points whose confidence scores do not meet the preset score threshold, and determining the target boxes, gesture boxes, and key points representing the relevant information of the target person.

[0337] After determining the target box and key points, the target person can be tracked through Kalman filtering. The determination of the target person can be based on the pixel area of each person object in the image to be detected. For example, the person object with the largest pixel area is determined as the target person.

[0338] After selecting the target person, the first distance is calculated according to the key point information of the target person. Thus, the wind speed control instruction can be determined based on the first distance. Through Kalman filtering, the target person can be tracked. When it is determined that the images collected before and after are of the same target person, the position change information of the target person can be calculated according to the collected images, and the adjustment instruction of the electronic device can be determined based on the position change information and the first distance.

[0339] When performing gesture recognition on a user, it is necessary to first obtain the gesture box of the target person obtained by the target detection network, extract the gesture map corresponding to the gesture from the image to be processed, and recognize the gesture in the gesture map. Before performing gesture recognition, the quality of the gesture map can be detected first. For example, the gesture map is sent into a quality filtering model to obtain the clarity score and blurriness score of the gesture map. Among them, the quality filtering model is the second neural network model pre-trained above.

[0340] It is judged whether the gesture map is clear through the clarity score and the blurriness score. If not, the gesture map is not processed and the judgment of the gesture map is directly ended; if so, the gesture map is input into the gesture prediction model for gesture prediction. Among them, the gesture prediction model is the third neural network model pre-trained in the above embodiment. According to the gesture classification result of the gesture map, it is judged whether the gesture in the gesture map is a pre-set business gesture. If so, the corresponding gesture adjustment instruction is determined according to the gesture type; if not, the judgment of the gesture map is ended without processing it.

[0341] In practical applications, the electronic device can also maintain the device adjustment method in the related technology. For example, the electronic device is provided with a button module, and the user can directly press the button to trigger the corresponding device adjustment instruction. Among them, there can be a priority between the device adjustment instructions triggered by different methods. For example, during the use of the electronic device, the electronic device adjusts the device in real time according to the user's position information. However, if the user triggers a gesture adjustment instruction through a gesture at this time, the electronic device preferentially executes the relevant adjustment operation according to the gesture adjustment instruction.

[0342] By applying the method of the embodiment of the present application, the position information and / or gesture information of the target person can be calculated according to the target box, the gesture box and the key points, and then the electronic device can be adjusted according to the position information and / or gesture information of the target person, providing a more convenient human-computer interaction method for the user and improving the user experience.

[0343] In the second aspect of the embodiment of the present application, an electronic device adjustment device is provided. The electronic device includes a fan and a camera. The device includes the following Figure 12 structure:

[0344] A first image to be processed acquisition module 1201, configured to acquire a first image to be processed collected by the camera; wherein, the first image to be processed includes a target person;

[0345] A key point detection module 1202, configured to perform human key point detection on the first image to be processed, and determine the shoulder key point and the hip key point of the target person in the first image to be processed;

[0346] The first distance calculation module 1203 is configured to calculate the shoulder-hip image distance based on the image coordinates of the shoulder key point and the hip key point in the first image to be processed, and convert the shoulder-hip image distance into a first distance, where the first distance represents the distance between the target person and the electronic device;

[0347] The target wind speed determination module 1204 is configured to query the preset wind speed correspondence relationship, determine the target wind speed corresponding to the first distance, and adjust the fan to the target wind speed, where the wind speed correspondence relationship is the pre-set correspondence relationship between the wind speed and the distance.

[0348] In a possible implementation manner, the first distance calculation module includes:

[0349] The key point coordinate calculation sub-module is specifically configured to calculate the coordinates of the shoulder key point and / or the hip key point through the following formula:

[0350] ;

[0351] ;

[0352] where k x represents the abscissa of the shoulder key point or the hip key point, k y represents the ordinate of the shoulder key point or the hip key point, and the abscissa and ordinate of the offset coordinates of the key point predicted by the first neural network model relative to the anchor point are t kx and t ky , and the abscissa and ordinate of the anchor point coordinates are c x and c y ; strides represents the downsampling scaling ratio;

[0353] The center point coordinate calculation sub-module is specifically configured to calculate the coordinates of the center point of the two shoulder key points and the coordinates of the center point of the two hip key points through the following formula:

[0354] ;

[0355] where k csx represents the abscissa of the center point of the two shoulder key points, k lsx represents the abscissa of the left shoulder key point of the target person, k rsx represents the abscissa of the right shoulder key point of the target person;

[0356] ;

[0357] where k csy represents the ordinate of the center point of the two shoulder key points, k lsy represents the ordinate of the left shoulder key point of the target person, k rsyrepresents the ordinate of the right shoulder key point of the target person;

[0358] ;

[0359] where k chx represents the abscissa of the center point of two hip key points, k lhx represents the abscissa of the left hip key point of the target person, k rhx represents the abscissa of the right hip key point of the target person;

[0360] ;

[0361] where k chy represents the ordinate of the center point of two hip key points, k lhy represents the ordinate of the left hip key point of the target person, k rhy represents the ordinate of the corresponding right hip key point of the target;

[0362] The shoulder-hip image distance calculation sub-module is specifically configured to calculate the shoulder-hip image distance according to the coordinates of the center points of two shoulder key points and the coordinates of the center points of two hip key points through the following formula;

[0363] ;

[0364] where d pix represents the shoulder-hip image distance;

[0365] The first distance calculation sub-module is specifically configured to convert the shoulder-hip image distance into a first distance through the following formula;

[0366] ;

[0367] where D represents the first distance, f represents the camera focal length, and d body represents the preset average shoulder-hip distance.

[0368] In a possible implementation manner, the key point detection module includes:

[0369] The first neural network model prediction sub-module is specifically configured to input the first image to be processed into a pre-trained first neural network model to obtain the shoulder key points and hip key points of the target person in the first image to be processed;

[0370] where the first neural network model also outputs a gesture box of the target person, and the device further includes:

[0371] The image cutting module is configured to cut the first image to be processed according to the gesture box to obtain a gesture image;

[0372] A gesture detection module, configured to input a gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, where the gesture classification result represents the business gesture category corresponding to the gesture in the gesture image;

[0373] A device control instruction determination module, configured to determine a device control instruction according to the gesture classification result so that the fan executes an adjustment operation corresponding to the device control instruction.

[0374] In a possible implementation manner, the apparatus further includes:

[0375] A clarity detection module, configured to input the gesture image into a pre-trained third neural network model to obtain a clarity score and a blurriness score of the gesture image; where the third neural network model is used to detect the clarity of the gesture image;

[0376] A blurred image processing module, configured to discard the gesture image when the blurriness score is greater than or equal to the clarity score;

[0377] A clear image processing module, configured to, when the clarity score is greater than the blurriness score, perform the steps of: inputting the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image.

[0378] In a possible implementation manner, the apparatus further includes:

[0379] A second image to be processed acquisition module, configured to acquire a second image to be processed collected by the fan, where the second image to be processed is an image collected after the first image to be processed;

[0380] A target box acquisition module, configured to input the first image to be processed and the second image to be processed into a first neural network model respectively to obtain a first target box and a second target box, where the first target box represents the position of the target person in the first image to be processed, and the second target box represents the position of the target person in the second image to be processed;

[0381] A third target box prediction module, configured to calculate a third target box of the target person according to the first target box, where the third target box represents the predicted position of the target person in the second image to be processed;

[0382] A position relationship matching module, configured to perform a position relationship matching between the second target box and the third target box to obtain a matching result;

[0383] A position transformation information calculation module, configured to calculate a difference between the first target box and the second target box to obtain position transformation information when the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person;

[0384] A device adjustment instruction determination module, configured to determine a device adjustment instruction according to position transformation information, so that a fan is adjusted according to the device adjustment instruction.

[0385] In a possible implementation manner, the third target box prediction module includes:

[0386] An initial state vector calculation sub-module, specifically configured to calculate an initial state vector of a target person in a first image to be processed according to a first target box;

[0387] A predicted vector calculation sub-module, specifically configured to calculate an error between a first predicted state vector and the first predicted state vector according to the initial state vector;

[0388] A measurement value calculation sub-module, specifically configured to calculate an error between a measurement value of a target person in a second image to be detected and the measurement value according to a measurement matrix of a measurement device;

[0389] A Kalman gain calculation sub-module, specifically configured to calculate a Kalman gain by combining the error of the measurement value and the error of the first predicted state vector;

[0390] A predicted vector correction sub-module, specifically configured to correct the first predicted state vector through the Kalman gain to obtain a first corrected state vector;

[0391] A third target box calculation sub-module, specifically configured to calculate a third target box according to the first corrected state vector.

[0392] In a possible implementation manner, a position relationship matching module; includes:

[0393] An intersection over union calculation sub-module, specifically configured to calculate the intersection over union of a second target box and a third target box;

[0394] An intersection over union judgment sub-module, specifically configured to determine that the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person when the intersection over union is greater than or equal to a preset threshold.

[0395] In a possible implementation manner, a key point detection module, includes:

[0396] An image processing sub-module, specifically configured to input the first image to be processed into a first neural network model, and a first network structure in the first neural network model performs feature extraction and downsampling processing on the first image to be processed to obtain a first image feature;

[0397] Input the first image feature into a second network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a second image feature; wherein, the network structure of the second network structure is the same as that of the first network structure;

[0398] Input the second image feature into the third network structure for feature enhancement processing to obtain the third image feature;

[0399] Input the third image feature into the fourth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the fourth image feature; wherein, the network structure of the fourth network structure is the same as that of the second network structure;

[0400] Input the fourth image feature into the fifth network structure for feature enhancement processing to obtain the fifth image feature; wherein, the network structure of the fifth network structure is the same as that of the third network structure;

[0401] Input the fifth image feature into the sixth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the sixth image feature, wherein the network structure of the sixth network structure is the same as that of the fourth network structure;

[0402] Input the sixth image feature into the seventh network structure for feature enhancement processing to obtain the seventh image feature, wherein the network structure of the seventh network structure is the same as that of the fifth network structure;

[0403] Input the seventh image feature into the eighth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the eighth image feature; wherein, the network structure of the eighth network structure is the same as that of the sixth network structure;

[0404] Input the eighth image feature into the ninth network structure for multi-scale pooling operation to fuse image features of different scales to obtain the ninth image feature;

[0405] Input the ninth image feature into the tenth network structure for convolution processing and feature enhancement processing to obtain the tenth image feature, wherein the sizes of the tenth image feature, the seventh image feature, and the fifth image feature are different;

[0406] Perform convolution processing and feature enhancement processing on the tenth image feature through the eleventh network structure to obtain the eleventh image feature, wherein the network structure of the eleventh network structure is the same as that of the eighth network structure;

[0407] Perform upsampling processing on the eleventh image feature through the first upsampling module to obtain the twelfth image feature, wherein the size of the twelfth image feature is the same as that of the seventh image feature;

[0408] Perform feature fusion on the twelfth image feature and the seventh image feature through the first feature fusion layer to obtain the thirteenth image feature;

[0409] Input the thirteenth image feature into the twelfth network structure for convolution processing and feature enhancement processing to obtain the fourteenth image feature, where the network structure of the twelfth network structure is the same as that of the tenth network structure;

[0410] Input the fourteenth image feature into the thirteenth network structure for convolution processing and feature enhancement processing to obtain the fifteenth image feature, where the network structure of the thirteenth network structure is the same as that of the eleventh network structure;

[0411] Perform upsampling processing on the fifteenth image feature through the second upsampling module to obtain the sixteenth image feature, where the size of the sixteenth image feature is the same as that of the fifth image feature;

[0412] Perform feature fusion on the fifth image feature and the sixteenth image feature through the second feature fusion layer to obtain the seventeenth image feature;

[0413] Input the seventeenth image feature into the fourteenth network structure for convolution processing and feature enhancement processing to obtain the eighteenth image feature; where the network structure of the fourteenth network structure is the same as that of the twelfth network structure;

[0414] Input the eighteenth image feature into the fifteenth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the nineteenth image feature, where the network structure of the fifteenth network structure is the same as that of the thirteenth network structure;

[0415] Perform feature fusion on the nineteenth image feature and the fifteenth image feature through the third feature fusion layer to obtain the twentieth image feature;

[0416] Input the twentieth image feature into the sixteenth network structure for convolution processing and feature enhancement processing to obtain the twenty - first image feature, where the network structure of the sixteenth network structure is the same as that of the fourteenth network structure;

[0417] Input the twenty - first image feature into the seventeenth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain the twenty - second image feature, where the network structure of the seventeenth network structure is the same as that of the fifteenth network structure;

[0418] Perform feature fusion on the twenty - second image feature and the eleventh image feature through the fourth feature fusion layer to obtain the twenty - third image feature;

[0419] Input the twenty - third image feature into the eighteenth network structure for convolution processing and feature enhancement processing to obtain the twenty - fourth image feature, where the network structure of the eighteenth network structure is the same as that of the sixteenth network structure;

[0420] The twenty-fourth image feature, the twenty-first image feature, and the eighteenth image feature are respectively input into the network head modules corresponding to their respective sizes for classification prediction to obtain the shoulder key points and hip key points of the target person.

[0421] By applying the device according to the embodiment of the present application, key point detection can be performed on the first image to be processed collected, the shoulder key points and hip key points corresponding to the target person can be determined, the shoulder-hip image distance can be calculated according to the coordinates of the shoulder key points and hip key points in the image, and the shoulder-hip image distance can be converted into a first distance representing the actual distance between the target person and the electronic device. Thus, the target wind speed can be determined according to the first distance, so that the electronic device can adjust the wind speed according to the distance between the user and the electronic device. Compared with the operation method in the related art, the device adjustment method according to the embodiment of the present application can eliminate the need for the user to perform key operations, which is more convenient. Moreover, the position information and / or gesture information of the target person can be calculated according to the target frame, gesture frame, and key points, and then the electronic device can be adjusted according to the position information and / or gesture information of the target person, providing a more convenient human-computer interaction method for the user and improving the user experience.

[0422] The embodiment of the present application also provides an electronic device, as Figure 13 shown, including a processor 1301, a communication interface 1302, a memory 1303, and a communication bus 1304. Among them, the processor 1301, the communication interface 1302, and the memory 1303 complete communication with each other through the communication bus 1304.

[0423] The memory 1303 is used to store computer programs;

[0424] When the processor 1301 is used to execute the program stored in the memory 1303, the following steps are implemented:

[0425] Obtain the first image to be processed collected by the camera; wherein, the first image to be processed includes the target person;

[0426] Perform human key point detection on the first image to be processed to determine the shoulder key points and hip key points of the target person in the first image to be processed;

[0427] According to the image coordinates of the shoulder key points and hip key points in the first image to be processed, calculate the shoulder-hip image distance, and convert the shoulder-hip image distance into a first distance, where the first distance represents the distance between the target person and the electronic device;

[0428] Query the preset wind speed correspondence relationship, determine the target wind speed corresponding to the first distance, and adjust the fan to the target wind speed, where the wind speed correspondence relationship is the pre-set correspondence relationship between the wind speed and the distance.

[0429] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0430] The communication interface is used for communication between the above electronic device and other devices.

[0431] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0432] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0433] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned electronic device adjustment methods are implemented.

[0434] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above-mentioned electronic device adjustment methods in the embodiments.

[0435] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0436] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0437] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.

[0438] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A method for adjusting an electronic device, characterized in that: The electronic device includes a fan and a camera, and the method includes: Acquire a first image to be processed captured by the camera; wherein the first image to be processed includes a target person; Performing human body key point detection on the first image to be processed to determine the shoulder key points and crotch key points of the target person in the first image to be processed; Calculating a shoulder-span image distance according to image coordinates of a shoulder key point and a crotch key point in the first image to be processed, and converting the shoulder-span image distance into a first distance, wherein the first distance represents a distance between the target person and the electronic device; Querying a preset wind speed correspondence relationship, determining a target wind speed corresponding to the first distance, and adjusting the fan to the target wind speed, wherein the wind speed correspondence relationship is a preset correspondence relationship between wind speed and distance; Acquire a second image to be processed captured by the fan, wherein the second image to be processed is an image captured after the first image to be processed; Inputting the first image to be processed and the second image to be processed into a first neural network model respectively to obtain a first target frame and a second target frame, wherein the first target frame represents the position of the target person in the first image to be processed, and the second target frame represents the position of the target person in the second image to be processed; Calculating a third target frame of the target person according to the first target frame, wherein the third target frame represents a predicted position of the target person in the second image to be processed; Matching the second target frame with the third target frame in positional relationship to obtain a matching result; When the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person, calculating the difference between the first target frame and the second target frame to obtain position transformation information; According to the position change information, a device adjustment instruction is determined to enable the fan to be adjusted according to the device adjustment instruction; wherein the device adjustment instruction is used to change the air supply direction of the fan so that the fan supplies air to the location of the target person.

2. The method according to claim 1, characterized in that calculating the shoulder-span image distance according to the image coordinates of the shoulder key points and the crotch key points in the first image to be processed, and converting the shoulder-span image distance into the first distance; include: The coordinates of the shoulder key points and / or crotch key points are calculated by the following formula: ; ; Among them, the k x represents the horizontal coordinate of the shoulder key point or the crotch key point, the k y The ordinate of the shoulder key point or the crotch key point is represented by the horizontal and vertical coordinates of the offset coordinates of the key point predicted by the first neural network model relative to the anchor point. kx and t ky The horizontal and vertical coordinates of the anchor point coordinates are c x and c y ; The strides represent the downsampling scaling ratio; The coordinates of the center points of the two shoulder key points and the center points of the two span key points are calculated using the following formula: ; Among them, the k csx represents the horizontal coordinate of the center point of the two shoulder key points, the k lsx represents the horizontal coordinate of the left shoulder key point of the target person, and the k rsx The horizontal coordinate of the key point of the right shoulder of the target person; ; Among them, the k csy represents the ordinate of the center point of the two shoulder key points, the k lsy represents the ordinate of the left shoulder key point of the target person, and the k rsy The vertical coordinate of the key point of the right shoulder of the target person; ; Among them, the k chx represents the horizontal coordinate of the center point of the two cross-part key points, the k lhx represents the horizontal coordinate of the left span key point of the target person, and the k rhx The horizontal coordinate of the right span key point of the target person; ; Among them, the k chy represents the ordinate of the center point of the two cross-part key points, the k lhy represents the ordinate of the left span key point of the target person, and the k rhy The ordinate of the right-span key point corresponding to the target; According to the coordinates of the center points of the two shoulder key points and the coordinates of the center points of the two span key points, the shoulder-span image distance is calculated by the following formula; ; Among them, the d pix represents the shoulder-span image distance; The shoulder-span image distance is converted into the first distance by the following formula; ; Wherein, D represents the first distance, f represents the focal length of the camera, and d body Indicates the preset average shoulder span distance.

3. The method according to claim 1, characterized in that The step of performing human body key point detection on the first image to be processed to determine the shoulder key points and crotch key points of the target person in the first image to be processed includes: Inputting the first image to be processed into a pre-trained first neural network model to obtain shoulder key points and crotch key points of the target person in the first image to be processed; Wherein, the first neural network model further outputs a gesture frame of the target person, and the method further includes: According to the gesture frame, the first image to be processed is segmented to obtain a gesture image; Inputting the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, wherein the gesture classification result indicates a business gesture category corresponding to the gesture in the gesture image; According to the gesture classification result, a device control instruction is determined to enable the fan to perform an adjustment operation corresponding to the device control instruction.

4. The method according to claim 3, characterized in that Before inputting the gesture image into a pre-trained second neural network model to obtain a gesture classification result corresponding to the gesture image, the method further includes: Inputting the gesture image into a pre-trained third neural network model to obtain a clarity score and a blur score of the gesture image; wherein the third neural network model is used to perform clarity detection on the gesture image; When the blur score is greater than or equal to the clarity score, discarding the gesture image; In the case where the clarity score is greater than the fuzziness score, the step of inputting the gesture image into a pre-trained second neural network model is performed to obtain a gesture classification result corresponding to the gesture image.

5. The method according to claim 1, characterized in that The step of calculating a third target frame of the target person according to the first target frame includes: Calculating an initial state vector of the target person in the first image to be processed according to the first target frame; Calculating, based on the initial state vector, an error between the first predicted state vector and the first predicted state vector; Calculating the error between the measurement value of the target person in the second to-be-detected image and the measurement value according to the measurement matrix of the measuring device; Calculating a Kalman gain by combining an error in the measurement value with an error in the first predicted state vector; Correcting the first predicted state vector by using the Kalman gain to obtain a first corrected state vector; The third target frame is calculated according to the first corrected state vector.

6. The method according to claim 1, characterized in that The step of matching the second target frame with the third target frame in positional relationship to obtain a matching result comprises: Calculating an intersection-over-union ratio of the second target frame and the third target frame; When the intersection-over-union ratio is greater than or equal to a preset threshold, it is determined that the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person.

7. The method according to claim 3, characterized in that The step of inputting the first image to be processed into a pre-trained first neural network model to obtain the shoulder key points and crotch key points of the target person in the first image to be processed includes: Inputting the first image to be processed into the first neural network model, wherein the first network structure in the first neural network model performs feature extraction and downsampling processing on the first image to be processed to obtain first image features; Inputting the first image feature into a second network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a second image feature; wherein the second network structure is the same as the first network structure; Inputting the second image feature into the third network structure for feature enhancement processing to obtain a third image feature; Inputting the third image feature into a fourth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a fourth image feature; wherein the fourth network structure is the same as the second network structure; Inputting the fourth image feature into a fifth network structure for feature enhancement processing to obtain a fifth image feature; wherein the fifth network structure is the same as the third network structure; Inputting the fifth image feature into a sixth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a sixth image feature, wherein the sixth network structure is the same as the fourth network structure; Inputting the sixth image feature into a seventh network structure for feature enhancement processing to obtain a seventh image feature, wherein the seventh network structure is the same as the fifth network structure; Inputting the seventh image feature into an eighth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain an eighth image feature; wherein the eighth network structure is the same as the sixth network structure; Inputting the eighth image feature into a ninth network structure to perform a multi-scale pooling operation, fusing image features of different scales, and obtaining a ninth image feature; Inputting the ninth image feature into a tenth network structure for convolution processing and feature enhancement processing to obtain a tenth image feature, wherein the tenth image feature, the seventh image feature, and the fifth image feature have different sizes; Performing convolution processing and feature enhancement processing on the tenth image feature through an eleventh network structure to obtain an eleventh image feature, wherein the eleventh network structure is the same as the network structure of the eighth network structure; Performing upsampling processing on the eleventh image feature through a first upsampling module to obtain a twelfth image feature, wherein the twelfth image feature has the same size as the seventh image feature; The twelfth image feature and the seventh image feature are subjected to feature fusion through a first feature fusion layer to obtain a thirteenth image feature; Inputting the thirteenth image feature into a twelfth network structure for convolution processing and feature enhancement processing to obtain a fourteenth image feature, wherein the twelfth network structure has the same network structure as the tenth network structure; Inputting the fourteenth image feature into the thirteenth network structure for convolution processing and feature enhancement processing to obtain a fifteenth image feature, wherein the thirteenth network structure is the same as the eleventh network structure; performing upsampling processing on the fifteenth image feature through a second upsampling module to obtain a sixteenth image feature, wherein the sixteenth image feature has the same size as the fifth image feature; The fifth image feature and the sixteenth image feature are subjected to feature fusion by a second feature fusion layer to obtain a seventeenth image feature; Inputting the seventeenth image feature into a fourteenth network structure for convolution processing and feature enhancement processing to obtain an eighteenth image feature; wherein the fourteenth network structure has the same network structure as the twelfth network structure; Inputting the eighteenth image feature into a fifteenth network structure for convolution processing, feature enhancement processing and downsampling processing to obtain a nineteenth image feature, wherein the fifteenth network structure is the same as the thirteenth network structure; The nineteenth image feature is fused with the fifteenth image feature through a third feature fusion layer to obtain a twentieth image feature; Inputting the 20th image feature into a 16th network structure for convolution processing and feature enhancement processing to obtain a 21st image feature, wherein the 16th network structure is the same as the 14th network structure; Inputting the twenty-first image feature into a seventeenth network structure for convolution processing, feature enhancement processing, and downsampling processing to obtain a twenty-second image feature, wherein the seventeenth network structure is the same as the fifteenth network structure; The twenty-second image feature and the eleventh image feature are fused through a fourth feature fusion layer to obtain a twenty-third image feature; Inputting the twenty-third image feature into an eighteenth network structure for convolution processing and feature enhancement processing to obtain a twenty-fourth image feature, wherein the eighteenth network structure is the same as the sixteenth network structure; The twenty-fourth image feature, the twenty-first image feature and the eighteenth image feature are respectively input into the network head modules corresponding to their respective sizes for classification prediction to obtain the shoulder key points and crotch key points of the target person.

8. An electronic equipment adjustment device, characterized in that: The electronic device includes a fan and a camera, and the device includes: A first image to be processed acquisition module, used to acquire a first image to be processed captured by the camera; wherein the first image to be processed includes a target person; A key point detection module, used to perform human key point detection on the first image to be processed, and determine the shoulder key points and crotch key points of the target person in the first image to be processed; A first distance calculation module, configured to calculate a shoulder-span image distance according to image coordinates of shoulder key points and crotch key points in the first image to be processed, and convert the shoulder-span image distance into a first distance, wherein the first distance represents a distance between the target person and the electronic device; a target wind speed determination module, used to query a preset wind speed correspondence relationship, determine a target wind speed corresponding to the first distance, and adjust the fan to the target wind speed, wherein the wind speed correspondence relationship is a preset correspondence relationship between wind speed and distance; A second image to be processed acquisition module, used to acquire a second image to be processed acquired by the fan, wherein the second image to be processed is an image acquired after the first image to be processed; a target frame acquisition module, used to input the first image to be processed and the second image to be processed into a first neural network model respectively, to obtain a first target frame and a second target frame, wherein the first target frame represents the position of the target person in the first image to be processed, and the second target frame represents the position of the target person in the second image to be processed; A third target frame prediction module, used to calculate a third target frame of the target person according to the first target frame, wherein the third target frame represents a predicted position of the target person in the second image to be processed; A position relationship matching module, used for matching the position relationship between the second target frame and the third target frame to obtain a matching result; a position transformation information calculation module, configured to calculate a difference between the first target frame and the second target frame to obtain position transformation information when the matching result indicates that the target person in the first image to be processed and the target person in the second image to be processed are the same target person; The device adjustment instruction determination module is used to determine the device adjustment instruction according to the position change information so that the fan is adjusted according to the device adjustment instruction; wherein the device adjustment instruction is used to change the air supply direction of the fan so that the fan supplies air to the location of the target person.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Target tracking method and device, computer device and computer storage medium

    CN109903310A

  • Method for adjusting air speed of fan, fan and computer readable storage medium

    CN112922875A

  • Household appliance control method, household appliance and computer readable storage medium

    CN113050434A