Human-computer interaction method and system based on artificial intelligence, storage medium and computer
By performing color space conversion and convolution network processing on historical image data sets, and introducing attention mechanisms and optimization activation functions into neural network models, the gesture interaction model is trained, which solves the problems of operator experience dependence and inefficiency in the existing technology, and achieves efficient human-computer interaction.
Patent Information
- Application Number
- CN202411889486.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-23
AI Technical Summary
The existing human-computer interaction methods based on artificial intelligence require operators to have some experience, and when the object category or position changes, data parameters need to be reprogrammed, resulting in increased workload, increased costs and inefficiency.
A human-computer interaction method based on artificial intelligence is proposed. By obtaining the historical image data set in the preset area for color space conversion, inputting it to the convolutional network model for data processing, and introducing attention mechanism and optimization activation function into the neural network model, training to obtain a gesture interaction model, which is used to calculate gesture position information for human-computer interaction.
It improves the accuracy and processing speed of the model, reduces the dependence on operator experience, enhances the anti-interference ability to light, improves gesture image processing efficiency and perception ability, and achieves more efficient human-computer interaction.
Smart Images

Figure CN120029446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a human-computer interaction method, system, storage medium and computer based on artificial intelligence. Background Art
[0002] With the rapid development of science and technology and the improvement of people's living standards, more and more people have begun to rely on computer devices, thus giving rise to human-computer interaction methods.
[0003] Gesture interaction, as a typical interaction mode of human-computer interaction, can reduce the limitations caused by regional and cultural differences, and thus convey emotional information intuitively and accurately. The existing human-computer interaction method based on artificial intelligence usually realizes human-computer interaction by having the operator input the corresponding data parameters into the control robot to grasp the object. However, this method requires the operator to have a certain amount of experience to avoid the reduction in efficiency due to unfamiliarity with the use of the control robot; and when the category or position of the object to be grasped changes, new data parameters need to be reprogrammed. As the amount of data continues to increase, the workload of this method is also increasing, resulting in a significant increase in costs, low work efficiency, and failure to achieve the expected results. Summary of the invention
[0004] Based on this, the purpose of the present invention is to provide a human-computer interaction method, system, storage medium and computer based on artificial intelligence to at least solve the deficiencies in the above-mentioned technology.
[0005] The present invention proposes a human-computer interaction method based on artificial intelligence, comprising: Acquire a historical image data set in a preset area, and perform color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; Inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; Constructing a neural network model, introducing an attention mechanism into the neural network model, and optimizing the activation function of the neural network model to obtain a neural network optimization model; Extracting some images from the output image set as a training set, defining a model learning rate, and inputting the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; Collecting image data of the preset area, and performing feature extraction on the image data to extract feature information from the image data; Determine the center point of the feature information, input the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and use the gesture position information for human-computer interaction.
[0006] Furthermore, the step of converting the historical image data set into a color space to obtain a preliminary processed image set containing hand information includes: Converting the RGB color space of the historical image data set into a color space, and extracting the red offset in the color space; The red offset is binarized using a maximum inter-class variance algorithm to obtain a binarized image of the historical image data set.
[0007] Furthermore, the step of inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set includes: Constructing a convolutional network, and performing convolution operations on the input feature dimensions of the convolutional network using a number of standard convolutional kernels to output corresponding feature dimensions; Using a plurality of point-by-point convolution kernels to perform point-by-point convolution on the feature dimension to construct a convolutional network model; The preliminary processed image set is input into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and the feature images are subjected to feature fusion to obtain a corresponding output image set.
[0008] Furthermore, the steps of constructing a neural network model, introducing an attention mechanism into the neural network model, and optimizing the activation function of the neural network model to obtain a neural network optimization model include: Construct a YOLOv7 network model, and input the input features of the YOLOv7 network model into a multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two feature vectors; The two feature vectors are merged, and after activation through an activation function, a channel attention matrix is obtained, and the channel attention matrix is fused with the input features of the YOLOv7 network model to obtain fused input features; The fused input features are subjected to global average pooling and global maximum pooling compression processing along the corresponding channel direction to obtain two output features, the two output features are spliced in the channel dimension, and the spliced output features are subjected to convolution processing to obtain a spatial attention matrix; The spatial attention matrix is introduced into the YOLOv7 network model, and the activation function of the YOLOv7 network model is optimized to obtain a neural network optimization model.
[0009] Furthermore, the step of collecting image data of the preset area and performing feature extraction on the image data to extract feature information from the image data includes: Collecting image data of the preset area ,in, , , Respectively represent the width, height and number of channels of the image data; The main features of the image data are extracted using a convolutional neural network algorithm, and the extracted feature information is fused to obtain regression features in the image data.
[0010] The present invention also proposes a human-computer interaction system based on artificial intelligence, comprising: A data conversion module, used to obtain a historical image data set in a preset area, and perform color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; A data processing module, used for inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; A model optimization module, used to construct a neural network model, introduce an attention mechanism into the neural network model, and optimize the activation function of the neural network model to obtain a neural network optimization model; A model training module, used to extract some images from the output image set as a training set, and define a model learning rate to input the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; A feature extraction module, used to collect image data of the preset area and perform feature extraction on the image data to extract feature information from the image data; The human-computer interaction module is used to determine the center point of the feature information, input the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and use the gesture position information for human-computer interaction.
[0011] Furthermore, the data conversion module includes: A data conversion unit, used to convert the RGB color space of the historical image data set into a color space, and extract a red offset in the color space; The binarization processing unit is used to perform binarization processing on the red offset by using a maximum inter-class variance algorithm to obtain a binarized image of the historical image data set.
[0012] Furthermore, the data processing module includes: A network construction unit, used to construct a convolutional network, and use a number of standard convolution kernels to perform convolution operations on the input feature dimensions of the convolutional network to output corresponding feature dimensions; A model building unit, used to perform point-by-point convolution on the feature dimension using a plurality of point-by-point convolution kernels to build a convolutional network model; A data processing unit is used to input the preliminary processed image set into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and to perform feature fusion on each of the feature images to obtain a corresponding output image set.
[0013] Furthermore, the model optimization module includes: A model processing unit, used for constructing a YOLOv7 network model, and inputting the input features of the YOLOv7 network model into a multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two feature vectors; A feature fusion unit is used to merge the two feature vectors, obtain a channel attention matrix after activation through an activation function, and fuse the channel attention matrix with the input features of the YOLOv7 network model to obtain fused input features; A convolution processing unit, used to perform global average pooling and global maximum pooling compression processing on the fused input features along the corresponding channel direction to obtain two output features, splice the two output features in the channel dimension, and perform convolution processing on the spliced output features to obtain a spatial attention matrix; A model optimization unit is used to introduce the spatial attention matrix into the YOLOv7 network model and optimize the activation function of the YOLOv7 network model to obtain a neural network optimization model.
[0014] Furthermore, the feature extraction module includes: A data acquisition unit, used to acquire image data of the preset area ,in, , , Respectively represent the width, height and number of channels of the image data; The feature fusion unit is used to extract the main features of the image data using a convolutional neural network algorithm, and fuse the extracted feature information to obtain the regression features in the image data.
[0015] The present invention also provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned human-computer interaction method based on artificial intelligence is implemented.
[0016] The present invention also proposes a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned human-computer interaction method based on artificial intelligence when executing the computer program.
[0017] The human-computer interaction method, system, storage medium and computer based on artificial intelligence in the present invention perform color space conversion on the historical image data set in a preset area, thereby reducing the influence of light intensity and improving the anti-interference ability of light; the preliminary processed image is input into a pre-learned convolutional network model for data processing to obtain an output image set, maximizing the retention of key feature information of the input image to prevent the loss of local information; the attention mechanism is input into the constructed neural network model, and the activation function of the neural network model is optimized, thereby improving the accuracy, processing speed and corresponding processing performance of the model; the output image set is divided into a training set to train the model, further improving the model's processing efficiency and perception ability for gesture images, using the obtained gesture interaction model to process the image data of the preset area, and performing human-computer interaction based on the obtained gesture position information. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flow chart of a human-computer interaction method based on artificial intelligence in a first embodiment of the present invention; Figure 2 for Figure 1 Detailed flow chart of step S101; Figure 3 for Figure 1 Detailed flow chart of step S102; Figure 4 for Figure 1 Detailed flow chart of step S103; Figure 5 is a structural block diagram of a human-computer interaction system based on artificial intelligence in a second embodiment of the present invention; Figure 6 FIG. 4 is a structural block diagram of a computer in a third embodiment of the present invention.
[0019] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0020] In order to facilitate understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0022] Embodiment 1 See also Figure 1 , which shows a human-computer interaction method based on artificial intelligence in a first embodiment of the present invention, and the method specifically includes steps S101 to S106: S101, obtaining a historical image data set in a preset area, and performing color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; For further information, see Figure 2 , the step S101 specifically includes steps S1011~S1012: S1011, converting the RGB color space of the historical image data set into a color space, and extracting a red offset in the color space; S1012: binarize the red offset using a maximum inter-class variance algorithm to obtain a binarized image of the historical image dataset.
[0023] In a specific implementation, a historical image data set within a preset area is obtained, wherein the preset area may be a specific area where the user needs to perform gesture interaction, or may be a collection area of an image acquisition device in a human-computer interaction system, and the RGB color space of the obtained historical image data set is converted into a color space. It can be understood that the historical image data set collected by the image acquisition device is stored in an RGB encoding format, and the RGB-based color space encoding is easily affected by light intensity and has weak generalization ability. Therefore, it is converted into a corresponding color space to improve the resistance to light interference and exhibit better clustering characteristics. Moreover, the conversion from RGB space to color space is linear, with a small amount of calculation and a good segmentation effect.
[0024] Furthermore, the red offset in the above-mentioned color space is extracted, and the red offset is binarized using the maximum inter-class variance algorithm, thereby obtaining a binarized image of the above-mentioned historical image data set. In this embodiment, the skin color of the human body in the image is taken as a significant feature. Therefore, the red offset is binarized using the maximum inter-class variance algorithm (the maximum inter-class variance algorithm adopts the OTSU algorithm), which can better resist the intervention of brightness and contrast, thereby improving the efficiency of data processing and reducing errors.
[0025] S102, inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; For further information, see Figure 3 , the step S102 specifically includes steps S1021 to S1023: S1021, constructing a convolutional network, and using a number of standard convolutional kernels to perform convolution operations on the input feature dimensions of the convolutional network to output corresponding feature dimensions; S1022, performing point-by-point convolution on the feature dimension using a plurality of point-by-point convolution kernels to construct a convolutional network model; S1023, inputting the preliminary processed image set into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and performing feature fusion on each of the feature images to obtain a corresponding output image set.
[0026] In the specific implementation, a convolutional network is constructed and several standard convolution kernels are used. Perform convolution operation on the input feature dimension of the convolutional network to output the corresponding feature dimension, where the size of the input feature dimension of the convolutional network is , the dimension of the standard convolution kernel is , the size of the output feature dimension is ; Specifically, several point-by-point convolution kernels are used to act on each input feature channel in the feature dimension of the above output, and the feature dimension is convolved point-by-point to construct a convolutional network model, wherein the number of channels of the output feature map of the point-by-point convolution is consistent with the input, and the sizes of the convolution kernels constructed in the constructed convolutional network model are 1*1, 3*3 and 5*5 respectively.
[0027] Furthermore, the above-obtained preliminary processed image set is input into the constructed convolutional network model to obtain feature images containing the above-mentioned different channel numbers 1*1, 3*3 and 5*5, and the above-obtained feature images are feature fused to obtain the corresponding output image set. It can be understood that feature fusion can maximize the retention of key feature information of the input image to prevent the loss of local information. The calculation formula of feature fusion is as follows: ; In the formula, represents the feature cascade operation function, represents the feature branch fusion operator, Represents the first 1*1 convolution channel feature images, Represents the first 3*3 convolution channel feature images, Represents the first 5*5 convolution channel feature image.
[0028] S103, constructing a neural network model, introducing an attention mechanism into the neural network model, and optimizing an activation function of the neural network model to obtain a neural network optimization model; For further information, see Figure 4 , the step S103 specifically includes steps S1031 to S1033: S1031, constructing a YOLOv7 network model, and inputting the input features of the YOLOv7 network model into a multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two feature vectors; S1032, merging the two feature vectors, activating them through an activation function to obtain a channel attention matrix, and fusing the channel attention matrix with the input features of the YOLOv7 network model to obtain fused input features; S1033, performing global average pooling and global maximum pooling compression processing on the fused input features along the corresponding channel direction to obtain two output features, splicing the two output features in the channel dimension, and performing convolution processing on the spliced output features to obtain a spatial attention matrix; S1034, introducing the spatial attention matrix into the YOLOv7 network model, and optimizing the activation function of the YOLOv7 network model to obtain a neural network optimization model.
[0029] In the specific implementation, a YOLOv7 network model is constructed, and the input feature map of the constructed YOLOv7 network model is compressed in spatial dimension, and the obtained results are respectively input into the multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two corresponding feature vectors; Specifically, the two feature vectors obtained above are merged and activated through an activation function (in this embodiment, the activation function includes a Tanh function, a Softmax function, and a Sigmoid function, and the activation function can increase the nonlinear factor of the model) to obtain the final channel attention matrix, and the obtained channel attention matrix is fused with the input features of the YOLOv7 network model to obtain the fused input features.
[0030] Furthermore, the fused input features obtained above are subjected to global average pooling and global maximum pooling compression along the corresponding channel direction to generate two output features of the same size, the two output features are concatenated in the channel dimension, and the concatenated output features are convolved to reduce the corresponding channels to one dimension, and activated using the activation function to obtain the final spatial attention matrix, where the calculation formula of the spatial attention matrix is: ; In the formula, represents the activation function, It represents the output feature obtained by averaging the fused input features along the corresponding channel direction. It represents the output feature obtained by performing global maximum pooling processing on the fused input features along the corresponding channel direction. Represents a convolution layer with a convolution kernel size of 7*7.
[0031] In this embodiment, the spatial attention matrix obtained above is introduced into the YOLOv7 network model, and the spatial attention matrix obtained above is introduced into the last layer of the basic convolution model in the YOLOv7 network model, so as to construct a neural network optimization model.
[0032] S104, extracting some images from the output image set as a training set, and defining a model learning rate to input the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; In the specific implementation, the output image set obtained above is divided into a training set and a validation set in a ratio of 8:2, and the model learning rate is defined as 0.002. The obtained training set is input into the neural network optimization model for iterative training to ensure the stability of the model and optimize the effect of the model training. During the training process, the model learning rate is gradually increased until the preset requirement is met (in this embodiment, the preset requirement is the preset learning rate threshold of 0.01).
[0033] After the model is trained, the corresponding evaluation index is constructed using the above-mentioned verification set, and the trained model is verified using the evaluation index. If the verification meets the requirements, the trained model, i.e., the gesture interaction model, is output.
[0034] S105, collecting image data of the preset area, and performing feature extraction on the image data to extract feature information from the image data; Furthermore, the step S105 specifically includes steps S1051-S1052: S1051, collecting image data of the preset area ,in, , , Respectively represent the width, height and number of channels of the image data; S1052, using a convolutional neural network algorithm to extract the main features of the image data, and fusing the extracted feature information to obtain regression features in the image data.
[0035] In the specific implementation, the image data of the above-mentioned preset area is collected ,in, , , They respectively represent the width, height and number of channels of the image data. The image data is the real-time image data of the preset area. The main features of the image data are extracted using the convolutional neural network algorithm to output three feature information. The three feature information respectively represent the type prediction information of the target, the size regression feature information of the target and the regression accuracy feature information of the target. The three feature information are fused to obtain the regression features in the image data.
[0036] S106, determining the center point of the feature information, inputting the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and using the gesture position information for human-computer interaction.
[0037] In the specific implementation, the center point of the above-mentioned feature information is determined: ; In the formula, Represents the number of categories of targets in the image data, Indicates the output step length, which is a constant set by the user or automatically generated by the system; Furthermore, the center points of the image data and feature information are input into the gesture interaction model obtained above, so that the gesture interaction model performs data processing. When the value calculated by the center point of the feature information in the gesture interaction model meets the requirements, it means that the corresponding target object is detected in the feature information. When the value calculated by the center point of the feature information in the gesture interaction model does not meet the requirements, it means that the corresponding target object is not detected in the feature information. The gesture interaction model is used to identify the real key points of the image data, and the real key points are smoothed by a Gaussian kernel and mapped to the output feature map, so as to locate the gesture position information of the target object in the image data, and use the gesture position information for human-computer interaction.
[0038] In summary, the human-computer interaction method based on artificial intelligence in the above-mentioned embodiments of the present invention performs color space conversion on the historical image data set in the preset area, thereby reducing the influence of light intensity and improving the anti-interference ability of light; the preliminary processed image is input into the pre-learned convolutional network model for data processing to obtain an output image set, and the key feature information of the input image is retained to the maximum extent to prevent the loss of local information; the attention mechanism is input into the constructed neural network model, and the activation function of the neural network model is optimized, so as to improve the accuracy, processing speed and corresponding processing performance of the model; the output image set is divided into a training set to train the model, so as to further improve the model's processing efficiency and perception ability of gesture images, and the obtained gesture interaction model is used to process the image data of the preset area, and human-computer interaction is performed according to the obtained gesture position information.
[0039] Embodiment 2 On the other hand, the present invention also proposes a human-computer interaction system based on artificial intelligence, please refer to Figure 5 , which shows a human-computer interaction system based on artificial intelligence in a second embodiment of the present invention, the system comprises: The data conversion module 11 is used to obtain a historical image data set in a preset area and perform color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; Furthermore, the data conversion module 11 includes: A data conversion unit, used to convert the RGB color space of the historical image data set into a color space, and extract a red offset in the color space; The binarization processing unit is used to perform binarization processing on the red offset by using a maximum inter-class variance algorithm to obtain a binarized image of the historical image data set.
[0040] A data processing module 12 is used to input the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; Furthermore, the data processing module 12 includes: A network construction unit, used to construct a convolutional network, and use a number of standard convolution kernels to perform convolution operations on the input feature dimensions of the convolutional network to output corresponding feature dimensions; A model building unit, used to perform point-by-point convolution on the feature dimension using a plurality of point-by-point convolution kernels to build a convolutional network model; A data processing unit is used to input the preliminary processed image set into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and to perform feature fusion on each of the feature images to obtain a corresponding output image set.
[0041] A model optimization module 13 is used to construct a neural network model, introduce an attention mechanism into the neural network model, and optimize the activation function of the neural network model to obtain a neural network optimization model; Furthermore, the model optimization module 13 includes: A model processing unit, used for constructing a YOLOv7 network model, and inputting the input features of the YOLOv7 network model into a multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two feature vectors; A feature fusion unit is used to merge the two feature vectors, obtain a channel attention matrix after activation through an activation function, and fuse the channel attention matrix with the input features of the YOLOv7 network model to obtain fused input features; A convolution processing unit, used to perform global average pooling and global maximum pooling compression processing on the fused input features along the corresponding channel direction to obtain two output features, splice the two output features in the channel dimension, and perform convolution processing on the spliced output features to obtain a spatial attention matrix; A model optimization unit is used to introduce the spatial attention matrix into the YOLOv7 network model and optimize the activation function of the YOLOv7 network model to obtain a neural network optimization model.
[0042] A model training module 14 is used to extract some images from the output image set as a training set, and define a model learning rate to input the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; A feature extraction module 15, used for collecting image data of the preset area and performing feature extraction on the image data to extract feature information from the image data; Furthermore, the feature extraction module 15 includes: A data acquisition unit, used to acquire image data of the preset area ,in, , , Respectively represent the width, height and number of channels of the image data; The feature fusion unit is used to extract the main features of the image data using a convolutional neural network algorithm, and fuse the extracted feature information to obtain the regression features in the image data.
[0043] The human-computer interaction module 16 is used to determine the center point of the feature information, input the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and use the gesture position information to perform human-computer interaction.
[0044] The functions or operation steps implemented when the above modules and units are executed are generally the same as those in the above method embodiments, and will not be repeated here.
[0045] The human-computer interaction system based on artificial intelligence provided by the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, reference can be made to the corresponding contents in the aforementioned method embodiment.
[0046] Embodiment 3 The present invention also provides a computer, see Figure 6 , shown is a computer in the third embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, the above-mentioned human-computer interaction method based on artificial intelligence is implemented.
[0047] The memory 10 includes at least one type of storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10 may be an internal storage unit of a computer, such as a hard disk of the computer. In other embodiments, the memory 10 may also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. Further, the memory 10 may also include both an internal storage unit of the computer and an external storage device. The memory 10 may be used not only to store application software and various types of data installed in the computer, but also to temporarily store data that has been output or is to be output.
[0048] Among them, in some embodiments, the processor 20 can be an electronic control unit (Electronic Control Unit, abbreviated as ECU, also known as a vehicle computer), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run the program code stored in the memory 10 or process data, such as executing access restriction programs, etc.
[0049] It should be pointed out that Figure 6 The structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0050] An embodiment of the present invention further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the human-computer interaction method based on artificial intelligence as described above is implemented.
[0051] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0052] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0053] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0054] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0055] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A human-computer interaction method based on artificial intelligence, characterized in that: include: Acquire a historical image data set in a preset area, and perform color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; Inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; Constructing a neural network model, introducing an attention mechanism into the neural network model, and optimizing the activation function of the neural network model to obtain a neural network optimization model; Extracting some images from the output image set as a training set, defining a model learning rate, and inputting the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; Collecting image data of the preset area, and performing feature extraction on the image data to extract feature information from the image data; Determine the center point of the feature information, input the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and use the gesture position information for human-computer interaction.
2. The human-computer interaction method based on artificial intelligence according to claim 1, characterized in that: The steps of converting the historical image data set into a color space to obtain a preliminary processed image set containing hand information include: Converting the RGB color space of the historical image data set into a color space, and extracting the red offset in the color space; The red offset is binarized using a maximum inter-class variance algorithm to obtain a binarized image of the historical image data set.
3. The human-computer interaction method based on artificial intelligence according to claim 1, characterized in that: The step of inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set includes: Constructing a convolutional network, and performing convolution operations on the input feature dimensions of the convolutional network using a number of standard convolutional kernels to output corresponding feature dimensions; Using a plurality of point-by-point convolution kernels to perform point-by-point convolution on the feature dimension to construct a convolutional network model; The preliminary processed image set is input into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and the feature images are subjected to feature fusion to obtain a corresponding output image set.
4. The human-computer interaction method based on artificial intelligence according to claim 1, characterized in that: The steps of constructing a neural network model, introducing an attention mechanism into the neural network model, and optimizing the activation function of the neural network model to obtain a neural network optimization model include: Construct a YOLOv7 network model, and input the input features of the YOLOv7 network model into a multi-layer perception network for processing, so as to predict the importance of each channel in the YOLOv7 network model and generate two feature vectors; The two feature vectors are merged, and after activation through an activation function, a channel attention matrix is obtained, and the channel attention matrix is fused with the input features of the YOLOv7 network model to obtain fused input features; The fused input features are subjected to global average pooling and global maximum pooling compression processing along the corresponding channel direction to obtain two output features, the two output features are spliced in the channel dimension, and the spliced output features are subjected to convolution processing to obtain a spatial attention matrix; The spatial attention matrix is introduced into the YOLOv7 network model, and the activation function of the YOLOv7 network model is optimized to obtain a neural network optimization model.
5. The human-computer interaction method based on artificial intelligence according to claim 1, characterized in that: The steps of collecting image data of the preset area and performing feature extraction on the image data to extract feature information from the image data include: Collecting image data of the preset area ,in, , , Respectively represent the width, height and number of channels of the image data; The main features of the image data are extracted using a convolutional neural network algorithm, and the extracted feature information is fused to obtain regression features in the image data.
6. A human-computer interaction system based on artificial intelligence, characterized in that: include: A data conversion module, used to obtain a historical image data set in a preset area, and perform color space conversion on the historical image data set to obtain a preliminary processed image set containing hand information; A data processing module, used for inputting the preliminary processed image set into a pre-learned convolutional network model for data processing to obtain a corresponding output image set; A model optimization module, used to construct a neural network model, introduce an attention mechanism into the neural network model, and optimize the activation function of the neural network model to obtain a neural network optimization model; A model training module, used to extract some images from the output image set as a training set, and define a model learning rate to input the training set into the neural network optimization model for training to obtain a gesture interaction model, wherein the model learning rate is gradually increased during the training process until the model learning rate reaches a preset requirement; A feature extraction module, used to collect image data of the preset area and perform feature extraction on the image data to extract feature information from the image data; The human-computer interaction module is used to determine the center point of the feature information, input the image data and the center point of the feature information into the gesture interaction model to calculate the gesture position information in the image data, and use the gesture position information for human-computer interaction.
7. The human-computer interaction system based on artificial intelligence according to claim 6 is characterized in that: The data conversion module comprises: A data conversion unit, used to convert the RGB color space of the historical image data set into a color space, and extract a red offset in the color space; The binarization processing unit is used to perform binarization processing on the red offset by using a maximum inter-class variance algorithm to obtain a binarized image of the historical image data set.
8. The human-computer interaction system based on artificial intelligence according to claim 6, characterized in that: The data processing module comprises: A network construction unit, used to construct a convolutional network, and use a number of standard convolution kernels to perform convolution operations on the input feature dimensions of the convolutional network to output corresponding feature dimensions; A model building unit, used to perform point-by-point convolution on the feature dimension using a plurality of point-by-point convolution kernels to build a convolutional network model; A data processing unit is used to input the preliminary processed image set into the convolutional network model for data processing to obtain feature images containing different numbers of channels, and to perform feature fusion on each of the feature images to obtain a corresponding output image set.
9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the human-computer interaction method based on artificial intelligence as described in any one of claims 1 to 5 is implemented.
10. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the human-computer interaction method based on artificial intelligence as described in any one of claims 1 to 5 is implemented.