A Human-Computer Interaction Optimization System and Method Based on Deep Learning

Through the deep learning intelligent evaluation mechanism and dynamic frame rate adjustment, the problem of insufficient recognition of fixed frame rate cameras when changing fast gestures is solved, achieving higher recognition accuracy and smooth user interaction experience.

CN119904734BActive Publication Date: 2025-07-22JOINT WARFARE COLLEGE NAT DEFENSE UNIV OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411964062.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-22
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In the prior art, a fixed frame rate camera cannot capture sufficient dynamic details when the user's gestures change rapidly, resulting in incomplete and errors in gesture recognition, affecting the user experience.

Method used

Through the deep learning intelligent evaluation mechanism, the jump changes and movement alternating frequency of key points in the hand are analyzed in real time, key factors are generated, and the camera frame rate is dynamically adjusted to capture more dynamic details, including using particle swarm optimization algorithm to optimize frame rate adjustment.

Benefits of technology

It improves the accuracy of gesture recognition and the real-time interactive response, provides a smoother and more accurate user experience, reduces misoperation, and improves user satisfaction and interactive perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904734B_ABST
    Figure CN119904734B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimized system and method for human-computer interaction based on deep learning, which relates to the technical field of human-computer interaction and includes the following steps: First, a camera captures the hand images of a user at a pre-set frame rate; the captured hand image frames are stored and organized into an analysis set, and key features reflecting rapid changes in gestures are extracted from the analysis set; the extracted key features are analyzed in detail within a detection window. The present invention accurately identifies rapid changes in user gestures through a deep learning intelligent evaluation mechanism. The system analyzes the jumping changes of hand key points and the action alternation frequency in real time, generates key factors, and evaluates gesture patterns. According to the evaluation results, the camera frame rate is dynamically adjusted to capture more dynamic details and avoid recognition errors or delays caused by insufficient frame rate. This mechanism improves the accuracy of gesture recognition and interaction response, provides a smoother and more accurate user experience, and optimizes the overall interaction perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a human-computer interaction optimization system and method based on deep learning. Background Art

[0002] The optimization of human-computer interaction based on deep learning refers to using deep learning algorithms to improve the quality and efficiency of the interaction between humans and computer systems. Traditional human-computer interaction mainly relies on rule-based and manually designed systems, while the optimization method based on deep learning automatically learns more intelligent interaction methods by analyzing and understanding input data such as user behavior, language, and actions. For example, through models such as deep neural networks (DNNs) or convolutional neural networks (CNNs), real-time recognition and prediction of user speech, gestures, expressions, etc. are performed, so as to achieve a more natural and personalized interaction experience. By continuously learning and adapting to user inputs, deep learning can optimize the interaction interface, improve the accuracy and speed of system response, and dynamically adjust the interaction mode according to user needs and preferences, achieving a more efficient and intelligent human-computer interaction.

[0003] The prior art has the following deficiencies:

[0004] In the prior art, gesture change information is usually obtained at a fixed frame rate through a camera in user gesture recognition. The camera captures hand images at a certain frame rate and transmits the image data to a computer or processing system in real time. Each frame of the image records the state or action information of the hand, and the system analyzes the change of the gesture based on these consecutive image frames. However, the fixed frame rate means that the camera can only capture a limited number of image frames per second. If the user's gesture changes rapidly in a short period of time, a low-frame-rate camera may not be able to capture enough details, resulting in serious consequences. When the gesture changes very rapidly, a low-frame-rate camera may not be able to capture all the key action details. For example, when the user waves quickly or makes complex gestures, the relatively long inter-frame time may cause some important dynamics to be missed. This will not only lead to incomplete gesture recognition, but also may cause the system to be unable to accurately identify the detailed changes in rapid actions, especially at the start, end, or turning points of the gesture. As a result, this inaccuracy in recognition may lead to misclassification, the system misinterpreting the user's intention, and then generating incorrect interaction responses, affecting the user experience.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a human-computer interaction optimization system and method based on deep learning. By introducing an intelligent evaluation mechanism based on deep learning, accurate recognition of rapid changes in user gestures can be achieved. When the gestures change rapidly, the system analyzes the jumping changes of hand key points and the action alternation frequency in real time, generates key factors, and evaluates the gesture pattern. Based on the evaluation results, the system dynamically adjusts the frame rate of the camera to capture more dynamic details, avoiding recognition errors or delays caused by insufficient frame rate. This mechanism improves the accuracy of gesture recognition and the real-time performance of interaction response, provides a smoother and more accurate experience for users, and optimizes the overall user interaction perception to solve the problems in the above-mentioned background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions: A human-computer interaction optimization method based on deep learning, comprising the following steps:

[0008] First, the camera captures the hand images of the user at a pre-set frame rate;

[0009] The captured hand image frames are stored and organized into an analysis set, and key features reflecting rapid gesture changes are extracted from the analysis set;

[0010] The extracted key features are analyzed in detail within the detection window, and the key features after analysis and processing are input into a pre-trained machine learning model, and the machine learning model is used to intelligently evaluate abnormal changes in user gestures;

[0011] Based on the evaluation results of the model, the gesture changes of the user are divided into two categories: normal gesture changes and rapid gesture changes;

[0012] For normal gesture changes, continue to capture hand images at a preset fixed frame rate, and use gesture recognition algorithms to identify actions and intentions;

[0013] For rapid gesture changes, according to the evaluation results of the machine learning model, dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and capture more dynamic details.

[0014] Preferably, key features reflecting rapid gesture changes are extracted from the analysis set. The extracted key features include the jumping changes of hand key points and the alternation frequency of actions in the gesture. Under the detection window, the jumping changes of the extracted hand key points and the alternation frequency of actions in the gesture are analyzed to generate a key point jumping factor and an alternation frequency factor respectively. The key point jumping factor quantifies the severity of the position change of hand key points between adjacent frames, reflecting the suddenness and discontinuity of the spatial position during the gesture movement; the alternation frequency factor quantifies the frequent switching of actions in the gesture, reflecting the alternation change intensity in the gesture dynamic pattern.

[0015] Preferably, after obtaining the key point jump factor and the alternating frequency factor generated by analyzing the extracted key features, the key point jump factor and the alternating frequency factor are input into a pre-trained machine learning model. A gesture change coefficient is generated through the machine learning model, and the intelligent evaluation of the abnormal change of the user's gesture is carried out through the gesture change coefficient.

[0016] Preferably, the gesture change coefficient generated by analyzing the extracted key features is compared and analyzed with a pre-set reference threshold of the gesture change coefficient to classify the user's gesture change. The classification steps are as follows:

[0017] If the gesture change coefficient is greater than or equal to the pre-set reference threshold of the gesture change coefficient, the user's gesture change is classified as a fast gesture change;

[0018] If the gesture change coefficient is less than the pre-set reference threshold of the gesture change coefficient, the user's gesture change is classified as a normal gesture change.

[0019] Preferably, for fast gesture changes, according to the evaluation results of the machine learning model, based on particle swarm optimization, the frame rate of the camera is dynamically adjusted to reduce the inter-frame time interval. The specific steps to capture more dynamic details are as follows:

[0020] When it is detected that the gesture is in a fast-changing state, the particle swarm optimization algorithm is started, and the frame rate increment of the particle is initialized. The initialization of the frame rate increment not only depends on the gesture change coefficient GVC, but also comprehensively considers the historical gesture change coefficient GVC value to more accurately adjust the frame rate subsequently. The initialized frame rate increment formula is as follows:

[0021] ΔFPS b (t)=(GVC(t)-GVC ref )+θ·GVC(t - 1)

[0022] , where ΔFPS b (t) is the frame rate increment of the b-th particle at time point t, GVC(t) is the current gesture change coefficient, θ is the historical weight factor reflecting the influence of historical changes, GVC ref is the reference threshold of the gesture change coefficient, and GVC(t - 1) is the gesture change coefficient at the previous moment;

[0023] The particle swarm optimization algorithm optimizes the optimal frame rate increment of each particle according to the change of the frame rate increment ΔFPS b (t). The calculation expression is as follows:

[0024] ΔFPS opt (t)=|GVC(t)-GVC ref |+τ·|ΔFPS b(t) - ΔFPS b (t - 1)|

[0025] , where ΔFPS opt (t) is the optimal frame rate increment, τ is the penalty coefficient, ΔFPS b (t - 1) is the frame rate increment at the previous moment, that is, the frame rate increment of the b-th particle at the moment t - 1;

[0026] After obtaining the optimal frame rate increment ΔFPS opt (t), the optimal frame rate increment ΔFPS opt (t) is combined with the initial preset frame rate to dynamically adjust the frame rate of the camera. The calculation expression of the optimized camera frame rate is as follows:

[0027]

[0028] , where FPS new (t) is the optimized camera frame rate, FPS0 is the preset camera frame rate, δ is the historical smoothing factor, which controls the influence degree of the historical frame rate change on the current frame rate adjustment, and B is the total number of particles.

[0029] Preferably, under the detection window, the specific steps for analyzing the jumping changes of the hand key points and generating the key point jumping factor are as follows:

[0030] Under the detection window, for each frame, use the hand tracking algorithm to extract the spatial coordinate data of the key points and calibrate it as P i (t), which represents the coordinate position of the i-th key point at the time point t. P i (t) = {x i (t), y i (t), z i (t)}, where x i (t), y i (t), z i (t) are the positions of the i-th key point on the x-axis, y-axis, and z-axis at the time point t respectively;

[0031] Calculate the change in the key point position between each pair of consecutive frames, that is, the displacement of the i-th key point between two adjacent frames. The expression is as follows:

[0032]

[0033] , where Δt is the time interval, and ΔP i (t) is the change amount of the key point position, that is, the spatial change of the i-th key point within the time interval Δt. x i (t + Δt) is the position of the i-th key point on the x-axis at the time point t + Δt, yi (t + Δt) is the position of the i-th key point on the y-axis at the time point t + Δt, z i (t + Δt) is the position of the i-th key point on the z-axis at the time point t + Δt;

[0034] To quantify the degree of change mutation of each hand key point, a non-linear weighting method is introduced to define the local key point jump factor, and the calculation expression is as follows:

[0035] J i (t) = (ΔP i (t)) α ·e -β·Δt

[0036] , where J i (t) is the local key point jump factor, α is the non-linear weighting exponent, e -β·Δt is the time decay factor, e is the natural base, and β is the time decay constant;

[0037] Integrate the jump factors J i (t) of all local key points to obtain the final key point jump factor, and the calculation expression is as follows:

[0038]

[0039] , where J(t) is the key point jump factor, w i is the weight of the i-th key point, and N is the total number of key points.

[0040] Preferably, under the detection window, the specific steps for analyzing the alternating frequency of actions in a gesture to generate an alternating frequency factor are as follows:

[0041] Under the detection window, divide the captured continuous gesture image frames into multiple segments, each segment representing the duration of a gesture action. According to the change of the hand key point positions between consecutive frames, divide each action segment, and the expression is as follows:

[0042]

[0043] , where A(t) is the gesture action change amount, representing the change amount of the gesture action at the time point t, P i (t) is the coordinate position of the i-th key point at the time point t, P i (t - 1) is the coordinate position of the i-th key point at the time point t - 1, and N is the total number of key points;

[0044] When the actions of each segment are divided, then calculate the alternation frequency between each action segment. The alternation frequency measures the rate of gesture action switching per unit time and reflects the dynamic changes of the gesture. The calculation expression is as follows:

[0045]

[0046] , where ω(t) is the alternation frequency, action j (t) is the eigenvalue of the j-th action segment at time point t, action j (t - 1) is the eigenvalue of the j-th action segment at time point t - 1, and Δt is the time interval;

[0047] To consider the influence of different action segments, the alternation frequency needs to be weighted and adjusted according to the intensity of different action segments. Based on the complexity and action intensity of each action segment, weights are used to adjust the calculation of the alternation frequency. The calculation expression of the weighted and adjusted alternation frequency is as follows:

[0048]

[0049] , where is the weighted alternation frequency, w j is the complexity weight of the j-th action segment, ω j (t) is the alternation frequency of the j-th action segment;

[0050] Finally, all the weighted alternation frequencies are further processed through a non-linear adjustment function to generate the final alternation frequency factor. The calculation expression is as follows:

[0051]

[0052] , where γ(t) is the alternation frequency factor, is the weighted alternation frequency of the j-th action segment, and λ j is the weighting factor of the j-th action segment.

[0053] A human-computer interaction optimization system based on deep learning includes a gesture image capture module, an image data storage and analysis set construction module, a feature extraction and analysis module, an intelligent evaluation and classification module, a normal gesture recognition module, and a dynamic frame rate adjustment module:

[0054] The gesture image capture module first captures the user's hand image by the camera at a preset frame rate;

[0055] The image data storage and analysis set construction module stores the captured hand image frames and organizes them into an analysis set, and extracts the key features reflecting the rapid changes of the gesture from the analysis set;

[0056] The feature extraction and analysis module analyzes the extracted key features in the detection window in detail, and inputs the key features after analysis and processing into a pre-trained machine learning model to intelligently evaluate the abnormal changes in the user's gestures through the machine learning model;

[0057] The intelligent evaluation and classification module divides the user's gesture changes into two categories: normal gesture changes and rapid gesture changes based on the evaluation results of the model;

[0058] The normal gesture recognition module continues to capture hand images at a preset fixed frame rate for normal gesture changes, and uses gesture recognition algorithms to identify actions and intentions;

[0059] The dynamic frame rate adjustment module dynamically adjusts the frame rate of the camera according to the evaluation results of the machine learning model for rapid gesture changes, reduces the inter-frame time interval, and captures more dynamic details.

[0060] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0061] By introducing an intelligent evaluation mechanism based on deep learning, the present invention can accurately identify rapid changes in the user's gestures. When the user's gestures change rapidly within a short period of time, traditional fixed-frame-rate cameras may miss key dynamic details, resulting in incomplete or incorrect gesture recognition. By analyzing the jumping changes of hand key points and the alternating frequency of actions in the gesture in real time, after generating the key point jumping factor and the alternating frequency factor, the machine learning model can intelligently evaluate the change pattern of the gesture. When it is detected that the user's gesture has changed rapidly, the system can dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and thus capture more dynamic details. This real-time and sensitive adjustment ensures that the system can more accurately identify rapidly changing gestures, avoid missing or misclassifying gesture actions, and thus improve the accuracy of gesture recognition. Users can obtain a more accurate and smooth interaction experience and reduce misoperations caused by inaccurate recognition.

[0062] The dynamic frame rate adjustment mechanism of the present invention combines real-time analysis of gesture changes and can flexibly adjust the working mode of the camera according to the changes in gestures. When the user makes rapid gesture movements, the system can timely adjust the frame rate to capture the key details of the gesture movements, avoiding interaction delays or missed information caused by insufficient frame rates. This mechanism ensures the real-time response of the interactive system to the user's intentions and avoids the system's misinterpretation of the user's behavior due to recognition errors. Further, the combination of intelligent evaluation and dynamic frame rate adjustment not only improves the system's sensitivity to rapid gestures but also makes the system's recognition of different types of gestures more natural and smooth. During the use process, users can obtain a smoother and delay-free interaction experience. This precise response and high sensitivity will greatly improve user satisfaction and interaction perception, optimizing the overall user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a method flow chart of a human-computer interaction optimization method based on deep learning according to the present invention.

[0065] Figure 2 It is a module schematic diagram of a human-computer interaction optimization system based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0067] The present invention provides a Figure 1 human-computer interaction optimization method based on deep learning as shown, including the following steps:

[0068] First, the camera captures the user's hand image at a pre-set frame rate;

[0069] The preset frame rate refers to the number of image frames captured by the camera within a certain period of time (usually per second) in a gesture recognition system. The setting of this frame rate is typically determined based on historical data or empirical data. For example, in past gesture recognition applications, through the dynamic analysis of the changes in user gestures, researchers or engineers may have found that the recognition accuracy is relatively high at a specific frame rate, and most key gesture actions can be captured. Therefore, based on this historical data, the frame rate is preset to a fixed value to ensure the efficient recognition of gestures during the operation of the system.

[0070] When setting the frame rate, it is necessary to ensure that this value can balance the consumption of computing resources and the detailed requirements of image capture. Generally, the set frame rate should be sufficient to capture the key detail changes in the user's gesture actions, while avoiding an overly high frame rate that may result in excessive data volume, thereby affecting the processing speed and the real-time response ability of the system. Therefore, the setting of the frame rate needs to balance the resolution of the camera, the computing power, and the speed of gesture changes, ensuring that images can be efficiently captured during common gesture executions without causing resource waste.

[0071] The captured hand image frames are stored and organized into an analysis set, and key features reflecting the rapid changes in the gesture are extracted from the analysis set;

[0072] First, the camera continuously captures image frames of the user's hand at the set frame rate. Each frame of the image contains the state and action information of the hand. These image frames are then transmitted to a computer or a processing system, and the system stores and organizes these images in chronological order into an analysis set. The analysis set is a data container that contains all the hand image frames captured within a certain time window and is arranged according to the chronological order of the gesture actions. This set provides the basic data for subsequent gesture analysis.

[0073] The role of the analysis set is to provide a complete and systematic data basis for subsequent feature extraction, processing, and gesture recognition. By organizing the image frames in chronological order, the analysis set can provide the system with continuous gesture change information, enabling the algorithm to identify the dynamic changes in the image sequence, especially the turning points, rapid changes, and key actions of the gesture. At the same time, it also facilitates the extraction of features reflecting the rapid changes in the gesture from it, helping the system to more accurately understand the intention and details of the gesture, providing high-quality input data for the machine learning model, and thus improving the recognition accuracy and real-time performance.

[0074] The key features extracted are analyzed in detail within the detection window, and the analyzed and processed key features are input into a pre-trained machine learning model to intelligently evaluate the abnormal changes in the user's gesture through the machine learning model;

[0075] Extract the key features that reflect the rapid changes of gestures from the analysis set. The extracted key features include the jumping changes of hand key points and the alternating frequency of actions in the gesture. Under the detection window, analyze the jumping changes of the extracted hand key points and the alternating frequency of actions in the gesture, and generate a key point jumping factor and an alternating frequency factor respectively. The key point jumping factor quantifies the intensity of the position change of hand key points between adjacent frames, reflecting the suddenness and discontinuity of the spatial position during the gesture movement; the alternating frequency factor quantifies the frequent switching of actions in the gesture, reflecting the alternating change intensity in the gesture dynamic pattern.

[0076] After obtaining the key point jumping factor and the alternating frequency factor generated after analyzing the extracted key features, input the key point jumping factor and the alternating frequency factor into a pre-trained machine learning model, and generate a gesture change coefficient through the machine learning model, and use the gesture change coefficient to intelligently evaluate the abnormal changes of the user's gesture.

[0077] The pre-trained machine learning model refers to the model obtained by training using a large amount of historical data and label information of gesture changes before the system is deployed. This model has learned different types of gesture data through certain algorithms (such as deep learning, support vector machines, random forests, etc.), so as to be able to identify and predict different gesture patterns and their change rules. Through training, the model can master the spatio-temporal features of gestures, understand which gesture changes are normal and which are abnormal, and even learn some subtle change patterns. During the training process, the dynamic features of gestures (such as the key point jumping factor and the alternating frequency factor) will be associated with the corresponding labels (such as whether the gesture is abnormal, the gesture type, etc.), so that the model can predict the results of unknown inputs based on the rules in the data.

[0078] This pre-trained machine learning model can perform intelligent evaluation in real-time applications. Specifically, once the key point jumping factor and the alternating frequency factor are extracted from the analysis set and input into the model, the model will evaluate the input data according to the patterns learned before and output a gesture change coefficient. The gesture change coefficient reflects the abnormal degree of the user's current gesture. If this coefficient exceeds a certain threshold, the model can judge that the gesture may be in a rapid change or abnormal state, and thus issue a corresponding warning or adjust the recognition strategy. Through continuous learning and adjustment, the pre-trained model can continuously optimize its judgment criteria to ensure that it can still make accurate recognition and evaluation when facing different scenarios and new gesture changes.

[0079] The sudden change in the jumping of hand key points indicates that the current user's gesture is in a state of rapid change. When there are significant jumping changes in hand key points between adjacent frames, it usually means that there has been a significant change in the spatial position of the gesture, which may be a rapid action transition or movement completed in a short period of time. For example, when the user waves their hand quickly or makes complex gestures, the positions of the key points on the hand may change significantly between each frame, and this sudden change reflects the rapid transformation of the gesture. This sudden change is usually related to a sharp change in the user's intention, such as suddenly starting or ending an action, or quickly switching the gesture type. Compared with smooth actions, jumping changes usually occur during dynamic acceleration or actions with a higher intensity of change. Therefore, when there are obvious jumping changes in hand key points, the system can infer that the user's gesture is undergoing a rapid and significant change state.

[0080] Under the detection window, the specific steps for analyzing the jumping changes of hand key points and generating the key point jumping factor are as follows:

[0081] Under the detection window, for each frame, use the hand tracking algorithm to extract the spatial coordinate data of the key points and label it as P i (t), representing the coordinate position of the i-th key point at time point t, P i (t) = {x i (t), y i (t), z i (t)}, x i (t), y i (t), z i (t) are respectively the positions of the i-th key point on the x-axis, y-axis, and z-axis at time point t;

[0082] The hand tracking algorithm is a computer vision technology used to real-time track and identify the position, posture, and movement of the hand in images or videos. It determines the position and movement trajectory of the hand in three-dimensional space by analyzing feature points, contours, or depth information in the image or video sequence. The hand tracking algorithm usually combines machine learning, deep learning, or traditional computer vision methods, uses a camera to capture continuous image frames of the hand, and infers the position and actions of the hand based on the dynamic changes between these frames.

[0083] The reason for using the hand tracking algorithm is that it can precisely capture the subtle movements and changes of the hand, thus providing a basis for gesture recognition and interaction. Through hand tracking, the system can real-time perceive the user's intention, posture, and movement amplitude during various actions, and then provide a more efficient and intuitive interaction method. For example, in applications such as virtual reality, augmented reality, and gesture control, the hand tracking algorithm can achieve a smooth and precise user interaction experience, avoiding the limitations of traditional input devices and enhancing the user's immersion and operation convenience.

[0084] Calculate the change in the key point positions between each pair of consecutive frames, that is, the displacement of the i-th key point between two adjacent frames. The expression is as follows:

[0085]

[0086] , where Δt is the time interval, and ΔP i (t) is the change in the key point position, that is, the spatial change of the i-th key point within the time interval Δt. x i (t + Δt) is the position of the i-th key point on the x-axis at the time point t + Δt, and y i (t + Δt) is the position of the i-th key point on the y-axis at the time point t + Δt, and z i (t + Δt) is the position of the i-th key point on the z-axis at the time point t + Δt;

[0087] The role of the change in the key point position is to quantify the dynamics of hand movement, especially to evaluate the change rate and suddenness of gestures. By analyzing the change in the key point positions between adjacent frames, we can determine whether the hand movement is smooth or has significant fluctuations. A large change usually means that the hand has experienced a large displacement in a short time, which may indicate that the gesture is changing rapidly or there is a sudden change. This information is crucial for identifying fast gestures or abnormal changes, enabling the system to accurately capture the key details in fast actions and avoid missing important action information due to frame rate limitations. Therefore, calculating the change in the key point position provides basic data support for further gesture classification, anomaly detection, and intelligent recognition, enhancing the response ability of the gesture recognition system in a dynamic environment.

[0088] To quantify the degree of change and mutation of each hand key point, a non-linear weighting method is introduced to define the local key point jump factor. The calculation expression is as follows:

[0089] J i (t) = (ΔP i (t)) α ·e -β·Δt

[0090] , where J i(t) is the local key point jump factor, α is the nonlinear weighting index, e -β·Δt is the time decay factor, e is the natural base, and β is the time decay constant;

[0091] The time decay factor is a mathematical term introduced when calculating dynamic changes. Its main function is to adjust the impact of a certain change on the final result according to the length of the time interval. Specifically, the time decay factor is usually expressed by an exponential decay function (such as e -β·Δt ), where β is the decay constant and Δt is the time interval. Its function is that when the time interval is long, the decay factor will reduce the impact of the change, and conversely, when the time interval is short, the change has a greater impact on the result. In this way, the time decay factor can effectively handle dynamic changes with long time intervals, so that the system pays more attention to rapid changes that occur in a short time span, while ignoring slow changes with a long time span, thereby more accurately reflecting the rapid changes in gestures or actions, and avoiding the unnecessary impact of slow changes caused by long time intervals on the results.

[0092] The jump factor J for all local key points i (t) is integrated to obtain the final key point jump factor, and the calculation expression is as follows:

[0093]

[0094] , where J(t) is the key point jump factor, w i is the weight of the i-th key point, N is the total number of key points;

[0095] Under the detection window, the larger the key point jump factor expression value generated after analyzing the jump changes of the hand key points, the more drastic the spatial position change of the hand key points between adjacent frames is, which reflects that the user's gesture is undergoing rapid and sudden changes. Therefore, a larger key point jump factor expression value indicates that the user's gesture is in a state of rapid change. On the contrary, if the jump factor expression value is small, it means that the change of the hand key points is relatively stable or continuous, indicating that the user's gesture changes are relatively slow or stable, which is a normal gesture change state.

[0096] The rapid increase in the alternation frequency of actions in a gesture typically indicates that the current user's gesture is in a state of rapid change, as the alternation frequency reflects the rate of switching or repetition between gesture actions. When the user makes rapid gesture changes, the transitions between each action become more frequent and rapid, resulting in a significant increase in the alternation frequency. This frequent switching or repeated actions are usually signs of a drastic change in the gesture state, which may be caused by actions such as rapid waving, shaking, or rapid clicking. Compared with smooth or slowly changing gestures, rapidly changing gestures have higher dynamic characteristics, and the transitions of actions are faster and more obvious. Therefore, the rapid increase in the alternation frequency directly indicates that the gesture has experienced more action switches or changes in a short period of time, reflecting an increase in the intensity of gesture changes. At this time, the gesture pattern becomes more complex and dynamic, and the system needs to capture more details to ensure the accuracy of recognition.

[0097] Under the detection window, the specific steps for analyzing the alternation frequency of actions in a gesture and generating the alternation frequency factor are as follows:

[0098] Under the detection window, the captured continuous gesture image frames are divided into multiple segments, each segment representing the duration of a gesture action. According to the change in the positions of hand key points between consecutive frames, each action segment is divided, and the expression is as follows:

[0099]

[0100] , where A(t) is the gesture action change amount, representing the change amount of the gesture action at time point t, P i (t) is the coordinate position of the i-th key point at time point t, P i (t - 1) is the coordinate position of the i-th key point at time point t - 1, that is, the coordinate of the key point at the previous moment, and N is the total number of key points;

[0101] Dividing the captured continuous gesture image frames into multiple segments, with each segment representing the duration of a gesture action, is to accurately identify and analyze the time span and motion characteristics of each action. This division method helps the system understand the start and end times of each action, clarify the various dynamic stages of the gesture, and thus provide a more detailed reference for subsequent action analysis. For example, when the hand key points change significantly between consecutive frames, the system can recognize a change in an action and start a new action segment; when the change is small or stable, the system can continue to analyze within the current segment. This division can help accurately capture details such as mutations, pauses, and smoothness in gesture actions, provide more accurate input for subsequent action recognition, pattern analysis, and anomaly detection, and thereby improve the accuracy and real-time response ability of the overall recognition system.

[0102] When the actions of each segment are divided and completed, then the alternation frequency between each action segment is calculated. The alternation frequency measures the rate of gesture action switching per unit time and reflects the dynamic changes of the gesture. The calculation expression is as follows:

[0103]

[0104] , where ω(t) is the alternation frequency, action j (t) is the eigenvalue of the j-th action segment at time point t, action j (t - 1) is the eigenvalue of the j-th action segment at time point t - 1, that is, the eigenvalue of the action segment at the previous moment, and Δt is the time interval;

[0105] To consider the influence of different action segments, the alternation frequency needs to be weighted and adjusted according to the intensity of different action segments. Based on the complexity and action intensity of each action segment, weights are used to adjust the calculation of the alternation frequency. The calculation expression of the weighted-adjusted alternation frequency is as follows:

[0106]

[0107] , where is the weighted alternation frequency, w j is the complexity weight of the j-th action segment, ω j (t) is the alternation frequency of the j-th action segment;

[0108] The significance of weighting the alternation frequency is that it can more accurately reflect the relative importance and change intensity of each action segment in the gesture. Different action segments may have different time lengths, dynamic change characteristics, and frequencies during the gesture process. Therefore, their contributions to the overall gesture change are uneven. By weighting, higher weights can be assigned to the key action segments that play a dominant role in the rapid change of the gesture, so that the finally calculated alternation frequency is more sensitive and representative, and can accurately capture the important change moments in the gesture. The weighting process helps to reduce the interference of unimportant or short-term changes, ensuring that the system can give priority to those action segments that have a greater impact on gesture recognition and classification, thereby improving the accuracy and real-time response ability of gesture recognition.

[0109] Finally, all weighted alternation frequencies are further processed through a non-linear adjustment function to generate the final alternation frequency factor. The calculation expression is as follows:

[0110]

[0111] , where γ(t) is the alternation frequency factor, is the weighted alternation frequency of the j-th action segment, λ jis the weighting factor of the j-th action segment;

[0112] The weighted alternating frequency is further processed by a non-linear adjustment function to enhance the system's sensitivity to dynamic changes in gestures and more accurately capture complex change patterns. The alternating frequency of a gesture typically exhibits non-linear characteristics during rapid changes, and simple linear weighting may not fully represent these complex changes. The non-linear adjustment function can amplify the rapidly changing parts and reduce the influence of relatively stable or repetitive actions, thereby more precisely reflecting the dynamic characteristics of the gesture. In this way, it is possible to better identify and distinguish gestures in rapid change and stable states, improve the system's response ability to rapidly changing gestures, and avoid misjudgments or omissions due to insufficient linear processing. In addition, the non-linear adjustment function can also enhance the robustness of the model, enabling it to adapt to different gesture change scenarios and improving the system's performance in diverse and dynamic environments.

[0113] Under the detection window, the larger the value of the alternating frequency factor generated after analyzing the alternating frequency of the actions in the gesture, the more rapidly changing the current user's gesture is. The alternating frequency factor quantifies the rate of frequent switching or repetition of actions in the gesture. When the user makes rapid gestures, the transitions between actions become more frequent, resulting in an increase in the alternating frequency, manifested as a higher value of the alternating frequency factor. In this case, a larger value of the alternating frequency factor reflects the drastic fluctuations in the gesture state, meaning that the gesture actions change rapidly and unstably, usually indicating rapid changes. On the contrary, a smaller alternating frequency factor indicates that the actions in the gesture are relatively stable, with infrequent switching, and the gesture is in a normal or slowly changing state, indicating that there is no significant acceleration or fluctuation in the current gesture change.

[0114] The machine learning model is not limited here. Any machine learning model that can comprehensively analyze the key point jump factor J(t) and the alternating frequency factor γ(t) to generate the gesture change coefficient GVC is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method:

[0115] The formula for generating the gesture change coefficient GVC is as follows:

[0116]

[0117] , where k1 and k2 are the preset proportionality coefficients of the key point jump factor J(t) and the alternating frequency factor γ(t) respectively, and both k1 and k2 are greater than 0.

[0118] As can be seen from the calculation expression of the gesture change coefficient, under the detection window, the larger the performance value of the key point jump factor generated by analyzing the jumping changes of the hand key points, and the larger the performance value of the alternating frequency factor generated by analyzing the alternating frequency of the actions in the gesture, the larger the performance value of the gesture change coefficient generated by analyzing the extracted key features under the detection window, indicating that the current user's gesture is in a state of rapid change. On the contrary, it indicates that the current user's gesture does not change rapidly and is in a normal change state.

[0119] The preset proportionality coefficients (k1 and k2) are important weight coefficients used to adjust the key point jump factor J(t) and the alternating frequency factor γ(t) when calculating the gesture change coefficient GVC. Their role is to reasonably adjust and optimize the contributions of the two factors according to the actual application requirements. Specifically, k1 and k2 are preset values used to represent the relative importance or influence degree of the two features (jumping changes and alternating frequencies). By adjusting these two coefficients, the influence of the two features on the final gesture change coefficient GVC can be flexibly controlled, so as to better adapt to the requirements of different gesture recognition scenarios or specific tasks. The preset proportionality coefficients are usually determined based on previous experimental data or domain knowledge, aiming to achieve the best recognition accuracy and stability in comprehensive analysis.

[0120] Based on the evaluation results of the model, the user's gesture changes are divided into two categories: normal gesture changes and rapid gesture changes;

[0121] The gesture change coefficient generated by analyzing the extracted key features is compared and analyzed with the preset reference threshold of the gesture change coefficient to divide the user's gesture changes. The division steps are as follows:

[0122] If the gesture change coefficient is greater than or equal to the preset reference threshold of the gesture change coefficient, the user's gesture change is divided into rapid gesture changes;

[0123] If the gesture change coefficient is less than the preset reference threshold of the gesture change coefficient, the user's gesture change is divided into normal gesture changes;

[0124] Normal gesture change refers to the gentle and gradual change of the user's gesture within a certain period of time. Its speed and amplitude are usually relatively uniform and stable. In this case, the action conversion of the gesture is relatively slow, and there will be no violent jumps or sudden changes; rapid gesture change refers to the rapid and violent change of the user's gesture within a short period of time, usually accompanied by a large action amplitude or frequent dynamic switching. Such changes are manifested as rapid changes in hand movements, such as rapid waving, sudden swinging, or rapid gesture alternating actions.

[0125] For normal gesture changes, continue to capture hand images at a preset fixed frame rate, and use gesture recognition algorithms to identify actions and intentions;

[0126] For situations classified as normal gesture changes, the system continues to capture hand images at a preset fixed frame rate and uses existing gesture recognition algorithms to identify actions and intentions. The purpose is to ensure the efficiency and stability of the system in regular gesture operations. Normal gesture changes usually refer to relatively smooth changes in the user's actions without drastic changes or irregular fluctuations. In this case, maintaining a fixed frame rate for image capture can fully meet the requirements of gesture recognition while avoiding excessive consumption of computing resources. By using existing gesture recognition algorithms, the system can efficiently classify and analyze common gesture actions without additional computational burdens or frequent adjustments. This not only ensures real-time performance and smoothness but also guarantees the reliability and user experience of the system in daily operations, suitable for most regular gesture input scenarios.

[0127] For rapid gesture changes, based on the evaluation results of the machine learning model and particle swarm optimization, dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and capture more dynamic details;

[0128] For rapid gesture changes, based on the evaluation results of the machine learning model and particle swarm optimization, the specific steps to dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and capture more dynamic details are as follows:

[0129] When it is detected that the gesture is in a rapid change state, start the particle swarm optimization (PSO) algorithm and initialize the frame rate increment of the particle. The initialization of the frame rate increment of the particle not only depends on the gesture change coefficient GVC but also comprehensively considers the historical gesture change coefficient GVC value for more accurate subsequent frame rate adjustment. The initialized frame rate increment formula is as follows:

[0130] ΔFPS b (t)=(GVC(t)-GVC ref )+θ·GVC(t - 1)

[0131] , where ΔFPS b (t) is the frame rate increment of the bth particle at time point t, GVC(t) is the current gesture change coefficient, θ is the historical weight factor reflecting the influence of historical changes, GVC ref is the gesture change coefficient reference threshold, and GVC(t - 1) is the gesture change coefficient at the previous moment;

[0132] The particle swarm optimization algorithm optimizes the optimal frame rate increment of each particle according to the change of the frame rate increment ΔFPS b (t), and the calculation expression is as follows:

[0133] ΔFPS opt (t) = |GVC(t) - GVC ref | + τ·|ΔFPS b (t) - ΔFPS b (t - 1)|

[0134] , where ΔFPS opt (t) is the optimal frame rate increment, τ is the penalty coefficient, and ΔFPS b (t - 1) is the frame rate increment at the previous moment, that is, the frame rate increment of the b-th particle at the moment t - 1;

[0135] The penalty coefficient τ is a parameter used to adjust and control the influence of the change in the past frame rate increment during the particle swarm optimization (PSO) algorithm. Its main role is to prevent the particles from over-adjusting the frame rate and avoid unnecessary fluctuations in the system. Specifically, the penalty coefficient imposes a certain penalty on the change of each particle in the particle swarm during historical frame rate adjustment, making the particles pay more attention to smooth and gradual frame rate adjustment rather than drastic changes during the optimization process. This helps to balance the dynamic adjustment of the frame rate and the system stability in the case of rapid gesture changes, ensuring that the system will not be unstable or waste resources due to rapid gesture changes. By adjusting the penalty coefficient, the system can maintain sufficient smoothness when responding to the needs of rapid gesture changes and avoid the side effects caused by over-adjustment.

[0136] After obtaining the optimal frame rate increment ΔFPS opt (t), the optimal frame rate increment ΔFPS opt (t) is combined with the initial preset frame rate to dynamically adjust the frame rate of the camera. The calculation expression of the optimized camera frame rate is as follows:

[0137]

[0138] , where FPS new (t) is the optimized camera frame rate, FPS0 is the preset camera frame rate, δ is the historical smoothing factor that controls the influence degree of historical frame rate change on the current frame rate adjustment, and B is the total number of particles.

[0139] In the case of rapid gesture changes, the system needs to be able to respond in a timely manner and capture gesture details to ensure accurate recognition of the user's actions and intentions. The role of this step is to dynamically adjust the frame rate of the camera using the Particle Swarm Optimization (PSO) algorithm based on the evaluation results of the machine learning model, thereby reducing the inter-frame time interval and increasing the capture frequency and accuracy of the system for gesture changes. Specifically, when the Gesture Variation Coefficient (GVC) is higher than the set reference threshold, it indicates that the gesture is undergoing rapid changes, and a higher frame rate is required to capture the gesture details meticulously. The Particle Swarm Optimization algorithm can efficiently explore and optimize the frame rate adjustment strategy by simulating the search and aggregation process of particles in nature. Through particle swarm optimization, the system can find the optimal solution among multiple possible frame rate increments and dynamically adjust the frame rate of the camera to adapt to the rapid changes of the gesture. Reducing the inter-frame time interval helps the system obtain more dynamic images in a shorter time, thereby enhancing the capture ability for rapid gesture changes and enabling the system to achieve higher recognition accuracy in a high-dynamic gesture environment. This adjustment strategy not only ensures that the system can handle different gesture changes but also avoids action omissions or blurs caused by too low a frame rate, thus enhancing the system's rapid response ability and accuracy to the user's intentions.

[0140] The present invention can accurately identify rapid changes in user gestures by introducing an intelligent evaluation mechanism based on deep learning. When the user's gesture undergoes rapid changes within a short period, a traditional fixed-frame-rate camera may miss key dynamic details, resulting in incomplete or incorrect gesture recognition. By analyzing the jumping changes of hand key points and the alternating frequency of actions in the gesture in real time, after generating the key point jumping factor and the alternating frequency factor, the machine learning model can intelligently evaluate the change pattern of the gesture. When it detects that the user's gesture has undergone rapid changes, the system can dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and thus capture more dynamic details. This real-time and sensitive adjustment ensures that the system can more accurately identify rapidly changing gestures, avoid omitting or misclassifying gesture actions, and thereby improve the accuracy of gesture recognition. The user can obtain a more precise and smooth interaction experience and reduce misoperations caused by inaccurate recognition.

[0141] The dynamic frame rate adjustment mechanism of the present invention combines real-time analysis of gesture changes and can flexibly adjust the working mode of the camera according to the changes of gestures. When the user makes rapid gesture movements, the system can timely adjust the frame rate to capture the key details of the gesture movements, avoiding interaction delays or missed information caused by insufficient frame rates. This mechanism ensures the real-time response of the interactive system to the user's intentions and avoids the system's misunderstanding of the user's behavior due to recognition errors. Further, the combination of intelligent evaluation and dynamic frame rate adjustment not only improves the sensitivity of the system to rapid gestures but also makes the recognition of different types of gestures by the system more natural and smooth. During the use process, users can obtain a smoother and delay-free interaction experience. This precise response and high sensitivity will greatly improve user satisfaction and interactive perception, optimizing the overall user experience.

[0142] The present invention provides a human-computer interaction optimization system based on deep learning as shown in Figure 2 Figure 5, including a gesture image capture module, an image data storage and analysis set construction module, a feature extraction and analysis module, an intelligent evaluation and classification module, a normal gesture recognition module, and a dynamic frame rate adjustment module:

[0143] The gesture image capture module first captures the user's hand image by the camera at a preset frame rate;

[0144] The image data storage and analysis set construction module stores the captured hand image frames and organizes them into an analysis set, and extracts the key features reflecting the rapid changes of the gestures from the analysis set;

[0145] The feature extraction and analysis module conducts a detailed analysis of the extracted key features within the detection window and inputs the key features after analysis and processing into a pre-trained machine learning model to conduct intelligent evaluation of the abnormal changes of the user's gestures through the machine learning model;

[0146] The intelligent evaluation and classification module divides the user's gesture changes into two categories: normal gesture changes and rapid gesture changes based on the evaluation results of the model;

[0147] The normal gesture recognition module continues to capture the hand image at a preset fixed frame rate for normal gesture changes and uses gesture recognition algorithms to recognize the actions and intentions;

[0148] The dynamic frame rate adjustment module dynamically adjusts the frame rate of the camera according to the evaluation results of the machine learning model for rapid gesture changes, reduces the inter-frame time interval, and captures more dynamic details.

[0149] An optimization method for human-computer interaction based on deep learning provided by an embodiment of the present invention is implemented through the above-mentioned optimization system for human-computer interaction based on deep learning. For the specific methods and processes of an optimization system for human-computer interaction based on deep learning, refer to the embodiments of the above-mentioned optimization method for human-computer interaction based on deep learning, which will not be elaborated here.

[0150] Only some exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, various different ways can be used to modify the described embodiments without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A human-computer interaction optimization method based on deep learning, characterized in that, It includes the following steps: First, the camera captures the user's hand image at a preset frame rate; The captured hand image frames are stored and organized into an analysis set, and key features reflecting rapid gesture changes are extracted from the analysis set; The extracted key features are analyzed in detail within the detection window, and the key features after analysis and processing are input into a pre-trained machine learning model, and the machine learning model is used to intelligently evaluate abnormal changes in the user's gestures; Based on the evaluation results of the model, the user's gesture changes are divided into two categories: normal gesture changes and rapid gesture changes; For normal gesture changes, continue to capture hand images at a preset fixed frame rate, and use gesture recognition algorithms to identify actions and intentions; For rapid gesture changes, according to the evaluation results of the machine learning model, dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and capture more dynamic details; Key features reflecting rapid gesture changes are extracted from the analysis set. The extracted key features include the jumping changes of hand key points and the alternating frequency of actions in the gesture. Under the detection window, the jumping changes of the extracted hand key points and the alternating frequency of actions in the gesture are analyzed to generate a key point jumping factor and an alternating frequency factor respectively. The key point jumping factor quantifies the degree of drastic change in the position of hand key points between adjacent frames, reflecting the mutability and discontinuity of the spatial position during the gesture movement process; the alternating frequency factor quantifies the frequent switching of actions in the gesture, reflecting the intensity of alternating changes in the gesture dynamic pattern; Under the detection window, the specific steps for analyzing the jumping changes of hand key points to generate the key point jumping factor are as follows: Under the detection window, for each frame, the spatial coordinate data of the key points are extracted using a hand tracking algorithm and labeled as P i (t), representing the coordinate position of the i-th key point at time point t, P i (t) = {x i (t), y i (t), z i (t)}, where x i (t), y i (t), z i (t) are the positions of the i-th key point on the x-axis, y-axis, and z-axis at time point t, respectively; Calculate the change in the position of key points between each pair of consecutive frames, that is, the displacement of the i-th key point between adjacent frames. The expression is as follows: , where Δt is the time interval, and ΔP i (t) is the change in the position of the key point, that is, the spatial change of the i-th key point within the time interval Δt, x i (t + Δt) is the position of the i-th key point on the x-axis at the time point t + Δt, y i (t + Δt) is the position of the i-th key point on the y-axis at the time point t + Δt, z i (t + Δt) is the position of the i-th key point on the z-axis at the time point t + Δt; To quantify the mutation degree of the change of each hand key point, a non-linear weighting method is introduced to define the local key point jumping factor. The calculation expression is as follows: J i (t) = (ΔP i (t)) α ·e -β·Δt , where J i (t) is the local key point jump factor, α is the nonlinear weighting exponent, e -β·Δt is the time decay factor, e is the natural base, and β is the time decay constant; The jump factor J i (t) for all local key points is synthesized to obtain the final key point jump factor, and the calculation expression is as follows: , where J(t) is the key point jump factor, w i is the weight of the i-th key point, and N is the total number of key points.

2. The human-computer interaction optimization method based on deep learning according to claim 1, wherein After obtaining the key point jumping factor and the alternating frequency factor generated after analyzing the extracted key features, input the key point jumping factor and the alternating frequency factor into the pre-learned machine learning model, generate a gesture change coefficient through the machine learning model, and use the gesture change coefficient to intelligently evaluate abnormal changes in the user's gestures.

3. The human-computer interaction optimization method based on deep learning according to claim 2, wherein Compare and analyze the gesture change coefficient generated after analyzing the extracted key features with a preset reference threshold of the gesture change coefficient to divide the user's gesture changes. The division steps are as follows: If the gesture change coefficient is greater than or equal to the preset reference threshold of the gesture change coefficient, divide the user's gesture change into a rapid gesture change; If the gesture change coefficient is less than the preset reference threshold of the gesture change coefficient, divide the user's gesture change into a normal gesture change.

4. The human-computer interaction optimization method based on deep learning according to claim 3, characterized in that, For rapid gesture changes, according to the evaluation results of the machine learning model, based on particle swarm optimization, the specific steps for dynamically adjusting the frame rate of the camera, reducing the inter-frame time interval, and capturing more dynamic details are as follows: When it is detected that the gesture is in a rapidly changing state, the particle swarm optimization algorithm is started, and the frame rate increment of the particles is initialized. The initialization of the frame rate increment of the particles not only depends on the gesture variation coefficient GVC, but also synthesizes the historical gesture variation coefficient GVC values to perform subsequent frame rate adjustment more precisely. The formula for the initialized frame rate increment is as follows: ΔFPS b (t) = (GVC(t) - GVC ref ) + θ·GVC(t - 1), Where, ΔFPS b (t) is the frame rate increment of the b-th particle at time point t, GVC(t) is the current gesture change coefficient, θ is the historical weight factor, reflecting the influence of historical changes, GVC ref is the reference threshold of the gesture change coefficient, and GVC(t - 1) is the gesture change coefficient at the previous moment; The particle swarm optimization algorithm optimizes the optimal frame rate increment of each particle according to the change of the frame rate increment ΔFPS b (t), and the calculation expression is as follows: ΔFPS opt (t) = |GVC(t) - GVC ref | + τ·|ΔFPS b (t) - ΔFPS b (t - 1)|, where ΔFPS opt (t) is the optimal frame rate increment, τ is the penalty coefficient, and ΔFPS b (t - 1) is the frame rate increment at the previous moment, that is, the frame rate increment of the b-th particle at the time point t - 1; Obtain the optimal frame rate increment ΔFPS opt After (t), the optimal frame rate increment ΔFPS opt (t) is combined with the initial preset frame rate to dynamically adjust the frame rate of the camera. The calculation expression of the optimized camera frame rate is as follows: , Wherein, FPS new (t) is the optimized camera frame rate, FPS0 is the preset camera frame rate, δ is the historical smoothing factor that controls the influence degree of the historical frame rate change on the current frame rate adjustment, and B is the total number of particles.

5. A human-computer interaction optimization method based on deep learning according to claim 1, characterized in that, Under the detection window, the specific steps for analyzing the alternating frequency of actions in the gesture are as follows: Under the detection window, the captured continuous gesture image frames are divided into multiple segments, each segment representing the duration of a gesture action. According to the change of the positions of the hand key points between consecutive frames, each segment of the action is divided. The expression is as follows: , Where, A(t) is the change amount of the gesture action, representing the change amount of the gesture action at time point t, and P i (t) is the coordinate position of the i-th key point at time point t, and P i (t - 1) is the coordinate position of the i-th key point at time point t - 1, and N is the total number of key points; When each segment of the action is divided, then calculate the alternating frequency between each action segment. The alternating frequency measures the rate of gesture action switching per unit time and reflects the dynamic change of the gesture. The calculation expression is as follows: , Where, ω(t) is the alternating frequency, and action j (t) is the eigenvalue of the j-th action segment at time point t, and action j (t - 1) is the eigenvalue of the j-th action segment at time point t - 1, and Δt is the time interval; To consider the influence of different action segments, the alternating frequency needs to be weighted and adjusted according to the intensity of different action segments. Based on the complexity and action intensity of each action segment, weights are used to adjust the calculation of the alternating frequency. The calculation expression of the weighted-adjusted alternating frequency is as follows: , In the formula, is the weighted alternating frequency, w j is the complexity weight of the j-th action segment, ω j (t) is the alternating frequency of the j-th action segment; Finally, all weighted alternating frequencies are further processed through a non-linear adjustment function to generate the final alternating frequency factor. The calculation expression is as follows: , where γ(t) is the alternating frequency factor, is the weighted alternating frequency of the j-th action segment, and λ j is the weighting factor of the j-th action segment.

6. A human-computer interaction optimization system based on deep learning, which is used to implement the human-computer interaction optimization method based on deep learning according to any one of the above claims 1-5, characterized in that, Including a gesture image capture module, an image data storage and analysis set construction module, a feature extraction and analysis module, an intelligent evaluation and classification module, a normal gesture recognition module, and a dynamic frame rate adjustment module: The gesture image capture module. First, the camera captures the user's hand image at a pre-set frame rate; The image data storage and analysis set construction module. The captured hand image frames are stored and organized into an analysis set, and the key features reflecting the rapid change of the gesture are extracted from the analysis set; The feature extraction and analysis module. The extracted key features are analyzed in detail within the detection window, and the key features after analysis and processing are input into a pre-trained machine learning model. Through the machine learning model, the abnormal change of the user's gesture is intelligently evaluated; The intelligent evaluation and classification module. Based on the evaluation results of the model, the gesture changes of the user are divided into two categories: normal gesture changes and rapid gesture changes; The normal gesture recognition module. For normal gesture changes, continue to capture the hand image at a pre-set fixed frame rate, and use the gesture recognition algorithm to recognize the action and intention; The dynamic frame rate adjustment module. For rapid gesture changes, according to the evaluation results of the machine learning model, dynamically adjust the frame rate of the camera, reduce the inter-frame time interval, and capture more dynamic details.

Citation Information

Patent Citations

  • Ultrasonic imaging system, control method, imaging controller and storage medium

    CN117045281A

  • Multi-screen display man-machine interaction method and system based on Leap Motion

    CN117093076A