An Adaptive Quick Answer Optimization Method and System Based on Reinforcement Learning

By applying reinforcement learning algorithms and adaptive optical technology in the classroom quiz system, the signal delay, fairness and safety problems in the classroom quiz system are solved, and efficient, fair and safe classroom interaction is achieved.

CN118780426BActive Publication Date: 2025-06-20GUANGDONG ACAD OF EDUCATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410822543.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-06-20
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

The existing classroom quiz system has problems such as high signal transmission delay, difficulty in ensuring fairness between equipment and insufficient security, which affects the quality and effectiveness of classroom interaction.

Method used

Adaptive quick-answer optimization method based on reinforcement learning is adopted to optimize the processing order and weight of quick-answer signals through reinforcement learning algorithms, combined with adaptive optical technology and dynamic wavefront correction technology, high-speed and low-latency signal transmission is achieved, and the signal security is ensured through optical data encryption technology.

Benefits of technology

It effectively overcomes the signal delay, fairness and security issues of existing systems, and provides an efficient, fair and safe classroom interaction solution, which significantly improves the system's real-time response capabilities and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118780426B_ABST
    Figure CN118780426B_ABST
Patent Text Reader

Abstract

The present invention proposes an adaptive quick-answer optimization method and system based on reinforcement learning. The method includes: using sensors to collect multi-modal quick-answer signals of students, performing preprocessing to obtain target signals, and performing feature extraction and fusion; the edge computing node collects the target signals, forms multi-modal feature data, performs feature fusion on the multi-modal feature data, and constructs a reinforcement learning model to rank the priority of the quick-answer signals; dynamically adjusts the processing order and weight of the quick-answer signals; optimizes the transmission path of the quick-answer signals, and encrypts and decrypts the optical signals generated by the optical adaptive optical technology for transmission; the central processing unit receives the decrypted optical signals and performs processing and analysis, and sends the analysis results to each edge computing node and user. The present invention proposes an adaptive optical quick-answer fairness optimization system based on reinforcement learning, which effectively solves the problems existing in the existing classroom quick-answer system in terms of signal transmission delay, fairness, and security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of reinforcement learning, and particularly relates to an adaptive quick answer optimization method and system based on reinforcement learning. Background Art

[0002] With the rapid development of information technology and educational technology, the classroom interactive teaching method has been continuously innovated, and the quick answer system has gradually become an important tool to improve classroom interactivity and student participation. However, there are still many problems in the actual application of the existing classroom quick answer system, which are difficult to meet the needs of the modern educational environment.

[0003] Traditional classroom quick answer systems usually rely on button devices or gesture recognition devices, which are connected to the central server by wired or wireless means, and the central server is responsible for signal reception, processing and feedback. The biggest drawback of such systems is that it is difficult to ensure the signal transmission delay and fairness between devices. Especially in a large-scale classroom environment, problems such as signal interference, network congestion and transmission delay between devices are particularly prominent, resulting in the accuracy and fairness of the quick answer results being affected. In addition, there are also potential safety hazards in the existing systems. The quick answer signals are easily interfered with or tampered with during the transmission process, affecting the quality and effect of classroom interaction.

[0004] Some existing improvement methods, such as using high-frequency signal transmission, adding signal amplifiers and optimizing the network architecture, although alleviating the signal transmission delay and interference problems to a certain extent, still cannot fundamentally solve the fairness and security problems of the system. On the other hand, the application of artificial intelligence technology has improved the recognition and processing efficiency of quick answer signals to a certain extent, but due to the lack of real-time optimization of network delay and transmission path, there are still problems such as high delay, slow response and poor fairness.

[0005] In summary, the current classroom quick answer system faces the following main problems:

[0006] 1) High signal transmission delay, resulting in slow quick answer response speed;

[0007] 2) It is difficult to ensure fairness between devices, affecting students' participation enthusiasm;

[0008] 3) Insufficient security during signal transmission, vulnerable to interference and tampering.

[0009] In view of these problems, there is an urgent need for a new solution that can improve signal transmission efficiency, ensure system fairness and security. Summary of the Invention

[0010] The object of the present invention is to design an adaptive rush-answer optimization method and system based on reinforcement learning, which realizes the fairness optimization of rush-answer signals through a reinforcement learning algorithm, ensuring that all students have equal rush-answer opportunities at any time point. At the same time, the use of adaptive optical technology realizes the high-speed and low-latency transmission of rush-answer signals, significantly improving the real-time response ability and transmission efficiency of the system. In addition, the application of optical data encryption technology ensures the security and integrity of the rush-answer signal transmission process, fundamentally solving the security hidden dangers of the existing system. These innovative points work together to enable the present invention to effectively overcome the shortcomings of the existing classroom rush-answer system and provide an efficient, fair and secure classroom interaction solution.

[0011] To achieve the above object, in the first aspect of the present invention, an adaptive rush-answer optimization method based on reinforcement learning is provided, and the method includes the following steps:

[0012] S1. Use sensors to collect multi-modal rush-answer signals of students, preprocess the rush-answer information to obtain target signals, and the sensors perform feature extraction and fusion on the target to obtain a unified data set; wherein, the multi-modal rush-answer signals include voice signals, video signals and key signals;

[0013] S2. The edge computing node collects the target signals to form multi-modal feature data, the edge computing node performs feature fusion on the multi-modal feature data to obtain a fused multi-modal feature vector, and constructs a reinforcement learning model to rank the priority of the rush-answer signals;

[0014] S3. Introduce a reinforcement learning algorithm to dynamically adjust the processing order and weight of the rush-answer signals;

[0015] S4. Use adaptive optical technology combined with dynamic wavefront correction technology to optimize the transmission path of the rush-answer signals, and perform encryption and decryption transmission on the optical signals generated by the adaptive optical technology;

[0016] S5. The central processing unit receives the decrypted optical signals, processes and analyzes them, and sends the analysis results to each edge computing node and user;

[0017] Among them, the specific content of S3 includes:

[0018] S301. Construct a reinforcement learning environment:

[0019] Construct the state space S(t) of the reinforcement learning environment, where the state space S(t) includes the reduced-dimensional feature vector F pca (t) and the delay information D latency (t) at the current time step, which is expressed as follows:

[0020] S(t) = [F pca (t), D latency(t)

[0021] Among them, D latency (t) represents the delay information at the current time step;

[0022] Define the action space A(t), which includes the priority sorting and weight adjustment of the rush-answer signals, as follows:

[0023] A(t) = [P priority (t), W weight (t)]

[0024] Among them, P priority (t) represents the priority sorting at the current time step, and W weight (t) represents the weight adjustment at the current time step;

[0025] S302. Design the reward function:

[0026] Let the response time of the i-th student be T i (t), and the device performance be P i (t). Define the response time fairness index FairTime(A(t)) as:

[0027]

[0028] Among them, represents the average response time of all students, and N represents the total number of students;

[0029] The device difference fairness index FairEquip(A(t)) is defined as:

[0030]

[0031] Among them, represents the average device performance of all students;

[0032] By adding the response time and device difference, define the total fairness index Fairness(A(t)) as:

[0033] Fairness(A(t)) = γ1·FairTime(A(t)) + γ2·FairEquip(A(t))

[0034] Among them, γ1 and γ2 represent the weight parameters for balancing the response time and device difference;

[0035] The real-time index RealTime(A(t)) is used to measure the response speed of the system. Define the real-time index as:

[0036] RealTime(A(t)) = -max i∈{1,2,…,N} (Ti (t) - T start (t))

[0037] Among them, T start (t) represents the start time of the current time step;

[0038] Adding fairness and real-time performance, the comprehensive reward function R(t) is designed as:

[0039] R(t) = α R ·Fairness(A(t)) + β R ·RealTime(A(t))

[0040] Among them, α R and β R represent weight parameters used to balance the impacts of fairness and real-time performance.

[0041] In one embodiment, the preprocessing includes signal synchronization processing, noise cancellation, and signal normalization processing.

[0042] In one embodiment, the signal synchronization processing is expressed as follows:

[0043] Define the timestamp T of each signal i , and align all signals to a unified time reference T through interpolation method sync , which is expressed as follows

[0044]

[0045] Among them, T sync represents the unified time reference after synchronization, represents the timestamp of the voice signal S v ; represents the timestamp of the video signal S img ; represents the timestamp of the key signal S k ;

[0046] Perform time synchronization on each signal:

[0047]

[0048] Among them, represents the time derivative of each signal, S v (t) represents the voice signal, S img (t) represents the video signal, S k (t) represents the key signal;

[0049] The noise cancellation is expressed as follows:

[0050] Let \(S(t)\) be the original signal and \(N(t)\) be the noise signal. The noise is filtered by the adaptive filter \(H adapt (t)\) to obtain the purified signal \(S clean (t)\), which is calculated as follows:

[0051] S clean (t)=S(t)-H adapt (t)\cdot N(t)

[0052] The update formula of the filter is:

[0053] H adapt (t + 1)=H adapt (t)+\mu H \cdot e H (t)\cdot N(t)

[0054] where \(\mu H represents the filter learning rate, and \(e H (t)\) represents the current error, \(e H (t)=S(t)-S clean (t)\) represents the error;

[0055] The signal normalization is expressed as follows:

[0056]

[0057] where \(\mu S and \(\sigma s are the mean and standard deviation of the signal respectively, and \(S norm (t)\) represents the normalized signal.

[0058] In one embodiment, the reinforcement learning environment of the reinforcement learning model of the edge computing node is constructed as follows:

[0059] The state space includes the feature vector \(F pca (t)\) after dimensionality reduction, and the action space includes the priority ranking of the signals;

[0060] Design the reward function \(r(t)\) to evaluate the effect of each action, which is expressed as follows:

[0061] r(t)=\alpha R \cdot Fairness(a(t))+\beta R \cdot RealTime(a(t))

[0062] where \(\alpha R and \(\beta R are weight parameters, \(Fairness(a(t))\) is the fairness index of the current action, and \(RealTime(a(t))\) is the real-time index;

[0063] Training is carried out using the deep Q - network algorithm. Define the Q - value function Q(s,a), which is expressed as follows:

[0064] Q(s(t),a(t)) = Q(s(t),a(t))+η Q [r(t)+γ Q max A′ Q(s(t + 1),a′)-Q(s(t),a(t))]

[0065] Among them, η Q is the learning rate, and γ Q is the discount factor.

[0066] In one embodiment, according to the action a(t) output by the reinforcement learning algorithm, the signals at the current moment are sorted by priority. Let the priority sorting function be Π, and the calculation is as follows:

[0067] F sorted (t)=Π(F pca (t),a(t))

[0068] Among them, F pca (t) represents the feature vector after dimensionality reduction, and F sorted (t) represents the output priority sequence.

[0069] In one embodiment, in the above - mentioned S3, it further includes:

[0070] S303. Dynamically adjust the signal processing order and weights using the reinforcement learning algorithm:

[0071] Define the Q - value function Q(S,A), which represents the expected reward obtained by taking action A in state S:

[0072]

[0073] Among them, γ R represents the discount factor, and A′ represents the action at the next time step;

[0074] Use the deep Q - network algorithm to train and update the Q - value function; among them, the update formula of the Q - value is:

[0075] Use the deep Q - network algorithm to train and update the Q - value function; among them, the update formula of the Q - value is:

[0076]

[0077] Among them, η Q represents the learning rate;

[0078] Use a DNN network to fit the Q-value function; where the input of the neural network is the state S(t), and the output is the Q-value of each action, expressed as follows:

[0079]

[0080] Among them, θ represents the parameters of the neural network, which are optimized through the backpropagation algorithm, represents the Q-value function fitted by the DNN network, that is, the fitted Q-value;

[0081] S304. Add device performance information to the state space and introduce a time window mechanism into the action space; where the state space S(t) is defined as:

[0082] S(t) = [F pca (t), D latency (t), P device (t)]

[0083] Among them, P device (t) represents the device performance;

[0084] The action space A(t) is defined as:

[0085] A(t) = [P priority (t), W weight (t), T window (t)]

[0086] Among them, among them, T window (t) represents the time window parameter;

[0087] S305. At each time step, use the ∈-greedy strategy to select an action; where an action is randomly selected with a probability of ∈, and an action with the largest current Q-value is selected with a probability of 1 - ∈, expressed as follows:

[0088]

[0089] According to the selected action A(t), prioritize and adjust the weights of the signals at the current time step, expressed as follows:

[0090] F sorted (t) = Π(F pca (t), P priority (t))

[0091] F weighted (t) = F sorted (t) ⊙ W weight (t)

[0092] Among them, Π represents the priority sorting function, and ⊙ represents the element-wise multiplication operation;

[0093] S306. Store (S(t), A(t), R(t), S(t + 1)) of each interaction into the experience replay pool . During each training, randomly extract a mini - batch (S i , A i , R i , S i+1 ) from the experience replay pool for training, which is expressed as follows:

[0094]

[0095] Among them, represents extracting samples (S , A i , R i , S i , S i+1 ) from the experience replay pool for expectation calculation. L(θ) is the loss function, and the network parameter θ is optimized by minimizing the loss function;

[0096] Update the network parameter θ using the stochastic gradient descent algorithm, which is expressed as follows:

[0097]

[0098] Among them, η θ is the learning rate of the stochastic gradient descent algorithm, is the gradient of the loss function L(θ) with respect to the network parameter θ.

[0099] In one embodiment, in the said S4, the adaptive optical technology is combined with the dynamic wavefront correction technology to optimize the transmission path of the rush - answer signal, which specifically includes:

[0100] Use an electro - optic modulator to convert the multi - modal feature data Features(t) output from the edge computing node into an optical signal for transmission. Real - time monitor the wavefront distortion of the optical signal through a wavefront sensor, use an adaptive optical controller to calculate the correction signal C(t) to correct the wavefront distortion, and then perform wavefront correction on the optical signal through a correction mirror. At the same time, apply the correction signal C(t) to the drive signal of the correction mirror to obtain the signal E corrected (t);

[0101] Select the best optical fiber transmission path according to the real - time network condition, design a path optimization algorithm to calculate the shortest transmission path in real - time, and then transmit the signal E corrected (t) through the optimized optical fiber path;

[0102] The receiving end of the central server receives the transmitted optical signal E received(t), and demodulate the received electrical signal to restore the original transmitted signal Features received (t), and perform preprocessing and feature extraction on the signal Features received (t).

[0103] In one embodiment, in the step S4, encrypt and decrypt the optical signal generated by the optical adaptive optical technology for transmission, specifically including:

[0104] Use a phase encoder to encrypt the optical signal, and at the same time generate a phase mask through a physical random number generator;

[0105] Transmit the encrypted optical signal to the receiving end through an optical fiber. The receiving end decrypts the received encrypted optical signal with the same phase mask to restore the original optical signal. At the same time, adopt time-division multiplexing technology to synchronize the phase mask through a synchronous optical signal;

[0106] Process the decrypted optical signal to restore the original electrical signal.

[0107] In one embodiment, the step S5 specifically includes:

[0108] S501. The central server receives the transmitted optical signal through a photodetector, converts the optical signal into an electrical signal, and then decodes the received electrical signal to restore the original multimodal feature signal;

[0109] S502. Perform denoising processing on the decoded multimodal feature signal and synchronize the system clock, and extract features from the processed multimodal feature signal to obtain a feature signal;

[0110] S503. Sort the feature signals according to the importance and priority of the signals, and process the sorted feature signals to generate a final response result;

[0111] S504. Generate a feedback response of the system according to the response result, and send the generated feedback response to each edge node and user through the network to complete the feedback process of the system.

[0112] In the second aspect of the present invention, an adaptive quick answer optimization system based on reinforcement learning is provided, including the following modules:

[0113] A signal collection module, which is used to collect multimodal quick answer signals of students by using sensors, preprocess the quick answer information to obtain a target signal, and perform feature extraction and fusion on the target by the sensors to obtain a unified data set; wherein, the multimodal quick answer signals include voice signals, video signals and key signals;

[0114] A signal processing module is used for an edge computing node to collect target signals, form multi-modal feature data, the edge computing node performs feature fusion on the multi-modal feature data to obtain a fused multi-modal feature vector, and constructs a reinforcement learning model to rank the priority of the rush-answer signals; a reinforcement learning algorithm is introduced to dynamically adjust the processing order and weights of the rush-answer signals;

[0115] A signal transmission module is used for optimizing the transmission path of the rush-answer signals by adopting adaptive optical technology combined with dynamic wavefront correction technology, and performing encrypted and decrypted transmission on the optical signals generated by the optical adaptive optical technology;

[0116] A signal sending module is used for enabling a central processing unit to receive the decrypted optical signals, perform processing and analysis, and send the analysis results to each edge computing node and user;

[0117] Among them, the reinforcement learning algorithm specifically includes:

[0118] S301. Construct a reinforcement learning environment:

[0119] Construct a state space S(t) of the reinforcement learning environment, where the state space S(t) includes the dimensionality-reduced feature vector F pca (t) and the delay information D latency (t) at the current time step, which is expressed as follows:

[0120] S(t) = [F pca (t), D latency (t)]

[0121] Among them, D latency (t) represents the delay information at the current time step;

[0122] Define an action space A(t), including the priority ranking and weight adjustment of the rush-answer signals, which is expressed as follows:

[0123] A(t) = [P priority (t), W weight (t)]

[0124] Among them, P [riority (t) represents the priority ranking at the current time step, and W weight (t) represents the weight adjustment at the current time step;

[0125] S302. Design a reward function:

[0126] Let the response time of the i-th student be T i (t), and the device performance be P i (t). Define the response time fairness index FairTime(A(t)) as:

[0127]

[0128] Among them, represents the average response time of all students, and N represents the total number of students;

[0129] The device difference fairness index FairEquip(A(t)) is defined as:

[0130]

[0131] Among them, represents the average device performance of all students;

[0132] By adding the response time and device difference, the total fairness index Fairness(A(t)) is defined as:

[0133] Fairness(A(t)) = γ1·FairTime(A(t)) + γ2·FairEquip(A(t))

[0134] Among them, γ1 and γ2 represent the weight parameters for balancing the response time and device difference;

[0135] The real-time index RealTime(A(t)) is used to measure the response speed of the system, and the real-time index is defined as:

[0136] RealTime(A(t)) = -max i∈{1,2,…,N} (T i (t) - T start (t))

[0137] Among them, T start (t) represents the start time of the current time step;

[0138] By adding fairness and real-time, the comprehensive reward function R(t) is designed as:

[0139] R(t) = α R ·Fairness(A(t)) + β R ·RealTime(A(t))

[0140] Among them, α R and β R represent the weight parameters for balancing the impacts of fairness and real-time.

[0141] The beneficial technical effects of the present invention are at least as follows:

[0142] (1) The present invention adopts a fairness optimization algorithm for reinforcement learning. By designing an adaptive strategy, the system can dynamically adjust the processing order and weight of the rush-answer signals according to the students' historical rush-answer data, real-time network conditions, and latency feedback, ensuring that all students have fair rush-answer opportunities. This algorithm can monitor the fairness of the system in real time and continuously adjust and optimize through a feedback mechanism to ensure long-term fairness. In addition, the application of multi-modal data fusion technology improves the accuracy and fairness of rush-answer signal recognition, effectively avoiding the bias that may be brought by a single data source.

[0143] (2) The present invention introduces an adaptive optical signal transmission system. Using adaptive optical communication technology, high-speed and low-latency rush-answer signal transmission is achieved through optical fibers. The dynamic wavefront correction technology can adjust the optical fiber transmission path in real time, reducing the loss and latency in signal transmission. Compared with traditional wireless or wired transmission methods, adaptive optical technology provides higher transmission speed and lower latency, which is particularly suitable for large-scale classroom environments. In addition, optical data encryption technology ensures the security and integrity of rush-answer signal transmission, preventing signals from being eavesdropped or tampered with, and improving the overall reliability of the system.

[0144] (3) The present invention realizes the fairness optimization of rush-answer signals through a reinforcement learning algorithm, ensuring that all students have equal rush-answer opportunities at any time point. At the same time, the use of adaptive optical technology realizes high-speed and low-latency transmission of rush-answer signals, significantly improving the real-time response ability and transmission efficiency of the system. In addition, the application of optical data encryption technology ensures the security and integrity during the transmission of rush-answer signals, fundamentally solving the security hidden dangers of existing systems. These innovative points work together, enabling the present invention to effectively overcome the shortcomings of existing classroom rush-answer systems and provide an efficient, fair, and secure classroom interaction solution. Description of the Drawings

[0145] The present invention is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the following drawings.

[0146] Figure 1 It is a flowchart of an adaptive rush-answer optimization method based on reinforcement learning of the present invention.

[0147] Figure 2 It is a framework diagram of an adaptive rush-answer optimization system based on reinforcement learning of the present invention. Detailed Embodiments

[0148] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0149] In one or more embodiments, as Figure 1 shown, an adaptive quick-answer optimization method based on reinforcement learning is disclosed. The method includes the following steps S1 - S5:

[0150] S1. Use sensors to collect multi-modal quick-answer signals of students, and preprocess the quick-answer information to obtain target signals. The sensors perform feature extraction and fusion on the target to obtain a unified data set; wherein, the multi-modal quick-answer signals include voice signals, video signals, and button signals.

[0151] S2. The edge computing node collects the target signals to form multi-modal feature data. The edge computing node performs feature fusion on the multi-modal feature data to obtain a fused multi-modal feature vector, and constructs a reinforcement learning model to rank the priority of the quick-answer signals.

[0152] S3. Introduce a reinforcement learning algorithm to dynamically adjust the processing order and weights of the quick-answer signals.

[0153] S4. Adopt adaptive optical technology combined with dynamic wavefront correction technology to optimize the transmission path of the quick-answer signals, and perform encryption and decryption transmission on the optical signals generated by the adaptive optical technology.

[0154] S5. The central processing unit receives the decrypted optical signals, processes and analyzes them, and sends the analysis results to each edge computing node and user.

[0155] Next, each step will be explained.

[0156] Specifically, in S1, a variety of sensors are arranged in the classroom, including: a microphone array (for collecting the voice signals of students), a camera (for capturing the gestures and facial expressions of students), and a quick-answer button (for recording the button-pressing time of students). These sensors are arranged at different positions in the classroom to ensure that all students' positions can be comprehensively covered. The microphone array is arranged at the four corners of the front, back, left, and right of the classroom to ensure that the voice signals of each student can be clearly collected. The camera is arranged in the front and on the ceiling of the classroom to ensure that the gestures and facial expressions of each student can be captured. The quick-answer button is placed on each student's desktop to ensure that students can conveniently perform quick-answer operations.

[0157] Through sensor collection, multi-modal data signals are obtained: voice signal S v (t), video signal Simg (t), the key signal S k (t); where t represents time. These signals are collected by sensors and transmitted to the edge computing node for processing in real time.

[0158] Furthermore, the preprocessing includes signal synchronization processing, noise cancellation, and signal normalization processing.

[0159] Among them, the signal synchronization processing is as follows:

[0160] Since there are time differences in the signals collected by different sensors, signal synchronization processing is required. Define the timestamp T of each signal i , and align all signals to a unified time reference T through interpolation sync , which is expressed as follows

[0161]

[0162] Perform time synchronization on each signal:

[0163]

[0164] Among them, represents the time derivative of each signal, S v (t) represents the voice signal, S img (t) represents the video signal, S k (t) represents the key signal; in this way, the signals collected by different sensors are aligned to the same time reference to ensure data consistency.

[0165] The noise cancellation is as follows:

[0166] Let S(t) be the original signal and N(t) be the noise signal. Filter the noise through the adaptive filter H adapt (t) to obtain the purified signal S clean (t), which is calculated as follows:

[0167] S clean (t) = S(t) - H adapt (t)·N(t)

[0168] The update formula of the filter is:

[0169] H adapt (t + 1) = H adapt (t) + μ H ·e H (t)·N(t)

[0170] Among them, μ H represents the filter learning rate, e H(t) represents the current error, e H (t) = S(t) - S clean (t) represents the error; in this way, the noise in the signal can be effectively eliminated and the signal quality can be improved.

[0171] The signal standardization is represented as follows:

[0172]

[0173] where, μ S and σ S are the mean and standard deviation of the signal respectively, and S norm (t) represents the standardized signal. Through the standardization process, the dimensional differences of different signals are eliminated, ensuring the consistency of the data.

[0174] Furthermore, key features are extracted from the standardized signal. For the speech signal S v (t), the MFCC (Mel Frequency Cepstral Coefficients) feature MFCC(t) is extracted. The specific formula is:

[0175]

[0176] where, K is the number of filters. For the video signal S img (t), the gesture and facial expression features CNN feat (t) are extracted using a Convolutional Neural Network (CNN). The specific steps are:

[0177] Then the video signal is decomposed into a series of frames: each frame is input into a pre-trained CNN model to extract feature vectors; the feature vectors of all frames are averaged to obtain the final feature vector.

[0178] For the key signal S k (t), the key press time and pressure features Key feat (t) are recorded. The specific steps are: record the key press time t k ; record the key press pressure p k .

[0179] Furthermore, the extracted features are fused to form a unified data set.

[0180] Features(t) = [MFCC(t), CNN feat (t), Key feat (t)]

[0181] The fused data is stored in the local memory of the edge computing node, providing a basis for subsequent signal processing and optimization.

[0182] In the step S2, local signal processing of the edge computing node is formed.

[0183] Specifically, the edge computing node receives in real time the multi-modal feature data collected and preprocessed from the sensors. Suppose the received data set is:

[0184] F(t) = [MFCC i (t), CNN feat,i (t), Key feat,i (t)] for i = 1, 2, …, N

[0185] where, MFCC i (t) represents the speech feature of the i-th student, CNN feat,i (t) represents the video feature of the i-th student, Key feat,i (t) represents the key-pressing feature of the i-th student, and N represents the total number of students.

[0186] To smooth the signal processing process, the received multi-modal data is cached in a circular buffer B(t) with a fixed length. The size of the buffer is L, and each time new data arrives, the buffer content is updated:

[0187] B(t) = [F(t - L + 1), F(t - L + 2), …, F(t)]

[0188] This caching mechanism allows the system to retain the data within the past period of time when processing the current data, provides historical data reference, and increases the stability of processing.

[0189] Furthermore, the multi-modal feature data in the buffer is fused to form a comprehensive feature vector. Suppose the fusion function is Φ, and the fusion process can be implemented through weighted average or through a deep neural network DNN network (deep neural network), where the weighted coefficients or network parameters are learned and optimized through the training data set:

[0190]

[0191] where, α i is the fusion weight, and F fusion (t) is the fused feature vector. The fusion process uses the weighted average of historical data, which can retain the time dependence of the signal and improve the stability and representativeness of the feature vector.

[0192] To reduce the computational complexity, dimensionality reduction processing is performed on the fused feature vector F fusion (t). A dimensionality reduction method based on singular value decomposition (SVD) is adopted to reduce the feature vector to d dimensions:

[0193] F pca (t) = SVD(Ffusion (t), d)

[0194] Among them, F pca (t) is the feature vector after dimensionality reduction, and SVD is the singular value decomposition function. The dimensionality reduction process reduces data redundancy and improves computational efficiency by retaining the most important components in the feature vector.

[0195] Furthermore, construct the reinforcement learning space of the reinforcement learning model. The reinforcement learning environment of the edge computing node's reinforcement learning model is constructed as follows:

[0196] The state space includes the feature vector F pca (t) after dimensionality reduction, and the action space includes the priority ranking of signals;

[0197] Design a reward function r(t) to evaluate the effect of each action, which is expressed as follows:

[0198] r(t) = α R ·Fairness(a(t)) + β R ·RealTime(a(t))

[0199] Among them, α R and β R are weight parameters, Fairness(a(t)) is the fairness index of the current action, and RealTime(a(t)) is the real-time index;

[0200] Use the deep Q-network algorithm for training, and define the Q-value function Q(s, a), which is expressed as follows:

[0201] Q(s(t), a(t)) = Q(s(t), a(t)) + η Q [r(t) + γ Q max A′ Q(s(t + 1), a′) - Q(s(t), a(t))]

[0202] Among them, η Q is the learning rate, and γ Q is the discount factor.

[0203] According to the action a(t) output by the reinforcement learning algorithm, perform priority ranking on the signals at the current moment. Let the priority ranking function be Π, and the calculation is as follows:

[0204] F sorted (t) = Π(F pca (t), a(t))

[0205] Among them, F pca (t) represents the feature vector after dimensionality reduction, F sorted(t) represents the output priority sequence.

[0206] Further, the sorted signals are transmitted to the central server through the adaptive optical signal transmission system. Let the transmission function be Γ:

[0207] F transmitted (t) = Γ(F sorted (t))

[0208] In the step S3, S3 specifically includes:

[0209] S301. Construct a reinforcement learning environment:

[0210] Construct the state space S(t) of the reinforcement learning environment, where the state space S(t) includes the feature vector F pca (t) after dimensionality reduction and the delay information D latency (t), which is expressed as follows:

[0211] S(t) = [F pca (t), D latency (t)]

[0212] Among them, D latency (t) represents the delay information of the current time step;

[0213] Define the action space A(t), including the priority sorting and weight adjustment of the rush-answer signals, which is expressed as follows:

[0214] A(t) = [P priority (t), W weight (t)]

[0215] Among them, P priority (t) represents the priority sorting of the current time step, and W weight (t) represents the weight adjustment of the current time step;

[0216] S302. Design a reward function:

[0217] Let the response time of the i-th student be T i (t), and the device performance be P i (t). Define the response time fairness index FairTime(A(t)) as:

[0218]

[0219] Among them, represents the average response time of all students, and N represents the total number of students;

[0220] The device difference fairness index FairEquip(A(t)) is defined as:

[0221]

[0222] Among them, represents the average device performance of all students;

[0223] By incorporating the response time and device differences, the overall fairness metric Fairness(A(t)) is defined as:

[0224] Fairness(A(t)) = γ1·FairTime(A(t)) + γ2·FairEquip(A(t))

[0225] where γ1 and γ2 represent the weight parameters for balancing the response time and device differences;

[0226] The real-time metric RealTime(A(t)) is used to measure the response speed of the system, and the real-time metric is defined as:

[0227] RealTime(A(t)) = -max i∈{1,2,…,N} (T i (t) - T start (t))

[0228] where T start (t) represents the start time of the current time step;

[0229] By incorporating fairness and real-time, the comprehensive reward function R(t) is designed as:

[0230] R(t) = α R ·Fairness(A(t)) + β R ·RealTime(A(t))

[0231] where α R and β R represent the weight parameters for balancing the impacts of fairness and real-time.

[0232] For example, assume the response times of three students are T1(t) = 2s, T2(t) = 3s, T3(t) = 1s respectively. Calculate the average response time:

[0233]

[0234] Calculate the fairness metric for response time:

[0235]

[0236] Suppose the device performances of three students are P1(t) = 0.9, P2(t) = 1.0, and P3(t) = 0.8 respectively. Calculate the average device performance:

[0237]

[0238] Calculate the device difference fairness index:

[0239]

[0240] Set the weights γ1 = 0.5 and γ2 = 0.5, and calculate the comprehensive fairness:

[0241]

[0242] Suppose the start time T start (t) = 0s of the current time step, and calculate the real-time performance:

[0243] RealTime(A(t)) = -max(2s - 0s, 3s - 0s, 1s - 0s) = -3s

[0244] Set the weights α R = 0.7 and β R = 0.3, and calculate the comprehensive reward:

[0245]

[0246] S303. Dynamically adjust the signal processing order and weights using a reinforcement learning algorithm:

[0247] Define the Q-value function Q(S, A), which represents the expected reward obtained by taking action A in state S:

[0248]

[0249] Among them, γ R represents the discount factor, and A′ represents the action in the next time step;

[0250] Use the deep Q-network algorithm to train and update the Q-value function; among them, the update formula for the Q-value is:

[0251]

[0252] Among them, η Q represents the learning rate of the signal processing order and weights;

[0253] Use a DNN network to fit the Q-value function; among them, the input of the neural network is the state S(t), and the output is the Q-value of each action, which is expressed as follows:

[0254]

[0255] Among them, θ represents the parameters of the neural network and is optimized through the backpropagation algorithm;

[0256] S304. When facing special scenarios, device performance information is added to the state space and a time window mechanism is introduced into the action space; among them, the state space S(t) is defined as:

[0257] S(t) = [F pca (t), D latency (t), P device (t)]

[0258] Among them, P device (t) represents the device performance;

[0259] The action space A(t) is defined as:

[0260] A(t) = [P priority (t), W weight (t), T window (t)]

[0261] Among them, among them, T window (t) represents the time window parameter;

[0262] S305. At each time step, the ∈-greedy policy is used to select an action; among them, an action is randomly selected with a probability of ∈, and the action with the largest current Q value is selected with a probability of 1 - ∈, which is expressed as follows:

[0263]

[0264] According to the selected action A(t), the signals at the current time step are sorted by priority and weighted, which is expressed as follows:

[0265] F sorted (t) = Π(F pca (t), P priority (t))

[0266] F weighted (t) = F sorted (t) ⊙ W weight (t)

[0267] Among them, Π represents the priority sorting function, and ⊙ represents the element-wise multiplication operation;

[0268] S306. Store (S(t), A(t), R(t), S(t + 1)) of each interaction into the experience replay pool During each training, randomly extract a small batch (S i , A i , Ri ,S i+1 ) for training, which is expressed as follows:

[0269]

[0270] Among them, L(θ) is the loss function, and the network parameter θ is optimized by minimizing the loss function;

[0271] The stochastic gradient descent algorithm is used to update the network parameters θ, which is expressed as follows:

[0272]

[0273] Among them, η θ is the learning rate of the stochastic gradient descent algorithm, is the gradient of the loss function L(θ) with respect to the network parameters θ.

[0274] In step S4, the use of adaptive optics technology combined with dynamic wavefront correction technology to optimize the transmission path of the response signal specifically includes:

[0275] The fused feature signal Features(t) output from the edge computing node needs to be converted into an optical signal for transmission. Let the optical signal be E(t), which is modulated with the electrical signal Features(t) through an electro-optical modulator:

[0276] E(t)=E0·cos(ωt+φ(t))

[0277] Where E0 is the amplitude of the optical signal, ω is the angular frequency of the light, and φ(t) is the phase modulated by the electrical signal Features(t);

[0278] Furthermore, the multimodal characteristic signals are integrated into a single optical signal for transmission. This step requires the use of an electro-optical modulator to map different types of characteristic signals (voice, video, keystrokes) to different parameters of the optical signal, such as phase, frequency, amplitude, etc., as shown below:

[0279] Speech signal characteristics S v (t) (such as MFCC features);

[0280] Video signal characteristics S img (t) (such as features extracted by CNN);

[0281] Key signal characteristics S k (t) (e.g. key press time and pressure).

[0282] The wavefront distortion of the optical signal is monitored in real time by the wavefront sensor. Assume that the output of the wavefront sensor is W(t), which represents the distortion information of the optical wavefront:

[0283] W(t) = Δx(t) + Δy(t) + Δz(t)

[0284] Where Δx(t), Δy(t), and Δz(t) are the aberrations of the wavefront in the three spatial coordinates respectively.

[0285] An adaptive optical controller is used to calculate the correction signal C(t) to correct the wavefront aberration. The correction calculation formula is as follows:

[0286] C(t) = -K·W(t) (13)

[0287] Where K is the correction coefficient matrix and W(t) is the aberration information measured by the wavefront sensor.

[0288] Then, the wavefront of the optical signal is corrected by the correction mirror, and at the same time, the correction signal C(t) is applied to the drive signal of the correction mirror to obtain the signal E corrected (t) in the ideal state, which is expressed as follows:

[0289] E corrected (t) = E(t) + C(t) (14)

[0290] Select the best optical fiber transmission path according to the real-time network condition; let the transmission characteristics of the optical fiber path be P fiber (t), including transmission loss and delay information:

[0291] P fiber (t) = {L fiber (t), D fiber (t)} (15)

[0292] Where L fiber (t) is the transmission loss and D fiber (t) is the transmission delay.

[0293] Design a path optimization algorithm to calculate the shortest transmission path in real time, minimize the transmission loss and delay, and then transmit the signal E corrected (t) in the ideal state through the optimized optical fiber path; among them, the designed optimization objective function is:

[0294]

[0295] Where α P and β P are weight coefficients used to balance the effects of transmission loss and delay.

[0296] The receiving end of the central server receives the transmitted optical signal E received (t) through an optical detector, which is expressed as follows:

[0297] F received (t) = Dopt (E received (t)) (18)

[0298] Demodulate the received electrical signal to recover the original transmitted signal Features received (t):

[0299] Features received (t) = Demod(F received (t)) (19)

[0300] where Demod is the demodulation function to recover the original multimodal feature signal.

[0301] To ensure signal consistency and high quality, the demodulated signal must be synchronized and key features extracted for subsequent processing and analysis.

[0302] Then, preprocess and extract features from the signal Features received (t):

[0303] Synchronize the demodulated signal with the system clock to ensure that all signals have a unified time reference. Let the synchronization function be H sync , and the synchronized signal be F synced (t):

[0304] F synced (t) = H sync (Features received (t)) (20)

[0305] where H sync is the synchronization function used to eliminate the time differences between different sensor data.

[0306] Extract key features from the synchronized signal for subsequent processing and analysis. Let the feature extraction function be K extract , and the extracted feature signal be F features (t):

[0307] F features (t) = K extract (F synced (t))(21)

[0308] where K extract is the feature extraction function.

[0309] Among them, encrypt and decrypt the optical signal generated by the optical adaptive optical technology for transmission, specifically including:

[0310] Encrypt the optical signal using a phase encoder. Let the input optical signal be E(t), and the phase encoder performs phase modulation on it. The encrypted optical signal is E enc (t):

[0311]

[0312] where φ enc (t) is a randomly generated phase mask, and j is the imaginary unit.

[0313] The phase mask φ enc (t) is generated by a physical random number generator to ensure the randomness and security of encryption:

[0314] φ enc (t) = 2π·rand(t)

[0315] where rand(t) is a uniformly distributed random number.

[0316] Transmit the encrypted optical signal E enc (t) through an optical fiber to the receiving end:

[0317] E transmitted (t) = E enc (t)

[0318] The decryption process is as follows:

[0319] At the receiving end, use the same phase mask φ enc (t) to decrypt the received encrypted optical signal E received (t) and recover the original optical signal E dec (t):

[0320]

[0321] Ensure that the same phase mask φ enc (t) is used at the sending end and the receiving end for synchronization. Adopt time-division multiplexing technology to achieve synchronization of the phase mask through the synchronization signal S sync (t):

[0322] S sync (t) = sync(φ enc (t))

[0323] where sync is a synchronization function to ensure that the phase masks at the sending end and the receiving end are consistent.

[0324] Process the decrypted optical signal E dec (t) to recover the original electrical signal F recovered (t):

[0325] Frecovered I(t)=D opt (E dec (t))

[0326] where D opt is the conversion function of the photodetector, which converts the optical signal into an electrical signal.

[0327] Specifically, a phase encoder and a physical random number generator are used to implement an optical encryption system to ensure the high randomness and security of the phase mask. A phase decoder and a synchronization system are used to implement an optical decryption system to ensure that the receiver can accurately recover the original signal. By optimizing the generation algorithm of the phase mask and the synchronization technology, the encryption and decryption efficiency of the system is improved, and the security and integrity of real-time transmitted data are ensured.

[0328] In the step S5, it specifically includes:

[0329] S501. The central server receives the transmitted optical signal E received (t) through a photodetector and converts it into an electrical signal F received (t):

[0330] F received (t)=D opt (E received (t))

[0331] where D opt is the conversion function of the photodetector; the received electrical signal F received (t) is decoded to recover the original multimodal feature signal F decoded (t):

[0332] F decoded (t)=Decode(F received (t))

[0333] where Decode is the decoding function used to recover the original signal;

[0334] S502. The decoded signal is denoised to remove the noise interference during transmission. Let the denoising function be G denoise , and the denoised signal is F denoised (t):

[0335] F denoised (t)=G denoise (F decoded (t))

[0336] where G denoise is the denoising function;

[0337] Synchronize the denoised signal with the system clock to ensure that all signals have a unified time reference. Let the synchronization function be H sync , and the synchronized signal be F synced (t):

[0338] F synced (t) = H sync (F denoised (t))

[0339] where H sync is the synchronization function;

[0340] Extract key features from the synchronized signal for subsequent processing and analysis. Let the feature extraction function be K extract , and the extracted feature signal be F features (t):

[0341] F features (t) = K extract (F synced (t))

[0342] where K extract is the feature extraction function;

[0343] S503. Sort the feature signals according to the importance and priority of the signals. Let the sorting function be P sort , and the sorted feature signal be F sorted (t):

[0344] F sorted (t) = P sort (F features (t))

[0345] where P sort is the priority sorting function;

[0346] Process the sorted signal to generate the final response result. Let the processing function be Q process , and the processed response result be R(t):

[0347] R(t) = Q process (F sorted (t))

[0348] where Q process is the signal processing function;

[0349] S504. Generate the feedback response of the system according to the processing result R(t). Let the response generation function be M response , and the generated feedback response be F response (t):

[0350] Fresponse M(t) = M response (R(t))

[0351] where M response is a response generation function;

[0352] Send the generated feedback response to each edge node and user through the network to complete the feedback process of the system. Let the sending function be T send , and the feedback signal sent is F sent (t):

[0353] F sent (t) = T send (F response (t))

[0354] where T send is the sending function.

[0355] In one or more embodiments, as Figure 2 shown, an adaptive quick-answer optimization system based on reinforcement learning is disclosed, including the following modules:

[0356] A signal collection module 101, configured to collect multi-modal quick-answer signals of students by using sensors, preprocess the quick-answer information to obtain a target signal, and the sensor extracts and fuses features of the target to obtain a unified data set; wherein, the multi-modal quick-answer signals include voice signals, video signals, and key signals;

[0357] A signal processing module 102, configured to collect the target signal by an edge computing node to form multi-modal feature data, the edge computing node performs feature fusion on the multi-modal feature data to obtain a fused multi-modal feature vector, and constructs a reinforcement learning model to sort the priorities of the quick-answer signals; introduce a reinforcement learning algorithm to dynamically adjust the processing order and weights of the quick-answer signals;

[0358] A signal transmission module 103, configured to optimize the transmission path of the quick-answer signal by using adaptive optical technology combined with dynamic wavefront correction technology, and perform encrypted and decrypted transmission on the optical signal generated by the optical adaptive optical technology;

[0359] A signal sending module 104, configured to enable the central processing unit to receive the decrypted optical signal, process and analyze it, and send the analysis result to each edge computing node and user;

[0360] wherein, the reinforcement learning algorithm specifically includes:

[0361] S301. Construct a reinforcement learning environment:

[0362] Construct the state space S(t) of the reinforcement learning environment, wherein the state space S(t) includes the reduced-dimensional feature vector Fpca (t) and the delay information D at the current time step latency (t), which is expressed as follows:

[0363] S(t) = [F pca (t), D latency (t)]

[0364] where D latency (t) represents the delay information at the current time step;

[0365] Define the action space A(t), including the priority sorting and weight adjustment of the rush-answer signals, which is expressed as follows:

[0366] A(t) = [P priority (t), W weight (t)]

[0367] where P priority (t) represents the priority sorting at the current time step, and W weight (t) represents the weight adjustment at the current time step;

[0368] S302. Design the reward function:

[0369] Let the response time of the i-th student be T i (t), and the device performance be P i (t). Define the response time fairness index FairTime(A(t)) as:

[0370]

[0371] where represents the average response time of all students, and N represents the total number of students;

[0372] Define the device difference fairness index FairEquip(A(t)) as:

[0373]

[0374] where represents the average device performance of all students;

[0375] Incorporate the response time and device difference, and define the overall fairness index Fairness(A(t)) as:

[0376] Fairness(A(t)) = γ1·FairTime(A(t)) + γ2·FairEquip(A(t))

[0377] where γ1 and γ2 represent the weight parameters for balancing the response time and device difference;

[0378] The real-time index RealTime(A(t)) is used to measure the response speed of the system, and the real-time index is defined as:

[0379] RealTime(A(t)) = -max i∈{1,2,…,N} (T i (t) - T start (t))

[0380] Wherein, T start (t) represents the start time of the current time step;

[0381] Adding fairness and real-time, the comprehensive reward function R(t) is designed as:

[0382] R(t) = α R ·Fairness(A(t)) + β R ·RealTime(A(t))

[0383] Wherein, α R and β R represent weight parameters, which are used to balance the influence of fairness and real-time.

[0384] It should be noted that the specific working process of the adaptive rush-answer optimization system based on reinforcement learning provided in the embodiments of the present invention is the same as the process of the adaptive rush-answer optimization method based on reinforcement learning described in the above embodiments, and will not be elaborated here.

[0385] Compared with the prior art, the adaptive rush-answer optimization system based on reinforcement learning provided in the embodiments of the present invention uses sensors to collect multi-modal rush-answer signals of students, preprocesses the rush-answer information to obtain target signals, and the sensors extract and fuse features of the target to obtain a unified data set; wherein, the multi-modal rush-answer signals include voice signals, video signals and key signals. The edge computing node collects the target signals to form multi-modal feature data, the edge computing node performs feature fusion on the multi-modal feature data to obtain a fused multi-modal feature vector, and constructs a reinforcement learning model to rank the priority of the rush-answer signals. The reinforcement learning algorithm is introduced to dynamically adjust the processing order and weight of the rush-answer signals. The adaptive optical technology is combined with the dynamic wavefront correction technology to optimize the transmission path of the rush-answer signals, and the optical signals generated by the adaptive optical technology are encrypted and decrypted for transmission. The central processing unit receives the decrypted optical signals, processes and analyzes them, and sends the analysis results to each edge computing node and user.

[0386] An embodiment of the present invention further provides an adaptive quick-answer optimization device based on reinforcement learning, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps in the embodiment of the adaptive quick-answer optimization method based on reinforcement learning as described above are implemented, such as Figure 1 the steps S1 to S5 described in

[0387] Exemplarily, the computer program may be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the adaptive quick-answer optimization device based on reinforcement learning.

[0388] The adaptive quick-answer optimization device based on reinforcement learning may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The adaptive quick-answer optimization device based on reinforcement learning may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the adaptive quick-answer optimization device based on reinforcement learning may further include input / output devices, network access devices, a bus, etc.

[0389] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the adaptive quick-answer optimization device based on reinforcement learning, and connects various parts of the entire adaptive quick-answer optimization device through various interfaces and lines.

[0390] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory, the processor can implement various functions of the adaptive rush-answer optimization device based on reinforcement learning. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the operation of the air-conditioning controller, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0391] Among them, if the module integrated in the adaptive rush-answer optimization device based on reinforcement learning is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0392] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, it can be completed by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, optical disc, read-only memory (ROM), or random access memory (RAM), etc.

[0393] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. An adaptive response optimization method based on reinforcement learning, characterized in that: The method comprises the following steps: S1. Using sensors to collect students' multimodal response signals, and preprocessing the response information to obtain target signals. The sensors extract and fuse features of the targets to obtain a unified data set; wherein the multimodal response signals include voice signals, video signals, and key signals; S2. The edge computing node collects the target signal to form multimodal feature data. The edge computing node performs feature fusion on the multimodal feature data to obtain the fused multimodal feature vector, and builds a reinforcement learning model to prioritize the response signals. S3, introduce reinforcement learning algorithm to dynamically adjust the processing order and weight of the response signal; S4. Adopt adaptive optics technology combined with dynamic wavefront correction technology to optimize the transmission path of the answer signal, and encrypt and decrypt the optical signal generated by the adaptive optics technology for transmission; S5, the central processor receives the decrypted optical signal and processes and analyzes it, and sends the analysis results to each edge computing node and user; Wherein, the S3 specifically includes: S301. Build a reinforcement learning environment: Construct the state space S(t) of the reinforcement learning environment, where the state space S(t) includes the feature vector F after dimensionality reduction pca (t) and the delay information D of the current time step latency (t), expressed as follows: S(t)=[F pca (t),D latency (t)] Define the action space A(t), including the priority sorting and weight adjustment of the response signals, as follows: A(t)=[P priority (t),W weight (t)] Among them, P priority (t) represents the priority order of the current time step, W weight (t) represents the weight adjustment of the current time step; S302. Design reward function: Let the response time of the i-th student be T i (t), define the response time fairness indicator FairTime(A(t)) as: in, represents the average response time of all students, and N represents the total number of students; The device difference fairness index FairEquip(A(t)) is defined as: in, represents the average device performance of all students, and the device performance is P i (t); Taking into account the response time and device differences, the overall fairness index Fairness(A(t)) is defined as: Fairness(A(t))=γ1·FairTime(A(t))+γ2·FairEquip(A(t)) Among them, γ1 and γ2 represent weight parameters to balance the response time and device differences; The real-time indicator RealTime (A(t)) is used to measure the response speed of the system. The real-time indicator is defined as: RealTime(A(t))=-max i∈{1,2,…,N} (T i (t)-T start (t)) Among them, T start (t) represents the start time of the current time step; Taking fairness and real-time into consideration, the comprehensive reward function R(t) is designed as follows: R(t)=α R ·Fairness(A(t))+β R ·RealTime(A(t)) Among them, α R and β R Represents the weight parameter, which is used to balance the impact of fairness and real-time performance; The reinforcement learning environment of the reinforcement learning model of the edge computing node is constructed as follows: Design the reward function r(t) to evaluate the effect of each action, which is expressed as follows: r(t)=α R ·Fairness(a(t))+β R ·RealTime(a(t)) Among them, α R and β R is the weight parameter, Fairness(a(t)) is the fairness index of the current action, and RealTime(a(t)) is the real-time index; The deep Q network algorithm is used for training, and the Q value function Q(s,a) is defined as follows: Q(s(t),a(t))=Q(s(t),a(t))+η Q [r(t)+γ Q max a′ Q(s(t+1),a′)-Q(s(t),a(t))] Among them, η Q is the learning rate, γ Q is the discount factor; According to the action a(t) output by the reinforcement learning algorithm, the signal at the current moment is prioritized. Assuming the priority ranking function is Π, the calculation is as follows: F sorted (t)=Π(F pca (t),a(t)) Among them, F pca (t) represents the feature vector after dimensionality reduction, F sorted (t) represents the output priority sequence; In the S3, it also includes: S303, using reinforcement learning algorithm to dynamically adjust the signal processing order and weight: Define the Q-value function Q(S,A), which represents the expected reward for taking action A in state S: Among them, γ R represents the discount factor, A′ represents the action at the next time step; The deep Q network algorithm is used to train and update the Q value function; the update formula of the Q value is: Q(S(t),A(t))=Q(S(t),A(t)) Among them, η Q represents the learning rate; Use the DNN network to fit the Q value function; the input of the neural network is the state S(t), and the output is the Q value of each action, which is expressed as follows: Among them, θ represents the parameters of the neural network, which is optimized by the back propagation algorithm. represents the Q value function obtained by fitting the DNN network, that is, the fitted Q value; S304, adding device performance information to the state space and introducing a time window mechanism to the action space; wherein the state space S(t) is defined as: S(t)=[F pca (t),D latency (t),P device (t)] Among them, P device (t) indicates equipment performance; The action space A(t) is defined as: A(t)=[P priority (t),W weight (t),T window (t)] Among them, T window (t) represents the time window parameter; S305. At each time step, an action is selected using the ∈-greedy strategy; wherein an action is randomly selected with a probability of ∈, and the action with the largest current Q value is selected with a probability of 1-∈, which is expressed as follows: According to the selected action A(t), the signal of the current time step is prioritized and weighted, as shown below: F weighted (t)=F sorted (t)′⊙W weight (t) Among them, Π represents the priority sorting function, ⊙ represents the element-by-element point-by-point multiplication operation; S306. Store (S(t), A(t), R(t), S(t+1)) of each interaction into the experience replay pool In each training, a small batch (S i ,A i ,R i ,S i+1 ) for training, which is expressed as follows: in, Indicates that the experience replay pool Samples are drawn from i ,A i ,R i ,S i+1 ) performs expected calculation, L(θ) is the loss function, and the network parameter θ is optimized by minimizing the loss function; The stochastic gradient descent algorithm is used to update the network parameters θ, which is expressed as follows: Among them, η θ is the learning rate of the stochastic gradient descent algorithm, is the gradient of the loss function L(θ) with respect to the network parameters θ.

2. The adaptive response optimization method based on reinforcement learning according to claim 1, characterized in that: The preprocessing includes signal synchronization processing, noise elimination and signal standardization processing.

3. The adaptive response optimization method based on reinforcement learning according to claim 2 is characterized in that: The signal synchronization processing is represented as follows: Define the timestamp T of each signal i , align all signals to a unified time base T by interpolation sync , which is expressed as follows Among them, T sync Indicates the unified time base after synchronization, Represents the speech signal S v timestamp, Represents the video signal S img timestamp, Indicates key signal S k timestamp; Time synchronize each signal: in, represents the time derivative of each signal, S v (t) represents the speech signal, S img (t) represents the video signal, S k (t) represents a key signal; The noise cancellation is expressed as follows: Let S(t) be the original signal and N(t) be the noise signal. adapt (t) Filter the noise to obtain the purified signal S clean (t), is calculated as follows: S clean (t)=S(t)-H adapt (t)·N(t) The filter update formula is: H adapt (t+1)=H adapt (t)+μ H ·e H (t)·N(t) Among them, μ H represents the filter learning rate, e H (t) represents the current error, e H (t) = S(t) - S clean (t) represents the error; The signal normalization is expressed as follows: Among them, μ S and σ S are the mean and standard deviation of the signal, S norm (t) represents the normalized signal.

4. The adaptive response optimization method based on reinforcement learning according to claim 1, characterized in that: In S4, the optimization of the transmission path of the response signal by using adaptive optical technology combined with dynamic wavefront correction technology specifically includes: The photoelectric modulator is used to convert the multimodal feature data Features(t) output from the edge computing node into an optical signal for transmission. The wavefront distortion of the optical signal is monitored in real time by the wavefront sensor. The adaptive optical controller is used to calculate the correction signal C(t) to correct the wavefront distortion. The optical signal is then corrected by the correction mirror. At the same time, the correction signal C(t) is applied to the driving signal of the correction mirror to obtain the ideal signal E. corrected (t); Select the best optical fiber transmission path according to the real-time network status, design a path optimization algorithm to calculate the shortest transmission path in real time, and then convert the ideal signal E corrected (t) transmission through optimized optical fiber paths; The central server receiving end receives the transmitted optical signal E through the optical detector received (t), and demodulate the received electrical signal to restore the original transmission signal Features received (t), and the signal Features received (t) Perform preprocessing and feature extraction.

5. The adaptive response optimization method based on reinforcement learning according to claim 1, characterized in that: In S4, encrypting and decrypting the optical signal generated by the optical adaptive optics technology for transmission specifically includes: The optical signal is encrypted using a phase encoder, while a phase mask is generated by a physical random number generator; The encrypted optical signal is transmitted to the receiving end through the optical fiber. The receiving end decrypts the received encrypted optical signal with the same phase mask to restore the original optical signal. At the same time, the time division multiplexing technology is used to synchronize the phase mask by synchronizing the optical signal. The decrypted optical signal is processed to restore the original electrical signal.

6. The adaptive response optimization method based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S501, the central server receives the transmitted optical signal through the photoelectric detector, converts the optical signal into an electrical signal, and then decodes the received electrical signal to restore the original multimodal characteristic signal; S502, performing denoising processing on the decoded multimodal feature signal and synchronization processing on the system clock, and extracting features from the processed multimodal feature signal to obtain a feature signal; S503, sorting the characteristic signals according to the importance and priority of the signals, processing the sorted characteristic signals, and generating a final response result; S504: Generate a system feedback response according to the response result, and send the generated feedback response to each edge node and user through the network to complete the system feedback process.

7. An adaptive response optimization system based on reinforcement learning, characterized in that: Includes the following modules: A signal collection module is used to collect students' multimodal response signals using sensors, and pre-process the response information to obtain target signals. The sensors extract and fuse features of the targets to obtain a unified data set; wherein the multimodal response signals include voice signals, video signals, and key signals; The signal processing module is used for edge computing nodes to collect target signals and form multimodal feature data. The edge computing node performs feature fusion on the multimodal feature data to obtain the fused multimodal feature vector, and builds a reinforcement learning model to prioritize the response signals. The reinforcement learning algorithm is introduced to dynamically adjust the processing order and weight of the response signals. A signal transmission module is used to optimize the transmission path of the answer signal by using adaptive optics technology combined with dynamic wavefront correction technology, and to encrypt and decrypt the optical signal generated by the optical adaptive optics technology for transmission; A signal sending module is used to enable the central processor to receive the decrypted optical signal, process and analyze it, and send the analysis results to each edge computing node and user; The reinforcement learning algorithm specifically includes: S301. Build a reinforcement learning environment: Construct the state space S(t) of the reinforcement learning environment, where the state space S(t) includes the feature vector F after dimensionality reduction pca (t) and the delay information D of the current time step latency (t), expressed as follows: S(t)=[F pca (t),D latency (t)] Define the action space A(t), including the priority sorting and weight adjustment of the response signals, as follows: A(t)=[P priority (t),W weight (t)] Among them, P priority (t) represents the priority order of the current time step, W weight (t) represents the weight adjustment of the current time step; S302. Design reward function: Let the response time of the i-th student be T i (t), define the response time fairness indicator FairTime(A(t)) as: in, represents the average response time of all students, and N represents the total number of students; The device difference fairness index FairEquip(A(t)) is defined as: in, represents the average device performance of all students, and the device performance is P i (t); Taking into account the response time and device differences, the overall fairness index Fairness(A(t)) is defined as: Fairness(A(t))=γ1·FairTime(A(t))+γ2·FairEquip(A(t)) Among them, γ1 and γ2 represent weight parameters to balance the response time and device differences; The real-time indicator RealTime (A(t)) is used to measure the response speed of the system. The real-time indicator is defined as: RealTime(A(t))=-max i∈{1,2,…,N} (T i (t)-T start (t)) Among them, T start (t) represents the start time of the current time step; Taking fairness and real-time into consideration, the comprehensive reward function R(t) is designed as follows: R(t)=α R ·Fairness(A(t))+β R ·RealTime(A(t)) Among them, α R and β R Represents the weight parameter, which is used to balance the impact of fairness and real-time performance; The reinforcement learning environment of the reinforcement learning model of the edge computing node is constructed as follows: Design the reward function r(t) to evaluate the effect of each action, which is expressed as follows: r(t)=α R ·Fairness(a(t))+β R ·RealTime(a(t)) Among them, α R and β R is the weight parameter, Fairness(a(t)) is the fairness index of the current action, and RealTime(a(t)) is the real-time index; The deep Q network algorithm is used for training, and the Q value function Q(s,a) is defined as follows: Q(s(t),a(t))=Q(s(t),a(t))+η Q [r(t)+γ Q max a′ Q(s(t+1),a′)-Q(s(t),a(t))] Among them, η Q is the learning rate, γ Q is the discount factor; According to the action a(t) output by the reinforcement learning algorithm, the signal at the current moment is prioritized. Assuming the priority ranking function is Π, the calculation is as follows: Among them, F pca (t) represents the feature vector after dimensionality reduction, F sorted (t) represents the output priority sequence; The signal processing module further includes: S303, using reinforcement learning algorithm to dynamically adjust the signal processing order and weight: Define the Q-value function Q(S,A), which represents the expected reward for taking action A in state S: Among them, γ R represents the discount factor, A′ represents the action at the next time step; The deep Q network algorithm is used to train and update the Q value function; the update formula of the Q value is: Q(S(t),A(t))=Q(S(t),A(t)) Among them, η Q represents the learning rate; Use the DNN network to fit the Q value function; the input of the neural network is the state S(t), and the output is the Q value of each action, which is expressed as follows: Among them, θ represents the parameters of the neural network, which is optimized by the back propagation algorithm. represents the Q value function obtained by fitting the DNN network, that is, the fitted Q value; S304, adding device performance information to the state space and introducing a time window mechanism to the action space; wherein the state space S(t) is defined as: S(t)=[F pca (t),D latency (t),P device (t)] Among them, P device (t) indicates equipment performance; The action space A(t) is defined as: A(t)=[P priority (t),W weight (t),T window (t)] Among them, T window (t) represents the time window parameter; S305. At each time step, an action is selected using the ∈-greedy strategy; wherein an action is randomly selected with a probability of ∈, and the action with the largest current Q value is selected with a probability of 1-∈, which is expressed as follows: According to the selected action A(t), the signal of the current time step is prioritized and weighted, as shown below: F weighted (t)=F sorted (t)′⊙W weight (t) Among them, Π represents the priority sorting function, ⊙ represents the element-by-element point-by-point multiplication operation; S306. Store (S(t), A(t), R(t), S(t+1)) of each interaction into the experience replay pool In each training, a small batch (S i ,A i ,R i ,S i+1 ) for training, which is expressed as follows: in, Indicates that the experience replay pool Samples are drawn from i ,A i ,R i ,S i+1 ) performs expected calculation, L(θ) is the loss function, and the network parameter θ is optimized by minimizing the loss function; The stochastic gradient descent algorithm is used to update the network parameters θ, which is expressed as follows: Among them, η θ is the learning rate of the stochastic gradient descent algorithm, is the gradient of the loss function L(θ) with respect to the network parameters θ.

Citation Information

Patent Citations

  • Quick answering interactive method and system

    CN104735609A

  • Method, device and system for realizing electronized classroom preemptive answering

    CN105448008A