Intelligent supervision method, device and equipment for third subject examination of driver, and storage medium
By using image and audio acquisition equipment in the driver's license subject three test, combined with convolutional neural networks and long short-term memory networks, the action and voice features of the safety officer are extracted, realizing automatic and intelligent supervision of the driver's license subject three test. This solves the problem of low efficiency in existing technologies and improves the supervision effect and test security.
Patent Information
- Application Number
- CN202511006150.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing methods for supervising the driver's license test (subject 3) rely on manual checks, which are inefficient and make it difficult to effectively monitor cheating, leading to an increased risk of traffic accidents.
During the driver's license test (subject 3), image data of the safety officer is collected by in-vehicle image acquisition equipment, and audio data is collected by in-vehicle audio acquisition equipment. Convolutional neural networks and long short-term memory networks are used to extract action and sound features to determine cheating behavior.
It has enabled automated and intelligent supervision of the driver's license test (subject 3), improving supervision efficiency and effectiveness, reducing cheating, and ensuring the fairness and security of the test.
Smart Images

Figure CN120877256A_ABST
Abstract
Description
Technical Field
[0001] This application relates to intelligent monitoring and testing technology, and in particular to an intelligent monitoring method, device, equipment, and storage medium for driver's license subject three examination. Background Technology
[0002] As motor vehicle driving skills tests gain increasing attention and popularity, tens of thousands of students obtain their driver's licenses every year. However, due to the large number of people needing to take the driving test, some cheating inevitably occurs.
[0003] Cheating can easily lead to traffic accidents while driving, which not only damage the vehicle itself but also pose a serious threat to people's economic well-being and safety. Therefore, it is necessary to supervise the driver's license test (Part 3).
[0004] The current regulatory method involves post-event manual inspection and verification. This method relies on manual labor, is inefficient, and makes it difficult to achieve good regulatory results. Summary of the Invention
[0005] To address one of the aforementioned technical deficiencies, this application provides a method, device, equipment, and storage medium for intelligent monitoring of driver's license subject three examinations.
[0006] The first aspect of this application provides an intelligent monitoring method for driver's license subject three examination, the method comprising:
[0007] During the driver's license test (subject 3), the in-vehicle image acquisition device collects the safety officer's image data; the in-vehicle audio acquisition device collects the in-vehicle audio data.
[0008] Extract the safety officer's motion characteristics from the image data;
[0009] Extract the sound features inside the vehicle based on the audio data;
[0010] Cheating behavior is determined based on movement and voice characteristics.
[0011] Optionally, the image acquisition device is located in a key position inside the vehicle, and the resolution of the image acquisition device is not less than 1080 pixels, and the frame rate is not less than 25 frames / second;
[0012] The audio acquisition device is located on the roof of the vehicle, and its frequency response range is 20 Hz to 20 kHz.
[0013] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture;
[0014] Based on the image data, extract the safety officer's motion characteristics, including:
[0015] Hand features are extracted from image data using a convolutional neural network; the convolutional neural network includes convolutional layers and pooling layers; the convolutional layers are processed using the formula... The processing is performed where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(·) be the deviation of the spatial dimension j of the input channel l, and let σ(·) be the Sigmoid activation function; the pooling layer is defined by the formula... Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R... j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l;
[0016] Based on hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture;
[0017] The distance between the safety officer's hand and the target component is determined by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the target component's position. h Let y be the vertical coordinate of the safety officer's hand. c The vertical coordinate represents the position of the target component; the velocity of the hand moving towards the target component is expressed by the formula... d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
[0018] Optionally, before determining the distance between the safety officer's hand and the target component, the speed at which the hand moves towards the target component, and the hand posture based on hand characteristics, the method further includes:
[0019] Identify target components inside the vehicle based on image data;
[0020] Determine multiple bounding boxes of the target component and the class probability corresponding to each bounding box;
[0021] Based on the category probability, using the formula Filter multiple bounding boxes to obtain the predicted bounding box of the target part; where B1 is the area of one bounding box during filtering and B2 is the area of another bounding box during filtering.
[0022] Kalman filtering is used to optimize the position of the target component; the position of the target component is determined based on the predicted bounding box of the target component; the measurement update equation during Kalman filtering optimization is: t1 is the number of times. For the position observation of the target component in the t1th optimization, To optimize the position observation matrix of the target component for the t1th iteration, Let t1 be the position state variable of the target component in the optimization step. The measurement noise of the target component is optimized for the t1th iteration.
[0023] Optionally, based on the audio data, the sound features inside the vehicle are extracted, including:
[0024] Through formula Noise reduction is applied to the audio data; where t2 is the time index of the audio data. y(t2-a1) is the output of the noise reduction process at time t2, a1 is the tap index of the filter used for noise reduction, A is the number of taps of the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 of the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2.
[0025] The denoised audio data is divided into frames; the duration of each frame is not less than 20 milliseconds and not more than 30 milliseconds; the frame shift is not less than 10 milliseconds and not more than 15 milliseconds.
[0026] Extract sound features based on frame segmentation;
[0027] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate;
[0028] Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, L be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g (a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame;
[0029] Short-time energy x e(g) represents the speech amplitude of the g-th frame of audio data;
[0030] Zero crossing rate sgn[·] is a symbolic function.
[0031] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture;
[0032] Based on movement and voice characteristics, cheating behavior is determined, including:
[0033] Based on sound characteristics, sound anomalies are classified to obtain sound anomaly results;
[0034] If the distance between the hand and the target component is less than the distance threshold, the speed at which the hand moves toward the target component is greater than the speed threshold, the hand posture matches the abnormal conditions, and the abnormal sound result is abnormal, then cheating is determined.
[0035] Optionally, sound anomalies can be classified based on sound characteristics, including:
[0036] Voice anomaly classification is performed using long short-term memory networks to identify sound features.
[0037] The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula:
[0038] f t =σ(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t ] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f For the offset of the forget gate;
[0039] i t =σ(W i [h t-1 ,x t ]+b i ), i t W is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b i This is the bias of the input gate;
[0040] For candidate memory cell states at time t, W C [h t-1 ,x t [] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C Bias for candidate memory cells;
[0041] C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1;
[0042] o t =σ(W o [h t-1 ,x t ]+b o ), o t W is the output gate activation value at time t. o [h t-1 ,x t ] represents the weight of the output gate, b O For the output gate bias;
[0043] h t =o t tanh(C t ), h t Let t be the hidden state at time t.
[0044] A second aspect of this application provides an intelligent monitoring device for driver's license subject three examination, the device comprising:
[0045] The data acquisition module is used to collect image data of the safety officer through the in-vehicle image acquisition device and audio data of the in-vehicle audio acquisition device during the driver's license test (subject 3).
[0046] The action feature extraction module is used to extract the action features of the safety officer based on the image data collected by the data acquisition module;
[0047] The sound feature extraction module is used to extract the sound features inside the vehicle based on the audio data collected by the data acquisition module.
[0048] The cheating behavior determination module is used to determine cheating behavior based on the action features obtained by the action feature extraction module and the sound features obtained by the sound feature extraction module.
[0049] A third aspect of this application provides an electronic device, comprising:
[0050] Memory;
[0051] Processor; and
[0052] Computer programs;
[0053] The computer program is stored in the memory and configured to be executed by the processor to implement the method described in the first aspect above.
[0054] In a fourth aspect, this application provides a computer-readable storage medium having a computer program stored thereon; the computer program is executed by a processor to implement the method described in the first aspect above.
[0055] This application provides a method, device, equipment, and storage medium for intelligent supervision of driver's license subject three examinations. The method includes: during the driver's license subject three examination, acquiring image data of the safety officer through an in-vehicle image acquisition device; acquiring in-vehicle audio data through an in-vehicle audio acquisition device; extracting the safety officer's motion characteristics based on the image data; extracting in-vehicle sound characteristics based on the audio data; and determining cheating behavior based on the motion and sound characteristics. This method obtains the safety officer's motion characteristics from the in-vehicle image data and in-vehicle sound characteristics from the in-vehicle audio data, and then determines cheating behavior based on the motion and sound characteristics. This eliminates the reliance on manual inspection and review for driver's license subject three examination supervision, achieving automated and intelligent supervision and ensuring both efficiency and effectiveness. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0057] Figure 1 A flowchart illustrating an intelligent monitoring method for driver's license subject three examination provided in this application embodiment;
[0058] Figure 2 A schematic diagram of the structure of an intelligent monitoring device for driver's license subject three examination provided in this application embodiment;
[0059] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0060] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0061] In developing this application, the inventors discovered that as motor vehicle driving skills tests gain increasing attention and popularity, tens of thousands of students obtain driver's licenses every year. However, due to the large number of people needing to take the driving test, cheating inevitably occurs. Cheating can easily lead to traffic accidents while driving, which not only affect the vehicle itself but also pose a serious threat to people's economic and personal safety. Therefore, it is necessary to supervise the driver's license test (subject three). The existing supervision method is post-test manual review. This method relies on manual labor, is inefficient, and struggles to achieve effective supervision.
[0062] To address the aforementioned issues, this application provides an intelligent monitoring method, device, equipment, and storage medium for driver's license subject three examinations. The method includes: during the driver's license subject three examination, acquiring image data of the safety officer using an in-vehicle image acquisition device; acquiring in-vehicle audio data using an in-vehicle audio acquisition device; extracting the safety officer's motion characteristics based on the image data; extracting in-vehicle sound characteristics based on the audio data; and determining cheating behavior based on the motion and sound characteristics. This method obtains the safety officer's motion characteristics from the in-vehicle image data and in-vehicle sound characteristics from the in-vehicle audio data, and then determines cheating behavior based on the motion and sound characteristics. This eliminates the reliance on manual inspection and verification for driver's license subject three examination monitoring, achieving automated and intelligent monitoring and ensuring both monitoring efficiency and effectiveness.
[0063] This embodiment provides an intelligent monitoring method for driver's license subject three examination. The method pre-installs image acquisition equipment and audio acquisition equipment in the vehicle. During the driver's license subject three examination, the image acquisition equipment in the vehicle will collect image data of the safety officer in real time, and the audio acquisition equipment in the vehicle will collect audio data in real time.
[0064] The image acquisition equipment is located in key positions inside the vehicle, and its resolution is no less than 1080p (pixels), with a frame rate of no less than 25 frames per second. In practice, the image acquisition equipment can consist of multiple high-definition cameras distributed in key locations inside the test vehicle. A high-definition wide-angle camera located at the top in front of the passenger seat, with a resolution of at least 1080p and a frame rate of no less than 25 frames per second, possesses good low-light shooting capabilities and dynamic range, and is used to capture the safety driver's upper body, especially the details of hand and arm movements. A camera above the driver's seat, near the roof, is used to supplement the capture of the safety driver's side and lower body movements, ensuring complete posture information; its resolution and frame rate match the former. A camera located in the middle of the rear seats captures the overall interior environment, including components such as the steering wheel, gear shift, and handbrake, as well as the relative position of the safety driver to these components; the shooting range covers the entire driving area, and it also has high resolution and frame rate. All these cameras are connected to the data transmission equipment via an in-vehicle Ethernet network to stably and reliably transmit the acquired image data.
[0065] The audio acquisition device is located on the vehicle roof, and its frequency response range is 20 Hz to 20 kHz. In practice, the audio acquisition device uses an array of high-sensitivity microphones, installed in a square arrangement in the center of the vehicle's roof to achieve omnidirectional sound acquisition within the vehicle. The microphones' frequency response range is [20Hz-20kHz], effectively capturing sound signals of various frequencies, including unusual sounds that might be used as cheating alerts. The audio acquisition device converts the acquired analog sound signals into digital signals using an audio amplifier and an analog-to-digital converter (ADC), performs preliminary noise reduction and filtering, and then transmits the audio data to the data processing center via a data transmission device.
[0066] In addition, after the image acquisition device and / or audio acquisition device collects data, it can transmit the data to the device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment via a transmission device. For example, the image acquisition device and / or audio acquisition device uses 4G wireless network or vehicle Ethernet technology to transmit the real-time data collected by the image acquisition device and / or audio acquisition device to the device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment (such as the monitoring center server of the examination site). Inside the vehicle, the transmission device is connected to the camera and audio acquisition device via a network cable or wireless access point, and is equipped with a high-performance signal conversion and transmission chip to convert the image data into a format suitable for network transmission (such as H.264 or H.265 video encoding format), and the audio data uses AAC or MP3 audio encoding format. The device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment (such as the monitoring center server of the examination site) is equipped with a corresponding 4G receiving device or Ethernet switch, as well as a storage medium for receiving and storing image and audio data from each examination vehicle. Meanwhile, to ensure the stability and security of data transmission, data encryption and verification technologies are adopted to prevent data from being lost, tampered with, or stolen during transmission.
[0067] Storage media, such as data storage media, can be stored in the device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment using an SD card. It supports storage cards of 64GB or more and can save image, short video, and long video data. Storage media, such as algorithm storage media, can store pre-trained human keypoint algorithm models, object classification algorithm models, and sound algorithm models on the system disk of the device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment. These models are loaded into memory and run when the system starts. Simultaneously, the relevant parameters and configuration files of these algorithm models are also stored on the system disk so that the models can be updated, optimized, and retrained as needed, ensuring that the system can continuously adapt to new examination scenarios and requirements, and improving the accuracy and reliability of cheating detection.
[0068] Furthermore, to improve processing efficiency, the device executing the intelligent monitoring method for the driver's license subject three examination provided in this embodiment can be located in the vehicle's trunk. This device is equipped with a high-performance processor, possessing powerful computing and multi-tasking capabilities. This ensures that the system can quickly store and retrieve image data, algorithm model parameters, and intermediate calculation results during operation. This accelerates the operation of the human body keypoint algorithm, object classification algorithm, and sound algorithm, significantly improving data processing efficiency and meeting the real-time processing needs of large amounts of image and audio data. After obtaining the result of the cheating behavior determination, it can be transmitted to the monitoring center at the examination site.
[0069] See Figure 1 This embodiment provides a method for intelligent supervision of driver's license test (subject three), and the implementation process is as follows:
[0070] 101. During the driver's license test (Part 3), image data of the safety officer is collected using in-vehicle image acquisition equipment. Audio data is collected using in-vehicle audio acquisition equipment.
[0071] The image acquisition device is located in a key position inside the vehicle, which is pre-determined by relevant supervisory personnel, such as in front of the passenger seat, above the driver's seat, and in the rear seats. After acquiring image data, the image acquisition device can transmit it via in-vehicle Ethernet to the device executing the intelligent supervision method for driver's license subject three examination provided in this embodiment.
[0072] The audio acquisition device is located on the vehicle roof; for example, it may be an array of microphones mounted on the roof. After acquiring audio data, the audio acquisition device can transmit it via in-vehicle Ethernet to the device executing the intelligent monitoring method for the driver's license test (subject three) provided in this embodiment.
[0073] 102. Extract the safety officer's motion features based on the image data.
[0074] The motion feature data includes: the distance between the hand and the target part, the speed at which the hand moves toward the target part, and the hand posture.
[0075] The implementation process for this step is as follows:
[0076] 102-1, using a convolutional neural network to extract hand features from image data.
[0077] Convolutional neural networks include convolutional layers and pooling layers.
[0078] Convolutional layers are obtained through the formula The processing is performed where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(·) be the deviation of the spatial dimension j of the input channel l, and let σ(·) be the Sigmoid activation function.
[0079] Pooling layers are obtained through formulas Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R... j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l.
[0080] 102-2. Based on hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0081] The distance between the safety officer's hand and the target component is determined by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the target component's position. h Let y be the vertical coordinate of the safety officer's hand. c The vertical coordinate represents the position of the target component.
[0082] The speed at which the hand moves toward the target part is expressed by the formula d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
[0083] In addition, the acceleration of the hand moving toward the target part can be obtained, which can be expressed by the formula... Where v2 is the speed at which the hand moves toward the target part in the current frame image data, and v1 is the speed at which the hand moves toward the target part in the previous frame image data.
[0084] Hand posture can be measured by finger joint angles, such as those calculated using the law of cosines.
[0085] The target components are internal parts of the vehicle, which are pre-determined by relevant regulatory personnel, such as the steering wheel, gear lever, and handbrake. Furthermore, there can be multiple target components. If there are multiple target components, the method provided in this embodiment will monitor each target component separately, and if any one target component is determined to be fraudulent, then fraud is confirmed.
[0086] The location of the target component is determined before step 102-2 is executed, and the process for determining this location is as follows:
[0087] A1 identifies target components inside the vehicle based on image data.
[0088] For example, target components in image data can be identified using deep learning-based object detection algorithms such as YOLOv5.
[0089] A2 determines multiple bounding boxes of the target component and the class probability corresponding to each bounding box.
[0090] For example, multiple bounding boxes of a target component and the class probability corresponding to each bounding box can be predicted using a relevant prediction model.
[0091] For example, the coordinates of a bounding box are determined by x. min y min x man y max It is determined that the category probability is p. c .
[0092] The category probability can be determined by a visual classification algorithm. This category probability is the confidence level that the content enclosed by the bounding box is the target part.
[0093] A3, based on category probability, using the formula Filter multiple bounding boxes to obtain the predicted bounding box of the target component.
[0094] Where B1 is the area of one bounding box during filtering, and B2 is the area of the other bounding box during filtering.
[0095] To improve execution efficiency, in practice, bounding boxes with class probabilities less than a preset probability threshold can be removed first. Then, the remaining bounding boxes are arranged in descending order, and the bounding box at the top of the list is determined as a candidate bounding box. This candidate bounding box serves as the basis for subsequent filtering (i.e., one of the area source bounding boxes in B1 or B2 during filtering). Finally, the other remaining bounding boxes are selected in turn, and the IoU between the selected remaining bounding boxes and the candidate bounding boxes is calculated. Then, the predicted bounding box of the target component is determined based on the IoU.
[0096] The above process uses Non-Maximum Suppression (NMS) IoU to determine the degree of overlap of bounding boxes and filters out low-confidence boxes.
[0097] A4 uses Kalman filtering to optimize the position of the target component.
[0098] The position of the target component is determined based on its predicted bounding box. For example, the center of the predicted bounding box of the target component is used to determine the position of the target component.
[0099] Furthermore, Kalman filter optimization is based on Kalman filtering, an algorithm that uses the state equations of a linear system to optimally estimate the system state using system input and output observation data. Since the observation data includes the effects of noise and interference in the system, the optimal estimation can also be viewed as a filtering process.
[0100] The measurement update equation during Kalman filter optimization is as follows: t1 is the number of times. For the position observation of the target component in the t1th optimization, To optimize the position observation matrix of the target component for the t1th iteration, Let t1 be the position state variable of the target component in the optimization step. The measurement noise of the target component is optimized for the t1th iteration.
[0101] It should be noted that in practical applications, after obtaining the image data, it can be preprocessed first, and then the preprocessed image data can be used as the image data in step 102. In step 102, the safety officer's action features can be extracted based on the preprocessed image data.
[0102] 103. Extract the sound features inside the vehicle based on the audio data.
[0103] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate.
[0104] The implementation process for this step is as follows:
[0105] 103-1, through formula Noise reduction for audio data.
[0106] Where t2 is the time index of the audio data. y(t2-a1) is the output of the noise reduction process at time t2, a1 is the tap index of the filter used for noise reduction, A is the number of taps of the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 of the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2.
[0107] 103-2, the noise-reduced audio data is divided into frames.
[0108] Each frame has a duration of no less than 20 milliseconds and no more than 30 milliseconds. The frame shift is no less than 10 milliseconds and no more than 15 milliseconds.
[0109] 103-3, Extract sound features based on frame segmentation.
[0110] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate.
[0111] Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, l be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g(a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame.
[0112] Short-time energy x e (g) represents the speech amplitude of the g-th frame of audio data.
[0113] Zero crossing rate sgn[·] is a symbolic function.
[0114] It should be noted that in practical applications, after obtaining the audio data, it can be preprocessed first, and then the preprocessed audio data can be used as the audio data in step 103. In step 103, the sound features inside the vehicle can be extracted based on the preprocessed audio data.
[0115] 104. Based on the characteristics of the movements and voice, determine the cheating behavior.
[0116] Since the motion feature data includes the distance between the hand and the target component, the speed at which the hand moves towards the target component, and the hand posture, the implementation process of this step is as follows:
[0117] 104-1, Classify sound anomalies based on sound characteristics to obtain sound anomaly results.
[0118] In practice, this step can use a long short-term memory network to classify sound features as sound anomalies.
[0119] The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula:
[0120] f t =σ(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t ] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f This is an offset for the forget gate.
[0121] i t =σ(W i [h t-1 ,x t ]+b i ), i tW is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b i This is the bias of the input gate.
[0122] For candidate memory cell states at time t, W C [h t-1 ,x t [] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C The bias is used to select candidate memory cells.
[0123] C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1.
[0124] o t =σ(W o [h t-1 ,x t ]+b o ), o t W is the output gate activation value at time t. o [h t-1 ,x t ] represents the weight of the output gate, b o This is the bias of the output gate.
[0125] h t =o t tanh(C t ), h t Let t be the hidden state at time t.
[0126] If the duration of a specific frequency exceeds the duration threshold, or the match degree with the cheating sound pattern is higher than the match degree threshold, the sound abnormal result is determined to be abnormal.
[0127] 104-2. If the distance between the hand and the target component is less than the distance threshold, the speed at which the hand moves toward the target component is greater than the speed threshold, the hand posture matches the abnormal conditions, and the abnormal sound result is abnormal, then cheating is determined.
[0128] In other words, when the safety officer's key hand points quickly approach and enter the target component, the hand posture matches the abnormal hand movement, and the abnormal sound result is abnormal (such as the duration of a specific frequency exceeding the duration threshold, or the matching degree with the cheating sound pattern being higher than the matching degree threshold), cheating is determined.
[0129] In practical implementation, by setting reasonable thresholds (such as duration threshold, matching degree threshold, distance threshold, speed threshold, etc.) and combining them with dynamic adjustment of the thresholds, normal and cheating behaviors can be effectively distinguished, and accurate cheating behavior judgment results can be obtained.
[0130] Additionally, if a large amount of data is collected in step 101, a time window can be set to sample the collected data, effectively reducing the amount of subsequent data processing and improving data processing efficiency. Furthermore, by setting a reasonable time window length and dynamically adjusting it, the effectiveness of the sampled data in supporting the determination of cheating behavior can be guaranteed while ensuring the amount of data processed, thus ensuring the accuracy of cheating behavior determination.
[0131] This embodiment provides an intelligent monitoring method for driver's license driving test (subject 3). During the test, image data of the safety officer is collected using in-vehicle image acquisition equipment; audio data is collected using in-vehicle audio acquisition equipment; motion characteristics of the safety officer are extracted from the image data; sound characteristics are extracted from the audio data; and cheating behavior is determined based on the motion and sound characteristics. This method obtains the safety officer's motion characteristics from the in-vehicle image data and the sound characteristics from the in-vehicle audio data, and then determines cheating behavior based on these characteristics. This eliminates the need for manual verification in monitoring the driver's license driving test (subject 3), achieving automated and intelligent monitoring and ensuring both efficiency and effectiveness.
[0132] Based on the same inventive concept as the intelligent monitoring method for driver's license subject three examination, this embodiment provides an intelligent monitoring device for driver's license subject three examination, see [link to relevant documentation]. Figure 2 The device includes:
[0133] The data acquisition module 201 is used to collect image data of the safety officer through in-vehicle image acquisition equipment during the driver's license test (subject three). It also collects in-vehicle audio data through in-vehicle audio acquisition equipment.
[0134] The motion feature extraction module 202 is used to extract the motion features of the safety officer based on the image data collected by the data acquisition module 201.
[0135] The sound feature extraction module 203 is used to extract the sound features inside the vehicle based on the audio data collected by the data acquisition module 201.
[0136] The cheating behavior determination module 204 is used to determine cheating behavior based on the action features obtained by the action feature extraction module 202 and the sound features obtained by the sound feature extraction module 203.
[0137] The image acquisition device is located in a key position inside the vehicle, and the resolution of the image acquisition device is no less than 1080 pixels, and the frame rate is no less than 25 frames per second.
[0138] The audio acquisition device is located on the roof of the vehicle, and its frequency response range is 20 Hz to 20 kHz.
[0139] The motion feature data includes: the distance between the hand and the target part, the speed at which the hand moves toward the target part, and the hand posture.
[0140] The action feature extraction module 202 is used to extract hand features from image data using a convolutional neural network. The convolutional neural network includes convolutional layers and pooling layers. The convolutional layers are processed using the formula... The processing is performed where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(·) be the deviation of the spatial dimension j of the input channel l, and let σ(·) be the Sigmoid activation function. The pooling layer is defined by the formula... Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R... j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l.
[0141] Based on hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0142] The distance between the safety officer's hand and the target component is determined by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the target component's position. h Let y be the vertical coordinate of the safety officer's hand. c The vertical coordinate represents the position of the target component. The velocity of the hand moving towards the target component is expressed by the formula... d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
[0143] The device further includes a processing module for identifying target components inside the vehicle based on image data. This involves determining multiple bounding boxes for the target components and the corresponding class probability for each bounding box. Based on the class probabilities, a formula is used to... Multiple bounding boxes are filtered to obtain the predicted bounding boxes of the target component. Here, B1 is the area of one bounding box during filtering, and B2 is the area of another bounding box during filtering. Kalman filtering is then used to optimize the position of the target component. The position of the target component is determined based on its predicted bounding box. The measurement update equation for Kalman filtering optimization is as follows: t1 is the number of times. For the position observation of the target component in the t1th optimization, To optimize the position observation matrix of the target component for the t1th iteration, Let t1 be the position state variable of the target component in the optimization step. The measurement noise of the target component is optimized for the t1th iteration.
[0144] The sound feature extraction module 203 is used to extract sound features using formulas. Noise reduction is performed on the audio data. Here, t2 is the time index of the audio data. Let y(t2-a1) be the output of the noise reduction process at time t2, where a1 is the tap index of the filter used for noise reduction, A is the number of taps in the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 in the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2. The denoised audio data is divided into frames. The duration of each frame is no less than 20 milliseconds and no more than 30 milliseconds. The frame shift is no less than 10 milliseconds and no more than 15 milliseconds. Sound features are extracted based on the frame divisions.
[0145] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate.
[0146] Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, L be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g (a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame.
[0147] Short-time energy x e (g) represents the speech amplitude of the g-th frame of audio data.
[0148] Zero crossing rate sgn[·] is a symbolic function.
[0149] The motion feature data includes: the distance between the hand and the target part, the speed at which the hand moves toward the target part, and the hand posture.
[0150] The cheating behavior determination module 204 is used to classify sound anomalies based on sound characteristics and obtain sound anomaly results. If the distance between the hand and the target part is less than a distance threshold, the speed at which the hand moves toward the target part is greater than a speed threshold, the hand posture matches the anomaly conditions, and the sound anomaly result is abnormal, then cheating is determined.
[0151] Among them, the classification of sound anomalies based on sound characteristics includes:
[0152] Voice anomalies are classified using long short-term memory networks.
[0153] The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula:
[0154] f t =σ(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t ] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f This is an offset for the forget gate.
[0155] i t =σ(W i [h t-1 ,x t ]+b i ), i t W is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b i This is the bias of the input gate.
[0156] For candidate memory cell states at time t, W C [h t-1 ,x t[] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C The bias is used to select candidate memory cells.
[0157] C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1.
[0158] o t =σ(W o [h t-1 ,x t ]+b o ), o t W is the output gate activation value at time t. o [h t-1 ,x t ] represents the weight of the output gate, b o This is the bias of the output gate.
[0159] h t =o t tanh(C t ), h t Let t be the hidden state at time t.
[0160] The device provided in this embodiment obtains the safety officer's action characteristics through in-vehicle image data and in-vehicle sound characteristics through in-vehicle audio data. Then, based on the action and sound characteristics, it determines cheating behavior, so that the supervision of driver's license test (subject 3) no longer relies on manual detection and review, realizing automatic and intelligent supervision, and ensuring supervision efficiency and effectiveness.
[0161] Based on the same inventive concept as the intelligent supervision method for driver's license subject three examination, this embodiment provides an electronic device, which is as follows: Figure 3 As shown, it includes: a memory 301, a processor 302, and a computer program.
[0162] The computer program is stored in memory 301 and configured to be executed by processor 302 to implement the above-mentioned intelligent supervision method for driver's license subject three examination.
[0163] Specifically,
[0164] During the driver's license test (Part 3), image data of the safety officer is collected using in-vehicle image acquisition equipment. Audio data is also collected using in-vehicle audio acquisition equipment.
[0165] Based on the image data, extract the safety officer's motion characteristics.
[0166] Extract the sound features inside the vehicle based on the audio data.
[0167] Cheating behavior is determined based on movement and voice characteristics.
[0168] Optionally, the image acquisition device is located in a key position inside the vehicle, and the resolution of the image acquisition device is not less than 1080 pixels, and the frame rate is not less than 25 frames per second.
[0169] The audio acquisition device is located on the roof of the vehicle, and its frequency response range is 20 Hz to 20 kHz.
[0170] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0171] Based on the image data, extract the safety officer's motion characteristics, including:
[0172] Hand features are extracted from image data using a convolutional neural network. This network consists of convolutional layers and pooling layers. The convolutional layers are defined using the formula... The processing is performed where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(j) be the deviation of the spatial dimension j of the input channel l, and let σ(j) be the Sigmoid activation function. The pooling layer is defined by the formula... Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R... j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l.
[0173] Based on hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0174] The distance between the safety officer's hand and the target component is determined by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the target component's position. h Let y be the vertical coordinate of the safety officer's hand. c The vertical coordinate represents the position of the target component. The velocity of the hand moving towards the target component is expressed by the formula... d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
[0175] Optionally, before determining the distance between the safety officer's hand and the target component, the speed at which the hand moves towards the target component, and the hand posture based on hand characteristics, the method further includes:
[0176] Identify target components inside the vehicle based on image data.
[0177] Determine multiple bounding boxes of the target component and the class probability corresponding to each bounding box.
[0178] Based on the category probability, using the formula Multiple bounding boxes are filtered to obtain the predicted bounding box of the target part. Here, B1 is the area of one bounding box during filtering, and B2 is the area of another bounding box during filtering.
[0179] Kalman filtering is used to optimize the position of the target component. The position of the target component is determined based on its predicted bounding box. The measurement update equation during Kalman filtering optimization is as follows: t1 is the number of times. For the position observation of the target component in the t1th optimization, To optimize the position observation matrix of the target component for the t1th iteration, Let t1 be the position state variable of the target component in the optimization step. The measurement noise of the target component is optimized for the t1th iteration.
[0180] Optionally, based on the audio data, the sound features inside the vehicle are extracted, including:
[0181] Through formula Noise reduction is performed on the audio data. Here, t2 is the time index of the audio data. y(t2-a1) is the output of the noise reduction process at time t2, a1 is the tap index of the filter used for noise reduction, A is the number of taps of the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 of the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2.
[0182] The denoised audio data is divided into frames. The duration of each frame is no less than 20 milliseconds and no more than 30 milliseconds. The frame shift is no less than 10 milliseconds and no more than 15 milliseconds.
[0183] Sound features are extracted based on frame segmentation.
[0184] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate.
[0185] Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, L be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g (a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame.
[0186] Short-time energy x e (g) represents the speech amplitude of the g-th frame of audio data.
[0187] Zero crossing rate sgn[·] is a symbolic function.
[0188] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0189] Based on movement and voice characteristics, cheating behavior is determined, including:
[0190] Based on sound characteristics, sound anomalies are classified to obtain sound anomaly results.
[0191] If the distance between the hand and the target component is less than the distance threshold, the speed at which the hand moves toward the target component is greater than the speed threshold, the hand posture matches the abnormal conditions, and the abnormal sound result is abnormal, then cheating is determined.
[0192] Optionally, sound anomalies can be classified based on sound characteristics, including:
[0193] Voice anomalies are classified using long short-term memory networks.
[0194] The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula:
[0195] f t =σ(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f This is an offset for the forget gate.
[0196] i t =σ(W i [h t-1 ,x t ]+b i ), i t W is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b o This is the bias of the input gate.
[0197] For candidate memory cell states at time t, W C [h t-1 ,x t [] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C The bias is used to select candidate memory cells.
[0198] C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1.
[0199] o t =σ(W o [h t-1 ,x t ]+b o ), o t W is the output gate activation value at time t. o [h t-1 ,x t ] represents the weight of the output gate, b o This is the bias of the output gate.
[0200] h t =o t tanh(C t ), h t Let t be the hidden state at time t.
[0201] The electronic device provided in this embodiment has a computer program executed by a processor to obtain the safety officer's action characteristics through in-vehicle image data and in-vehicle sound characteristics through in-vehicle audio data. Based on the action and sound characteristics, cheating behavior is determined, so that the supervision of driver's license test (subject 3) no longer relies on manual detection and review, realizing automatic and intelligent supervision and ensuring supervision efficiency and effectiveness.
[0202] Based on the same inventive concept as the intelligent monitoring method for driver's license subject three examination, this embodiment provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the aforementioned intelligent monitoring method for driver's license subject three examination.
[0203] Specifically,
[0204] During the driver's license test (Part 3), image data of the safety officer is collected using in-vehicle image acquisition equipment. Audio data is also collected using in-vehicle audio acquisition equipment.
[0205] Based on the image data, extract the safety officer's motion characteristics.
[0206] Extract the sound features inside the vehicle based on the audio data.
[0207] Cheating behavior is determined based on movement and voice characteristics.
[0208] Optionally, the image acquisition device is located in a key position inside the vehicle, and the resolution of the image acquisition device is not less than 1080 pixels, and the frame rate is not less than 25 frames per second.
[0209] The audio acquisition device is located on the roof of the vehicle, and its frequency response range is 20 Hz to 20 kHz.
[0210] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0211] Based on the image data, extract the safety officer's motion characteristics, including:
[0212] Hand features are extracted from image data using a convolutional neural network. This network consists of convolutional layers and pooling layers. The convolutional layers are defined using the formula... The processing is performed where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(·) be the deviation of the spatial dimension j of the input channel l, and let σ(·) be the Sigmoid activation function. The pooling layer is defined by the formula... Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R... j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l.
[0213] Based on hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0214] The distance between the safety officer's hand and the target component is determined by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the target component's position. h Let y be the ordinate of the safety officer's hand. c The vertical coordinate represents the position of the target component. The velocity of the hand moving towards the target component is expressed by the formula... d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
[0215] Optionally, before determining the distance between the safety officer's hand and the target component, the speed at which the hand moves towards the target component, and the hand posture based on hand characteristics, the method further includes:
[0216] Identify target components inside the vehicle based on image data.
[0217] Determine multiple bounding boxes of the target component and the class probability corresponding to each bounding box.
[0218] Based on the category probability, using the formula Multiple bounding boxes are filtered to obtain the predicted bounding box of the target part. Here, B1 is the area of one bounding box during filtering, and B2 is the area of another bounding box during filtering.
[0219] Kalman filtering is used to optimize the position of the target component. The position of the target component is determined based on its predicted bounding box. The measurement update equation during Kalman filtering optimization is as follows: t1 is the number of times. For the position observation of the target component in the t1th optimization, To optimize the position observation matrix of the target component for the t1th iteration, Let t1 be the position state variable of the target component in the optimization step. The measurement noise of the target component is optimized for the t1th iteration.
[0220] Optionally, based on the audio data, the sound features inside the vehicle are extracted, including:
[0221] Through formula Noise reduction is performed on the audio data. Here, t2 is the time index of the audio data. y(t2-a1) is the output of the noise reduction process at time t2, a1 is the tap index of the filter used for noise reduction, A is the number of taps of the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 of the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2.
[0222] The denoised audio data is divided into frames. The duration of each frame is no less than 20 milliseconds and no more than 30 milliseconds. The frame shift is no less than 10 milliseconds and no more than 15 milliseconds.
[0223] Sound features are extracted based on frame segmentation.
[0224] Among these, sound characteristics include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate.
[0225] Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, L be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g (a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame.
[0226] Short-time energy x e (g) represents the speech amplitude of the g-th frame of audio data.
[0227] Zero crossing rate sgn[·] is a symbolic function.
[0228] Optionally, the motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture.
[0229] Based on movement and voice characteristics, cheating behavior is determined, including:
[0230] Based on sound characteristics, sound anomalies are classified to obtain sound anomaly results.
[0231] If the distance between the hand and the target component is less than the distance threshold, the speed at which the hand moves toward the target component is greater than the speed threshold, the hand posture matches the abnormal conditions, and the abnormal sound result is abnormal, then cheating is determined.
[0232] Optionally, sound anomalies can be classified based on sound characteristics, including:
[0233] Voice anomalies are classified using long short-term memory networks.
[0234] The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula:
[0235] f t =v(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t ] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f This is an offset for the forget gate.
[0236] i t =σ(W i [h t-1 ,x t ]+b i ), i t W is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b i This is the bias of the input gate.
[0237] For candidate memory cell states at time t, W C [h t-1 ,x t [] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C The bias is used to select candidate memory cells.
[0238] C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1.
[0239] o t =σ(W o [h t-1 ,x t ]+b o ), o t W is the output gate activation value at time t. o [h t-1 ,xt ] represents the weight of the output gate, b o This is the bias of the output gate.
[0240] h t =o t tanh(C t ), h t Let t be the hidden state at time t.
[0241] The computer-readable storage medium provided in this embodiment has a computer program thereon that is executed by a processor to obtain the safety officer's action characteristics through in-vehicle image data and the in-vehicle sound characteristics through in-vehicle audio data. Based on the action and sound characteristics, cheating behavior is determined, so that the supervision of the driver's subject test no longer relies on manual detection and review, realizing automatic and intelligent supervision and ensuring supervision efficiency and effectiveness.
[0242] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0243] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0244] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0245] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0246] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0247] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for intelligent supervision of driver's license test (subject three), characterized in that, The method includes: During the driver's license test (subject 3), the in-vehicle image acquisition device collects the safety officer's image data; the in-vehicle audio acquisition device collects the in-vehicle audio data. Based on the image data, extract the safety officer's motion characteristics; Based on the audio data, extract the sound features inside the vehicle; Cheating behavior is determined based on the described action characteristics and the described sound characteristics.
2. The method according to claim 1, characterized in that, The image acquisition device is located in a key position inside the vehicle, and the resolution of the image acquisition device is not less than 1080 pixels, and the frame rate is not less than 25 frames / second. The audio acquisition device is located on the roof of the vehicle, and the frequency response range of the audio acquisition device is 20 Hz to 20 kHz.
3. The method according to claim 1, characterized in that, The motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture; The step of extracting the safety officer's motion features based on the image data includes: Hand features are extracted from the image data using a convolutional neural network; wherein the convolutional neural network includes convolutional layers and pooling layers; the convolutional layers are processed using a formula... The processing is performed, where i and j are the spatial dimension indices of the convolutional kernel in the convolutional layer, and l is the input channel index of the convolutional layer. M is the convolution output of input channel l with spatial dimension j. j A collection of spatial dimension indexes. The input is a convolution with spatial dimension j of input channel l. For convolution kernel, Let σ(·) be the deviation of the spatial dimension j of the input channel l, and let σ(·) be the Sigmoid activation function; the pooling layer is defined by the formula... Processing is performed, where j′ is the spatial dimension index of the pooling layer, and R j′ Let be the pooling window corresponding to spatial dimension j′, and m and n be the row and column indices within the pooling window. For the input of the spatial dimension j′ of the input channel l, The output is the spatial dimension j′ of the input channel l; Based on the hand characteristics, determine the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture; The distance between the safety officer's hand and the target component is expressed by the formula... Confirmed, x h Let x be the x-coordinate of the safety officer's hand. c Let y be the x-coordinate of the position of the target component. h Let y be the ordinate of the safety officer's hand. c The vertical coordinate of the target component is given; the speed at which the hand moves towards the target component is expressed by the formula... Determine that d2 is the distance between the safety officer's hand and the target component in the current frame image data, d1 is the distance between the safety officer's hand and the target component in the previous frame image data, and Δt is the acquisition time difference between the current frame image data and the previous frame image data.
4. The method according to claim 3, characterized in that, Before determining the distance between the safety officer's hand and the target component, the speed at which the hand moves toward the target component, and the hand posture based on the hand characteristics, the method further includes: Based on the image data, identify the target components inside the vehicle; Determine multiple bounding boxes of the target component and the category probability corresponding to each bounding box; Based on the category probability, using the formula Multiple bounding boxes are filtered to obtain the predicted bounding box of the target component; where B1 is the area of one bounding box during filtering, and B2 is the area of another bounding box during filtering. The position of the target component is optimized using Kalman filtering; wherein the position of the target component is determined based on the predicted bounding box of the target component; the measurement update equation during Kalman filtering optimization is as follows: t1 is the number of times. For the t1th optimization of the position observation of the target component, To optimize the position observation matrix of the target component for the t1th time, For the t1th optimization, the position state quantity of the target component is... The measurement noise of the target component is optimized for the t1th time.
5. The method according to claim 1, characterized in that, The step of extracting the sound features inside the vehicle based on the audio data includes: Through formula Noise reduction is applied to the audio data; where t2 is the time index of the audio data. y(t2-a1) is the output of the noise reduction process at time t2, a1 is the tap index of the filter used for noise reduction, A is the number of taps of the filter used for noise reduction, h(a1) is the tap coefficient of tap a1 of the filter used for noise reduction, and y(t2-a1) is the input of the noise reduction process at time t2. The denoised audio data is divided into frames; the duration of each frame is not less than 20 milliseconds and not more than 30 milliseconds; the frame shift is not less than 10 milliseconds and not more than 15 milliseconds. Extract sound features based on the frame segmentation; The sound features include Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate; Mel cepstral coefficients e is the identifier for the Mel-frequency cepstral coefficient, C e Let be the e-th coefficient of the Mel-spectral coefficients, g be the frame number identifier, G be the total number of audio data frames, L be the Mel-order, and M(g) be the Mel-filter value of the g-th frame of audio data. a2 is the identifier for the Mel filter, F is the number of Fourier transform points, X(a2) is the cosine transform value of the audio data input to Mel filter a2, and H... g (a2) is the frequency response of Mel filter a2 to the audio data of the g-th frame; Short-time energy x e (g) represents the speech amplitude of the g-th frame of audio data; Zero crossing rate sgn[·] is a symbolic function.
6. The method according to claim 1, characterized in that, The motion feature data includes: the distance between the hand and the target component, the speed at which the hand moves toward the target component, and the hand posture; The step of determining cheating behavior based on the action characteristics and the voice characteristics includes: Based on the aforementioned sound characteristics, sound anomalies are classified to obtain sound anomaly results; If the distance between the hand and the target component is less than a distance threshold, the speed at which the hand moves toward the target component is greater than a speed threshold, the hand posture matches the abnormal conditions, and the abnormal sound result is abnormal, then cheating is determined.
7. The method according to claim 6, characterized in that, The step of classifying sound anomalies based on the sound features includes: The sound features are classified as sound anomalies using a long short-term memory network. The Long Short-Term Memory (LSTM) network updates its hidden state using the following formula: f t =σ(W f [h t-1 ,x t ]+b f ), t is the time identifier, f t h is the forget gate activation value at time t. t-1 Let x be the hidden state at time t-1. t W is the input to the forget gate at time t. f [h t-1 ,x t ] represents the weights of the forget gate, σ(·) represents the Sigmoid activation function, and b f For the offset of the forget gate; i t =σ(W i [h t-1 ,x t ]+b i ), i t W is the input gate activation value at time t. i [h t-1 ,x t ] represents the weights of the input gate, b i For the input gate bias; For candidate memory cell states at time t, W C [h t-1 ,x t [] represents the weights of the candidate memory units, tanh(·) is the Tanh activation function, and b C Bias for candidate memory cells; C t Let C be the state of the memory cell at time t. t-1 The state of memory cells at time t-1; o t =σ(W o [h t-1 ,x t ]+b o ), o t w is the output gate activation value at time t. o [h t-1 ,x t ] represents the weight of the output gate, b o For the output gate bias; h t =o t tanh(C t ), h t Let t be the hidden state at time t.
8. A smart monitoring device for driver's license subject three examination, characterized in that, The device includes: The data acquisition module is used to collect image data of the safety officer through the in-vehicle image acquisition device and audio data of the in-vehicle audio acquisition device during the driver's license test (subject 3). The action feature extraction module is used to extract the action features of the safety officer based on the image data collected by the data acquisition module; The sound feature extraction module is used to extract the sound features inside the vehicle based on the audio data collected by the data acquisition module. The cheating behavior determination module is used to determine cheating behavior based on the action features obtained by the action feature extraction module and the sound features obtained by the sound feature extraction module.
9. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program thereon; the computer program is executed by a processor to implement the method as described in any one of claims 1-7.