Method and system for automatic detection of drunk driving based on deep learning model and robot system
By constructing a dynamic interaction framework for detection scenarios using deep learning models and robotic systems, the problems of low efficiency and environmental interference in existing drunk driving detection methods are solved, achieving efficient and accurate automatic drunk driving detection.
Patent Information
- Application Number
- CN202511433276.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing drunk driving detection methods are inefficient, struggle to cope with disguises and environmental interference, and lack intelligent interaction and closed-loop processing mechanisms, resulting in inaccurate test results and missed detections.
By constructing a dynamic interaction framework for detection scenarios based on deep learning models and robot systems, integrating behavior guidance logic, temporal collaborative rules for multi-source information acquisition, and adaptive adjustment mechanisms for environmental parameters, the framework collects and corrects respiratory gas and facial dynamic feature data, inputs them into a pre-trained model for judgment, and triggers a closed-loop processing flow.
It improves the accuracy and efficiency of drunk driving detection, realizes an intelligent detection system, and can accurately identify drunk driving behavior and quickly record and reset it under different environmental conditions.
Smart Images

Figure CN120894671B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning, in particular to a drunk driving automatic detection method and system based on a deep learning model and a robot system. BACKGROUND
[0002] In the field of traffic safety, drunk driving detection is an important link to ensure road safety and reduce traffic accidents. Traditional drunk driving detection methods mainly rely on manual interception and on-site detection, which has many limitations.
[0003] On the one hand, manual interception detection is low in efficiency, especially during periods and sections with heavy traffic flow, making it difficult to comprehensively detect all passing vehicles, which can lead to missed detection of some drunk drivers. Moreover, manual detection requires a large number of police forces, increasing law enforcement costs.
[0004] On the other hand, existing detection devices have relatively simple functions and can usually only detect alcohol content in respiratory gases, making it difficult to effectively respond to situations where some people try to evade detection through disguising, interference, etc. For example, some people may use mouthwash and other items to cover up the smell of alcohol after drinking, affecting the accuracy of the detection results. In addition, existing detection devices lack the ability to adapt to the detection environment, and under different temperature, humidity, and other environmental conditions, the detection results may be disturbed, leading to false positives or false negatives. At the same time, the detection process lacks intelligent interaction and closed-loop processing mechanisms, and cannot achieve functions such as rapid recording, synchronization, and scene resetting of detection results, which is not conducive to subsequent law enforcement and management. SUMMARY
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a drunk driving automatic detection method based on a deep learning model and a robot system, which comprises:
[0006] A detection scene dynamic interaction framework is constructed by a scene interaction module of the robot system, which includes behavior guidance logic of the object to be detected, time sequence coordination rules of multi-source information collection, and an environment parameter adaptive adjustment mechanism;
[0007] Based on the detection scene dynamic interaction framework, a multi-modal collection unit of the robot system is started, and the respiratory gas data stream and the facial dynamic feature stream of the object to be detected are collected according to the time sequence coordination rules, and the detection environment interference is compensated through the environment parameter adaptive adjustment mechanism, to obtain a synchronous and associated respiratory gas data set and facial dynamic feature set after interference correction;
[0008] The breath gas data set and the facial dynamic feature set are input into a pre-trained drunk driving detection deep learning model, a space-time mapping relationship of the breath gas data set and the facial dynamic feature set is established, a fusion detection feature vector is generated through bidirectional feature conduction and feature interaction processing, and the fusion feature validity is verified;
[0009] The fusion detection feature vector is subjected to drunk driving judgment processing, and a drunk driving comprehensive judgment result containing a detection confidence, a feature correlation strength and a judgment basis is generated.
[0010] According to the drunk driving comprehensive judgment result, a closed-loop processing flow of the detection scene is triggered, the closed-loop processing flow contains an encryption result record, multi-terminal information synchronization and detection scene reset, and the closed-loop processing flow needs to drive the robot components to cooperatively act based on the output result of the drunk driving detection deep learning model.
[0011] In still another aspect, the embodiment of the present application also provides a drunk driving automatic detection system based on a deep learning model and a robot system, characterized in that it comprises:
[0012] A processor; a machine readable storage medium for storing machine executable instructions of the processor; wherein the processor is configured to execute the machine executable instructions to perform the above-mentioned drunk driving automatic detection method based on a deep learning model and a robot system.
[0013] In still another aspect, the embodiment of the present application also provides a computer program product, which comprises machine executable instructions stored in a computer readable storage medium, a processor of a computer device reads the machine executable instructions from the computer readable storage medium, and the processor executes the machine executable instructions, so that the computer device executes the above-mentioned drunk driving automatic detection method based on a deep learning model and a robot system.
[0014] Based on the above aspects, a detection scene dynamic interaction framework is constructed through a scene interaction module of a robot system, which integrates behavior guidance logic of a to-be-detected object, time sequence coordination rules of multi-source information collection, and an environment parameter adaptive adjustment mechanism. Based on the detection scene dynamic interaction framework, a multi-modal collection unit is started, which can accurately collect respiratory gas data flow and facial dynamic feature flow of the to-be-detected object according to the time sequence coordination rules, and effectively compensate for detection environment interference through the environment parameter adaptive adjustment mechanism, to obtain a high-quality data set after synchronization association and interference correction. The respiratory gas data set and the facial dynamic feature set are input into a pre-trained drunk driving detection deep learning model to establish a space-time mapping relationship and generate a fusion detection feature vector, while verifying the effectiveness of the fusion feature, thereby improving the accuracy and reliability of detection. The fusion detection feature vector is subjected to drunk driving judgment processing to generate a comprehensive judgment result containing detection confidence, feature correlation strength, and judgment basis. The detection scene closed-loop processing procedure triggered according to the judgment result realizes functions such as encrypted result recording, multi-terminal information synchronization, and detection scene resetting. Moreover, the robot component is driven to cooperate based on the model output result, forming a complete and intelligent drunk driving automatic detection system, which greatly improves the efficiency, accuracy, and intelligent level of drunk driving detection. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 is an execution flow schematic diagram of a drunk driving automatic detection method based on a deep learning model and a robot system provided by an embodiment of the present application.
[0016] Figure 2 is a schematic diagram of exemplary hardware and software components of a drunk driving automatic detection system based on a deep learning model and a robot system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0017] The present application will be described in detail below with reference to the accompanying drawings, Figure 1 is a flow schematic diagram of a drunk driving automatic detection method based on a deep learning model and a robot system provided by an embodiment of the present application. The drunk driving automatic detection method based on a deep learning model and a robot system will be described in detail below.
[0018] Step S110: A detection scene dynamic interaction framework is constructed through a scene interaction module of a robot system, which contains behavior guidance logic of a to-be-detected object, time sequence coordination rules of multi-source information collection, and an environment parameter adaptive adjustment mechanism.
[0019] The embodiment takes the urban road traffic safety inspection station as the application scene, and a robot system integrating multi-modal collection functions is deployed at the station to automatically detect suspected drunk drivers. As the core control unit of the robot system, the scene interaction module needs to first integrate the spatial information, equipment characteristics and environmental factors of the detection scene to build an interactive framework that can dynamically adapt to the detection process. The interactive framework needs to meet the behavior specification guidance of the object to be detected, the collaborative work of multiple types of collection equipment and the real-time compensation of complex environmental interference.
[0020] Step S111: Call the environment perception unit of the scene interaction module to collect spatial parameters and environmental interference parameters through a multi-type sensor array in the detection area. The spatial parameters include the three-dimensional boundary coordinates of the detection area, the equipment layout position and the detection point space size. The environmental interference parameters include light intensity distribution, air flow rate, environmental temperature change and background noise intensity.
[0021] After the scene interaction module is started, the environment perception unit is activated first. The environment perception unit communicates with the preset sensor array in the detection area through wired connection. The sensor array includes laser radar, infrared range finder, light sensor, anemometer, temperature and humidity sensor and noise meter. These sensors are distributed at the top, around and ground of the detection area according to the preset density. The laser radar and infrared range finder are used to collect spatial parameters. The laser radar performs three-dimensional scanning on the detection area and outputs point cloud data. After coordinate conversion, the three-dimensional boundary coordinates of the detection area are obtained. The infrared range finder is installed on the fixed bracket of each component of the robot system to measure the relative position between the components and calculate the equipment layout position combined with the initial installation coordinates. The spatial size of the detection point is collected by the infrared range finder group deployed around the detection core area. The long, wide and high parameters of the detection point are obtained by fitting the multi-point ranging data.
[0022] The light sensors are uniformly distributed on different height planes of the detection area. Each sensor collects the light intensity at the location at a fixed period to form light intensity distribution data. The anemometer is installed above and on the side of the detection core area to collect air flow rate and wind direction data. The temperature and humidity sensor is distributed at multiple points in the detection area to collect environmental temperature change data in real time. The noise meter is installed at the entrance of the detection area and the edge of the core area to collect background noise intensity. The data collected by all sensors is transmitted to the cache area of the environment perception unit through the data bus. The cache area stores the data by type.
[0023] Step S112: Based on the spatial parameters, a three-dimensional modeling tool is used to build a basic space model of the detection scene. The basic space model divides the activity prohibited area, the guide path area and the detection core area of the object to be detected, and labels the effective working range boundary of each component of the multi-modal collection unit.
[0024] The environmental perception unit packages the collected spatial parameters and sends them to the three-dimensional modeling tool. The three-dimensional modeling tool first checks the spatial parameters, removes outliers, and then constructs the overall profile of the scene based on the three-dimensional boundary coordinates of the detection area as the framework. Within this profile, the functional areas are further divided according to the device layout position and the spatial size of the detection points. The activity prohibited area includes the device installation area, cable wiring area, and sensor deployment area, which is marked by red solid blocks in the model; the guide path area extends from the entrance of the detection area to the core detection area, with the path width set according to ergonomic parameters, marked by a yellow dashed line on the ground layer of the model; the detection core area is the area where the object to be detected collects respiratory gas and facial features, which is marked by a green solid line frame according to the spatial size of the detection points.
[0025] The effective working range boundary of the multi-modal acquisition unit needs to be marked in combination with the technical parameters of each component. The effective working range boundary of the gas acquisition component is calculated according to the inhalation radius of its sampling port and the gas diffusion model, and a translucent sphere is drawn in the model with the sampling port as the sphere center; the effective working range boundary of the image acquisition component is calculated according to the field of view angle, focal length, and resolution of the camera, and a four-pyramid region is drawn in the model with the camera optical center as the vertex. All functional areas and effective working range boundaries are displayed in layers in the basic space model, which can be viewed by switching through the interactive interface.
[0026] Step S113: Combine the basic space model with the behavior habit template of the object to be detected to construct a behavior guidance logic that includes a staged path planning from the entrance of the detection area to the core detection area, posture holding specifications in the core detection area, interaction timing requirements with the acquisition device, and guidance correction rules for abnormal behaviors.
[0027] The behavior habit template is stored in the database of the scene interaction module and contains typical behavior characteristic data of different genders and age groups in similar detection scenarios. The basic space model and the behavior habit template are fused through a data interface. The staged path planning is generated based on the guide path area in the basic space model and the walking speed distribution data in the behavior habit template. The guide path is divided into an entrance guide segment, a transition segment, and a core preparation segment. The entrance guide segment sets a prominent ground mark to guide the object to be detected into the detection area; the transition segment guides the object to be detected to adjust the walking direction through ground arrows and voice prompts; the length of the core preparation segment is set according to the deceleration distance parameter in the behavior habit template, guiding the object to be detected to gradually reduce the walking speed and smoothly enter the core detection area.
[0028] The posture keeping specification is made based on the spatial size of the core detection area and the effective working range boundary of the multi-modal acquisition unit. The object to be detected needs to stand at the center mark point of the core detection area, with both feet as wide as the shoulders, straight torso, face towards the image acquisition assembly, head kept in a natural state, and eyes level with the camera. The timing requirements for interaction with the acquisition device stipulate that after the object to be detected enters the core detection area, it needs to wait for the voice prompt, and after hearing the "please blow" instruction, blow into the sampling port of the gas acquisition assembly until hearing the "blow end" instruction. The facial feature acquisition is carried out synchronously during the blowing process, and the object to be detected needs to keep the facial expression natural and avoid large-scale movements.
[0029] The guiding correction rules of abnormal behavior include behavior recognition conditions and correction measures. The behavior recognition conditions are determined by comparing the real-time position and posture data of the object to be detected with the preset threshold, such as entering the activity prohibited area, deviating from the guide path by more than a preset distance, and the posture not meeting the specification for more than a preset time. The correction measures include voice prompt, light guidance and mechanical arm assistance. The voice prompt module plays the preset guide statement, such as "please return to the guide path" and "please face the camera". The ground guide light strip lights up the corresponding area according to the real-time position of the object to be detected, forming a visual guide path. When the voice and light guidance are ineffective, the mechanical arm assistance is started, and the flexible guide device installed at the end of the mechanical arm touches the arm of the object to be detected, guiding it to adjust the position and posture.
[0030] Step S114: Through multiple pre-acquisition experiments, record the start-up delay parameters, data acquisition cycle parameters, data transmission delay parameters and acquisition accuracy fluctuation range of the gas acquisition assembly and the image acquisition assembly in the multi-modal acquisition unit, establish a component working characteristic database, and based on the component working characteristic database, generate timing coordination rules including the start trigger time difference setting of the gas acquisition assembly and the image acquisition assembly, the acquisition cycle synchronization calibration mechanism and the data transmission timing alignment requirements, so that the acquisition synchronization error of the respiratory gas data stream and the facial dynamic feature stream in the time dimension is controlled within the preset synchronization error range.
[0031] The pre-acquisition experiment is carried out in an interference-free environment, and the experimental samples are simulated dummies of different sizes, and each size of dummy is subjected to multiple repeated experiments. The gas acquisition assembly of the multi-modal acquisition unit includes a gas sensor array and a sampling pump, and the image acquisition assembly includes a high-definition camera and a fill light. During the experiment, the control module sends a start instruction to the gas acquisition assembly and the image acquisition assembly, and records the instruction sending time and the timestamp of the actual start of data acquisition of the assembly, and the difference between the two is the start delay parameter. The data acquisition period parameter is calculated by the timestamp sequence of continuous data acquisition, and the timestamp difference between two adjacent data frames is the acquisition period. The data transmission delay parameter is determined by recording the time difference from the component output port to the system receiving buffer area. The acquisition accuracy fluctuation range is obtained by statistical analysis of multiple acquisition data of the same standard sample, and the standard deviation and the range are calculated.
[0032] All experimental data are classified and stored according to component type and experimental times to form a component working characteristic database. Based on the database, the start trigger time difference is calculated according to the average start delay parameter difference of the gas acquisition assembly and the image acquisition assembly. If the gas acquisition assembly has a longer start delay, the gas acquisition assembly is triggered in advance, and the advance time is equal to the average start delay parameter difference of the two. The acquisition cycle synchronization calibration mechanism adopts a master-slave synchronization mode, taking the acquisition cycle of the image acquisition assembly as the reference, adjusting the sampling pump control signal frequency of the gas acquisition assembly, so that the acquisition cycle deviation of the two is controlled within a preset range. The data transmission timing alignment requirement is realized by adding a timestamp to each data frame, and the system receiving end interpolates or resamples the respiratory gas data stream and the facial dynamic feature stream according to the timestamp, so that the sampling points on the time axis are one-to-one corresponding, and the synchronization error is calculated by the absolute value of the timestamp difference, and it is ensured that the absolute value does not exceed the preset synchronization error threshold.
[0033] Step S115: According to the environmental interference parameters collected by the environment perception unit, an environment parameter adaptive adjustment mechanism including light intensity compensation rules, air flow interference suppression rules, temperature influence correction rules and background noise filtering rules is constructed.
[0034] The construction of the environmental parameter adaptive adjustment mechanism needs to analyze the relevance between the environmental interference parameters and the acquisition accuracy of the multi-modal acquisition unit. The light intensity compensation rule is formulated according to the light intensity distribution data collected by the light sensor and the exposure characteristic curve of the image acquisition component. When the light intensity is lower than the preset lower threshold, the fill light is started, and the brightness level of the fill light is determined according to the difference between the light intensity and the lower threshold; when the light intensity is higher than the preset upper threshold, the aperture size and shutter speed of the camera are adjusted to reduce the exposure amount. The air flow interference suppression rule is formulated based on the air flow rate data collected by the anemometer and the gas diffusion model, and when the air flow rate exceeds the preset threshold, the wind shield of the gas acquisition component is started, and the opening angle of the wind shield is adjusted according to the wind direction data; at the same time, the sampling time is extended, and the sampling time extension ratio is proportional to the multiple of the air flow rate exceeding the threshold.
[0035] The temperature influence correction rule is formulated according to the environmental temperature change data collected by the temperature and humidity sensor and the temperature characteristic model of the gas sensor, a temperature compensation coefficient table is established, different compensation coefficients correspond to different temperature intervals, and the output value of the gas sensor is multiplied by the compensation coefficient to obtain the corrected gas concentration data. The background noise filtering rule is based on the background noise intensity data collected by the noise meter, when the noise intensity exceeds the preset threshold, the noise reduction algorithm of the voice prompt module is started, the noise frequency band is identified through spectrum analysis, and the corresponding frequency band of the voice signal is attenuated; at the same time, the volume of the voice prompt is adjusted, and the volume adjustment amplitude is proportional to the decibel number of the noise intensity exceeding the threshold.
[0036] Step S116: Establish the mutual influence relationship between the behavior guidance logic, the time sequence coordination rule and the environmental parameter adaptive adjustment mechanism, and fuse to generate a detection scene dynamic interaction framework containing space constraints, time constraints and environmental constraints.
[0037] The behavior guidance logic, the time sequence coordination rule and the environmental parameter adaptive adjustment mechanism output control instructions to different execution units, and there is mutual influence among them. The posture maintenance specification in the behavior guidance logic affects the acquisition cycle setting of the time sequence coordination rule, when the posture change frequency of the object to be detected is high, the time sequence coordination rule needs to increase the acquisition cycle to capture more details; the start trigger time difference setting of the time sequence coordination rule affects the preheating time of the environmental parameter adaptive adjustment mechanism, the components that are triggered in advance need to be monitored in advance; the light intensity compensation rule of the environmental parameter adaptive adjustment mechanism affects the visual guidance effect of the behavior guidance logic, the opening of the fill light may change the visibility of the ground mark, and the identification brightness of the guidance path area needs to be adjusted synchronously.
[0038] The interaction weight between the three is quantified by constructing an influence relationship matrix, and the weight value is obtained by training historical data. The functional area in the space parameter is divided as a space constraint, the synchronization error requirement in the time sequence coordination rule is taken as a time constraint, and each threshold in the environment parameter adaptive adjustment mechanism is taken as an environment constraint. The detection scene dynamic interaction framework is generated by fusion. The framework is stored in the form of data structure, including behavior guidance logic, time sequence coordination rule, specific parameters and execution priority of environment parameter adaptive adjustment mechanism, and the execution priority is dynamically adjusted according to the stage of the detection process.
[0039] Step S120: based on the detection scene dynamic interaction framework, starting the multi-modal acquisition unit of the robot system, collecting the respiratory gas data stream and the facial dynamic feature stream of the object to be detected according to the time sequence coordination rule, compensating the detection environment interference through the environment parameter adaptive adjustment mechanism, obtaining the synchronized and associated respiratory gas data set and facial dynamic feature set after interference correction.
[0040] After the detection scene dynamic interaction framework is constructed, the robot system enters the detection state. When the object to be detected enters the detection area entrance, the infrared sensing device is triggered, and the multi-modal acquisition unit is started. The multi-modal acquisition unit includes a gas acquisition assembly and an image acquisition assembly. The gas acquisition assembly includes a gas sensor array, a sampling pump and a wind shield; the image acquisition assembly includes a high-definition camera, a fill light and an image processor. The system controls the starting time sequence and acquisition parameters of the multi-modal acquisition unit according to the time sequence coordination rule in the detection scene dynamic interaction framework, and adjusts the acquisition strategy in real time according to the environment parameter adaptive adjustment mechanism to eliminate environmental interference. During the acquisition process, the respiratory gas data stream and the facial dynamic feature stream are associated by timestamp, and form a structured data set after interference correction.
[0041] Step S121: according to the behavior guidance logic in the detection scene dynamic interaction framework, the initial path guidance signal is sent at the detection area entrance through the guidance execution component of the robot system, the initial path guidance signal includes the lighting sequence of the ground embedded visual guidance light belt and the phased prompt instruction of the voice broadcast module, which is used to guide the object to be detected to move to the preset detection point of the detection core area.
[0042] The guiding execution component includes a ground-embedded visual guidance light strip, a voice broadcast module, and a mechanical guiding arm. The ground-embedded visual guidance light strip is laid along the guiding path area and is composed of variable color LED lamp beads arranged at a preset interval. The generation of the initial path guiding signal is based on the phased path planning in the behavior guiding logic. When the object to be detected enters the detection area entrance, the infrared sensing device outputs a trigger signal, and after the guiding execution component receives the trigger signal, the visual guidance light strip is controlled to light up in a preset light sequence, with the light sequence being to light up sequentially from the entrance to the detection core area, forming a flowing light effect; at the same time, the voice broadcast module is started to play the first stage content of the phased prompt instruction, prompting the object to be detected to move following the light strip.
[0043] The color and flashing frequency of the visual guidance light strip can be adjusted according to the moving speed of the object to be detected. When the object to be detected moves too slowly, the light strip flashing frequency is accelerated; when the object to be detected deviates from the guiding path, the light strip at the deviated position flashes red. The voice broadcast module uses a high-fidelity loudspeaker, and the playback volume is automatically adjusted according to the background noise intensity to ensure the clarity of the prompt instruction. The preset detection point has a clear identification pattern on the ground in the detection core area, and the end point of the visual guidance light strip coincides with the identification pattern.
[0044] Step S122: In the moving process of the object to be detected, the position information of the object to be detected is obtained in real time by the infrared sensors arranged along the path by the guiding execution component, and the light-up area of the visual guidance light strip and the content of the voice prompt instruction are dynamically adjusted in combination with the path planning in the basic space model, so that the object to be detected moves along the preset path.
[0045] The infrared sensors along the path are arranged in an array, and each sensor is responsible for monitoring a specific area. When the object to be detected enters the monitoring area of a certain sensor, the sensor outputs a high-level signal, and the system calculates the real-time position information of the object to be detected according to the sensor number triggered. The position information is compared with the guiding path area in the basic space model to determine whether the object to be detected is on the preset path and the deviation distance from the path center line.
[0046] When the object to be detected is on the preset path, the light-up area of the visual guidance light strip is a preset length range in front of the current position, and the light-up color is green; the voice prompt instruction is switched to the second stage content, prompting the object to be detected to maintain the current direction. When the object to be detected deviates from the preset path, the system calculates the deviation direction and the deviation distance, controls the light strip on the deviation direction side to light up red, and at the same time, the voice broadcast module plays the direction correction instruction, and the instruction content includes the deviation direction and the correction suggestion. The cycle of dynamic adjustment is consistent with the acquisition cycle of the infrared sensor, ensuring the real-time of the guidance.
[0047] Step S123: When the to-be-detected object reaches the preset detection point, the guidance execution component sends a posture adjustment signal according to the posture keeping specification in the behavior guidance logic, and starts the visual acquisition unit of the robot system to capture the current posture image of the to-be-detected object. The current posture image is compared with the preset standard posture template at the pixel level, the head angle deviation, the body posture deviation, and the distance deviation from the acquisition device are calculated respectively, and are standardized and mapped to the same scale range and then weighted and summed to obtain a posture deviation parameter. The fine degree of the voice prompt instruction is adjusted according to the posture deviation parameter, and the current posture image is continuously captured and compared until the posture deviation parameter of the to-be-detected object is within the preset deviation threshold interval, and the posture adjustment is completed.
[0048] The ground identification pattern of the preset detection point is provided with a pressure sensor. When the to-be-detected object stands on the identification pattern, the pressure sensor outputs a trigger signal. After the guidance execution component receives the trigger signal, a posture adjustment signal is sent according to the posture keeping specification. The posture adjustment signal includes color change of the visual guidance light belt and switching of the voice prompt instruction. The visual guidance light belt changes from flowing light effect to stable green constant, and the voice prompt instruction is switched to posture adjustment prompt content.
[0049] The visual acquisition unit includes high-definition cameras arranged around the detection core area. The cameras are installed at preset angles to ensure that the front, side, and top images of the to-be-detected object can be captured. The current posture image is synchronously acquired by the cameras, and is compared with the standard posture template after image preprocessing. The standard posture template is a standard human posture image in the front and side stored in the system. The head angle deviation is calculated by extracting the coordinate difference of the head feature points in the current posture image and the standard posture template, including the eye corner, nose tip, and chin tip. The body posture deviation is calculated by extracting the coordinate difference of the torso feature points, including the shoulder peak, anterior superior iliac spine, and ankle joint. The distance deviation from the acquisition device is calculated by the focal length and imaging scale of the camera.
[0050] The calculated head angle deviation, body posture deviation and distance deviation are respectively normalized to map the deviation values to a range of zero to one. The normalized deviation values are multiplied by preset weight coefficients and summed to obtain a posture deviation parameter. The weight coefficients are set according to the influence of each deviation on the collection accuracy, and the weight coefficient of the head angle deviation is the highest. When the posture deviation parameter is greater than the upper limit of the preset deviation threshold, the voice broadcast module plays detailed posture adjustment instructions, such as "please turn your head to the left" and "please separate your feet to the width of your shoulders"; when the posture deviation parameter is between the upper and lower limits of the preset deviation threshold, a simplified adjustment instruction is played; when the posture deviation parameter is less than the lower limit of the preset deviation threshold, it is determined that the posture adjustment is completed, and the voice prompt "correct posture, please prepare for collection" is given. During the posture adjustment process, the camera collects images at a fixed period and calculates the posture deviation parameter until the parameter is within the preset deviation threshold interval.
[0051] Step S124: Start the synchronous trigger module of the multi-modal collection unit, call the time sequence coordination rule in the detection scene dynamic interaction framework, obtain the start trigger time difference and collection cycle parameters of the gas collection component and the image collection component, synchronously trigger the start of the gas collection component and the image collection component, and monitor the working states of the gas collection component and the image collection component in real time during the collection process. The time stamps of the collected data are calibrated through data transmission time sequence alignment, so that the respiratory gas data and the facial dynamic feature data at the same collection time have the same time identifier.
[0052] The synchronous trigger module is the control core of the multi-modal collection unit, and is internally provided with a high-precision clock chip and a trigger signal generator. When started, the synchronous trigger module reads the time sequence coordination rule from the detection scene dynamic interaction framework through a data interface, and parses the start trigger time difference and the collection cycle parameters of the gas collection component and the image collection component. The trigger signal generator generates two trigger signals according to the start trigger time difference, and sends them to the start control ends of the gas collection component and the image collection component respectively, to ensure that the two components start according to the preset time sequence.
[0053] During the collection process, the synchronous trigger module reads the working state words of the gas collection component and the image collection component in real time through a state monitoring interface. The working state words include power state, sensor state and data transmission state. When an abnormal state is detected, the synchronous trigger module sends a reset signal to restart the components. The data transmission time sequence alignment is realized by embedding a time stamp in the trigger signal. Each time the collection is triggered, the clock chip outputs the current time stamp, which is transmitted to the system receiving end together with the collected data. The system receiving end compares the time stamps of the respiratory gas data and the facial dynamic feature data, and if there is a deviation, interpolation processing is performed according to the data transmission time sequence alignment requirement, so that the data at the same collection time have the same time identifier. The time identifier adopts the UTC time format unified by the system, and is accurate to the millisecond level.
[0054] Step S125: Call the environmental parameter adaptive adjustment mechanism in the scene dynamic interaction framework, and dynamically adjust the sampling flow of the gas collection component and the exposure parameter of the image collection component according to the real-time collected environmental interference parameters: when the light intensity is outside the preset light threshold interval, the light compensation intensity of the image collection component is adjusted; when the air flow rate is outside the preset flow rate threshold interval, the sampling port orientation of the gas collection component is adjusted and the sampling time length is adjusted.
[0055] The calling of the environmental parameter adaptive adjustment mechanism is realized through the real-time data interaction of the environment perception unit and the multi-modal collection unit. The environment perception unit sends the environmental interference parameters to the control module of the multi-modal collection unit at a fixed period, and the control module calculates the adjustment parameters according to the rules in the environmental parameter adaptive adjustment mechanism. The sampling flow of the gas collection component is realized by controlling the rotating speed of the sampling pump, and the rotating speed of the sampling pump is proportional to the sampling flow, and the control module adjusts the rotating speed instruction according to the temperature compensation coefficient in the temperature influence correction rule.
[0056] The exposure parameters of the image collection component include aperture size, shutter speed and ISO sensitivity, and the control module calculates the optimal exposure parameter combination corresponding to the current light intensity according to the light intensity compensation rule. When the light intensity is lower than the lower limit of the preset light threshold interval, the control module calculates the light compensation intensity requirement, which is proportional to the difference between the light intensity and the lower limit threshold, and the driving current of the light compensation lamp is adjusted according to the light compensation intensity; when the light intensity is higher than the upper limit of the preset light threshold interval, the control module preferentially reduces the aperture, and when the aperture is adjusted to the minimum and still cannot meet the requirements, the shutter speed is shortened.
[0057] The air flow rate is monitored in real time by an anemometer, and when the air flow rate is higher than the upper limit of the preset flow rate threshold interval, the control module drives the steering motor of the gas collection component to adjust the sampling port orientation, so that the sampling port faces away from the wind direction; at the same time, the sampling time length adjustment value is calculated according to the air flow interference suppression rule, and the new sampling time length is obtained by adding the adjustment value to the original sampling time length. The calculation of the adjustment value is based on the proportion of the air flow rate exceeding the threshold, and the larger the proportion, the larger the adjustment value. The adjusted sampling flow and exposure parameters are sent to the execution unit of the corresponding component through the control bus, and the execution unit configures the hardware circuit according to the new parameters.
[0058] Step S126: Collect the respiratory gas samples of the to-be-detected object through the gas collection component according to the adjusted sampling parameters, form an initial respiratory gas data stream by arranging the gas composition response value, sampling flow parameter and environmental temperature parameter of each collection time in chronological order through the gas sensor array.
[0059] The gas collection assembly comprises a sampling port, a sampling pump, a gas sensor array and a data preprocessing circuit. The sampling port is detachable and has a dustproof filter screen inside. The sampling pump is a miniature diaphragm pump with adjustable rotating speed. The gas sensor array comprises an alcohol sensor, a carbon dioxide sensor and a humidity sensor, and outputs analog signals. The data preprocessing circuit comprises a signal amplification, filtering and A / D conversion module. According to the adjusted sampling parameters, the sampling pump operates at the set rotating speed, and the respiratory gas of the object to be detected is sucked in through the sampling port and detected by the gas sensor array.
[0060] Each sensor in the gas sensor array responds to different gas components and outputs an analog signal proportional to the gas concentration. The analog signal is converted into a digital signal by the data preprocessing circuit, which is the gas component response value. At each collection time, the system simultaneously records the gas component response value, the current rotating speed of the sampling pump (corresponding to the sampling flow parameter) and the current environmental temperature parameter sent by the environmental perception unit. The above data are packaged in time sequence, each data packet contains a timestamp, an array of gas component response values, a sampling flow parameter and an environmental temperature parameter, multiple data packets are arranged in time sequence to form an initial respiratory gas data stream, and the data stream is stored in a ring buffer. The size of the buffer is set according to the preset collection time.
[0061] Step S127: The face image of the object to be detected is continuously captured by the image collection assembly according to the adjusted exposure parameters, the eye dynamic features, facial muscle movement features and skin color change features in each face image are extracted, the illumination intensity parameter at the time of capturing is recorded, and the initial face dynamic feature stream is formed by combining in time sequence.
[0062] The image collection assembly comprises a multi-angle camera group, a fill light and an image processor. The multi-angle camera group is composed of high-definition cameras arranged on the front, left and right sides of the object to be detected, and each camera is equipped with an adjustable focal length lens. The fill light is a ring-shaped LED lamp installed around the front camera. The image processor comprises an FPGA and a GPU for image collection and feature extraction. According to the adjusted exposure parameters, the camera group starts shooting synchronously, the image sensor outputs raw image data at the set frame rate, and the raw image data is preprocessed by the FPGA module of the image processor, including denoising, contrast enhancement and distortion correction.
[0063] The preprocessed face image is input into a GPU module for feature extraction. The extraction of eye dynamic features is achieved by locating the pupil and eyelid contours. An edge detection algorithm is used to extract the circumscribed ellipse of the pupil, and the center coordinates and long axis direction of the ellipse are calculated over time. Meanwhile, the upper and lower eyelid contour curves are extracted, and the eyelid opening degree is calculated. The extraction of facial muscle movement features is achieved by tracking facial feature points. Key feature points such as the corners of the mouth, the nasal ala, and the brow bone are selected, and the displacement and movement speed of the feature points are calculated. The extraction of skin color change features is achieved by analyzing the color space component changes in the face region. The RGB image is converted to the HSV color space, and the mean and variance of the H and S components over time are extracted.
[0064] The eye dynamic features, facial muscle movement features, and skin color change features extracted at each collection time form a feature vector. The illumination intensity parameter at the time of shooting is recorded in real time by the environment perception unit. The feature vector and the illumination intensity parameter are combined in chronological order. Each time point corresponds to a feature data block. Multiple feature data blocks are arranged in chronological order to form an initial face dynamic feature stream. The initial face dynamic feature stream is stored in a video stream packaging format. Each key frame contains complete feature vector data.
[0065] Step S128: Based on the timestamp information of the collection time, the data of the same time node in the initial respiratory gas data stream and the initial face dynamic feature stream are associated and marked to generate a pair of associated data. The initial respiratory gas data in the pair of associated data is removed from the temperature variation interference according to the temperature influence correction rule in the environmental parameter adaptive adjustment mechanism, obtaining the corrected respiratory gas data. The initial face dynamic feature data in the pair of associated data is removed from the light variation interference according to the illumination intensity compensation rule, obtaining the corrected face dynamic feature data. All corrected respiratory gas data are integrated in chronological order to form a set of synchronous and interference-corrected respiratory gas data. All corrected face dynamic feature data are integrated in chronological order to form a set of synchronous and interference-corrected face dynamic features.
[0066] The timestamp information is the UTC time generated by the synchronous triggering module, accurate to the millisecond level. Each data unit in the initial respiratory gas data stream and the initial face dynamic feature stream contains a timestamp field. The association marking process is achieved by traversing the timestamp fields of the two data streams. When the timestamp difference of two data units is less than a preset matching threshold (the preset matching threshold is set to 50 ms, which is verified by multiple collection experiments, and this preset matching threshold can ensure the accuracy of the data of the same time node), it is determined that the data is of the same time node, and a pair of associated data is generated, which contains the initial respiratory gas data and the initial face dynamic feature data.
[0067] The removal of temperature change interference is based on a temperature influence correction rule, a compensation coefficient corresponding to the current environmental temperature is read from an environmental parameter self-adaptive adjustment mechanism, the gas component response value in the initial respiratory gas data is multiplied by the compensation coefficient to obtain a corrected gas component response value, the sampling flow parameter and the environmental temperature parameter remain unchanged, and a corrected respiratory gas data is formed. The removal of illumination change interference is based on an illumination intensity compensation rule, an illumination intensity parameter I is extracted from the initial face dynamic feature data, and an illumination compensation factor K is calculated. The specific calculation method is as follows: first, a standard illumination intensity I0 (the value is 500 lux, which meets the standard detection illumination environment requirement of the indoor and outdoor transition area of the urban road traffic safety inspection station) is set, when I≠0, K=I0 / I; when I=0 (extreme low light environment), K=10 (through experimental calibration, the value can avoid excessive correction of features in a low light environment); at the same time, an illumination response correction coefficient a (the value is 0.95, which is obtained through 1000 group face feature collection experiments under different illumination conditions, and is used to compensate the nonlinear response of the image sensor) is introduced, and finally the illumination compensation factor K_final=K×a. The H component and the S component in the skin color change feature are corrected, and the correction formula is as follows: corrected H component=original H component×K_final, corrected S component=original S component×K_final; the eye dynamic feature and the facial muscle movement feature are not affected by illumination, and remain unchanged, and a corrected face dynamic feature data is formed.
[0068] The corrected respiratory gas data is arranged in ascending order of timestamp, each data unit contains the corrected gas component response value, the sampling flow parameter, the environmental temperature parameter and the timestamp, all data units are combined to form a respiratory gas data set, the data set is stored in JSON format, contains metadata and a data list, and the metadata records the start time, the collection time length and the number of data units. The corrected face dynamic feature data is also arranged in ascending order of timestamp, forming a face dynamic feature data set, which is stored in the same JSON format as the respiratory gas data set to ensure the consistency of the data structure.
[0069] Step S130: input the respiratory gas data set and the face dynamic feature set into the pre-trained drunk driving detection deep learning model, establish the space-time mapping relationship of the respiratory gas data set and the face dynamic feature set, generate a fusion detection feature vector through the feature interaction processing of bidirectional feature conduction, and verify the validity of the fusion feature.
[0070] The drunk driving detection deep learning model is deployed in the edge computing unit of the robot system, which is a multi-modal fusion model based on the Transformer architecture, including a gas feature extraction layer, a face feature extraction layer, a feature correlation layer and a feature fusion layer. The construction and training process of the model is as follows: 1. Model structure design: the gas feature extraction layer adopts a 3-layer 1D convolutional network, and the convolution kernel size is 3, 5 and 7 respectively (used to capture the gas concentration change features of different time scales), and each layer is followed by batch normalization and ReLU activation function, and finally the global average pooling outputs a gas feature vector with a dimension of 256; the face feature extraction layer adopts 2 layers of 2D convolution (capture spatial features) + 2 layers of Temporal Convolutional Network (TCN, capture time features), the 2D convolution kernel size is 3x3, and the TCN convolution kernel size is 3, and the output dimension is a face feature vector with a dimension of 256; the feature correlation layer adopts 8 heads of multi-head self-attention mechanism, and the attention weight is calculated by the covariance matrix of the gas and face features; the feature fusion layer adopts a fully connected network, which concatenates the correlated features into a 512-dimensional vector, and outputs a fusion detection feature vector through 2 layers of fully connected (hidden layer dimension is 1024) and Sigmoid activation function. 2. Training data preparation: collect 100,000+ sets of labeled data, each set of labeled data contains known alcohol concentration (0-200mg / 100ml, covering non-drunk driving, drunk driving, drunk driving scenarios) breath gas data, corresponding face dynamic feature data, and artificial labeled drunk driving state label (non-drunk driving / drunk driving / drunk driving); the data is enhanced: the gas data is added with ±5% Gaussian noise (to simulate sensor error), and the face data is rotated by ±10° and scaled by 0.8-1.2 times (to simulate posture changes). 3. Training process: adopt end-to-end training method, the loss function is cross entropy loss (weight 0.7) and contrast loss (weight 0.3, used to enhance the consistency of multi-modal features); the optimizer adopts Adam, the initial learning rate is 1e-4, and every 10 rounds is decayed to 0.1 times of the original learning rate according to the cosine annealing strategy; the training rounds are 100 rounds, 5-fold cross validation is adopted, and the training is stopped when the verification set accuracy reaches 98.2%; after the training is completed, 30,000 sets of test set (not involved in the training) are used for model evaluation to ensure the generalization ability of the model, and the test set accuracy is not less than 97.5%. The breath gas data set and the face dynamic feature set are used as the input of the model, and the gas feature vector sequence and the face feature vector sequence are obtained by feature encoding through the gas feature extraction layer and the face feature extraction layer respectively. The feature correlation layer establishes the spatio-temporal mapping relationship between the two through the attention mechanism, realizes the bidirectional feature conduction, and the feature fusion layer performs concatenation and non-linear transformation on the conducted features to generate a fusion detection feature vector.The fusion feature effectiveness is verified by calculating the information entropy and dispersion of the feature, to ensure that the feature contains sufficient discriminant information.
[0071] Step S131: input the set of respiratory gas data into the gas feature extraction layer of the drunk driving detection deep learning model to obtain a sequence of gas feature vectors.
[0072] Step S1311: receive the set of respiratory gas data through the preprocessing sublayer of the gas feature extraction layer, and extract the gas component response value, sampling flow parameter and environmental temperature parameter in each respiratory gas data.
[0073] The preprocessing sublayer parses the input set of respiratory gas data and extracts key parameters from each data unit. The gas component response value reflects the detection results of the gas sensor array on various components in the respiratory gas; the sampling flow parameter represents the inhalation rate during gas collection; and the environmental temperature parameter is used for subsequent temperature compensation calculation. The above parameters are extracted in sequence and organized into a structured data list.
[0074] Step S1312: remove abnormal response values in the gas component response value that exceed the range of the gas sensor, fill in the missing points of the remaining response values using the mean filling method to obtain the preliminary processed gas component response value, convert the response values under different sampling flow rates to the response values under the standard sampling flow rate based on the flow rate-response characteristic curve of the gas sensor, and correct the response value offset caused by temperature changes based on the temperature-response characteristic model of the gas sensor to obtain the temperature-compensated gas component response value.
[0075] Abnormal response value removal is achieved by comparing the gas component response value with the sensor range threshold. Values exceeding the range are marked as abnormal and removed. For missing points in the data sequence, the mean filling method is used, that is, the arithmetic mean of a predetermined number of valid data before and after the missing point is calculated to fill in the missing point, to ensure the continuity of the data sequence. In the flow conversion process, the response value under the actual sampling flow rate is mapped to the theoretical response value corresponding to the standard sampling flow rate according to the flow rate-response characteristic curve calibrated when the gas sensor is shipped. Temperature compensation uses the pre-established temperature-response characteristic model of the gas sensor to calculate the compensation coefficient by substituting the environmental temperature parameter into the model, and corrects the response value after flow conversion to eliminate the influence of temperature changes on the detection accuracy of the sensor.
[0076] Step S1313: standardize and map the temperature-compensated gas component response value to obtain a standardized gas component response value, and combine the standardized gas component response value, the corrected sampling flow parameter and the compensated environmental temperature parameter to form a standardized respiratory gas data.
[0077] The standardized mapping adopts a min-max normalization method to compress the temperature-compensated gas component response value to a specific numerical interval, and the relative size in the interval is used to represent the strength of the gas component response. The corrected sampling flow parameter is compared with the standard sampling flow to obtain a relative flow coefficient, which is used to reflect the deviation of the actual sampling flow from the standard flow. The compensated ambient temperature parameter is converted into a deviation value from the standard temperature to reflect the influence of the ambient temperature on the detection process. The standardized gas component response value, the relative flow coefficient and the temperature deviation value are combined in a predetermined order to form standardized respiratory gas data containing multi-dimensional information.
[0078] Step S1314: calling the convolution filtering unit of the gas feature extraction layer, loading the pre-trained multi-scale convolution kernel parameters containing different size convolution windows, each convolution window corresponding to different convolution kernel weights, arranging the standardized gas component response values in the standardized respiratory gas data in time sequence into a one-dimensional data sequence, and inputting them into different size convolution windows for convolution operation respectively.
[0079] The convolution filtering unit loads the pre-trained multi-scale convolution kernel parameters, which have different size convolution windows to capture features in different time ranges. The standardized gas component response values in the standardized respiratory gas data are arranged in time sequence to form a one-dimensional data sequence. Then the one-dimensional data sequence is input into different size convolution windows respectively, and the data sequence is filtered by convolution operation, and each convolution window extracts specific feature patterns in the data according to the corresponding convolution kernel weights.
[0080] Step S1315: different range convolutions are performed on the one-dimensional data sequence by different size convolution windows to capture local peak value features, slope change features between local peak values and overall concentration distribution trend features respectively and output corresponding feature maps, and the feature maps are spliced in channel dimension to obtain gas original features.
[0081] Different size convolution windows perform sliding convolution operation on the one-dimensional data sequence. Smaller size convolution windows are used for local convolution of the data sequence to capture local peak value features in a short time, which may correspond to mutation points of gas components; medium size convolution windows perform medium range convolution to capture slope change features between local peak values, reflecting the change rate of gas component concentration; larger size convolution windows perform larger range convolution to capture overall concentration distribution trend features, reflecting the overall change law of gas components over time. Each convolution window outputs a corresponding feature map, and these feature maps are spliced in channel dimension to integrate different types of features together to form gas original features containing multi-aspect information.
[0082] Step S1316: input the gas original features into the nonlinear transformation sublayer of the gas feature extraction layer, call the pre-trained activation function to perform nonlinear transformation on each feature value in the gas original features, strengthen the response signal related to the alcohol molecule features, suppress the response signal related to the interfering gas components, and obtain the strengthened gas features.
[0083] After the nonlinear transformation sublayer receives the gas original features, the pre-trained activation function is called to perform nonlinear processing on each feature value in the features. The activation function can perform nonlinear mapping on the input feature values, so that the differences between the feature values are more obvious. Through the nonlinear transformation, the response signal related to the alcohol molecule features can be enhanced, the key features of alcohol detection can be highlighted, the response signal related to the interfering gas components can be suppressed, and the influence of irrelevant factors on subsequent feature extraction can be reduced, thereby obtaining the strengthened gas features.
[0084] Step S1317: performing feature dimension analysis on the strengthened gas features through the feature integration sublayer of the gas feature extraction layer, identifying the feature dimensions related to the molecular structure and the feature dimensions related to the concentration distribution, extracting the feature peak position, the feature peak width and the feature peak intensity from the feature dimensions related to the molecular structure to form the molecular structure response features, and calculating the concentration uniformity parameter and the concentration gradient parameter from the feature dimensions related to the concentration distribution to form the concentration distribution features.
[0085] The feature integration sublayer first performs feature dimension analysis on the strengthened gas features, identifies the feature dimensions related to the molecular structure of the gas and the feature dimensions related to the concentration distribution of the gas according to the physical meaning of the features and the pre-defined rules. For the feature dimensions related to the molecular structure, the feature peak position, the feature peak width and the feature peak intensity are extracted by analyzing the morphological parameters of the feature peaks, and these information can reflect the structural characteristics of the gas molecules, which are combined to form the molecular structure response features. For the feature dimensions related to the concentration distribution, the concentration uniformity parameter is calculated to describe the uniformity of the gas concentration in space or time, and the concentration gradient parameter is calculated to reflect the rate of concentration change, and the two parameters are combined to form the concentration distribution features.
[0086] Step S1318: arranging the molecular structure response features and the concentration distribution features in a pre-set dimension order to form a gas feature vector of a single breath gas data, and arranging the gas feature vectors of all single breath gas data in sequence according to the collection time sequence of the breath gas data in the breath gas data set, to generate a gas feature vector sequence.
[0087] The molecular structure response feature and the concentration distribution feature are arranged according to a preset dimension order, and combined into a gas feature vector capable of comprehensively representing a single respiratory gas data feature. Then, according to the collection time order of the respiratory gas data in the respiratory gas data set, the gas feature vectors corresponding to all single respiratory gas data are sequentially arranged to form a gas feature vector sequence.
[0088] Step S132: inputting the face dynamic feature set into a face feature extraction layer of the drunk driving detection deep learning model to obtain a face feature vector sequence.
[0089] Step S1321: extracting, by a feature segment division sub-layer of the face feature extraction layer, a face feature frame corresponding to each face dynamic feature in the face dynamic feature set and a shooting time stamp, and calculating a feature difference degree between adjacent two face feature frames by sum of squares of normalized inter-frame pixel differences.
[0090] The feature segment division sub-layer first extracts the face feature frame corresponding to each face dynamic feature in the face dynamic feature set and the shooting time stamp corresponding to the frame, and establishes the correspondence between the face feature frame and the time. Then, for the adjacent two face feature frames, the feature difference degree between them is calculated. In the calculation process, first, the two frames of images are normalized to eliminate the influence of image size or brightness difference, and then the sum of squares of the normalized inter-frame pixel differences is calculated, which is taken as the feature difference degree for measuring the difference degree between the adjacent face feature frames.
[0091] Step S1322: determining the feature change frequency according to the size of the feature difference degree, dividing the face feature frames in a corresponding time period into feature segments of a first preset time length interval when the continuous multiple feature difference degrees are all within a preset difference threshold interval, and dividing the face feature frames in the time period into feature segments of a second preset time length interval when the continuous multiple feature difference degrees are all outside the preset difference threshold interval, so that the number of face feature frames contained in each feature segment is kept within a preset frame number range, and the face feature frames less than the lower limit of the preset frame number range are merged with the feature frames of adjacent time periods, and the face feature frames exceeding the upper limit of the preset frame number range are split into multiple feature segments.
[0092] The variation frequency of the facial features is determined according to the calculated feature difference degree. A difference threshold interval is preset. When the feature difference degrees of the continuous multiple feature difference degrees are all within the interval, it is indicated that the facial features are relatively gentle in the time period, and the facial feature frames in the corresponding time period are divided into feature segments of a first preset time length interval with a relatively long time length. When the feature difference degrees of the continuous multiple feature difference degrees are all outside the preset difference threshold interval, it is indicated that the facial features are relatively intense in the time period, and the facial feature frames in the time period are divided into feature segments of a second preset time length interval with a relatively short time length. In the division process, it is necessary to ensure that the number of facial feature frames included in each feature segment is within a preset frame number range. If the number of frames included in the divided feature segment is less than a preset lower limit value, the feature segment is combined with the facial feature frames of the adjacent time period; if the number of frames exceeds a preset upper limit value, the feature segment is split into multiple feature segments to ensure the effectiveness of subsequent feature processing.
[0093] Step S1323: The feature point matching sub-layer of the facial feature extraction layer adopts a pre-trained facial feature point detection model to locate feature points in the first frame of each feature segment, and marks the coordinate positions of the eye feature points, facial muscle feature points and skin color feature points in the frame image.
[0094] The feature point matching sub-layer first processes the first frame of each feature segment. A pre-trained facial feature point detection model is adopted, which can automatically identify and locate key feature points in a facial image. In the frame image, eye feature points such as pupil edge points and eyelid contour points are marked; facial muscle feature points such as mouth corner contour points and cheek contour points are marked; skin color feature points such as forehead region points and cheek region points are marked, and the coordinate positions of these feature points in the image coordinate system are recorded.
[0095] Step S1324: For the subsequent facial feature frames in the feature segment, a light flow tracking algorithm is adopted to track the located feature points to determine the coordinate positions of each feature point in the subsequent frames, a template matching-based method is adopted to reposition the feature point loss in the tracking process, and a feature point coordinate mapping table of all facial feature frames in each feature segment is established.
[0096] For subsequent face feature frames after the first frame in the feature segment, a light flow tracking algorithm is used to track the feature points that have been located in the first frame. The light flow tracking algorithm determines the new coordinate positions of the feature points in the subsequent frames by calculating the motion vectors of the pixels between adjacent frames. When the feature points are lost due to occlusion, blur, or other reasons during the tracking process, a template matching-based method is used to reposition. The image in the neighborhood of the feature point before loss is taken as a template, and template matching is performed in the corresponding search region of the current frame to find the most similar region as the new position of the feature point. The coordinates of the feature points of all face feature frames in each feature segment are sorted in frame order to establish a feature point coordinate mapping table, which records the dynamic changes of the feature points in different frames.
[0097] Step S1325: The timing convolution network of the face feature extraction layer is called to convert the feature point coordinate mapping table of the feature segment into feature point motion trajectory data containing coordinate sequences changing over time, and the motion trajectory data is input into the convolution layers of the timing convolution network for convolution operation to extract motion features of different time scales.
[0098] The timing convolution network receives the feature point coordinate mapping table of the feature segment and converts it into feature point motion trajectory data, which records the sequence of changes of each feature point coordinate over time with time as the axis. The motion trajectory data is input into the timing convolution network, which contains multiple convolution layers, each of which uses a convolution kernel of different size. The first convolution layer uses a smaller size convolution kernel to perform convolution operation on the motion trajectory data to extract the local rapid motion features of the feature points; the middle convolution layer uses a medium size convolution kernel to extract the medium scale motion features of the feature points; and the last convolution layer uses a larger size convolution kernel to extract the overall motion trend features of the feature points.
[0099] Step S1326: The outputs of the convolution layers of the timing convolution network are pooled using the max pooling method to retain the key feature values in each trajectory feature, obtain the simplified trajectory features, calculate the total amount of position change of each feature point within the time length of the feature segment based on the trajectory features, divide the time length to obtain the average change rate, extract the coordinate difference value between the current frame and the previous frame of each time node as the instantaneous change rate and arrange it in time sequence to form the instantaneous change rate sequence, and combine the average change rate and the instantaneous change rate sequence to form the dynamic change rate feature.
[0100] The feature maps output by each convolutional layer of the time convolutional network are processed by a max-pooling method, that is, the maximum value in each preset-size pooling window is selected as the output of the window to retain key feature values in the trajectory features, realize feature dimension reduction and data simplification, and obtain simplified trajectory features. According to the simplified trajectory features, the total amount of position change of each feature point in the entire feature segment time length is calculated, and the average change rate of the feature point is obtained by dividing the total amount of position change by the time length. At the same time, at each time node, the difference between the current frame feature point coordinates and the previous frame feature point coordinates is calculated as the instantaneous change rate, and the instantaneous change rates of all time nodes are arranged in time sequence to form an instantaneous change rate sequence. The average change rate and the instantaneous change rate sequence are combined to form a dynamic change rate feature that can describe the speed and change trend of the feature point motion.
[0101] Step S1327: Analyzing the change correlation between different types of feature points, selecting the position change amount of the eye feature points and the position change amount of the facial muscle feature points in the same time interval to calculate the correlation coefficient of the two as the change synchronization rate, determining the time difference between the facial muscle feature point change and the skin color feature point change as the change lag time through cross-correlation analysis, and combining the change synchronization rate and the change lag time to form a feature correlation feature.
[0102] The change correlation between different types of feature points such as eye feature points, facial muscle feature points and skin color feature points is analyzed. In the same time interval, the position change amount of the eye feature points and the position change amount of the facial muscle feature points are calculated respectively, and the correlation coefficient between the two is calculated to measure the synchronization degree of their changes, and the correlation coefficient is the change synchronization rate. The cross-correlation analysis method is used to analyze the time delay relationship between the facial muscle feature point change sequence and the skin color feature point change sequence, and the time difference between the facial muscle feature point change and the skin color feature point change is determined as the change lag time. The change synchronization rate and the change lag time are combined to form a feature correlation feature that can reflect the dynamic correlation relationship between different types of feature points.
[0103] Step S1328: The feature fusion sub-layer of the face feature extraction layer unifies the dimensions of the dynamic change rate feature and the feature correlation feature, and uses a weighted fusion algorithm to weight and fuse the dynamic change rate feature and the feature correlation feature after dimension unification according to the pre-trained feature weight parameters to obtain a fusion feature vector of each feature segment. All fusion feature vectors are arranged in time sequence according to the time sequence of the feature segments in the face dynamic feature set to generate a face feature vector sequence.
[0104] The feature fusion sublayer first performs dimension unification processing on the dynamic change rate feature and the feature correlation feature, and converts the two into feature vectors of the same dimension. Then, a weighted fusion algorithm is used to perform weighted summation on the dynamic change rate feature and the feature correlation feature after dimension unification according to the feature weight parameters learned in the pre-training process, to obtain a fusion feature vector of each feature segment. The fusion feature vector comprehensively reflects the motion rate information of the feature points and the correlation information between different feature points. Finally, the fusion feature vectors of all feature segments are arranged in sequence according to the time sequence of the feature segments in the facial dynamic feature set, to form a facial feature vector sequence, which can reflect the overall dynamic change of the facial features over time.
[0105] Step S133: Calculate the normalized correlation coefficient of each feature dimension in the gas feature vector sequence and the facial feature vector sequence respectively as the first type of correlation parameter, calculate the matching degree of the feature dimension change rate as the second type of correlation parameter, generate the cross-correlation degree parameter after weighted fusion of the first type of correlation parameter and the second type of correlation parameter, and construct the space-time mapping matrix based on the cross-correlation degree parameter. According to the size of the cross-correlation degree parameter, the dynamic weight of each element in the space-time mapping matrix is allocated through the attention mechanism, the feature weight of the region where the cross-correlation degree parameter is in the preset correlation threshold interval is strengthened, the feature weight of the region where the cross-correlation degree parameter is outside the preset correlation threshold interval is weakened, and the weighted space-time mapping matrix is generated.
[0106] The lengths of the gas feature vector sequence and the facial feature vector sequence may be different. First, the two are adjusted to the same length through linear interpolation, and the adjusted length is set as N, the dimension of the gas feature vector is Dg, and the dimension of the facial feature vector is Df. When calculating the first type of correlation parameter, the normalized correlation coefficient of each dimension of the gas feature vector sequence and each dimension of the facial feature vector sequence is calculated, the normalized correlation coefficient is obtained by calculating the Pearson correlation coefficient and taking the absolute value, and a Dg×Df correlation coefficient matrix, i.e., the first type of correlation parameter matrix, is formed.
[0107] When calculating the second type of correlation parameter, first, difference operation is performed on each dimension of the gas feature vector sequence and the facial feature vector sequence to obtain a change rate sequence, and then the matching degree of each dimension combination of the gas change rate sequence and the facial change rate sequence is calculated. The matching degree is obtained by calculating the mutual information of the change rate sequence, and a Dg×Df matching degree matrix, i.e., the second type of correlation parameter matrix, is formed.
[0108] The first type of correlation parameter matrix and the second type of correlation parameter matrix are element-wise weighted summed, and the weight coefficient is determined through cross-validation in the training data to obtain a cross-correlation degree parameter matrix. The dimension of the space-time mapping matrix is N*N, and the matrix element represents the correlation strength of the gas feature vector sequence at time t and the face feature vector sequence at time t'. The correlation strength is initialized as the average value of the cross-correlation degree parameter matrix. The attention mechanism adjusts the weight of the space-time mapping matrix according to the element value of the cross-correlation degree parameter matrix. For the dimension pair with a cross-correlation degree parameter greater than a preset correlation threshold, the weight of the corresponding space-time position is increased; for the dimension pair with a cross-correlation degree parameter less than the preset correlation threshold, the weight of the corresponding space-time position is reduced. The adjusted space-time mapping matrix is the weighted space-time mapping matrix, and the matrix element value is between zero and one.
[0109] Step S134: According to the weighted space-time mapping matrix, bidirectional feature conduction is performed to establish the conduction path of the gas feature vector to the face feature vector and the face feature vector to the gas feature vector. In the first conduction path, the molecular structure response feature in the gas feature vector is conducted to the face feature vector to correct the dynamic change feature related to alcohol influence in the face feature. In the second conduction path, the feature correlation feature in the face feature vector is conducted to the gas feature vector to correct the concentration distribution feature related to the breathing state in the gas feature. The dynamic change feature and the concentration distribution feature after bidirectional conduction are analyzed to extract a cooperative feature. The cooperative feature is integrated with the key features in the original gas feature vector sequence and the face feature vector sequence to generate a fusion detection feature vector.
[0110] The weighted space-time mapping matrix is used to guide bidirectional feature conduction, and the conduction process is realized in the feature correlation layer. The first conduction path is the gas feature vector to the face feature vector conduction. First, the molecular structure response feature is extracted from the gas feature vector sequence. The molecular structure response feature corresponds to the response mode of the alcohol sensor in the gas sensor array, which is realized by selecting the feature dimension related to the alcohol molecular structure in the gas feature vector. The molecular structure response feature is multiplied by the weighted space-time mapping matrix to obtain a space-time weighted molecular structure response feature, which represents the influence intensity of the molecular structure response at different time points on each face feature point. The space-time weighted molecular structure response feature and the face feature vector sequence are element-wise added to realize the correction of the dynamic change feature related to alcohol influence, such as the slowing down of eye movement speed and the relaxation of facial muscles.
[0111] The second conduction path is from the face feature to the gas feature, and a feature correlation feature is extracted from the face feature vector sequence, which represents the cooperative change relationship between each dynamic feature of the face. The feature correlation feature is obtained by calculating the covariance matrix of the face feature vector sequence. The feature correlation feature is multiplied by the transpose of the weighted space-time mapping matrix to obtain a space-time weighted feature correlation feature, which represents the influence of the face feature correlation at different time points on the gas concentration distribution. The space-time weighted feature correlation feature and the gas feature vector sequence are added element by element to realize the correction of the concentration distribution features related to the respiratory state, such as the concentration fluctuation caused by the respiratory rhythm.
[0112] The face feature vector sequence and the gas feature vector sequence after bidirectional conduction are input into a cooperative analysis module, which extracts a cooperative feature using the mutual information maximization criterion. The cooperative feature represents the part that changes together in the two modal features. The dimension of the cooperative feature is a preset fixed value, and dimension reduction is realized through principal component analysis. The original gas feature vector sequence and the face feature vector sequence obtain a key feature vector through global average pooling. The key feature vector and the cooperative feature are spliced in the feature dimension to form a fusion detection feature vector, and the vector dimension is the sum of the cooperative feature dimension and the key feature vector dimension.
[0113] Step S135: Calculate the information entropy value and the feature dispersion of the fusion detection feature vector, compare the information entropy value with the preset information entropy threshold, compare the feature dispersion with the preset dispersion threshold, if the information entropy value is within the preset information entropy threshold interval and the feature dispersion is within the preset dispersion threshold interval, determine that the fusion detection feature vector is valid, if the information entropy value is outside the preset information entropy threshold interval or the feature dispersion is outside the preset dispersion threshold interval, trigger the feature reconstruction mechanism, return to adjust the weight distribution of the space-time mapping matrix, until an effective fusion detection feature vector is generated.
[0114] The effectiveness verification of the fusion detection feature vector is realized by calculating the information entropy value and the feature dispersion. The calculation of the information entropy value is based on the probability distribution estimation of the fusion detection feature vector. The kernel density estimation method is used to estimate the probability density function of each dimension of the feature vector, and then the entropy value of each dimension is calculated and summed to obtain the total information entropy value. The information entropy value reflects the uncertainty of the feature vector, and too high represents feature confusion and too low represents feature singleness.
[0115] The calculation of the feature dispersion degree uses the sum of the standard deviations of each dimension of the feature vector. The greater the standard deviation, the more dispersed the feature distribution, and the more information it contains. The smaller the standard deviation, the more concentrated the feature distribution, and there may be information redundancy. The preset information entropy threshold interval and the preset dispersion degree threshold interval are obtained by statistical analysis of the effective feature vectors in the training data. The 95% confidence interval of the information entropy values and the feature dispersion degrees in the training data is taken as the preset threshold interval.
[0116] If the information entropy value and the feature dispersion degree of the fusion detection feature vector are both within the preset threshold interval, it is determined to be valid, and the fusion detection feature vector is output. Otherwise, the feature reconstruction mechanism is triggered. The feature reconstruction mechanism adjusts the weight distribution of the weighted spatiotemporal mapping matrix, increasing the weight in the direction away from the threshold in the information entropy value and the feature dispersion degree. For example, if the information entropy value is too low, the weight of the weak correlation region is increased to introduce more changes. If the feature dispersion degree is too high, the weight of the strong correlation region is increased to enhance the consistency of the features. The adjusted weighted spatiotemporal mapping matrix is used for bidirectional feature conduction again, and steps S134 to S135 are repeated until a valid fusion detection feature vector is generated. The maximum number of iterations is preset to ten. If the iteration number exceeds the preset number, a warning message is output, and the fusion detection feature vector generated in the last iteration is used.
[0117] Step S140: Perform drunk driving judgment processing on the fusion detection feature vector to generate a comprehensive drunk driving judgment result containing detection confidence, feature correlation strength, and judgment basis.
[0118] The drunk driving judgment processing is implemented in the classification decision layer of the model, which includes a first-level judgment module and a second-level judgment module, adopting a two-level cascade structure. The first-level judgment module performs preliminary classification and outputs a suspected state category and a preliminary confidence. The second-level judgment module performs fine judgment on the suspected drunk driving state, calculates the detection confidence and the feature correlation strength. The judgment basis includes the state category, the confidence, the feature correlation strength, and the feature difference analysis result, which are output in the form of structured data.
[0119] Step S141: Input the fusion detection feature vector into the first-level judgment module, load the pre-trained classification model parameters through the pre-trained binary classifier included in the first-level judgment module, and convert the fusion detection feature vector into a feature point in the classification space.
[0120] The binary classifier of the first level decision module is a support vector machine-based classification model, and model parameters include support vectors, classification hyperplane coefficients, and bias terms, which are obtained by training a training data set and stored in a model parameter file. After the fusion detection feature vector is input into the binary classifier, it is first mapped to a high-dimensional classification space by a kernel function, and the kernel function adopts a radial basis function. The kernel function parameters are optimized and determined in the training process. The inner product operation is performed between the mapped feature vector and the support vector to obtain the coordinates of the feature points in the classification space, and the dimension of the feature points is equal to the number of support vectors.
[0121] Step S142: Calculate the distance between the feature point and the center of the normal state category and the center of the suspected drunk driving state category. If the distance between the feature point and the center of the normal state category is less than the distance between the feature point and the center of the suspected drunk driving state category, the suspected state category is determined to be the normal state. If the distance between the feature point and the center of the suspected drunk driving state category is less than the distance between the feature point and the center of the normal state category, the suspected state category is determined to be the suspected drunk driving state. Meanwhile, the normalized distance ratio of the feature point to the corresponding category center is calculated as the confidence parameter of the first level decision.
[0122] The category center in the classification space is calculated based on the category samples in the training data. The center of the normal state category is the mean vector of the mapped feature points of all normal samples, and the center of the suspected drunk driving state category is the mean vector of the mapped feature points of all drunk driving samples. The distance between the feature point and the category center is calculated by using the Euclidean distance. The smaller the distance value is, the closer the feature point is to the category.
[0123] The determination of the suspected state category is realized by comparing two distance values. If the distance between the feature point and the center of the normal state category is less than the distance between the feature point and the center of the suspected drunk driving state category, the suspected state category is determined to be the normal state. Otherwise, the suspected state category is determined to be the suspected drunk driving state. The confidence parameter of the first level decision is calculated as 1 minus the distance between the feature point and the corresponding category center divided by the distance between the two category centers. The closer the value is to 1, the more reliable the determination is. The closer the value is to 0, the more ambiguous the determination is.
[0124] Step S143: When the suspected state category is the normal state, integrate the suspected state category and the confidence parameter of the first level decision to generate a preliminary drunk driving decision result. If the confidence parameter of the first level decision is within a preset normal confidence threshold interval, directly integrate the suspected state category, the confidence parameter of the first level decision, and the basic correlation degree between the gas feature component and the face feature component as the feature correlation strength to generate a drunk driving comprehensive decision result. If the confidence parameter of the first level decision is outside the preset normal confidence threshold interval, trigger the re-determination mechanism to return to the feature correlation layer to re-generate the fusion detection feature vector and perform two-level hierarchical decision again until a determination result with a confidence parameter within the preset normal confidence threshold interval is obtained.
[0125] The preliminary drunk driving determination result is a dictionary structure containing a suspected state category and a confidence parameter. The preset normal confidence threshold interval is set according to the confidence distribution of normal samples in the training data, and is usually 0.7 to 1.0. The basic correlation degree of the gas feature component and the face feature component is obtained by calculating the cosine similarity of the gas feature part and the face feature part in the fusion detection feature vector, which reflects the consistency of the two modal features.
[0126] When the confidence parameter is within the preset normal confidence threshold interval, the feature correlation strength is the basic correlation degree, the drunk driving comprehensive determination result contains the suspected state category, the confidence parameter, the feature correlation strength and the determination basis, and the determination basis explains the reason for determining the normal state, such as “the alcohol concentration in the gas feature is lower than the threshold, and the face feature has no abnormal dynamics”. When the confidence parameter is lower than the lower limit of the preset normal confidence threshold interval, the re-determination mechanism is triggered, which sends a signal to the feature correlation layer to adjust the weight parameter of the attention mechanism, increases the weight of the gas feature component, and then regenerates the fusion detection feature vector. Repeat steps S133 to S143, and the maximum number of re-determination is three. If the confidence requirement cannot be met, a low-confidence warning is output, and the current determination result is used.
[0127] Step S144: When the suspected state category is a suspected drunk driving state, the fusion detection feature vector and the confidence parameter of the first-level determination are input into the second-level determination module to calculate the detection confidence.
[0128] Step S1441: The fusion detection feature vector and the confidence parameter of the first-level determination are input into the second-level determination module, and the second-level determination module loads a preset drunk driving feature template library. The drunk driving feature template library is generated by training a plurality of standard fusion feature templates, and each standard fusion feature template contains a template identifier, a drunk driving degree category corresponding to the standard fusion feature template, a template feature vector, and a feature dimension weight table.
[0129] After receiving the fusion detection feature vector and the confidence parameter of the first-level determination, the second-level determination module first loads the preset drunk driving feature template library. The drunk driving feature template library is generated by feature extraction and training on a large number of standard samples labeled with a drunk driving degree category, and contains a plurality of standard fusion feature templates. Each standard fusion feature template has a unique template identifier and corresponds to a specific drunk driving degree category, such as no drunk driving, slight drunk driving, moderate drunk driving, and severe drunk driving. At the same time, each template also contains a template feature vector, which is obtained by feature fusion on the standard sample and can represent the typical features of the drunk driving degree category, and a feature dimension weight table, which records the importance weight of each feature dimension in the matching process of the standard fusion feature template.
[0130] Step S1442: The fusion detection feature vector is compared with each standard fusion feature template in the drunk driving feature template library one by one, the template feature vector and the feature dimension weight table of the standard fusion feature template are extracted, the normalized absolute difference value of the fusion detection feature vector and the template feature vector in each feature dimension is calculated, the normalized absolute difference value of each feature dimension is weighted according to the weight value in the feature dimension weight table, the weighted difference values of all feature dimensions are counted and the sum is obtained to obtain the total weighted difference value, and the total weighted difference value is converted into a matching degree parameter.
[0131] The second level judgment module compares the input fusion detection feature vector with each standard fusion feature template in the drunk driving feature template library one by one. For each standard fusion feature template, its template feature vector and feature dimension weight table are extracted. The absolute difference value of the fusion detection feature vector and the template feature vector in each feature dimension is calculated, and the absolute difference value is normalized to eliminate the influence of different feature dimension scales. According to the weight value of each feature dimension in the feature dimension weight table, the corresponding normalized absolute difference value is weighted, that is, multiplied by the weight value. Then the weighted difference values of all feature dimensions are accumulated to obtain the total weighted difference value. The smaller the total weighted difference value is, the more similar the fusion detection feature vector is to the standard fusion feature template. The total weighted difference value is converted into a matching degree parameter through a preset conversion function, and the larger the matching degree parameter is, the higher the matching degree is.
[0132] Step S1443: The matching degree parameters of all standard fusion feature templates are sorted, the standard fusion feature template with the largest matching degree parameter is selected as the optimal matching template, the template identification and the corresponding drunk driving degree category of the optimal matching template are recorded, the feature dimension weight table of the optimal matching template is extracted, the number of feature dimensions in the feature dimension difference value between the complete fusion detection feature vector and the optimal matching template that are in the preset difference threshold interval is counted, the proportion of the feature dimension number in the total feature dimension number is calculated to obtain an auxiliary matching degree parameter, if the auxiliary matching degree parameter is outside the preset auxiliary threshold interval, the standard fusion feature template with the second largest matching degree parameter is selected as a candidate optimal matching template and the auxiliary matching degree parameter is repeatedly calculated until the standard fusion feature template with the auxiliary matching degree parameter in the preset auxiliary threshold interval is found as the optimal matching template.
[0133] The matching degree parameters of all standard fusion feature templates are sorted in descending order, the standard fusion feature template with the largest matching degree parameter is preliminarily determined as the optimal matching template, and the template identifier and the corresponding drunk driving degree category are recorded. Then the feature dimension weight table of the optimal matching template is extracted, the feature dimension difference values of the fusion detection feature vector and the template feature vector of the optimal matching template in each feature dimension are calculated, and it is judged whether the difference values are within the preset difference threshold interval. The number of feature dimensions within the interval is counted, which is divided by the total number of feature dimensions to obtain an auxiliary matching degree parameter, which is used to measure the overall matching of the fusion detection feature vector and the optimal matching template in the feature dimension. If the auxiliary matching degree parameter is within the preset auxiliary threshold interval, the template is confirmed as the optimal matching template; if it is outside the interval, the standard fusion feature template with the second largest matching degree parameter is selected as the candidate optimal matching template, and the above process of calculating the auxiliary matching degree parameter is repeated until the optimal matching template with the auxiliary matching degree parameter meeting the requirements is found, so as to improve the reliability of the matching result.
[0134] Step S1444: The confidence degree parameter of the first level determination, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template are normalized, and the final detection confidence is calculated by using the product correction method, and the variance of the confidence degree parameter, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template is calculated, if the variance is within the preset variance threshold interval, the initial product result is directly used as the detection confidence; if the variance is outside the preset variance threshold interval, the weight adjustment mechanism is triggered, different correction weights are assigned to the confidence degree parameter, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template according to the pre-trained parameter weight model, and the detection confidence is recalculated.
[0135] First, the confidence parameter of the first level determination, the core matching degree parameter of the optimal matching template (i.e. the matching degree parameter obtained in step S1442) and the auxiliary matching degree parameter are normalized to compress them to the interval [0, 1] to ensure that the parameters are operated in the same dimension. Then, the product correction method is used to multiply the three normalized parameters to obtain the initial detection confidence. At the same time, the variance of the three parameters is calculated, which is used to measure the dispersion degree of the values of the three parameters. If the variance is within the preset variance threshold interval, it means that the values of the three parameters are consistent and mutually confirmatory, and the initial product result is directly taken as the final detection confidence. If the variance is outside the preset variance threshold interval, it means that the dispersion degree of the values of the three parameters is large, and there may be some unreliable parameters, at which time the weight adjustment mechanism is triggered. According to the pre-trained parameter weight model, the parameter weight model learns the importance of each parameter under different conditions according to historical data, and assigns different correction weights to the confidence parameter, the core matching degree parameter and the auxiliary matching degree parameter. Then, the normalized parameters are multiplied by the corresponding correction weights to obtain the adjusted detection confidence.
[0136] Step S145: Extract the gas feature components corresponding to the key features of the original gas feature vector sequence and the face feature components corresponding to the key features of the original face feature vector sequence in the fusion detection feature vector, and calculate the mutual influence coefficient of the gas feature components and the face feature components obtained through the trace value of the covariance matrix as the feature correlation strength.
[0137] The extraction of the gas feature components and the face feature components is realized through feature segmentation. The first half of the fusion detection feature vector is the gas feature components, which correspond to the key features of the original gas feature vector sequence; and the second half is the face feature components, which correspond to the key features of the original face feature vector sequence. The length of the feature components is pre-set during model training.
[0138] The calculation of the mutual influence coefficient is realized by constructing the covariance matrix of the gas feature components and the face feature components. The size of the covariance matrix is (gas feature component dimension + face feature component dimension) x (gas feature component dimension + face feature component dimension), and the trace value of the matrix is the sum of the diagonal elements, which reflects the overall correlation of the two feature components. The trace value of the covariance matrix is normalized to the range of zero to one to obtain the mutual influence coefficient, i.e. the feature correlation strength. The larger the value is, the stronger the mutual influence of the two feature components is, and the more reliable the determination result is.
[0139] Step S146: integrate the suspected state category, detection confidence, feature correlation strength, optimal matching template identifier and feature dimension difference distribution to form a judgment basis. If the detection confidence is within the preset alcohol driving confidence threshold interval, integrate the suspected state category, detection confidence, feature correlation strength and judgment basis to generate an alcohol driving comprehensive judgment result. If the detection confidence is outside the preset alcohol driving confidence threshold interval, trigger a supplementary collection mechanism to supplement the collection of the respiratory gas data and facial dynamic feature data of the to-be-detected object by the multi-modal collection unit of the robot system, regenerate the fusion detection feature vector, and then perform two-level hierarchical judgment again until the detection confidence is within the preset alcohol driving confidence threshold interval.
[0140] The optimal matching template identifier is obtained by comparing the fusion detection feature vector with a preset alcohol driving feature template library. The alcohol driving feature template library contains standard feature vectors of different alcohol driving degree categories. The Euclidean distance is used for matching. The template with the smallest distance is the optimal matching template, and its identifier is recorded. The feature dimension difference distribution is the difference between each feature dimension of the fusion detection feature vector and the optimal matching template. The difference vector is arranged in dimension order to reflect the deviation of the current feature from the standard template.
[0141] The judgment basis is a JSON object containing the suspected state category, detection confidence, feature correlation strength, optimal matching template identifier and feature dimension difference distribution. The preset alcohol driving confidence threshold interval is set according to the legal alcohol concentration threshold. The specific conversion method is as follows: 1. Establish a mapping relationship between alcohol concentration and model confidence: Through the 100,000+ sets of labeled data used in the above model training process, extract the corresponding relationship between the alcohol concentration value C (mg / 100ml) of each set of data and the detection confidence P output by the model, and use the Platt scaling method for probability calibration to fit the mapping function P=f(C), where f(C) is an S-shaped function: f(C)=1 / (1+e^(-aC+b)), the fitting parameters are a=0.032 and b=-0.64 (fitting error R 2 =0.96, ensuring mapping accuracy). 2. Determine the confidence threshold interval: According to the relevant provisions, alcohol concentration C<20mg / 100ml is not alcohol driving, 20mg / 100ml≤C<80mg / 100ml is alcohol driving, and C≥80mg / 100ml is drunk driving; Substitute the legal threshold into the mapping function to calculate the corresponding confidence: when C=20mg / 100ml, P=f(20)=0.6; when C=80mg / 100ml, P=f(80)=0.8; therefore, set the preset alcohol driving confidence threshold interval: alcohol driving (20-80mg / 100ml) corresponds to 0.6-0.8, drunk driving (≥80mg / 100ml) corresponds to 0.8-1.0, and not alcohol driving (<20mg / 100ml) corresponds to <0.6.
[0142] When the detection confidence is within the preset drunk driving confidence threshold interval, the drunk driving comprehensive determination result includes a suspected state category (refined into non-drunk driving / drunk driving / drunk driving), detection confidence, feature correlation strength, and determination basis. When the detection confidence is lower than the lower limit of the preset drunk driving confidence threshold interval (for example, the determination is drunk driving but the confidence <0.6), the supplementary collection mechanism is triggered. The supplementary collection mechanism issues a supplementary collection instruction through the guidance execution component of the robot system, the voice broadcast module prompts the to-be-detected object to "please blow again and look at the camera", and at the same time, the multi-modal collection unit starts a new round of collection, and the collection time is half of the initial collection time (the initial collection time is 10s, and the supplementary collection time is 5s, which balances efficiency and data volume). After the collection data are processed through steps S125 to S128, the collection data are merged with the original data set (using data splicing, the time stamp continuity is preserved), a fusion detection feature vector is regenerated, steps S133 to S146 are repeated, and the maximum number of supplementary collection is twice. If the confidence requirement cannot be met, a high-risk warning is output, and manual review is recommended (the manual review process includes reusing a portable alcohol detector to detect, and viewing a dynamic video recording of a face).
[0143] Step S150: According to the drunk driving comprehensive determination result, a closed-loop processing flow of the detection scene is triggered, the closed-loop processing flow includes encrypted result recording, multi-terminal information synchronization, and detection scene resetting, and the closed-loop processing flow needs to be driven by the output result of the drunk driving detection deep learning model to drive the robot component to act cooperatively.
[0144] The closed-loop processing flow is executed by the central control unit of the robot system, and after the central control unit receives the drunk driving comprehensive determination result, the corresponding processing flow is triggered according to the type of the drunk driving comprehensive determination result. The encrypted result recording ensures the security and non-tamperability of the detection data; the multi-terminal information synchronization realizes real-time sharing of the detection result; and the detection scene resetting restores the system to the initial state, preparing for the next detection. The state machine controls between the flows to ensure sequential execution.
[0145] Based on the same inventive concept, please refer to Figure 2 , a structure schematic block diagram of a drunk driving automatic detection system 100 based on a deep learning model and a robot system for executing the above-mentioned drunk driving automatic detection method based on a deep learning model and a robot system is shown, the drunk driving automatic detection system 100 based on the deep learning model and the robot system can include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0146] The machine readable storage medium 120 is configured to store machine executable instructions for implementing the scheme of the present application, and the processor 130 is configured to execute the machine executable instructions stored in the machine readable storage medium 120 to implement the deep learning model-based and robot system-based drunk driving automatic detection method provided by the foregoing method embodiment.
[0147] It should be noted that, in order to simplify the expression of the present disclosure and to help the understanding of one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A drunk driving automatic detection method based on a deep learning model and a robot system, characterized by, The method comprises: A detection scene dynamic interaction framework is constructed by a scene interaction module of the robot system, and the detection scene dynamic interaction framework comprises behavior guidance logic of a to-be-detected object, time sequence coordination rules of multi-source information collection, and an environment parameter self-adaptive adjustment mechanism; Based on the detection scene dynamic interaction framework, a multi-modal collection unit of the robot system is started, and according to the time sequence coordination rules, respiratory gas data flow and facial dynamic feature flow of the to-be-detected object are collected, and through the environment parameter self-adaptive adjustment mechanism, detection environment interference is compensated, so as to obtain a synchronous and associated respiratory gas data set and facial dynamic feature set after interference correction. The breath gas data set and the facial dynamic feature set are input into a pre-trained drunk driving detection deep learning model, a space-time mapping relationship of the breath gas data set and the facial dynamic feature set is established, a fusion detection feature vector is generated through bidirectional feature conduction and feature interaction processing, and the fusion feature effectiveness is verified, specifically including: inputting the breath gas data set into a gas feature extraction layer of the drunk driving detection deep learning model to obtain a gas feature vector sequence; inputting the breath gas data set into a facial feature extraction layer of the drunk driving detection deep learning model to obtain a facial feature vector sequence; calculating a normalized correlation coefficient of each feature dimension in the gas feature vector sequence and the facial feature vector sequence as a first type of correlation parameter, and calculating a matching degree of a feature dimension change rate as a second type of correlation parameter, weighting and fusing the first type of correlation parameter and the second type of correlation parameter to generate a cross-correlation degree parameter, and constructing a space-time mapping matrix based on the cross-correlation degree parameter, assigning a dynamic weight to each element in the space-time mapping matrix according to the size of the cross-correlation degree parameter through an attention mechanism, strengthening the feature weight of the region where the cross-correlation degree parameter is in a preset correlation threshold interval, and weakening the feature weight of the region where the cross-correlation degree parameter is outside the preset correlation threshold interval, to generate a weighted space-time mapping matrix; performing bidirectional feature conduction according to the weighted space-time mapping matrix, establishing a conduction path of the gas feature vector to the facial feature vector and the facial feature vector to the gas feature vector, and in the first conduction path, conducting a molecular structure response feature in the gas feature vector to the facial feature vector to correct a dynamic change feature related to alcohol influence in the facial feature vector, and in the second conduction path, conducting a feature correlation feature in the facial feature vector to the gas feature vector to correct a concentration distribution feature related to a breath state in the gas feature vector, and performing collaborative analysis on the dynamic change feature and the concentration distribution feature after bidirectional conduction to extract a collaborative feature, integrating the collaborative feature with key features in the original gas feature vector sequence and the facial feature vector sequence to generate a fusion detection feature vector; calculating an information entropy value and a feature dispersion degree of the fusion detection feature vector, comparing the information entropy value with a preset information entropy threshold value and comparing the feature dispersion degree with a preset dispersion threshold value, if the information entropy value is in a preset information entropy threshold interval and the feature dispersion degree is in a preset dispersion threshold interval, determining that the fusion detection feature vector is effective, if the information entropy value is outside the preset information entropy threshold interval or the feature dispersion degree is outside the preset dispersion threshold interval, triggering a feature reconstruction mechanism, returning to adjust the weight distribution of the space-time mapping matrix, and generating an effective fusion detection feature vector until the effective fusion detection feature vector is generated; drunk driving judgment processing is performed on the fusion detection feature vector to generate a drunk driving comprehensive judgment result containing a detection confidence, a feature correlation strength and a judgment basis; based on the drunk driving comprehensive judgment result, a closed-loop processing flow of the detection scene is triggered, the closed-loop processing flow includes an encryption result record, a multi-terminal information synchronization and a detection scene reset, and the closed-loop processing flow needs to drive a robot component to act cooperatively based on the output result of the drunk driving detection deep learning model. 2.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 1, wherein, The scene interaction module of the robot system constructs a detection scene dynamic interaction framework, comprising: The environment perception unit of the scene interaction module detects the space parameters and environmental interference parameters in the detection area through a multi-type sensor array, the space parameters include the three-dimensional boundary coordinates of the detection area, the device layout position, and the detection point space size, and the environmental interference parameters include the light intensity distribution, the air flow rate, the environmental temperature change, and the background noise intensity; Based on the space parameters, a three-dimensional modeling tool is used to construct a basic space model of the detection scene, which divides the activity prohibited area, the guide path area, and the detection core area of the object to be detected, and labels the effective working range boundary of each component of the multi-modal acquisition unit; The behavior guidance logic is constructed by combining the basic space model and the behavior habit template of the object to be detected, which includes the stage path planning from the detection area entrance to the detection core area, the posture keeping specification in the detection core area, the interaction timing requirement with the acquisition device, and the abnormal behavior guidance correction rule; Through multiple pre-acquisition experiments, the startup delay parameters, data acquisition period parameters, data transmission delay parameters, and acquisition accuracy fluctuation range of the gas acquisition component and the image acquisition component in the multi-modal acquisition unit are recorded to establish a component working characteristic database, based on which the timing coordination rules including the startup trigger time difference setting of the gas acquisition component and the image acquisition component, the acquisition cycle synchronization calibration mechanism, and the data transmission timing alignment requirement are generated to control the acquisition synchronization error of the respiratory gas data stream and the facial dynamic feature stream in the time dimension within the preset synchronization error range; According to the environmental interference parameters collected by the environment perception unit, an environmental parameter adaptive adjustment mechanism is constructed, which includes the light intensity compensation rule, the air flow interference suppression rule, the temperature influence correction rule, and the background noise filtering rule; The mutual influence relationship among the behavior guidance logic, the timing coordination rule, and the environmental parameter adaptive adjustment mechanism is established to generate a detection scene dynamic interaction framework that integrates space constraints, time constraints, and environmental constraints. 3.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 1, wherein, Based on the detection scene dynamic interaction framework, the multi-modal acquisition unit of the robot system is started, and the respiratory gas data stream and the facial dynamic feature stream of the object to be detected are acquired according to the timing coordination rule, and the detection environment interference is compensated through the environmental parameter adaptive adjustment mechanism to obtain the synchronized and correlated respiratory gas data set and facial dynamic feature set after interference correction, comprising: According to the behavior guidance logic in the detection scene dynamic interaction framework, the initial path guidance signal is sent at the detection area entrance through the guidance execution component of the robot system, which includes the light-on sequence of the ground embedded visual guidance light belt and the stage prompt instruction of the voice broadcast module, used to guide the object to be detected to move to the preset detection point of the detection core area. In the process of moving the to-be-detected object, the infrared sensor deployed by the guiding execution component along the path acquires the position information of the to-be-detected object in real time, and the lighting area of the visual guide light belt and the content of the voice prompt instruction are dynamically adjusted in combination with the path planning in the basic space model, so that the to-be-detected object moves along the preset path; When the to-be-detected object reaches the preset detection point, the guiding execution component sends a posture adjustment signal according to the posture keeping specification in the behavior guiding logic, and simultaneously starts the visual acquisition unit of the robot system to capture the current posture image of the to-be-detected object, compares the current posture image with the preset standard posture template at the pixel level, respectively calculates the head angle deviation, body posture deviation and distance deviation from the acquisition device, respectively standardizes and maps them to the same scale range, and then weightedly sums them to obtain a posture deviation parameter, adjusts the refinement degree of the voice prompt instruction according to the posture deviation parameter, continuously captures the current posture image and compares it, until the posture deviation parameter of the to-be-detected object is within the preset deviation threshold interval, and the posture adjustment is completed. The synchronous triggering module of the multi-modal acquisition unit is started, the time sequence coordination rule in the detection scene dynamic interaction framework is called, the starting trigger time difference and the acquisition cycle parameters of the gas acquisition component and the image acquisition component are obtained, and the gas acquisition component and the image acquisition component are synchronously triggered to start. In the acquisition process, the working states of the gas acquisition component and the image acquisition component are monitored in real time, the timestamps of the acquisition data are calibrated through data transmission time sequence alignment requirements, so that the respiratory gas data and the facial dynamic feature data at the same acquisition time have the same time identifier. The environment parameter self-adaptive adjustment mechanism in the detection scene dynamic interaction framework is called, and the sampling flow of the gas acquisition component and the exposure parameter of the image acquisition component are dynamically adjusted according to the real-time collected environmental interference parameters: when the light intensity is outside the preset light threshold interval, the light intensity of the image acquisition component is adjusted; when the air flow rate is outside the preset flow rate threshold interval, the sampling port of the gas acquisition component is adjusted and the sampling time is adjusted; The respiratory gas samples of the to-be-detected object are continuously acquired by the gas acquisition component through the gas sensor array according to the adjusted sampling parameters, and the gas component response value, sampling flow parameter and environmental temperature parameter at each acquisition time are arranged in time sequence to form an initial respiratory gas data stream; The facial images of the to-be-detected object are continuously shot by the image acquisition component through the multi-angle camera according to the adjusted exposure parameters, the eye dynamic features, facial muscle movement features and skin color change features in each facial image are extracted, and the light intensity parameter at the shooting time is recorded, and the initial facial dynamic feature stream is formed by combining in time sequence. The data of the same time node in the initial respiratory gas data stream and the initial face dynamic feature stream are associated and marked based on timestamp information of a collection time, to generate an associated data pair, the initial respiratory gas data in the associated data pair is removed from temperature variation interference according to a temperature influence correction rule in an environmental parameter adaptive adjustment mechanism, to obtain corrected respiratory gas data, the initial face dynamic feature data in the associated data pair is removed from illumination variation interference according to an illumination intensity compensation rule, to obtain corrected face dynamic feature data, all the corrected respiratory gas data are integrated in time sequence to form a set of respiratory gas data that are synchronous, associated and interference-corrected, and all the corrected face dynamic feature data are integrated in time sequence to form a set of face dynamic feature data that are synchronous, associated and interference-corrected. 4.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 1, wherein, The gas feature extraction layer of the breath alcohol detection deep learning model receives the set of respiratory gas data, extracts the gas component response value, the sampling flow parameter and the environmental temperature parameter in each respiratory gas data through the preprocessing sublayer of the gas feature extraction layer; The abnormal response values exceeding the range of the gas sensor are removed from the gas component response values, the mean filling method is used to fill in the data missing points of the remaining response values to obtain the preliminary processed gas component response values, the response values under different sampling flows are converted into the response values under the standard sampling flow based on the sampling flow parameter and the flow-response characteristic curve of the gas sensor, and the response value offset caused by the temperature variation is corrected based on the temperature-response characteristic model of the gas sensor to obtain the temperature-compensated gas component response values; The temperature-compensated gas component response values are normalized and mapped to obtain the normalized gas component response values, and the normalized gas component response values, the corrected sampling flow parameter and the compensated environmental temperature parameter are combined to form the standardized respiratory gas data; The convolution filtering unit of the gas feature extraction layer is called to load the pre-trained multi-scale convolution kernel parameters including the first size convolution window, the second size convolution window and the third size convolution window, each convolution window corresponds to different convolution kernel weights, the normalized gas component response values in the standardized respiratory gas data are arranged in time sequence into a one-dimensional data sequence, and are input into the three different size convolution windows for convolution operation respectively; The local convolution is performed on the one-dimensional data sequence through the first size convolution window to capture the local peak value features and output the local peak value feature map, the medium range convolution is performed on the one-dimensional data sequence through the second size convolution window to capture the slope change features between the local peak values and output the slope change feature map, and the larger range convolution is performed on the one-dimensional data sequence through the third size convolution window to capture the overall concentration distribution trend features and output the trend feature map, and the local peak value feature map, the slope change feature map and the trend feature map are spliced in the channel dimension to obtain the gas original features. The gas original feature is input into a nonlinear transformation sublayer of a gas feature extraction layer, a pre-trained activation function is called to perform nonlinear transformation on each feature value in the gas original feature, a response signal related to an alcohol molecule feature is strengthened, a response signal related to an interfering gas component is suppressed, and a strengthened gas feature is obtained; The strengthened gas feature is subjected to feature dimension analysis by a feature integration sublayer of the gas feature extraction layer, a feature dimension related to a molecular structure and a feature dimension related to a concentration distribution are identified, a feature peak position, a feature peak width and a feature peak intensity are extracted from the feature dimension related to the molecular structure to form a molecular structure response feature, and a concentration uniformity parameter and a concentration gradient parameter are calculated from the feature dimension related to the concentration distribution to form a concentration distribution feature; The molecular structure response feature and the concentration distribution feature are arranged in a preset dimension sequence to form a gas feature vector of a single breath gas data, and the gas feature vectors of all the single breath gas data are sequentially arranged in a collection time sequence of the breath gas data in a breath gas data set to generate a gas feature vector sequence. 5.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 1, wherein, The face feature extraction layer of the drunk driving detection deep learning model is used to input the breath gas data set, and a face feature vector sequence is obtained, including: Each face dynamic feature corresponding to a face feature frame and a shooting time stamp in a face dynamic feature set is extracted by a feature segment division sublayer of the face feature extraction layer, and a feature difference degree obtained by squaring a normalized interframe pixel difference between two adjacent face feature frames is calculated; A feature change frequency is determined according to the size of the feature difference degree, face feature frames in a corresponding time period are divided into feature segments of a first preset time length interval when a plurality of continuous feature difference degrees are in a preset difference threshold interval, and the face feature frames in the time period are divided into feature segments of a second preset time length interval when the plurality of continuous feature difference degrees are out of the preset difference threshold interval, so that the number of face feature frames contained in each feature segment is kept within a preset frame number range, and the face feature frames are merged with adjacent time period feature frames when the number of face feature frames is less than a lower limit value of the preset frame number range, and the face feature frames are split into a plurality of feature segments when the number of face feature frames exceeds an upper limit value of the preset frame number range; A pre-trained face feature point detection model is used to locate feature points in a first frame face feature frame in each feature segment by a feature point matching sublayer of the face feature extraction layer, and coordinates of eye feature points, face muscle feature points and skin color feature points are marked in the frame image, the eye feature points include pupil edge points containing a preset number of points uniformly distributed around a pupil contour and eyelid contour points of upper and lower eyelid contour features, the face muscle feature points include mouth corner contour points containing points on both sides of a mouth corner and upper and lower lip edge points and cheek contour points containing contour feature points on both sides of a cheek, and the skin color feature points include forehead region points containing a center point and surrounding points of a forehead and cheek region points containing center points and surrounding points of left and right cheeks. For the subsequent face feature frames in the feature segment, a optical flow tracking algorithm is used to track the located feature points to determine the coordinate positions of each feature point in the subsequent frames, a feature point prediction algorithm based on the motion trajectory of the feature points in the previous N frames is used to supplement the lost feature point coordinates in the tracking process, the deviation between the tracked feature point coordinates and the predicted feature point coordinates is calculated, if the deviation is within a preset tracking deviation threshold interval, the tracked coordinates are used, if the deviation is outside the preset tracking deviation threshold interval, the feature point detection model is called again to locate the feature points, and a feature point coordinate mapping table of all face feature frames in each feature segment is established; The feature point coordinate mapping table of the feature segment is converted into feature point motion trajectory data containing coordinate sequences changing over time by calling the time sequence convolution network of the face feature extraction layer, the motion trajectory data is input into the first convolution layer of the time sequence convolution network, a first size convolution kernel is used for convolution operation to extract the local curvature change feature of the trajectory, the output of the first convolution layer is input into the second convolution layer, a second size convolution kernel is used for convolution operation to extract the displacement distance feature of the trajectory, and the output of the second convolution layer is input into the third convolution layer, a third size convolution kernel is used for convolution operation to extract the motion direction feature of the trajectory; The output of the third convolution layer is processed by a max-pooling method to retain the key feature values in each trajectory feature, and a simplified trajectory feature is obtained, the total position change of each feature point in the time length of the feature segment is calculated based on the trajectory feature, and the average change rate is obtained by dividing the time length, the coordinate difference between the current frame and the previous frame of each time node is extracted as the instantaneous change rate and arranged in time sequence to form an instantaneous change rate sequence, and the average change rate and the instantaneous change rate sequence are combined to form a dynamic change rate feature; The change correlation between different types of feature points is analyzed, the position change amount of the eye feature points and the position change amount of the facial muscle feature points in the same time interval are selected to calculate the correlation coefficient of the two as the change synchronization rate, the time difference between the change of the facial muscle feature points and the change of the skin color feature points is determined as the change lag time through cross-correlation analysis, and the change synchronization rate and the change lag time are combined to form a feature correlation feature; The dynamic change rate feature and the feature correlation feature are dimensionally unified through the feature fusion sub-layer of the face feature extraction layer, the dynamic change rate feature and the feature correlation feature after dimensional unification are weighted and fused by using a weighted fusion algorithm according to pre-trained feature weight parameters to obtain a fusion feature vector of each feature segment, and all the fusion feature vectors are arranged in sequence according to the time sequence of the feature segments in the face dynamic feature set to generate a face feature vector sequence. 6.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 1, wherein, The fusion detection feature vector is subjected to drunk driving judgment processing to generate a drunk driving comprehensive judgment result containing a detection confidence, a feature correlation strength and a judgment basis, including: The fusion detection feature vector is input into a first-level judgment module, the pre-trained classifier included in the first-level judgment module loads the pre-trained classification model parameters, the fusion detection feature vector is converted into a feature point in the classification space through feature mapping, The distance between the feature point and the center of the normal state category and the center of the suspected drunk driving state category in the classification space is calculated. If the distance between the feature point and the center of the normal state category is less than the distance between the feature point and the center of the suspected drunk driving state category, the suspected state category is determined to be the normal state. If the distance between the feature point and the center of the suspected drunk driving state category is less than the distance between the feature point and the center of the normal state category, the suspected state category is determined to be the suspected drunk driving state. Meanwhile, the normalized distance ratio of the feature point to the corresponding category center is calculated as the confidence parameter of the first-level determination. When the suspected state category is the normal state, the suspected state category and the confidence parameter of the first-level determination are integrated to generate a preliminary drunk driving determination result. If the confidence parameter of the first-level determination is within a preset normal confidence threshold interval, the suspected state category, the confidence parameter of the first-level determination, and the basic correlation degree between the gas feature component and the face feature component are directly integrated as the feature correlation strength to generate a drunk driving comprehensive determination result. If the confidence parameter of the first-level determination is outside the preset normal confidence threshold interval, a re-determination mechanism is triggered to return to the feature correlation layer to re-generate a fusion detection feature vector, and two-level hierarchical determinations are performed again until a determination result with a confidence parameter within the preset normal confidence threshold interval is obtained. When the suspected state category is the suspected drunk driving state, the fusion detection feature vector and the confidence parameter of the first-level determination are input into a second-level determination module to calculate a detection confidence. The gas feature component corresponding to the key feature of the original gas feature vector sequence and the face feature component corresponding to the key feature of the original face feature vector sequence in the fusion detection feature vector are extracted. The mutual influence coefficient obtained by the covariance matrix trace value of the gas feature component and the face feature component is taken as the feature correlation strength. The suspected state category, the detection confidence, the feature correlation strength, the optimal matching template identifier, and the feature dimension difference distribution are integrated to form a determination basis. If the detection confidence is within a preset drunk driving confidence threshold interval, the suspected state category, the detection confidence, the feature correlation strength, and the determination basis are integrated to generate a drunk driving comprehensive determination result. If the detection confidence is outside the preset drunk driving confidence threshold interval, a supplementary collection mechanism is triggered to supplement the collection of the breath gas data and the face dynamic feature data of the to-be-detected object through the multi-modal collection unit of the robot system, and the fusion detection feature vector is re-generated to perform two-level hierarchical determinations again until a determination result with a detection confidence within the preset drunk driving confidence threshold interval is obtained. 7.The drunk driving automatic detection method based on a deep learning model and a robot system according to claim 6, characterized in that, The second-level determination module loads a preset drunk driving feature template library. The drunk driving feature template library is generated by training a plurality of standard fusion feature templates. Each standard fusion feature template includes a template identifier, a drunk driving degree category corresponding to the standard fusion feature template, a template feature vector, and a feature dimension weight table recording the importance weight of each feature dimension in the matching process of the standard fusion feature template. The second-level determination module loads a preset drunk driving feature template library. The drunk driving feature template library is generated by training a plurality of standard fusion feature templates. Each standard fusion feature template includes a template identifier, a drunk driving degree category corresponding to the standard fusion feature template, a template feature vector, and a feature dimension weight table recording the importance weight of each feature dimension in the matching process of the standard fusion feature template. The fusion detection feature vector is compared with each standard fusion feature template in the drunk driving feature template library one by one, a template feature vector and a feature dimension weight table of the standard fusion feature template are extracted, a normalized absolute difference value of the fusion detection feature vector and the template feature vector in each feature dimension is calculated, the normalized absolute difference value of each feature dimension is weighted according to the weight value in the feature dimension weight table, the weighted difference values of all feature dimensions are counted and a total weighted difference value is obtained by summing, and the total weighted difference value is converted into a matching degree parameter; The matching degree parameters of all standard fusion feature templates are sorted, the standard fusion feature template with the largest matching degree parameter is selected as the optimal matching template, the template identifier and the corresponding drunk driving degree category of the optimal matching template are recorded, the feature dimension weight table of the optimal matching template is extracted, the number of feature dimensions in the feature dimension difference value between the complete fusion detection feature vector and the optimal matching template that are within a preset difference threshold interval is counted, a proportion of the number of feature dimensions to the total number of feature dimensions is calculated to obtain an auxiliary matching degree parameter, if the auxiliary matching degree parameter is outside a preset auxiliary threshold interval, the standard fusion feature template with the second largest matching degree parameter is selected as a candidate optimal matching template and the auxiliary matching degree parameter is repeatedly calculated until the standard fusion feature template with the auxiliary matching degree parameter within the preset auxiliary threshold interval is found as the optimal matching template; The confidence parameter of the first level judgment, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template are normalized, and the final detection confidence is calculated by using the product correction method, and the variances of the confidence parameter, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template are calculated, if the variances are within a preset variance threshold interval, the initial product result is directly used as the detection confidence; if the variances are outside the preset variance threshold interval, a weight adjustment mechanism is triggered, different correction weights are assigned to the confidence parameter, the core matching degree parameter and the auxiliary matching degree parameter of the optimal matching template according to the pre-trained parameter weight model, and the detection confidence is recalculated. 8.A drunk driving automatic detection system based on a deep learning model and a robot system, characterized by, It comprises: a processor; a machine readable storage medium for storing machine executable instructions of the processor; wherein the processor is configured to execute the machine executable instructions to perform the drunk driving automatic detection method based on the deep learning model and the robot system in any one of claims 1 to 7.
9. A computer program product, characterised in that, The computer program product comprises machine executable instructions stored in a computer readable storage medium, a processor of a computer device reads the machine executable instructions from the computer readable storage medium, and the processor executes the machine executable instructions, so that the computer device performs the drunk driving automatic detection method based on the deep learning model and the robot system in any one of claims 1 to 7.
Citation Information
Patent Citations
Alcohol concentration detection method and device, wearable equipment and storage medium
CN115032250A
Detection unit position regulation and control method and track processing system
CN119811102A