Plateau intelligent medical care cabin and man-machine interaction method thereof
Through multi-microphone arrays and intelligent noise reduction technology, user voice is recognized in a plateau environment, and combined with emotional scoring and large language models to generate control instructions, the convenience and intelligence of traditional interaction methods in high-altitude environments are solved, and efficient and safe voice interaction control is achieved.
Patent Information
- Application Number
- CN202510859373.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-12
AI Technical Summary
In high-altitude environments, traditional touch screen or button-type interaction methods are difficult to meet users' high-standard needs for convenience and intelligence. Low temperature and hypoxia lead to a decrease in touch screen sensitivity and slow response of mechanical buttons, which brings additional burden to plateau personnel and unstable interactions.
Multi-microphone arrays are used to combine intelligent noise reduction technology and deep neural network to identify user voice signals, and structured control instructions are generated through emotional scoring models and large language models to realize voice interaction without touch and adjust device power according to emotional scoring.
In a plateau environment, it significantly reduces the physical burden of users, improves interaction comfort and success rate, ensures equipment response speed and energy efficiency, and provides a safe and labor-saving intelligent interactive experience.
Smart Images

Figure CN120472902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a plateau smart medical cabin and a human-computer interaction method thereof. Background Art
[0002] The Smart Medical and Nursing Cabin is a closed cabin system that integrates medical monitoring, maintenance and adjustment, and intelligent control. It is designed to safeguard personnel health and work efficiency in extreme environments such as high altitudes. It uses multi-channel sensors to monitor key environmental parameters such as oxygen concentration, air pressure, temperature, and humidity in real time, and uses intelligent algorithms to automatically adjust oxygen supply, cabin pressure, and maintain constant temperature and humidity.
[0003] In high-altitude environments, traditional touch screens or button-based interaction methods are difficult to meet users' high standards for convenience and intelligence. Specifically: low temperatures can significantly reduce the sensitivity of touch screens, making it even more difficult to click accurately when wearing thick gloves; mechanical buttons respond slowly in the cold and require greater force to trigger, placing an additional burden on people in the plateau who are already tired and have stiff muscles. At the same time, lack of oxygen causes people to often experience symptoms such as dizziness, palpitations, and shortness of breath. Finger movements become slow and unstable, and they often only want to minimize unnecessary limb movements, but they have to flex, focus, and even search back and forth for button positions in order to complete interactive tasks. Summary of the Invention
[0004] The purpose of the present invention is to provide a plateau smart medical cabin and a human-computer interaction method thereof to solve the above-mentioned technical problems.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] A plateau smart medical cabin and a human-computer interaction method thereof, comprising the following steps:
[0007] Set up multiple microphone arrays to collect sound signals in the smart medical cabin based on the microphone array, and combine intelligent noise reduction technology to capture the user's voice signal in the sound signal;
[0008] Inputting the speech signal into a pre-trained emotion scoring model and outputting an emotion score of the speech signal;
[0009] The voice signal is converted into text format to obtain voice text, and the voice text is input into a pre-integrated large language model. The large language model gives a corresponding response, and the corresponding equipment in the medical cabin is controlled by the response;
[0010] The power of the device is adjusted according to the emotion score.
[0011] As a further solution of the present invention, capturing the user's voice signal in the sound signal includes:
[0012] Each channel of the microphone array synchronously acquires sound signals, completes channel delay correction through a time domain alignment algorithm, and uses adaptive beamforming technology to suppress interference from noise sources inside and outside the cabin;
[0013] The processed sound signal is input into the deep neural network noise reduction module. The deep neural network noise reduction module uses the pre-trained noise feature dictionary and speech feature embedding to recognize and separate background noise and user speech in real time.
[0014] For the separated user voice signal, a method combining spectrum subtraction and periodic spectrum estimation is used to eliminate the residual noise component to obtain the user's voice signal.
[0015] As a further solution of the present invention: training the sentiment scoring model includes:
[0016] Establish a database to store speech signals annotated with emotion scores;
[0017] A sentiment scoring model is established based on deep learning, and the sentiment scoring model is trained and verified based on the database to obtain a pre-trained sentiment scoring model.
[0018] As a further solution of the present invention, labeling the speech signal with an emotion score includes:
[0019] Set the emotional level, which includes acute and slow levels;
[0020] The closer the emotional label of the speech signal is to acuteness, the higher the emotional score; the closer the emotional label of the speech signal is to slowness, the lower the emotional score;
[0021] The emotional labels of speech signals are manually assigned.
[0022] As a further solution of the present invention: controlling the corresponding equipment in the medical cabin includes:
[0023] The large language model performs intent analysis and contextual reasoning on the input speech text, and generates structured control instructions based on the current environmental status of the medical cabin and historical interaction records.
[0024] Input the control instructions into the actuator of the corresponding device to control the corresponding device.
[0025] As a further solution of the present invention, the power of the device during operation is corrected according to the emotion score, including:
[0026] Obtain the current power P of the device, calculate the corrected power P1 = (1 + K)P, and use the corrected power P1 as the power when the device is working.
[0027] As a further solution of the present invention: a power upper limit Pmax is set, and if P1>Pmax, the power upper limit Pmax is used as the power when the device is working.
[0028] A plateau smart medical cabin, used to implement the human-computer interaction method of a plateau smart medical cabin, comprising:
[0029] Acquisition module: multiple microphone arrays are set up to collect sound signals in the smart medical cabin based on the microphone array, and the user's voice signal is captured in the sound signal by combining intelligent noise reduction technology;
[0030] Scoring module: inputs the speech signal into a pre-trained emotion scoring model and outputs the emotion score of the speech signal;
[0031] Control module: converts the voice signal into text format to obtain voice text, inputs the voice text into a pre-integrated large language model, and the large language model gives a corresponding response, which controls the corresponding equipment in the medical cabin;
[0032] Control optimization module: Modifies the power consumption of the device during operation based on the emotion score.
[0033] The beneficial effects of the present invention are as follows:
[0034] 1) Through multi-microphone array channel synchronization, time domain alignment, adaptive beamforming, and deep neural network noise reduction, this invention can still extract user voice with high fidelity in high-altitude cabins with high wind noise and low air pressure. This entire process is completely independent of touch screens and mechanical buttons, and is not restricted by low temperatures, thick gloves, and hand stiffness. It can significantly reduce the extra physical effort and pain caused by reaching and locating buttons, and avoid repeated interactions and frustration caused by movement errors in hypoxic environments. It allows personnel in high altitude areas to easily complete command input even in cold and fatigued conditions, thereby improving overall interaction comfort and success rate.
[0035] 2) The system maps speech signal features such as speech rate, pitch, and energy into an emotion score, and collaborates with the semantic reasoning of a large language model to generate structured control instructions. When the emotion score indicates a strong sense of urgency, the algorithm immediately increases the output power of core equipment such as oxygen supply, heating, and lighting to ensure that physiological support and environmental adjustment are quickly implemented. When the tone returns to a stable state, the power is not significantly increased, thereby saving energy, suppressing noise, and extending component life. This "emotion-power" adaptive pathway ensures both emergency response speed and long-term safety and economy. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described below with reference to the accompanying drawings.
[0037] Figure 1It is a flow chart of a human-computer interaction method for a plateau smart medical cabin according to the present invention. DETAILED DESCRIPTION
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] See also Figure 1 As shown, the present invention is a plateau smart medical cabin and a human-computer interaction method thereof, comprising the following steps:
[0040] Three groups of eight-channel microphone ring arrays are arranged on the top, side walls and head of the bed (other locations are acceptable). The element spacing of each array is designed according to the half-wavelength principle to take into account the changes in sound speed under high altitude and low pressure. For example, the top array is responsible for omnidirectional sound pickup, the side wall array focuses on the direction of the rest position, and the small array at the head of the bed is used to listen to whispers. An anti-frost microporous wind shield is installed in front of the array to prevent low-temperature condensation from clogging the holes. All channel signals are first clock synchronized and time-domain aligned by the local FPGA, and then adaptive beamforming based on minimum variance distortion-free response is executed in the edge computing unit to suppress the interference sound outside the main lobe. If a strong wind is detected outside the cabin, the weight direction pattern is dynamically adjusted to ensure that the main lobe still locks the user's mouth and nose area. After digitization by a first-order high-speed ADC, multiple raw waveforms are packaged and sent to an off-board cloud-based noise reduction server via a 5G module with low latency. Simultaneously, a deep neural network noise reduction backup link runs on a local GPU to ensure continuous processing even when the cloud link jitters. The noise reduction network uses gated convolution combined with an attention mask to distinguish speech from background sound, then uses spectral subtraction to compensate for residual noise, ultimately outputting a clear user voice stream, providing high signal-to-noise input for subsequent emotion recognition and command parsing.
[0041] Understandably, this design allows plateau personnel to "control the cabin by speaking" even when wearing thick gloves, with stiff fingers, or even in a semi-recumbent position, eliminating the physical effort of searching for buttons. The multi-array layout and beam adaptation enable the system to firmly lock on to the speaker even when noisy equipment is running in the cabin or the cabin door is open, improving interaction reliability. 5G's high-speed backhaul and large cloud models work together to achieve almost simultaneous noise reduction, emotion recognition, and intent analysis. When an urgent tone is detected, the oxygen supply and heating power can be immediately increased to ensure a timely life support response. At the same time, when the tone slows down, the power is automatically reduced to save energy and reduce noise and vibration, thus providing a labor-saving and safe full-voice adaptive control experience for the smart medical cabin.
[0042] In a preferred embodiment of the present invention, capturing the user's voice signal in the sound signal includes:
[0043] During the acquisition phase, each channel of the three microphone arrays is first triggered and activated synchronously by a unified clock pulse. A cross-correlation peak alignment algorithm is used in the local digital signal processor to measure relative delays and interpolate to compensate, ensuring that the waveforms of all channels remain synchronized at the sampling point level. The aligned multi-channel data is then fed into a minimum variance distortionless response adaptive beamformer. The system dynamically updates weight coefficients based on the array geometry and real-time estimated sound source orientations. If the system detects wind whistling through the door gap outside the cabin or the low-frequency hum of the oxygen concentrator inside the cabin, it automatically constructs deep traps in these directions to suppress interference. The beamformed signals enter a deep convolutional neural network noise reduction module. This network pre-builds a noise dictionary based on soundscapes such as high-altitude wind noise and the whistling of medical equipment. It also extracts speech embeddings for joint training using a mixed corpus of Mandarin and various dialects. During inference, it uses time-frequency masks to distinguish between background noise and user speech in real time. Spectral subtraction removes broadband random noise from the separated speech traces. Periodic spectrum estimation then compensates for periodic residues to eliminate artifacts. The resulting output is a continuous, smooth, and rhythmically characterized user speech signal for use by the sentiment analysis and command parsing modules.
[0044] It is important to note that this approach enables stable capture of clear speech in the noisy and changeable acoustic environment of the plateau cabin. Time delay correction is used to ensure that beamforming effectively focuses on the speaker's direction. Adaptive weights are used to suppress external wind noise and equipment operation sounds to reduce false triggering. A two-stage noise reduction method combining a deep neural network with frequency domain post-processing is then used to completely remove small noises. This allows the signals sent to the emotion discrimination and semantic understanding modules to maintain their complete voiceprint and rhythm, thereby improving the accuracy of emotion assessment and intent recognition. This further makes the closed-loop control that adjusts device power based on the degree of emotion more sensitive and safer, providing plateau personnel with an intelligent interactive experience that allows them to quickly obtain environmental support without touching.
[0045] Inputting the speech signal into a pre-trained emotion scoring model and outputting an emotion score of the speech signal;
[0046] In another preferred embodiment of the present invention, training the sentiment scoring model includes:
[0047] During the sample collection phase, standard Mandarin, dialects, and plateau interview materials are first recorded in a low-altitude laboratory. Then, in a real cabin environment at an altitude of more than 4,000 meters, additional speech containing situations such as rapid instructions, slow narration, groaning, and coughing is added, and the cabin temperature, air pressure, and blood oxygen saturation are simultaneously recorded for subsequent horizontal retrieval. After the recording is completed, it is annotated and scored for emotion. All labeled results, together with the original waveform and the extracted Mel spectrum, are stored in a hybrid database of PostgreSQL and object storage. Multidimensional indexes are established according to the speaker's gender, scenario type, emotional range, and altitude level. The calling interface can be backhauled to the cloud for backup via the 5G private network. The above data is then used to train an emotion scoring network composed of dual-channel time-frequency convolution blocks, gated recurrent units, and multi-head attention on a local GPU cluster. Random truncation, emotional balance sampling, and mixed precision strategies are used during training. The loss and weighted F1 index on the validation set are monitored and the weights are frozen when the early stopping condition is triggered. Finally, a pre-trained model file with a quantized suffix is exported for the embedded inference engine to call.
[0048] It is important to note that by building a database covering multiple altitudes and situations with reliable emotional labeling consistency, the model can fully learn the unique patterns of human speech changes under conditions of hypoxia and cold. Subsequently, a small computing unit deployed in the cabin can use the lightweight model to quickly determine the user's level of urgency. In this way, the system can instantly increase oxygen supply, heating, or lighting power when a high emotional score is detected, and automatically reduce power consumption when emotions are stable, thereby achieving adaptive environmental control that is both energy-saving and silent and meets physiological needs. This provides a human-computer interaction experience for plateau personnel that provides timely and safe support without any hands-on efforts.
[0049] In a preferred embodiment of the present invention, labeling the speech signal with an emotion score includes:
[0050] When building the labeling system, an emotional scale is first set on the annotation platform, gradually shifting from "slow" to "acute." While playing the recording, annotators can drag the slider to select a position for each speech segment. For example, when hearing "Please give me oxygen immediately," the slider will be pushed toward the "acute" end to generate a high score, while when encountering "Dim the lights a little," the slider will stay in the "slow" area, resulting in a low score. The platform interface also provides sample waveforms and timbre prompts to guide annotators in making judgments based on speech speed, pitch, loudness, breathing pauses, and semantic content. After completion, a second annotator will review the results. If the two results differ significantly, an expert audit will be automatically triggered. Finally, the determined scale position is transferred to the emotional label and attached to the metadata of the corresponding audio file.
[0051] It is important to note that manually labeling on a one-dimensional sliding scale eliminates the psychological cues associated with specific scores, allowing annotators to focus on the core dimension of "urgency" rather than complex emotion classification. This ensures comparability between labels and reduces subjective bias in the training data. The model then learns the implicit rule that the closer to the "acute" end, the faster and stronger the action feedback required. This allows the system to immediately increase oxygen supply or heating power when high emotion levels are detected, and to adjust smoothly when low emotion levels are detected. This ensures that device output aligns with the user's current physiological and emotional needs, providing energy-efficient and user-friendly adaptive control for the high-altitude smart medical cabin.
[0052] The voice signal is converted into text format to obtain voice text, and the voice text is input into a pre-integrated large language model. The large language model gives a corresponding response, and the corresponding equipment in the medical cabin is controlled by the response;
[0053] It is important to note that the use of speech-to-text conversion followed by a large language model eliminates the need for buttons and menus, allowing plateau personnel experiencing hypoxia or hand stiffness to complete multi-step operations simply by stating their needs. The simultaneous delivery of emotional tags and historical context to the model enables the system to understand the implicit urgency of instructions, avoiding misjudgments and further ensuring the accuracy and security of device calls through structured responses. 5G transmission enables near-real-time synchronization of parsing and execution, reducing anxiety caused by waiting. The overall process not only improves interactive fluency but also ensures that equipment adjusts power on demand, providing key support for the realization of true "speaking as a service" in the plateau smart medical cabin.
[0054] In another preferred embodiment of the present invention, controlling corresponding equipment in the medical cabin includes:
[0055] After receiving the text after speech recognition, the text is first packaged into a multimodal context with the last five rounds of conversation records, sentiment scores, real-time device status, and cabin environmental sensor data, and sent to the large language model on the edge server; the intention parsing unit inside the model first locks the action target based on keyword matching and dependency syntax tree, and then uses the attention network to compare historical conversations to eliminate ambiguity. For example, when the user says "the oxygen seems to be insufficient" and the oxygen concentration was previously reduced, the model will infer that the current requirement is to "increase the oxygen concentration"; then the policy generator combines the current cabin oxygen concentration, temperature, humidity, and equipment working gear according to the preset JSONScheme a. Generates structured control instructions, which include fields such as device ID, adjustment amplitude, gradient duration, and safety threshold. If an instruction involves multiple devices, it is automatically split into parallel sub-instructions. These instructions are written to the execution controller via the MQTT message bus. The controller's built-in Lua sandbox verifies the legitimacy and permissions of the fields before calling the underlying driver. For example, it sends a command to the oxygen supply machine to increase the target concentration by two levels and complete it smoothly within ten seconds, or to the heater to maintain the current power to prevent a sudden temperature rise. After each device completes the action, it writes the execution receipt and the new status parameters back through the same bus, triggering the local speech synthesis module to generate a voice feedback prompt to the user, saying "the oxygen concentration has been increased."
[0056] It is understandable that by integrating intent analysis, contextual reasoning, and real-time environmental data into a large language model, the system can accurately understand the implicit needs behind user requests, avoiding miscontrol of equipment due to ambiguous spoken expressions. At the same time, a structured instruction format is used to ensure that downstream execution control is verifiable and traceable. The multi-device splitting and safety threshold verification mechanism enables each subsystem to collaborate within the same control link without interfering with each other, reducing the risk of device conflicts and malfunctions. The two-way receipt of the message bus allows users to know the execution results without waiting or guessing in states such as hypoxia and palpitations. The overall process ensures that the life support equipment in the medical cabin responds promptly while taking into account energy consumption and temperature stability, providing a reliable closed loop for voice emotion-driven adaptive control in plateau environments.
[0057] Correct the power of the device according to the emotion score;
[0058] It is worth noting that mapping the voice emotion score to real-time correction of device power allows the system to use the most intuitive physiological clues to judge the user's urgency, avoiding the need for plateau personnel to specify the required force when they suffer from hypoxia or dizziness. The power automatically increases and decreases with the degree of urgency. On the one hand, it can quickly increase oxygen supply, heating or ventilation when anxiety or breathing difficulties are identified, promptly relieving discomfort. On the other hand, it can moderately reduce output when emotions are stable, reducing energy consumption, noise and overheating risks. This ensures rapid response to life support while taking into account device life and energy reserves. This builds an agile and frugal closed-loop control mechanism for the smart medical cabin, giving users a safer, more labor-saving and humane environmental protection experience.
[0059] Correcting the power consumption of the device based on the emotion score includes:
[0060] Before executing a control instruction, the system first reads the current power P through the device's power acquisition module. This module can call the device's built-in current and voltage monitoring chip or obtain real-time power data via the Modbus bus, and perform a sliding window average on instantaneous jitter. The scheduler then extracts the gain coefficient K from the emotion score mapping table and calculates the target power. For example, if the current oxygen supply power is P, the emotion score corresponding to K represents "increase by two levels", which automatically results in a higher P1. The calculation result is encapsulated into an adjustment instruction and issued via the message bus, driving the controller to smoothly transition the output to P1 within the set time using a linear gradient. At the same time, the user is prompted in the interface and voice feedback that "power has been adjusted to adaptive mode" and continuously monitors the new power and environmental changes for the next round of iteration.
[0061] It is important to note that by reading power in real time and accurately correcting it according to the emotional gain formula, the user's urgency can be directly mapped to the device output, avoiding excessive or insufficient response caused by fixed gears or absolute thresholds. The smooth transition reduces the mechanical and thermal shock to components such as motors and power supplies, and also makes the changes in physical sensation more gentle. This mechanism enables the medical cabin to quickly provide stronger support during high-risk moments such as anxiety and rapid breathing. When emotions ease, the power is automatically reduced to save energy and reduce noise and vibration. This achieves sensitive adjustment of environmental control and resource optimization, ultimately ensuring that people in the plateau receive a timely and comfortable life support experience.
[0062] It should be noted that a power upper limit Pmax is set. If P1>Pmax, the power upper limit Pmax is used as the power when the device is working.
[0063] It is worth noting that data is stored in an internal encrypted server to effectively ensure information security. In addition, a strain gauge deformation data monitoring module has been specially added to the cabin to capture subtle changes in the cabin structure in real time. Combined with an intelligent early warning mechanism, it comprehensively guarantees the safe and stable operation of the equipment. In response to the extreme environment of the plateau, a GPS and Beidou positioning module has been installed, and a current acquisition module has been deployed to determine the operating status of the equipment.
[0064] At the same time, it adopts an integrated design, highly integrated photovoltaic energy management (energy storage cabinet equipped with RS232 serial port, industrial control host transmits photovoltaic mains power, mains output voltage, daily power generation, monthly power generation, etc. through 4G module), pressurization and oxygenation control (independently designed and programmed plateau ecological cabin smart system, through PLC automation control technology to achieve the above mode of pressurization and oxygenation system), environmental perception monitoring (outdoor environment (0 2 , temperature and humidity, air pressure, CO 2); buffer room air pressure; personnel dynamics) and infrared remote control function to realize closed-loop management (control the home appliances in the cabin by learning the infrared remote control signal. The transmitter is installed on the ceiling of the cabin to receive the instructions of the infrared remote control controller);
[0065] It should be noted that the present invention can also realize functions such as control of high-pressure oxygen enrichment chambers and conventional smart home control. For high-pressure oxygen enrichment chambers, the oxygen enrichment parameters can be accurately adjusted according to the user's health status data (such as blood oxygen saturation, etc.) and voice commands to ensure the user's respiratory health. In terms of conventional smart home control, intelligent control of lights, curtains, home appliances, etc. can be achieved. At the same time, with the help of the five-bureau DeepSeek large model and local Internet of Things and Internet Internet of Things technologies, various types of equipment in the smart medical cabin can achieve interconnection and intelligent collaboration. For example, when it is detected that the light in the cabin is too strong and the user is in a resting state, the system can automatically link the curtain control device to close the curtains and adjust the light brightness to a soft state; when the user issues a "prepare to rest" voice command, the system can simultaneously control the air conditioner to adjust to a suitable temperature, turn off unnecessary electrical equipment, etc., to create a comfortable rest environment.
[0066] A plateau smart medical cabin, used to implement the human-computer interaction method of a plateau smart medical cabin, comprising:
[0067] Acquisition module: multiple microphone arrays are set up to collect sound signals in the smart medical cabin based on the microphone array, and the user's voice signal is captured in the sound signal by combining intelligent noise reduction technology;
[0068] Scoring module: inputs the speech signal into a pre-trained emotion scoring model and outputs the emotion score of the speech signal;
[0069] Control module: converts the voice signal into text format to obtain voice text, inputs the voice text into a pre-integrated large language model, and the large language model gives a corresponding response, which controls the corresponding equipment in the medical cabin;
[0070] Control optimization module: Modifies the power consumption of the device during operation based on the emotion score.
[0071] The above is a detailed description of an embodiment of the present invention. However, the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.
Claims
1. A human-computer interaction method for a plateau smart medical cabin, characterized in that: The following steps are involved: Set up multiple microphone arrays to collect sound signals in the smart medical cabin based on the microphone array, and combine intelligent noise reduction technology to capture the user's voice signal in the sound signal; Inputting the speech signal into a pre-trained emotion scoring model and outputting an emotion score of the speech signal; The voice signal is converted into text format to obtain voice text, and the voice text is input into a pre-integrated large language model. The large language model gives a corresponding response, and the corresponding equipment in the medical cabin is controlled by the response; The power of the device is adjusted according to the emotion score.
2. The human-computer interaction method for a plateau smart medical cabin according to claim 1 is characterized in that: Capturing the user's voice signal in the sound signal includes: Each channel of the microphone array synchronously acquires sound signals, completes channel delay correction through a time domain alignment algorithm, and uses adaptive beamforming technology to suppress interference from noise sources inside and outside the cabin; The processed sound signal is input into the deep neural network noise reduction module. The deep neural network noise reduction module uses the pre-trained noise feature dictionary and speech feature embedding to recognize and separate background noise and user speech in real time. For the separated user voice signal, a method combining spectrum subtraction and periodic spectrum estimation is used to eliminate the residual noise component to obtain the user's voice signal.
3. The human-computer interaction method for a plateau smart medical cabin according to claim 1 is characterized in that: Training the sentiment scoring model includes: Establish a database to store speech signals annotated with emotion scores; A sentiment scoring model is established based on deep learning, and the sentiment scoring model is trained and verified based on the database to obtain a pre-trained sentiment scoring model.
4. The human-computer interaction method for a plateau smart medical cabin according to claim 3 is characterized in that: The emotional scoring of speech signals includes: Set the emotional level, which includes acute and slow levels; The closer the emotional label of the speech signal is to acuteness, the higher the emotional score; the closer the emotional label of the speech signal is to slowness, the lower the emotional score; The emotional labels of speech signals are manually assigned.
5. The human-computer interaction method for a plateau smart medical cabin according to claim 1 is characterized in that: The corresponding equipment in the medical cabin includes: The large language model performs intent analysis and contextual reasoning on the input speech text, and generates structured control instructions based on the current environmental status of the medical cabin and historical interaction records. Input the control instructions into the actuator of the corresponding device to control the corresponding device.
6. The human-computer interaction method for a plateau smart medical cabin according to claim 1 is characterized in that: Correcting the power consumption of the device based on the emotion score includes: Obtain the current power P of the device, calculate the corrected power P1 = (1 + K)P, and use the corrected power P1 as the power when the device is working.
7. The human-computer interaction method for a plateau smart medical cabin according to claim 1 is characterized in that: Set the power upper limit Pmax. If P1>Pmax, the power upper limit Pmax is used as the operating power of the device.
8. A plateau smart medical cabin, the claim being used to implement a human-computer interaction method for a plateau smart medical cabin according to any one of claims 1 to 7, characterized in that: include: Acquisition module: multiple microphone arrays are set up to collect sound signals in the smart medical cabin based on the microphone array, and the user's voice signal is captured in the sound signal by combining intelligent noise reduction technology; Scoring module: inputs the speech signal into a pre-trained emotion scoring model and outputs the emotion score of the speech signal; Control module: converts the voice signal into text format to obtain voice text, inputs the voice text into a pre-integrated large language model, and the large language model gives a corresponding response, which controls the corresponding equipment in the medical cabin; Control optimization module: Modifies the power consumption of the device during operation based on the emotion score.
Citation Information
Patent Citations
Speech recognition method suitable for noise environment
CN110148420A
System and method for generating virtual engine sound for vehicle
CN115309361A
Training method and device for voice emotion interaction model and electronic equipment
CN118711572A
Environment adjusting system with emotion intelligence
CN119277605A
Digital human system based on interactive artificial intelligence
CN120012817A
Cited By
Off-line voice and cup state feedback method and system
CN121096338A