Elevator identity instruction correction method and system based on audio recognition and storage medium
By improving the quality of elevator voice signals through multi-microphone arrays and adaptive noise reduction algorithms, and combining voiceprint features and intent understanding technology, high-precision identity authentication and multi-type command processing for elevator operation have been achieved. This solves the problems of poor signal quality and imperfect command processing in elevator voice control, and improves the safety and intelligence level of elevators.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI CHANGYI ELECTROMECHANICAL TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing elevator voice control and identity authentication technologies are difficult to adapt to complex usage scenarios. They suffer from poor voice signal quality, disconnect between identity authentication and command processing, weak command recognition and classification capabilities, lack of reasonable division of labor, and imperfect execution and feedback, resulting in insufficient safety and intelligence levels.
The system employs a multi-microphone array combined with an adaptive noise reduction algorithm to improve the signal-to-noise ratio. It converts voiceprint features and speech recognition into text commands, uses an intent understanding engine to classify command types, and executes intelligent decisions based on user identity and command type to collaboratively control elevator operation and environmental parameters. It also builds a user preference library and a blockchain-based evidence storage system.
It achieves high-precision voice recognition and identity verification in complex noisy environments, accurately classifies multiple types of commands, improves the safety, intelligence level and user experience of elevator operation, and ensures safety level and personalized service.
Smart Images

Figure CN122035666A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio recognition technology, and in particular to an elevator identity instruction correction method, system, and storage medium based on audio recognition. Background Technology
[0002] With the development of intelligent technology, elevators, as core equipment in vertical transportation, are upgrading their control methods from traditional manual buttons to intelligent methods such as voice control. Voice control, because it eliminates the need for users to manually touch the control panel, offers significant advantages in convenience and hygiene, and has gradually become an important direction for the intelligent upgrading of elevators. Meanwhile, to address the safety management needs of elevators, identity authentication technology is also widely used to restrict unauthorized users' access to specific floors, such as equipment floors, machine room floors, and private floors, to prevent safety incidents.
[0003] However, existing elevator voice control and authentication technologies still have many shortcomings and are difficult to adapt to the complex usage scenarios of elevator cars. For example, the preprocessing effect of voice acquisition is poor, the coverage of a single microphone is limited, and fixed noise reduction algorithms cannot cope with dynamic noise and echoes, resulting in poor signal quality. Identity authentication and command processing are disconnected; traditional authentication operations are cumbersome, and voiceprint recognition has low accuracy due to its single feature and rudimentary comparison methods. Command recognition and classification capabilities are weak, supporting only single elevator riding commands, with large classification biases and the inability to filter non-command inputs. Furthermore, command execution and feedback are imperfect; independent hardware drivers are inefficient, command processing lacks reasonable division of labor, and there is a lack of status recording and real-time feedback, making it difficult to form a closed-loop control system.
[0004] Therefore, there is an urgent need for an elevator identity command correction method, system, and storage medium based on audio recognition to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide an elevator identity command correction method based on audio recognition, comprising the following steps: User voice input is collected, and the voice signal is acquired through a multi-microphone array set in the elevator car. An adaptive noise reduction algorithm is used to preprocess the voice signal to improve the signal-to-noise ratio. The preprocessed speech signal is subjected to voiceprint feature extraction. The extracted voiceprint features are compared with the pre-stored voiceprint database to determine the user's identity. At the same time, the speech command is converted into a text command through the speech recognition engine. The text instructions are semantically analyzed, and the intent understanding engine is used to classify the instructions into elevator access instructions, environmental control instructions, or scenario-based composite instructions. Intelligent decisions are made based on user identity and command type. When the command is an elevator ride command, the user's access rights to the target floor are verified. When the command is an environmental control command, the user preference library is queried and an environmental parameter adjustment command is generated. When the command is a scenario-based composite command, the predefined scenario rules are parsed and the elevator ride operation and environmental control operation are executed in a coordinated manner. Based on the intelligent decision-making results, the system controls the elevator operation system to register the target floor or controls the intelligent lighting system to adjust the car environment parameters and updates the personalized settings data in the user preference database.
[0006] Furthermore, the steps of collecting user voice input, acquiring voice signals through a multi-microphone array installed in the elevator car, and preprocessing the voice signals using an adaptive noise reduction algorithm to improve the signal-to-noise ratio include: User voice signals inside the elevator car are collected using a multi-microphone array arranged in a ring. An adaptive noise reduction algorithm based on the least mean square filter is used to process the collected speech signal to suppress elevator fan noise and mechanical vibration noise. The denoised speech signal is framed and windowed, and a short-time Fourier transform is performed after the Hamming window function is applied to reduce spectral leakage. The clarity of the speech signal is enhanced by dynamic range compression and echo cancellation modules, and the pre-processed speech data is output to the voiceprint recognition module.
[0007] Furthermore, the steps of performing voiceprint feature extraction on the preprocessed speech signal, comparing the extracted voiceprint features with a pre-stored voiceprint database to determine the user's identity, and converting the speech command into a text command using a speech recognition engine include: The fundamental frequency profile, formant frequencies, and Mel frequency cepstral coefficients are extracted from the preprocessed speech signal as speakerprint features. The dynamic time warping algorithm is used to calculate the minimum distance between the extracted features and the feature sequences in the voiceprint database. When the distance is lower than a preset threshold, the user's identity is confirmed. The speech command is converted into a text command by a speech recognition engine based on a hidden Markov model, and the conversion result is processed by grammar correction. By combining the voiceprint recognition results with the analysis of the speech segments before and after the instruction words, the validity of the instruction is verified and non-instructional speech input is filtered out.
[0008] Furthermore, the step of performing semantic analysis on the text instructions and classifying them into elevator access instructions, environmental control instructions, or scenario-based composite instructions using an intent understanding engine includes: Natural language processing technology is used to parse text instructions, extract key verbs and nouns, and match them with predefined instruction templates; The intent understanding engine categorizes instructions that match the elevator riding template as elevator riding instructions, instructions that match the environmental control template as environmental control instructions, and instructions that match the scene template as scene-based composite instructions. The scene rules corresponding to the scenario-based composite instructions are parsed, and the scene rules include target floor parameters, lighting parameters, and temperature and humidity parameters. When multiple conflicting user commands are detected, the execution order of the commands is determined based on user identity priority and timestamp sorting.
[0009] Furthermore, the steps of performing intelligent decision-making based on user identity and instruction type, including verifying the user's access rights to the target floor when the instruction is an elevator ride instruction, querying the user preference library and generating an environmental parameter adjustment instruction when the instruction is an environmental control instruction, and parsing predefined scenario rules and coordinating the execution of elevator ride and environmental control operations when the instruction is a scenario-based composite instruction, include: For elevator access commands, query the user-floor mapping relationship in the permission database. If the user does not have permission to access the target floor, refuse to execute the command and generate a voice prompt. For environmental control commands, the corresponding user's lighting color temperature and brightness parameters are obtained from the user preference library and an environmental adjustment command is generated. For scenario-based composite commands, the target floor registration operation and the car environment parameter adjustment operation are executed simultaneously. When a conflict of instructions is detected, the execution order is determined based on user identity priority and the urgency of the scenario. For high-risk commands, initiate cloud-based security authentication and record operation logs to the blockchain evidence storage system.
[0010] Furthermore, the steps of controlling the elevator operation system to register the target floor or controlling the intelligent lighting system to adjust the car environment parameters based on the intelligent decision-making results, and updating the personalized settings data in the user preference database, include: The elevator control unit executes the target floor registration operation, triggering the elevator car to move to the designated floor. The intelligent LED driver module controls the color temperature and brightness of the car lights; When an environment fine-tuning instruction is received, the environment parameters in the user preference library are updated according to a preset step size and bound to the user identity; The offline-cloud collaborative computing framework handles routine instructions and sensitive operations, with routine instructions processed by the local FPGA and sensitive operations uploaded to the cloud for processing. Record the execution status of instructions to the log database and return the operation results to the user through the voice feedback module.
[0011] Furthermore, the present invention also discloses an elevator identity command correction system based on audio recognition, comprising: The acquisition module is used to acquire user voice input. It obtains voice signals through a multi-microphone array set in the elevator car and uses an adaptive noise reduction algorithm to preprocess the voice signals to improve the signal-to-noise ratio. The extraction module is used to extract voiceprint features from the preprocessed speech signal, compare the extracted voiceprint features with the pre-stored voiceprint database to determine the user's identity, and convert the speech commands into text commands through the speech recognition engine. The analysis module is used to perform semantic analysis on the text instructions and use the intent understanding engine to classify the instructions into elevator instructions, environmental control instructions, or scenario-based composite instructions. The decision-making module is used to make intelligent decisions based on user identity and command type. When the command is an elevator ride command, it verifies the user's access rights to the target floor. When the command is an environmental control command, it queries the user preference library and generates environmental parameter adjustment commands. When the command is a scenario-based composite command, it parses predefined scenario rules and coordinates the execution of elevator ride and environmental control operations. The control module is used to control the elevator operation system to register the target floor or control the intelligent lighting system to adjust the car environment parameters based on the intelligent decision results, and update the personalized settings data in the user preference library.
[0012] Furthermore, the acquisition module includes: The acquisition unit is used to acquire user voice signals inside the car through a multi-microphone array arranged in a ring; The noise reduction unit is used to process the acquired speech signal using an adaptive noise reduction algorithm based on the least mean square filter to suppress elevator fan noise and mechanical vibration noise. The transform unit is used to perform framing and windowing processing on the denoised speech signal, and then performs short-time Fourier transform after applying the Hamming window function to reduce spectral leakage. The output unit is used to enhance the clarity of the speech signal through dynamic range compression and echo cancellation modules, and output the pre-processed speech data to the voiceprint recognition module.
[0013] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described elevator identity instruction correction method based on audio recognition.
[0014] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described elevator identity instruction correction method based on audio recognition.
[0015] The beneficial effects of this application are as follows: Firstly, this invention achieves full voice coverage in the car through a ring-shaped multi-microphone array, and combines adaptive noise reduction, Hamming windowing and other algorithms to effectively suppress dynamic noise and echo, and significantly improve the signal-to-noise ratio and clarity of the voice signal.
[0016] Secondly, this invention can accurately distinguish between elevator ride, environmental control, and scenario-based composite commands by using natural language processing and template matching mechanisms, filter non-command inputs, reduce the risk of misoperation, and extract multi-dimensional voiceprint features and use dynamic time warping algorithms to improve the accuracy of identity verification and realize the linkage control of identity and commands.
[0017] Third, this invention can build a user preference library to achieve personalized adjustment of environmental parameters, adjudicate command conflicts according to identity priority and scenario urgency, and enhance the security level of high-risk commands through cloud security authentication.
[0018] Fourth, this invention can improve execution efficiency through hardware-driven collaboration, balance response speed and safety through offline and cloud-based collaborative processing, record execution status and provide voice feedback, forming a complete control closed loop, and comprehensively improving the safety, intelligence level and user experience of elevator operation. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a method flow proposed in an embodiment of this application.
[0020] Figure 2 This is a schematic diagram of the system structure proposed in an embodiment of the present invention.
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0023] like Figure 1 As shown, this application provides an elevator identity instruction correction method based on audio recognition, including the following steps: S1, collect user voice input, acquire voice signal through a multi-microphone array set in the elevator car, and preprocess the voice signal using an adaptive noise reduction algorithm to improve the signal-to-noise ratio; S2, perform voiceprint feature extraction on the preprocessed speech signal, compare the extracted voiceprint features with the pre-stored voiceprint database to determine the user's identity, and at the same time convert the speech command into a text command through the speech recognition engine. S3, perform semantic analysis on the text instructions, and use the intent understanding engine to classify the instructions into elevator instructions, environmental control instructions, or scenario-based composite instructions; S4 executes intelligent decisions based on user identity and instruction type. When the instruction is an elevator ride instruction, it verifies the user's access rights to the target floor. When the instruction is an environmental control instruction, it queries the user preference library and generates an environmental parameter adjustment instruction. When the instruction is a scenario-based composite instruction, it parses the predefined scenario rules and coordinates the execution of elevator ride and environmental control operations. S5, based on the results of intelligent decision-making, controls the elevator operation system to register the target floor or controls the intelligent lighting system to adjust the car environment parameters, and updates the personalized settings data in the user preference library.
[0024] As described in steps S1-S5 above, elevators, as enclosed and dynamically operating spaces, have two core requirements: first, security control, where unauthorized personnel must be restricted from accessing certain floors to prevent accidents; and second, user experience, as different users' elevator operation needs extend beyond simply reaching a specific floor, potentially involving adjusting environmental parameters such as the color temperature and brightness of the car's lighting. Furthermore, the continuous noise from the fans and mechanical vibrations within the car can superimpose on the user's voice signal, reducing the accuracy of voice command recognition. Failure to effectively address these issues will directly impact elevator safety and user experience. Therefore, a technical solution is needed that adapts to the complex environment of the elevator car, integrates identity verification and multi-type command processing, to meet the specific needs of elevator scenarios.
[0025] Traditional elevator control methods, when using button input, lack an identity authentication process, allowing anyone to operate the elevator to any floor, posing a security vulnerability. When using ordinary voice control, it is not optimized for the noise environment inside the elevator car, resulting in a low signal-to-noise ratio and large recognition errors. Furthermore, it can only process single elevator commands and cannot respond to environmental control needs. It also lacks a mechanism for recording user personalized preferences, requiring users to repeat commands for each environmental adjustment, leading to a poor user experience. This invention, however, addresses these shortcomings through a comprehensive design encompassing voice preprocessing, voiceprint recognition and command conversion, semantic analysis and classification, intelligent decision-making, execution, and preference updates, achieving end-to-end control from voice signal to precise execution.
[0026] This invention achieves safe, personalized, and intelligent elevator operation through a complete closed-loop process from voice input to command execution and preference update. It integrates voiceprint authentication, semantic command classification, and intelligent decision control through audio recognition technology. Ultimately, it solves the technical problems of lack of identity authentication, low voice recognition accuracy in complex noise environments, single command type, and insufficient personalized services in traditional elevator command control.
[0027] In one embodiment, the steps of acquiring user voice input, obtaining voice signals through a multi-microphone array installed in the elevator car, and preprocessing the voice signals using an adaptive noise reduction algorithm to improve the signal-to-noise ratio include: S11 uses a multi-microphone array in a ring layout to collect user voice signals in the car, ensuring that voice collection covers all directions in the car. S12 uses an adaptive noise reduction algorithm based on the minimum mean square filter to process the collected voice signal and suppress elevator fan noise and mechanical vibration noise. S13, the noise-reduced speech signal is framed and windowed, and a short-time Fourier transform is performed after the Hamming window function is applied to reduce spectral leakage. S14 enhances the clarity of the speech signal through dynamic range compression and echo cancellation modules, and outputs the pre-processed speech data to the voiceprint recognition module.
[0028] As described in steps S11-S14 above, the above four steps together constitute a complete process of voice signal acquisition and preprocessing. A ring-shaped multi-microphone array is used to achieve full coverage acquisition of voice signals inside the car. Combined with technologies such as adaptive noise reduction based on the least mean square filter, frame windowing, short-time Fourier transform, dynamic range compression, and echo cancellation, the acquired voice signals are systematically preprocessed to ultimately improve the signal-to-noise ratio of the voice signal, providing high-quality voice data support for subsequent voiceprint feature extraction, user identity verification, and voice command conversion.
[0029] As an enclosed and dynamically operating space, the elevator car presents an uncertain user position. Users may stand in any direction within the car, and if the coverage of the voice acquisition equipment is limited, incomplete voice signal acquisition from some locations may occur. Noise interference also exists; the continuous fan noise and mechanical vibration noise generated during car operation can superimpose with the user's voice signal. Furthermore, the smooth interior walls of the car can easily generate voice echoes, leading to a decrease in the signal-to-noise ratio and clarity of the voice signal. If these problems are not addressed, subsequent voiceprint recognition may result in misidentification, and voice command conversion may lead to textual deviations, directly affecting the reliability of the entire elevator command correction method.
[0030] A ring-shaped multi-microphone array is used to collect user voice signals inside the elevator car. Considering that the elevator car's interior is typically a circular or square space with a diameter of 1.5m-2.5m, the multi-microphone array is designed in a ring layout. Four to six microphones are evenly installed on the top or side walls of the car, with a microphone spacing of 0.5m-0.8m. This ensures that the collection range covers all areas inside the car, so that at least two microphones can capture the user's voice signal regardless of whether the user is standing in the center, a corner, or near the door, avoiding signal loss due to different user positions. For example, when a user is standing in a corner, the two microphones closest to that corner can simultaneously collect their voice signal, and the signal superposition initially enhances the voice intensity, laying the foundation for subsequent processing.
[0031] An adaptive noise reduction algorithm based on a minimum mean square filter is used to process the acquired speech signal. The algorithm constructs a noise reference signal by real-time acquisition of noise samples from inside the elevator car, obtained from fan and mechanical vibration noise collected by the microphone when there is no user speech. The algorithm then dynamically adjusts the filter coefficients with the minimum mean square error as the objective, setting a convergence factor of 0.01 to ensure that the coefficient adjustment speed adapts to noise changes, minimizing the noise component in the filter output signal. Targeting the low-frequency characteristics of elevator fan noise (mainly concentrated in the 50Hz-200Hz range) and the mid-frequency characteristics of mechanical vibration noise (mainly concentrated in the 200Hz-1000Hz range), the algorithm can specifically suppress noise signals in these two frequency bands, improving the speech signal-to-noise ratio from the initial 10dB-15dB to 25dB-30dB, effectively preserving the high-frequency characteristics of the user's speech, and retaining key information for subsequent voiceprint feature extraction.
[0032] The denoised speech signal is framed and windowed. A Hamming window function is applied to reduce spectral leakage before a short-time Fourier transform is performed. Speech signals are non-stationary, and their frequency characteristics change over time. Direct Fourier transform cannot accurately capture instantaneous frequency characteristics. Therefore, the denoised speech signal is framed in 20ms-30ms frames with a 50% overlap to ensure smooth transitions between frames. The signal continuity between intervals is ensured. Windowing processing uses the Hamming window function, whose expression is: w(n) = 0.54 - 0.46cos( ) ; in, The index of the sampling point within the frame is represented, ranging from 0 to N-1. w(n) represents the Hamming window weight value of the nth sampling point in a single frame of speech signal, and N represents the total number of sampling points in a single frame of speech signal (i.e., the frame length). This is a framing parameter preset by the system based on the short-time stationarity of the speech signal. This function can reduce the abrupt changes at the edges of the signal frames, reduce spectral leakage caused by framing, and make the frequency characteristics after frequency domain transformation more accurate. Subsequently, a short-time Fourier transform is performed on each frame of the windowed signal to convert the time-domain signal into the amplitude and phase spectra in the frequency domain, providing frequency domain data support for subsequent speaker feature extraction, such as the extraction of fundamental frequency contours and formant frequencies.
[0033] The clarity of the voice signal is enhanced through dynamic range compression and echo cancellation modules. The amplitude of a user's voice signal can vary significantly depending on the volume of their voice; for example, the amplitude might be 0.1V when speaking softly and 1V when speaking loudly. The dynamic range compression module compresses the signal amplitude to a fixed range of -1dB to 1dB, preventing weak signals from being ignored and strong signals from being distorted in subsequent processing due to amplitude differences. The echo cancellation module collects echo signals reflected from the elevator car's inner wall; the echo delay time is typically 50ms-100ms. An echo estimation model is constructed, and the estimated echo component is subtracted from the voice signal to eliminate speech repetition. After these two processing steps, the clarity of the voice signal is significantly improved. Finally, the preprocessed voice data is output to the voiceprint recognition module, ensuring that the voiceprint recognition module can extract accurate voiceprint features based on the high-quality signal, reducing identity verification errors.
[0034] In one embodiment, the steps of performing voiceprint feature extraction on the preprocessed speech signal, comparing the extracted voiceprint features with a pre-stored voiceprint database to determine the user's identity, and simultaneously converting the speech command into a text command using a speech recognition engine include: S21, extract the fundamental frequency profile, formant frequency and Mel frequency cepstral coefficients from the preprocessed speech signal as voiceprint features; S22, the dynamic time warping algorithm is used to calculate the minimum distance between the extracted features and the feature sequences in the voiceprint database. When the distance is lower than the preset threshold, the user's identity is confirmed. S23, converts speech commands into text commands using a speech recognition engine based on a hidden Markov model, and performs grammatical correction processing on the conversion results; S24. Combine the voiceprint recognition results to analyze the speech segments before and after the instruction word, verify the validity of the instruction, and filter out non-instructional speech input.
[0035] As described in steps S21-S24 above, these steps together constitute the core process of voiceprint identity verification and voice command conversion. By extracting multi-dimensional voiceprint features from the preprocessed speech signal and combining them with the dynamic time warping algorithm, user identity comparison is completed. At the same time, a speech recognition engine based on a hidden Markov model is used to convert speech commands into standardized text commands, and non-command inputs are filtered through validity verification. Finally, accurate user identity verification and accurate voice command conversion are achieved, providing a reliable identity basis and command text foundation for subsequent semantic analysis and intelligent decision-making.
[0036] Elevator operation safety relies on the binding of user identity and floor access permissions, while accurate execution of user commands requires converting voice signals into machine-readable text. Although preprocessed voice signals improve the signal-to-noise ratio, it's still necessary to extract voiceprint features that uniquely identify the user. Voiceprint features are determined by the user's physiological structure, such as vocal cords and resonance cavities, and are unique, serving as the core basis for identity authentication. Furthermore, unconverted voice signals cannot be directly used for command classification; they must be converted to text for semantic analysis. Incomplete voiceprint feature extraction can lead to misidentification, such as misidentifying an ordinary user as an administrator. Inaccurate voice command conversion or the inclusion of non-command input can cause subsequent command classification errors, such as misinterpreting casual conversation as elevator commands, directly impacting the safety and accuracy of elevator operation. Therefore, a technical solution that balances accurate identity verification with reliable command conversion is needed to address these issues.
[0037] Traditional processing methods often employ only single-dimensional features for voiceprint extraction, such as extracting only the fundamental frequency contour. This results in insufficient feature information, low accuracy in identity verification, and the use of simple algorithms like Euclidean distance for identity comparison. These algorithms fail to account for inconsistencies in feature sequence lengths caused by variations in user speaking speed, leading to significant comparison errors. Voice command conversion utilizes simple template matching, which cannot handle the temporal characteristics of speech signals, resulting in high conversion error rates. Furthermore, the lack of grammatical correction leads to inconsistent text formatting. Finally, the absence of non-command input filtering mechanisms makes it prone to misinterpreting casual conversation and ambient noise as valid commands.
[0038] Fundamental frequency profile, formant frequencies, and Mel-frequency cepstral coefficients are extracted from preprocessed speech signals as voiceprint features. The fundamental frequency profile reflects the basic frequency of vocal cord vibration; for adult males, it ranges from 85Hz to 180Hz, and for adult females, from 165Hz to 255Hz. This feature can be extracted by calculating the zero-crossing rate and autocorrelation function of the speech signal frame by frame, and it can initially distinguish voiceprint differences between different users. Formant frequencies reflect the structural characteristics of the resonance cavity; the first formant frequency is typically in the 500Hz-1000Hz range. It can be extracted by finding the extreme points of the frequency domain amplitude spectrum of the speech signal, and it can further supplement the physiological characteristic information of the user's voiceprint. Mel-frequency cepstral coefficients simulate the sensitivity of the human ear to different frequencies of sound. By mapping the frequency domain signal to the Mel scale, calculating the logarithmic energy, and performing a discrete cosine transform, 12-16 dimensional coefficients can be extracted, which can capture the detailed spectral features of the speech signal. The combination of these three types of features forms a multi-dimensional voiceprint feature set, which improves information coverage compared to a single feature, providing sufficient basis for accurate identity verification.
[0039] A dynamic time warping algorithm is used to calculate the minimum distance between the extracted features and the feature sequences in the voiceprint database. The pre-stored voiceprint database is pre-loaded with voiceprint samples from authorized users through an elevator management system. Each user has 5-10 voiceprint data entries for different scenarios, such as normal speaking and slightly louder speaking. Each data entry contains three types of feature sequences corresponding to the first step. Because different users speak at different speeds, even the same user speaking at different times will have different feature sequence lengths. For example, user A takes 1.2 seconds to say "to the 5th floor," while user B takes 1.5 seconds to say the same thing. The dynamic time warping algorithm stretches or compresses the feature sequences, setting a maximum stretch / compression ratio of 1.5 times, to find the optimal alignment path between the two sequences. The sum of the feature distances along the path is then calculated as the minimum distance. The preset distance threshold is 0.3. When the calculated minimum distance is lower than 0.3, such as when the minimum distance between the extracted features of user C and the features of its samples in the database is 0.25, the user's identity is confirmed to match. If it is higher than 0.3, such as when the minimum distance between the extracted features of user D and all samples in the database is 0.4, the identity confirmation fails to avoid misjudgment.
[0040] A speech recognition engine based on a Hidden Markov Model (HMM) converts speech commands into text commands and performs grammatical correction on the conversion results. The HMM constructs a Markov chain of 3-5 states, each state corresponding to a short-term feature of the speech signal. A forward algorithm is used to calculate the matching probability between the speech signal and each text sequence, selecting the text with the highest probability as the conversion result. Considering the characteristics of speech commands in elevator scenarios, such as containing floor numbers and adjustment actions, a dedicated elevator speech dataset containing several elevator-related speech samples is used during model training to improve conversion accuracy. After conversion, grammatical correction is performed, correcting the text according to predefined command formats, such as "go to floor XX" or "adjust parameter XX to a higher value." For example, "go to the 10th floor" is corrected to "go to the 10th floor," and "turn the lights up" is corrected to "adjust the lighting brightness to a higher value," ensuring a consistent text command format and laying the foundation for template matching in subsequent semantic analysis.
[0041] By combining voiceprint recognition results with analysis of the speech segments before and after the command word, the validity of the command is verified and non-command speech input is filtered out. After the voiceprint recognition results confirm the existence of a valid user voice, the system locates the core command word in the voice command, such as "go to floor XX" or "adjust the lights," and extracts speech segments of 500ms-1000ms before and after the command word for analysis. If the segment has no continuous semantics, such as only containing "the weather is nice today" or "casual conversation," it is determined to be non-command input and directly filtered out. If the segment forms a complete semantic meaning around the command word, such as "go to floor 8, please hurry," it is determined to be a valid command, and the converted text command is retained for the next step of semantic analysis. This step can reduce the false trigger rate of non-command input and avoid elevator misoperation caused by irrelevant voice.
[0042] In one embodiment, the step of performing semantic analysis on the text instruction and classifying the instruction into elevator access instructions, environmental control instructions, or contextualized composite instructions using an intent understanding engine includes: S31 uses natural language processing technology to parse text instructions, extract key verbs and nouns, and match them with predefined instruction templates; S32, through the intent understanding engine, classifies instructions that match the elevator template as elevator instructions, instructions that match the environmental control template as environmental control instructions, and instructions that match the scene template as scene-based composite instructions; S33, parse the scene rules corresponding to the scene-based composite instructions, the scene rules including target floor parameters, lighting parameters and temperature and humidity parameters; S34, when multiple user command conflicts are detected, the command execution order is determined according to user identity priority and timestamp sorting.
[0043] As described in steps S31-S34 above, the above steps together constitute the core process of text instruction semantic analysis and classification. By extracting key information of text instructions through natural language processing technology and matching it with predefined templates, the instruction type is classified in combination with the intent understanding engine. The scenario rules of scenario-based compound instructions are parsed simultaneously. It can also determine the execution order for multi-user instruction conflicts, and finally achieve accurate classification and orderly processing of text instructions, providing a clear basis for subsequent intelligent decision-making based on instruction type.
[0044] The converted text commands contain diverse user needs. Some commands only require the elevator to register floors, some require adjusting car environmental parameters, and others require both elevator use and environmental adjustment simultaneously. If these command types cannot be distinguished, subsequent decision-making will lead to operational errors. For example, environmental control commands might be misinterpreted as elevator use commands, resulting in incorrect floor registration. Furthermore, multiple users may issue commands simultaneously inside the elevator car; failure to handle command conflicts can cause elevator operation chaos, impacting operational efficiency and user experience. Therefore, a technical solution is needed that can accurately identify command intent, classify command types, and resolve conflicts to ensure the accuracy and orderliness of subsequent decision-making and execution.
[0045] Traditional processing solutions can only recognize single elevator access commands and cannot distinguish between environmental control or combined commands. They also fail to respond to environmental commands such as temperature and lighting adjustments. Command classification relies on simple keyword matching, lacks standardized templates, has low classification accuracy, does not support scenario-based parsing of combined commands, and cannot handle combined needs of "elevator access + environmental adjustment." When faced with conflicting commands from multiple users, there is no clear rule for the execution order; commands are executed only in the order they are received, easily overlooking the needs of high-priority users.
[0046] Natural Language Processing (NLP) technology is used to parse text commands, extracting key verbs and nouns for matching against predefined command templates. NLP employs word segmentation algorithms and a dictionary-based forward maximum matching method. The dictionary contains commonly used elevator-related terms such as "to," "adjust," "floor," and "lights." The text commands are then broken down and further processed using part-of-speech tagging algorithms with a conditional random field model, resulting in higher tagging accuracy and the ability to identify key verbs and nouns. For example, the text command "adjust car lights to 4000K" is broken down into the key verb "adjust" and the key nouns "car lights" and "4000K." The text command "to the 10th floor" extracts the key verb "to" and the key noun "10th floor." Predefined command templates are pre-defined and stored in the system based on common elevator operation scenarios. These include elevator riding templates containing keyword combinations such as "to the XX floor" and "go to the XX floor"; environmental control templates containing keyword combinations such as "adjust lights," "adjust brightness," and "adjust temperature"; and scenario templates containing keyword combinations involving both elevator riding and environmental control, such as "to the XX floor and adjust XX." The extracted key information is compared with each template, laying the foundation for subsequent classification.
[0047] The intent understanding engine categorizes instructions matching different templates into corresponding types. This engine has a built-in classification model trained on an elevator scene instruction dataset containing samples of elevator ride, environmental control, and scenario-based composite instructions. The model's input is the first step's matching result, and its output is the instruction type. When key information matches the elevator ride template (e.g., "to the 10th floor" matches the "to the XXth floor" template), the engine classifies the instruction as an elevator ride instruction. When key information matches the environmental control template (e.g., "adjust the car lights to 4000K" matches the "adjust the lights" template), it is classified as an environmental control instruction. When key information matches the scene template (e.g., "to the 15th floor and adjust the lights to 4000K" matches the "to the XXth floor and adjust XX" template), it is classified as a scenario-based composite instruction. This classification method improves accuracy compared to traditional keyword matching, ensuring precise instruction type classification.
[0048] The system parses scenario-based composite commands according to corresponding scenario rules. These rules are predefined and stored in the system based on common user needs. Each rule includes target floor parameters, lighting parameters, and temperature and humidity parameters. For example, the office scenario rule corresponds to the 10th floor, a lighting color temperature of 4500K, brightness of 80%, and a temperature of 25℃. The rest scenario rule corresponds to the 20th floor, a lighting color temperature of 3000K, brightness of 50%, and a temperature of 24℃. During parsing, the intent understanding engine first identifies scenario-related terms in the composite command, such as "office mode" and "rest mode," then calls the corresponding scenario rule to extract the target floor and environmental parameters, clarifying the specific parameters of the two operations that need to be executed simultaneously in the composite command, providing data support for subsequent collaborative execution.
[0049] When multiple conflicting user commands are detected, the execution order is determined based on user priority and timestamp. User priority is pre-defined in the permission database, divided into levels: Administrator > Employee > Visitor. Each level corresponds to a fixed priority value: Administrator has a priority value of 3, Employee has 2, and Visitor has 1. The timestamp is the precise time automatically recorded by the system when it receives the command. For example, if User A (Administrator, priority value 3) issues the command "Go to the 5th floor," the system records the timestamp as 14:00:00; if User B (Employee, priority value 2) issues the command "Go to the 10th floor," the timestamp is 14:00:02. In this case, User A's command "Go to the 5th floor" is executed first because of its higher priority. If two users have the same priority (both are employees), the command received first is executed according to the timestamp order. This processing method ensures orderly elevator operation in multi-command scenarios, avoids chaos, and improves operational efficiency.
[0050] In one embodiment, the steps of performing intelligent decision-making based on user identity and instruction type, including verifying the user's access rights to the target floor when the instruction is an elevator ride instruction, querying the user preference library and generating an environmental parameter adjustment instruction when the instruction is an environmental control instruction, and parsing predefined scenario rules and coordinating the execution of elevator ride and environmental control operations when the instruction is a scenario-based composite instruction, include: S41, for elevator access commands, query the mapping relationship between users and floors in the permission database. If the user does not have permission to access the target floor, refuse to execute the command and generate a voice prompt. S42, for environmental control commands, retrieve the corresponding user's lighting color temperature and brightness parameters from the user preference library and generate environmental adjustment commands; S43, for scenario-based composite instructions, simultaneously executes the target floor registration operation and the car environment parameter adjustment operation; S44, when a conflict of instructions is detected, the execution order is determined based on the user's identity priority and the urgency of the scenario; S45 initiates cloud-based security authentication for high-risk commands and records operation logs to the blockchain evidence storage system.
[0051] As described in steps S41-S45 above, the above steps together constitute an intelligent decision-making process based on user identity and instruction type. By performing floor permission verification for elevator ride instructions, calling user preference parameters for environmental control instructions, and simultaneously executing dual operations for scenario-based composite instructions, the process also resolves instruction conflicts according to rules and strengthens security authentication and log storage for high-risk instructions. Ultimately, this process achieves the safety, personalization, and orderliness of elevator instruction execution, providing accurate decision-making basis for subsequent instruction implementation.
[0052] Elevator operation needs to balance safety control and personalized service. On the one hand, some floors, such as machine room floors and equipment floors, pose security risks, requiring restricted access for unauthorized users. Without proper authorization verification, this can easily lead to safety incidents. On the other hand, different users have different preferences for car environment parameters. If users have to repeat commands for every adjustment, it will degrade the experience. Furthermore, scenario-based compound commands need to simultaneously complete both elevator riding and environment adjustment operations; executing them separately will prolong the operation time. In addition, multiple users issuing commands simultaneously may cause conflicts, and high-risk commands, such as access to sensitive floors, require additional security measures. Failure to address these issues will result in elevator operation that is neither safe nor flexible.
[0053] Traditional solutions lack access control for elevator commands, allowing any user to access any floor, creating a security vulnerability. Environmental control commands lack user preference invocation functionality, requiring users to repeatedly input parameters for each adjustment, making the process cumbersome. Contextualized compound commands cannot be executed synchronously, requiring two separate operations, resulting in low efficiency. In case of command conflicts, they are executed only in the order they are received, neglecting user priority and scenario urgency, easily overlooking critical needs. High-risk commands lack additional security authentication, relying solely on basic identity verification, resulting in insufficient security levels. Furthermore, operation logs lack tamper-proof mechanisms, making traceability difficult.
[0054] For elevator access commands, the system queries the user-floor mapping in the access control database. This database is pre-built using the elevator management system, recording the identity information of all authorized users and their accessible floor mappings. For example, user A (administrator) can access floors 1-20 and the machine room floor; user B (employee) can only access floors 1-10; and user C (visitor) can only access floor 1 and a specified floor (e.g., floor 5). When a user issues an elevator access command, such as user B issuing a "to the 12th floor" command, the system first verifies the user's identity (user B) using permission 3, then queries the access control database for that user's floor mapping. If it finds that user B does not have access to the 12th floor, it immediately refuses to process the floor registration and generates a voice feedback module prompting "You do not have access to the 12th floor," preventing unauthorized access from causing security risks.
[0055] For environmental control commands, the system retrieves the corresponding lighting color temperature and brightness parameters from the user preference library and generates an environmental adjustment command. The user preference library is built using historical user operation data. The system automatically records the parameters after each environmental adjustment by the user. For example, if user D repeatedly adjusts the lighting color temperature to 4000K and the brightness to 80%, these parameters are stored and bound to the user's identity. When user D issues an environmental control command, such as adjusting the lighting, the system, based on the confirmed user identity (user D), retrieves the corresponding color temperature of 4000K and brightness of 80% from the preference library and directly generates the environmental adjustment command "Adjust the car lighting color temperature to 4000K and brightness to 80%", eliminating the need for the user to repeatedly input parameters and improving operational convenience.
[0056] For scenario-based composite commands, the system simultaneously executes the target floor registration operation and the car environment parameter adjustment operation. After the scenario-based composite command is classified and confirmed, the system first parses the pre-stored scenario rules. For example, the office scenario rule corresponds to the target floor 10, the lighting color temperature 4500K, and the brightness 80%. Then, based on the user authentication and floor access permissions confirmed above, such as confirming that user E has access to the 10th floor, two operations are triggered simultaneously: registering the 10th floor through the elevator control unit, and simultaneously sending a command to the intelligent LED driver module to adjust the lighting to 4500K and 80%. Compared with the traditional two-stage operation, simultaneous execution can reduce the time for command implementation by more than 50%, improving operational efficiency.
[0057] When command conflicts are detected, the execution order is determined based on user priority and scenario urgency. User priority is pre-defined through the permission database, divided into administrator > employee > visitor, with corresponding priority values of 3, 2, and 1 respectively. Scenario urgency is divided into emergency scenarios (e.g., medical treatment, fire) > ordinary scenarios. Emergency scenarios require the user command to contain specific keywords (e.g., "urgent to the 5th floor"). For example, if user F (administrator, priority 3) issues the command "to the 8th floor" and user G (employee, priority 2) issues the command "to the 12th floor," user F's command "to the 8th floor" will be executed first because of its higher priority. If user H (employee, priority 2) issues the command "urgent to the 5th floor" and user I (employee, priority 2) issues the command "to the 7th floor," user H's command is an emergency scenario, so the command "urgent to the 5th floor" will be executed first, ensuring that critical needs are met first.
[0058] High-risk commands trigger cloud-based security authentication and are logged to a blockchain-based evidence storage system. High-risk commands are identified using predefined rules and include commands to access sensitive floors such as the server room floor and equipment floor. When user J issues a command to access the server room floor, the system first verifies user J's identity as an authorized maintenance personnel before initiating cloud-based security authentication. This is done through encrypted communication between the elevator and the cloud server (using AES-256 encryption), sending user J's identity information and the command content to the cloud. Once the cloud verifies the information and provides feedback, the system executes the command. Simultaneously, the operation log (including user J's identity, command content, and execution time) is recorded in real-time to the blockchain-based evidence storage system. This system uses a consortium blockchain architecture, where each node stores a copy of the logs, ensuring the logs are immutable and can be quickly retrieved for subsequent traceability, thus improving the security and traceability of high-risk command operations.
[0059] In one embodiment, the steps of controlling the elevator operation system to register the target floor or controlling the intelligent lighting system to adjust the car environment parameters based on the intelligent decision-making results, and updating the personalized settings data in the user preference database, include: S51 executes the target floor registration operation through the elevator control unit, triggering the car to run to the designated floor; S52 controls the intelligent LED driver module to adjust the color temperature and brightness of the car lights; S53, when receiving an environment fine-tuning instruction, update the environment parameters in the user preference library according to a preset step size and bind them to the user identity; S54 processes routine instructions and sensitive operations through an offline-cloud collaborative computing framework. Routine instructions are processed by the local FPGA, while sensitive operations are uploaded to the cloud for processing. S55 records the command execution status to the log database and returns the operation results to the user through the voice feedback module.
[0060] As described in steps S51-S55 above, the above steps together constitute the implementation and data update process of intelligent decision results. By performing floor registration through the elevator control unit, adjusting environmental parameters through the intelligent LED drive module, combining preference updates after environmental fine-tuning, offline and cloud collaborative instruction processing division, and instruction status recording and voice feedback, the elevator instructions are ultimately executed accurately, user preferences are dynamically optimized, and the operation process is traceable. This ensures that the entire elevator instruction correction method forms a closed loop, meeting the user's operational effectiveness and personalized needs.
[0061] After intelligent decision-making, the results need to be translated into actual actions by the elevator hardware. Elevator riding decisions require triggering the car to move to the target floor, while environmental control decisions require adjusting the lighting parameters inside the car. If the hardware cannot be effectively driven to execute these decisions, the decisions become meaningless. Simultaneously, users may have fine-tuning needs during environmental adjustments, such as further brightening the already adjusted area. This requires updating the preference library to adapt to subsequent operations. Furthermore, command processing must balance response speed and safety. Routine commands need to be executed quickly, while sensitive operations require enhanced safety measures. Users also need to be promptly informed of the operation results. The lack of these steps can lead to operational gaps, inconsistent preferences, and a degraded user experience. Therefore, an execution scheme covering hardware driving, preference updates, command processing, and result feedback is needed to ensure effective decision implementation and optimize subsequent services.
[0062] Traditional solutions separate floor registration from environmental adjustment, requiring manual hardware operation for each step, resulting in low efficiency. There's no preference update mechanism after environmental fine-tuning, requiring users to repeat commands after each adjustment, leading to a poor personalized experience. Command processing lacks division of labor, with all commands processed through a single local or cloud channel. Routine commands are slow due to cloud latency, while sensitive operations suffer from insecurity due to insufficient local computing power. Furthermore, the lack of command status logging and real-time feedback prevents users from confirming successful operations and makes subsequent tracking of the process impossible.
[0063] The elevator control unit performs the target floor registration operation. The elevator control unit is directly connected to the drive module of the elevator operating system and pre-stores the corresponding logic between floor registration and car drive. For example, when registering the 8th floor, the control unit sends a control signal to the drive module to "move from the current floor to the 8th floor." When the decision result is a boarding instruction, such as user A's target floor being the 8th floor and the permission verification passing, the elevator control unit receives the decision result and immediately executes the floor registration, marking the 8th floor as the target floor in the floor display area of the elevator control panel. Simultaneously, it sends a drive signal to the drive module, triggering the car to start and travel from the current 3rd floor to the 8th floor at a preset speed, ensuring that the boarding decision is quickly translated into car action.
[0064] The intelligent LED driver module controls the color temperature and brightness of the car lights. The intelligent LED driver module is electrically connected to the LED lights inside the car and pre-stores the correspondence between color temperature, brightness parameters, and drive current. For example, a color temperature of 4000K corresponds to a drive current of 200mA, and 80% brightness corresponds to a drive current of 180mA. When the decision result is an environmental control command, such as user B's preferred parameters of 4000K color temperature and 80% brightness, the system sends this parameter command to the intelligent LED driver module. The driver module adjusts the output current according to the preset correspondence, controlling the LED lights to stabilize the color temperature at 4000K and the brightness at 80%, achieving precise adjustment of the car's environmental parameters.
[0065] When an environmental fine-tuning command is received, the environmental parameters in the user preference library are updated according to a preset step size. Environmental fine-tuning commands originate from user-added commands after the initial environmental adjustment. For example, if user C, after adjusting the brightness to 80%, issues a command to "brighten it further," the preset step size is pre-set in the system, such as a 10% step for brightness and a 500K step for color temperature, ensuring smooth adjustment. The system first parses the parameter changes corresponding to the fine-tuning command. For example, if user C's fine-tuning command corresponds to an increase in brightness from 80% to 90%, it then retrieves the user's historical environmental parameters from the user preference library, updates them to 90% brightness in 10% steps, and rebinds the updated parameters to user C's identity. This provides the latest preference basis for subsequent environmental control commands from user C, avoiding duplicate input.
[0066] A collaborative offline-cloud computing framework handles both routine commands and sensitive operations. This framework pre-divides command processing tasks. Routine commands, including simple elevator access and routine environmental adjustments (such as adjusting to preferred parameters), have simple logic and low security requirements, and are processed directly by the local FPGA to ensure rapid response. Sensitive operations, such as user identity modification and changes to permissions for high-risk floors, have high security requirements and must be uploaded to the cloud for processing. The cloud verifies the results through secondary identity verification and security algorithms before sending them back to the local processor for execution. This division of labor improves the execution speed of routine commands while simultaneously enhancing the security level of sensitive operations.
[0067] The system records the execution status of commands to a log database and returns the operation results to the user via a voice feedback module. The log database has a pre-defined storage format and includes fields such as user identity, command type, execution time, and execution result. When a command is completed, such as successful floor registration or brightness adjustment, the above information is automatically entered into the database for easy future tracking. The voice feedback module is a 2W speaker installed inside the car. It pre-stores voice templates for operation results. When a command is executed, the module calls the corresponding template and plays the feedback voice message to the user, allowing them to promptly confirm the operation result and avoid repeated operations due to unclear information.
[0068] like Figure 2 As shown, the present invention also discloses an elevator identity command correction system based on audio recognition, comprising: Acquisition module 1 is used to acquire user voice input. It obtains voice signals through a multi-microphone array set in the elevator car and uses an adaptive noise reduction algorithm to preprocess the voice signals to improve the signal-to-noise ratio. Extraction module 2 is used to extract voiceprint features from the preprocessed speech signal, compare the extracted voiceprint features with the pre-stored voiceprint database to determine the user's identity, and convert the speech commands into text commands through the speech recognition engine. Analysis module 3 is used to perform semantic analysis on the text instructions and classify the instructions into elevator instructions, environmental control instructions or scenario-based composite instructions using an intent understanding engine. Decision module 4 is used to make intelligent decisions based on user identity and instruction type. When the instruction is an elevator ride instruction, it verifies the user's access rights to the target floor. When the instruction is an environmental control instruction, it queries the user preference library and generates an environmental parameter adjustment instruction. When the instruction is a scenario-based composite instruction, it parses the predefined scenario rules and coordinates the execution of elevator ride and environmental control operations. Control module 5 is used to control the elevator operation system to register the target floor or control the intelligent lighting system to adjust the car environment parameters based on the intelligent decision results, and update the personalized settings data in the user preference library.
[0069] In one embodiment, the acquisition module includes: The acquisition unit is used to acquire user voice signals inside the car through a multi-microphone array arranged in a ring; The noise reduction unit is used to process the acquired speech signal using an adaptive noise reduction algorithm based on the least mean square filter to suppress elevator fan noise and mechanical vibration noise. The transform unit is used to perform framing and windowing processing on the denoised speech signal, and then performs short-time Fourier transform after applying the Hamming window function to reduce spectral leakage. The output unit is used to enhance the clarity of the speech signal through dynamic range compression and echo cancellation modules, and output the pre-processed speech data to the voiceprint recognition module.
[0070] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described elevator identity instruction correction method based on audio recognition.
[0071] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described elevator identity instruction correction method based on audio recognition.
[0072] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0074] The above description is merely a preferred embodiment of the present invention and does not limit the scope of this application. Any equivalent results or equivalent process transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.
Claims
1. A method for correcting elevator identity commands based on audio recognition, characterized in that, Includes the following steps: User voice input is collected, and the voice signal is acquired through a multi-microphone array set in the elevator car. An adaptive noise reduction algorithm is used to preprocess the voice signal to improve the signal-to-noise ratio. The preprocessed speech signal is subjected to voiceprint feature extraction. The extracted voiceprint features are compared with the pre-stored voiceprint database to determine the user's identity. At the same time, the speech command is converted into a text command through the speech recognition engine. The text instructions are semantically analyzed, and the instructions are classified into elevator instructions, environmental control instructions, or scenario-based composite instructions using an intent understanding engine. Natural language processing technology is used to parse the text instructions and extract key verbs and nouns to match them with predefined instruction templates. The intent understanding engine categorizes instructions that match the elevator riding template as elevator riding instructions, instructions that match the environmental control template as environmental control instructions, and instructions that match the scene template as scene-based composite instructions. The scene rules corresponding to the scenario-based composite instructions are parsed, and the scene rules include target floor parameters, lighting parameters, and temperature and humidity parameters. When multiple conflicting user commands are detected, the execution order of the commands is determined based on user identity priority and timestamp sorting. Intelligent decisions are made based on user identity and command type. When the command is an elevator ride command, the user's access rights to the target floor are verified. When the command is an environmental control command, the user preference library is queried and an environmental parameter adjustment command is generated. When the command is a scenario-based composite command, the predefined scenario rules are parsed and the elevator ride operation and environmental control operation are executed in a coordinated manner. For elevator ride commands, the mapping relationship between users and floors is queried in the permission database. If the user does not have permission to access the target floor, the execution is refused and a voice prompt is generated. For environmental control commands, the corresponding user's lighting color temperature and brightness parameters are obtained from the user preference library and an environmental adjustment command is generated. For scenario-based composite commands, the target floor registration operation and the car environment parameter adjustment operation are executed simultaneously. When a conflict of instructions is detected, the execution order is determined based on user identity priority and the urgency of the scenario. For high-risk commands, initiate cloud-based security authentication and record operation logs to the blockchain evidence storage system; Based on the intelligent decision-making results, the system controls the elevator operation system to register the target floor or controls the intelligent lighting system to adjust the car environment parameters and updates the personalized settings data in the user preference database.
2. The elevator identity command correction method based on audio recognition according to claim 1, characterized in that, The steps of collecting user voice input, acquiring voice signals through a multi-microphone array installed in the elevator car, and preprocessing the voice signals using an adaptive noise reduction algorithm to improve the signal-to-noise ratio include: User voice signals inside the elevator car are collected using a multi-microphone array arranged in a ring. An adaptive noise reduction algorithm based on the least mean square filter is used to process the collected speech signal to suppress elevator fan noise and mechanical vibration noise. The denoised speech signal is framed and windowed, and a short-time Fourier transform is performed after the Hamming window function is applied to reduce spectral leakage. The clarity of the speech signal is enhanced by dynamic range compression and echo cancellation modules, and the pre-processed speech data is output to the voiceprint recognition module.
3. The elevator identity command correction method based on audio recognition according to claim 1, characterized in that, The steps of extracting voiceprint features from the preprocessed speech signal, comparing the extracted voiceprint features with a pre-stored voiceprint database to determine the user's identity, and converting speech commands into text commands using a speech recognition engine include: The fundamental frequency profile, formant frequencies, and Mel frequency cepstral coefficients are extracted from the preprocessed speech signal as speakerprint features. The dynamic time warping algorithm is used to calculate the minimum distance between the extracted features and the feature sequences in the voiceprint database. When the distance is lower than a preset threshold, the user's identity is confirmed. The speech command is converted into a text command by a speech recognition engine based on a hidden Markov model, and the conversion result is processed for grammatical correction. By combining the voiceprint recognition results with the analysis of the speech segments before and after the instruction words, the validity of the instruction is verified and non-instructional speech input is filtered out.
4. The elevator identity command correction method based on audio recognition according to claim 1, characterized in that, The steps of controlling the elevator operation system to register the target floor or controlling the intelligent lighting system to adjust the car environment parameters based on the intelligent decision-making results, and updating the personalized settings data in the user preference database include: The elevator control unit executes the target floor registration operation, triggering the elevator car to move to the designated floor. The intelligent LED driver module controls the color temperature and brightness of the car lights; When an environment fine-tuning instruction is received, the environment parameters in the user preference library are updated according to a preset step size and bound to the user identity; The offline-cloud collaborative computing framework handles routine instructions and sensitive operations, with routine instructions processed by the local FPGA and sensitive operations uploaded to the cloud for processing. Record the execution status of instructions to the log database and return the operation results to the user through the voice feedback module.
5. An elevator identity command correction system based on audio recognition, characterized in that, include: The acquisition module is used to acquire user voice input. It obtains voice signals through a multi-microphone array set in the elevator car and uses an adaptive noise reduction algorithm to preprocess the voice signals to improve the signal-to-noise ratio. The extraction module is used to extract voiceprint features from the preprocessed speech signal, compare the extracted voiceprint features with the pre-stored voiceprint database to determine the user's identity, and convert the speech commands into text commands through the speech recognition engine. The analysis module is used to perform semantic analysis on the text instructions and classify the instructions into elevator instructions, environmental control instructions, or scenario-based composite instructions using an intent understanding engine. Natural language processing technology is used to parse the text instructions and extract key verbs and nouns to match them with predefined instruction templates. The intent understanding engine categorizes instructions that match the elevator riding template as elevator riding instructions, instructions that match the environmental control template as environmental control instructions, and instructions that match the scene template as scene-based composite instructions. The scene rules corresponding to the scenario-based composite instructions are parsed, and the scene rules include target floor parameters, lighting parameters, and temperature and humidity parameters. When multiple conflicting user commands are detected, the execution order of the commands is determined based on user identity priority and timestamp sorting. The decision-making module is used to make intelligent decisions based on user identity and command type. When the command is an elevator ride command, it verifies the user's access rights to the target floor. When the command is an environmental control command, it queries the user preference library and generates an environmental parameter adjustment command. When the command is a scenario-based composite command, it parses the predefined scenario rules and coordinates the execution of elevator ride and environmental control operations. Specifically, for elevator ride commands, it queries the user-floor mapping relationship in the permission database. If the user does not have permission to access the target floor, it refuses to execute the command and generates a voice prompt. For environmental control commands, the corresponding user's lighting color temperature and brightness parameters are obtained from the user preference library and an environmental adjustment command is generated. For scenario-based composite commands, the target floor registration operation and the car environment parameter adjustment operation are executed simultaneously. When a conflict of instructions is detected, the execution order is determined based on user identity priority and the urgency of the scenario. For high-risk commands, initiate cloud-based security authentication and record operation logs to the blockchain evidence storage system; The control module is used to control the elevator operation system to register the target floor or control the intelligent lighting system to adjust the car environment parameters based on the intelligent decision results, and update the personalized settings data in the user preference library.
6. The elevator identity command correction system based on audio recognition according to claim 5, characterized in that, The acquisition module includes: The acquisition unit is used to acquire user voice signals inside the car through a multi-microphone array arranged in a ring; The noise reduction unit is used to process the acquired speech signal using an adaptive noise reduction algorithm based on the least mean square filter to suppress elevator fan noise and mechanical vibration noise. The transform unit is used to perform framing and windowing processing on the denoised speech signal, and then performs short-time Fourier transform after applying the Hamming window function to reduce spectral leakage. The output unit is used to enhance the clarity of the speech signal through dynamic range compression and echo cancellation modules, and output the pre-processed speech data to the voiceprint recognition module.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.