Elevator intercom system based on AI real-time identification

By using AI real-time recognition technology, based on three-dimensional location tags and hierarchical channel control, the voice task mapping and resource distribution of the elevator intercom system are optimized, solving the problems of low efficiency of manual response and channel interference in traditional elevator intercom systems, and realizing the continuity of voice interaction and the stability of emergency response.

CN121565170APending Publication Date: 2026-02-24SHENZHEN LINGYUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511713531.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional elevator intercom systems rely on manual response, resulting in low communication efficiency, a lack of emotion recognition and task allocation capabilities, and severe channel interference in multi-elevator scenarios, which affects emergency rescue efficiency.

Method used

By using AI real-time recognition technology, three-dimensional location tags are generated based on building number, unit number, and floor number, enabling effective mapping of voice tasks in the spatial domain, hierarchical control of voice channels, identification of overlapping conflict intervals, optimization of voice task resource distribution, and formation of intelligent control scheme.

Benefits of technology

Ensure accurate correspondence between voice recognition commands and location, reduce information interference, achieve call continuity and semantic integrity, and improve the stability of emergency response and the efficiency of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565170A_ABST
    Figure CN121565170A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent man-machine interaction, in particular to an elevator intercom system based on AI real-time identification, which comprises an address index module, a channel layering module, a mute segment reconstruction module, an access freezing module and a resource isolation module. According to the invention, logic mapping is realized in a spatial domain through a voice task, a voice instruction and position corresponding relation is more accurate, instruction identification and task distribution are processed in a structured manner, concurrent conflicts and channel interference are reduced, conversation contents are dynamically reconstructed during an alternating period of manual talkback and intelligent voice accompanying, and semantic coherence and smooth communication are maintained. The rhythm is regulated and controlled according to the state lock value in semantic task execution, instruction dislocation and response lag are avoided, multi-car voice resource distribution is spatially optimized, communication channel interference is restrained, voice interaction forms a collaborative mechanism in safety, response efficiency and semantic consistency, and emergency response stability and interaction continuous experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent human-computer interaction technology, and in particular to an elevator intercom system based on real-time AI recognition. Background Technology

[0002] The field of intelligent human-computer interaction technology mainly involves the methods of information exchange between computers and users, covering key aspects such as speech recognition, image recognition, natural language processing, biometric recognition, emotion recognition, multimodal perception, and interaction mechanisms. It aims to enable machines to understand and respond to human language, expressions, actions, and psychological states through software and hardware means, forming systematic solutions in user interface design, interactive behavior recognition, and intelligent control response. These solutions are widely applied in various industrial scenarios such as smart terminals, service robots, autonomous driving systems, and smart homes. Traditional elevator intercom systems, for example, involve passengers triggering intercom communication via an emergency button inside the elevator. A human operator answers the call, communicates via voice, and initiates rescue. An audio channel is established between the elevator intercom terminal and the control room or service platform, allowing the operator to comfort and guide the trapped passenger and contact property maintenance and other rescue resources. This communication process relies heavily on manual response, has low information transmission efficiency, and lacks capabilities such as emotion recognition and language analysis.

[0003] Traditional elevator intercom systems rely on manual response during communication. After a passenger triggers the intercom, they need to wait for a human to answer. If the staff member is handling multiple requests for help at the same time, the voice channel is easily occupied, leading to response delays. The system only establishes a single audio channel and lacks semantic recognition and task diversion capabilities. It is difficult to judge the status of trapped persons based on the on-site environment or voice characteristics. Voice communication is limited to basic reassurance and information transmission, and fails to achieve dynamic semantic feedback and intelligent judgment. In multi-elevator scenarios, when multiple intercom terminals communicate concurrently, channel interference is obvious, and audio crosstalk and command confusion are likely to occur, affecting the efficiency of emergency rescue and the continuity of voice communication. Summary of the Invention

[0004] To address the technical problems existing in the prior art, this invention provides an elevator intercom system based on real-time AI recognition. The technical solution is as follows: On the one hand, an elevator intercom system based on AI real-time recognition is provided, the system including: The address index module generates three-dimensional location labels based on the triplet of building number, unit number and floor number. Based on the preset logical partitioning rules, it determines whether the labels meet the requirements. By adjusting the label structure, it realizes the effective mapping of the voice task in the spatial domain and associates it with the voice task to obtain the voice task binding number. Based on the voice task binding number, the channel layering module associates and extracts voice command channel information, attaches the user-triggered AI voice companionship task to the corresponding logical partition, generates command queue node number, and analyzes the task concurrency within the partition according to the number to obtain the voice channel layering control result. Based on the hierarchical control results of the voice channel, the silent segment reconstruction module identifies the intersection and conflict intervals between AI companion voice and human intercom voice during the call, extracts the semantic chain structure matching potential empty window intervals, and obtains the call content reconstruction results. Based on the reconstructed call content, the access freeze module calculates the semantic task stack depth and remaining execution cycle, forms a state lock value, determines whether the task is in a critical execution segment, and obtains an intelligent control scheme for elevator voice interaction.

[0005] As a further embodiment of the present invention, the voice task binding number includes a triplet label of building number, unit number, and floor number, logical partition mapping result, and adjusted label structure; the voice channel layer control result includes instruction queue node number, partition task concurrency record, and channel layer status; the call content reconstruction result includes signal segment replacement logic, semantic chain matching result, and empty window interval identification data; and the elevator voice interaction intelligent control scheme includes state lock value generation logic, key execution segment judgment result, and access freeze mark.

[0006] As a further aspect of the present invention, the address indexing module includes: The label generation submodule extracts the main control board number and floor parameters based on the triplet of building number, unit number and floor number, identifies the initial three-dimensional position label, determines whether the label conforms to the logical partitioning rules, and obtains the label structure adjustment requirements. The tag optimization submodule calls the tag structure adjustment requirements, analyzes the redundant parts of the tags, optimizes the number of tags and their arrangement order, and generates three-dimensional location tags that conform to the partitioning rules. The logical partition mapping submodule extracts partition mapping rules based on the three-dimensional location labels that conform to the partitioning rules, compares the matching of labels with the rules, determines the validity of the labels, and obtains the voice task binding number.

[0007] As a further aspect of the present invention, the channel layering module includes: The instruction mounting submodule extracts the voice instruction channel information based on the voice task binding number, mounts the user-triggered AI voice companion task to the corresponding logical partition, and generates an instruction queue node number. The partition concurrency analysis submodule calls the instruction queue node number to analyze the task concurrency within the partition, determine whether there is a risk of resource preemption, and obtain the partition task concurrency record; The channel layering control submodule extracts channel layering rules and analyzes channel layering status based on the concurrent records of the partitioned tasks to obtain the voice channel layering control results.

[0008] As a further aspect of the present invention, the voice command channel information refers to, when receiving an AI voice companion task triggered by a user, performing multi-dimensional feature analysis on the input voice signal based on the recognition confidence threshold of the voice input signal not being lower than a preset threshold, and storing the voice task identifier and the voice task binding number in association when the semantic matching degree parameter is greater than the preset threshold according to the analysis semantic matching degree parameter. The concurrent task status within the partition refers to the recording of resource usage information for the corresponding task when the resource load rate within the partition exceeds a preset threshold, based on the partition resource load rate parameter when the instruction queue node number is used.

[0009] As a further aspect of the present invention, the silent segment reconstruction module includes: Based on the hierarchical control results of the voice channel, the empty window interval identification submodule extracts the output time series and response time slice series of AI companion voice and human intercom voice, identifies potential voice empty window intervals, and obtains the empty window interval identification results. The signal segment replacement submodule calls the empty window interval identification result, matches the semantic chain structure, identifies the signal segment replacement logic, fills in the missing content, and obtains the preliminary reconstructed call content; The semantic continuity verification submodule verifies the semantic continuity of the call content based on the initially reconstructed call content, judges the reconstruction effect, optimizes the semantic chain matching logic, and obtains the call content reconstruction result.

[0010] As a further aspect of the present invention, the access freezing module includes: Based on the reconstructed call content, the state lock value calculation submodule calculates the semantic task stack depth and remaining execution cycle, identifies the state lock value, determines whether the task is in a critical execution segment, and obtains the critical execution segment judgment result. The channel admission control submodule calls the judgment result of the key execution section, freezes the edge channel admission, puts the request into the queuing buffer, performs AI task release locking, and obtains the admission freeze mark; Based on the admission freeze flag, the scheduling optimization submodule optimizes the voice channel scheduling logic, analyzes task independence and signal scheduling efficiency, and obtains an intelligent control scheme for elevator voice interaction.

[0011] As a further aspect of the present invention, the system also includes a resource isolation module: Based on the aforementioned intelligent control scheme for elevator voice interaction, the resource isolation module extracts the physical space overlap information of the multi-car system, optimizes the distribution of voice task resources, adjusts the activation logic of the backup channel of the edge gateway, and obtains the voice task resource isolation result.

[0012] As a further aspect of the present invention, the voice task resource isolation result includes physical space overlap information, backup channel activation records, and resource distribution optimization values.

[0013] As a further aspect of the present invention, the resource isolation module includes: The spatial overlap analysis submodule, based on the aforementioned elevator voice interaction intelligent control scheme, extracts the physical space overlap information of the multi-car system, analyzes the impact of the overlap area on the voice task, optimizes the resource distribution logic, and obtains the physical space overlap information. The backup channel activation submodule calls the physical space overlap information, adjusts the backup channel activation logic of the edge gateway, optimizes the independence of voice tasks, and obtains the backup channel activation record. The resource distribution optimization submodule optimizes the distribution of voice task resources based on the backup channel activation record, adjusts the signal scheduling efficiency in multi-task concurrent scenarios, and obtains the voice task resource isolation result.

[0014] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following: By implementing logical mapping in the spatial domain through voice tasks, the correspondence between voice recognition commands and specific locations is ensured to be more accurate. Command recognition and task allocation in voice interaction are processed in a structured manner, reducing information interference caused by concurrent conflicts and channel overlap. Voice content is dynamically reconstructed during the alternation of human intercom and intelligent voice accompaniment, so that the call process remains coherent and semantically complete. During the execution of semantic tasks, the task rhythm is dynamically adjusted according to the state lock value, effectively avoiding command misalignment and response lag. The distribution of voice resources in multiple car cabins is spatially optimized, and interference between communication channels is significantly suppressed. Voice interaction forms a coordination mechanism in terms of safety, response efficiency and semantic consistency, improving the stability of emergency response and the continuous experience of voice interaction. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of an elevator intercom system based on real-time AI recognition provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the system framework of the present invention; Figure 3 This is a flowchart of the address indexing module in this invention; Figure 4 This is a flowchart of the channel layering module in this invention; Figure 5 This is a flowchart of the silent segment reconstruction module in this invention; Figure 6 This is a flowchart of the admission freezing module in this invention; Figure 7 This is a flowchart of the resource isolation module in this invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] This invention provides an elevator intercom system based on real-time AI recognition, such as... Figure 1-2 The diagram shown illustrates an elevator intercom system based on real-time AI recognition. The system includes: The address index module generates three-dimensional location labels based on the triplet of building number, unit number and floor number. Based on the preset logical partitioning rules, it determines whether the labels meet the requirements. By adjusting the label structure, it realizes the effective mapping of the voice task in the spatial domain and associates it with the voice task to obtain the voice task binding number. The channel layering module associates and extracts voice command channel information based on the voice task binding number, attaches the AI ​​voice companion task triggered by the user to the corresponding logical partition, generates command queue node number, and analyzes the task concurrency within the partition according to the number to obtain the voice channel layering control result. The silent segment reconstruction module, based on the hierarchical control results of the voice channel, identifies the intersection and conflict intervals between AI companion voice and human intercom voice during the call, extracts the semantic chain structure matching potential empty window intervals, and obtains the call content reconstruction results. Based on the call content reconstruction results, the access freeze module calculates the semantic task stack depth and remaining execution cycle to form a state lock value, determines whether the task is in a critical execution segment, and obtains an intelligent control scheme for elevator voice interaction. The resource isolation module is based on the elevator voice interaction intelligent control scheme. It extracts the physical space overlap information of the multi-car system, optimizes the distribution of voice task resources, and adjusts the backup channel activation logic of the edge gateway to obtain the voice task resource isolation result.

[0023] The voice task binding number includes the triplet label of building number, unit number, and floor number, the logical partition mapping result, and the adjusted label structure. The voice channel layer control result includes the instruction queue node number, partition task concurrency record, and channel layer status. The call content reconstruction result includes the signal segment replacement logic, semantic chain matching result, and empty window interval identification data. The elevator voice interaction intelligent control scheme includes the state lock value generation logic, key execution section judgment result, and access freeze mark. The voice task resource isolation result includes physical space overlap information, backup channel activation record, and resource distribution optimization value.

[0024] Specifically, such as Figure 2 , 3 As shown, the address index module includes: The label generation submodule extracts the main control board number and floor parameters based on the triplet of building number, unit number and floor number, identifies the initial three-dimensional position label, determines whether the label conforms to the logical partitioning rules, and obtains the label structure adjustment requirements. The system receives a triplet of information consisting of building number "Building A", unit number "Unit 2", and floor number "Floor 25". First, it extracts the main control board number and floor parameters from the configuration data table. For example, based on "Building A", the main control board number is matched as "MCB-001"; based on "Unit 2", the sub-board area number is matched as "SUB-002"; and based on "Floor 25", the floor parameters are matched as "floor height coefficient 0.8, physical floor 25". Then, this extracted information is combined to generate the original form of the 3D location label: "MCB-001-SUB-002-F25". Next, the compliance of the label is determined according to the set logical partitioning rules. The logical partitioning rules stipulate that "the main control board number begins with MCB, the sub-board area number begins with SUB, and the floor number..." The label “MCB-001-SUB-002-F25” is followed by two digits, ranging from 01 to 30. A sub-board area can support a maximum of 30 physical layer stations. By checking the format and value of the label “MCB-001-SUB-002-F25” item by item, it is confirmed that “MCB-001” meets the “MCB” rule, “SUB-002” meets the “SUB” rule, and “F25” meets the “F followed by two digits ranging from 01 to 30” rule. In addition, the number of layer stations currently allocated in this sub-board area is 18, which is lower than the maximum limit of 30 layer stations. When all the checks pass, the label is determined to meet the logical partitioning rules. If there are any non-compliant items, they are marked as non-compliant. In this example, the label is determined to meet the logical partitioning rules, and there is no label structure adjustment instruction.

[0025] The label optimization submodule calls the label structure adjustment requirements, analyzes the redundant parts of the labels, optimizes the number of labels and their arrangement order, and generates three-dimensional position labels that conform to the partitioning rules. Based on the aforementioned determination of the unlabeled structure adjustment instruction, the data redundancy items of the label "MCB-001-SUB-002-F25" were analyzed. By comparing the label with the minimum valid label length specification, it was found that the main control board number occupies 6 characters, the sub-board area number occupies 5 characters, and the floor / station number occupies 3 characters, for a total length of 14 characters. However, the current label "MCB-001-SUB-002-F25" has a total length of 19 characters. Therefore, it was discovered that the "MCB-001" part of the label can be shortened. Written as "M1", "SUB-002" can be abbreviated as "S2", and the floor number "F25" remains unchanged. The number of digits in the label is adjusted accordingly. At the same time, according to the principle of best access efficiency, the label order is rearranged as "floor number - main control board number - sub-board area number". For example, floor information is accessed frequently in the elevator system, so "F25" is placed first, followed by main control board information, and then sub-board area information. Through this optimization operation, the three-dimensional location label "F25-M1-S2" that conforms to the zoning rules is finally generated.

[0026] The logical partition mapping submodule extracts partition mapping rules based on 3D location labels that conform to partitioning rules, compares the matching of labels with the rules, determines the validity of the labels, and obtains the voice task binding number. Based on the 3D location label "F25-M1-S2" conforming to the partitioning rules, rule items are first extracted from the logical partitioning mapping rule database. For example, the rule is set as "labels with headers of F01 to F10 are mapped to the low partition L1, labels with headers of F11 to F20 are mapped to the middle partition M1, and labels with headers of F21 to F30 are mapped to the high partition H1". After extraction, the label "F25-M1-S2" is compared with the rules, and the label header "F25" is extracted. Then, "F25" is compared with the range of each partitioning rule. The numerical comparison identifies "F25" as falling within the range of "F21" to "F30" and matches the mapping rule "High-level partition H1". Based on the matching result, the validity status of the label "F25-M1-S2" is determined. If the label successfully matches a unique logical partition, it is considered a valid label. If it fails to match or matches multiple partitions, it is considered invalid. In this example, the label successfully matches "High-level partition H1" and is considered valid. The voice task binding number "VTB-H1-003" associated with "High-level partition H1" is obtained.

[0027] Specifically, such as Figure 2 , 4 As shown, the channel layering module includes: The instruction mounting submodule extracts voice instruction channel information based on the voice task binding number, mounts the user-triggered AI voice companion task to the corresponding logical partition, and generates an instruction queue node number. Voice command channel information refers to the process of performing multi-dimensional feature analysis on the input voice signal when a user-triggered AI voice companion task is received, based on the recognition confidence threshold of the voice input signal not being lower than a preset threshold. According to the analyzed semantic matching degree parameter, when the semantic matching degree parameter is greater than the preset threshold, the voice task identifier and the voice task binding number are associated and stored. Based on the voice task binding number "VTB-H1-003", the system receives the user-triggered AI voice companionship task. First, it extracts the voice command channel information. Upon initiation, it monitors the recognition confidence level of the voice input signal, setting a confidence level threshold of 0.80. When the recognition confidence level of the user's voice "Please play background music" reaches 0.92, exceeding the set threshold of 0.80, it performs multi-dimensional feature analysis on the input voice signal, including voice waveform features, spectrogram features, and prosodic features. By training a semantic parsing model, it analyzes its semantic matching degree parameter. This parameter quantifies the degree of fit between the voice command and the semantic template of the AI ​​companionship task. The semantic matching degree threshold is set to... The semantic matching degree parameter of the current voice command is calculated to be 0.088, which is higher than the set threshold of 0.075. The voice task identifier "PlayMusic_A001" is associated with the voice task binding number "VTB-H1-003" and stored together, indicating that the task has passed the validity confirmation and is ready to be mounted. Then, based on the binding number "VTB-H1-003", the task "PlayMusic_A001" is located in "high-level partition H1" and mounted to the task queue of this logical partition. After the mounting is completed, the instruction queue node number "QNode-H1-003-20251010-001" is generated for the mounting operation.

[0028] The partition concurrency analysis submodule calls the instruction queue node number to analyze the task concurrency within the partition, determine whether there is a risk of resource preemption, and obtain the partition task concurrency record; The concurrency status of tasks within a partition refers to the resource usage information of the corresponding task when the resource load rate within the partition exceeds a preset threshold, based on the partition resource load rate parameter when the instruction queue node number is used. The command queue node number "QNode-H1-003-20251010-001" is invoked to analyze the task concurrency within its logical partition "High-Level Partition H1". First, the resource load rate parameter of "High-Level Partition H1" is monitored. This parameter considers CPU utilization, memory usage, and network bandwidth utilization. CPU utilization is set as the percentage of time the CPU spends processing tasks per unit of time; memory usage is set as the percentage of memory required by the current task queue relative to the total memory of the partition; and network bandwidth utilization is set as the percentage of the current partition's data transfer rate relative to its maximum rate. The current partition resource load rate is... The source load rate threshold is 0.75. When the CPU utilization rate within the partition is 0.65, the memory utilization rate is 0.70, and the network bandwidth utilization rate is 0.60, the calculated comprehensive load rate is (0.65+0.70+0.60) / 3=0.65. This value is lower than the set threshold of 0.75. If the calculated result is higher than the threshold, the resource usage information of the task is recorded, such as task ID, peak CPU usage, and peak memory usage. Based on the results of this analysis, there is no risk of resource preemption in the current partition "High-level partition H1". The partition task concurrency record is displayed as "Partition H1 has no risk of preemption, load rate 0.65".

[0029] The channel layering control submodule extracts channel layering rules and analyzes channel layering status based on concurrent recording of partitioned tasks to obtain voice channel layering control results; Based on the concurrent record of partitioned tasks, "Partition H1 has no preemption risk and a load rate of 0.65", the pre-stored items of the channel layering rules are first extracted. These rules specify the layering strategy for voice channels under different load states. For example, "When the partition load rate is below 0.70, the primary channel maintains priority 1 and the backup channel maintains priority 2", "When the partition load rate is between 0.70 and 0.90, the primary channel priority is reduced to 2 and the backup channel priority is reduced to 3", and "When the partition load rate is above 0.90, the primary channel priority is reduced to 3 and the backup channel is suspended". Then, the current channel layering status is analyzed. The current load rate of "high-level partition H1" is 0.65. According to the channel layering rules, this load rate is below the threshold of 0.70. The primary channel maintains priority 1 and the backup channel maintains priority 2. Through this comparison and matching operation, it is determined that the current voice channel does not need to be layered and adjusted. The voice channel layering control result is "primary channel priority 1, backup channel priority 2, no adjustment required".

[0030] Specifically, such as Figure 2 , 5 As shown, the silent segment reconstruction module includes: The empty window interval identification submodule extracts the output time series and response time slice series of AI companion voice and human intercom voice based on the hierarchical control results of the voice channel, identifies potential voice empty window intervals, and obtains the empty window interval identification results. Based on the voice channel hierarchical control result "primary channel priority 1, backup channel priority 2, no adjustment required", the output time sequence of the AI ​​companion voice and the response time slice sequence of the human intercom voice are first extracted from the real-time communication stream. For example, the AI ​​companion voice is sent at 10:05:32.120, lasts for 1.50 seconds, and ends at 10:05:33.620; the human intercom voice starts at 10:05:34.250, lasts for 2.00 seconds, and ends at 10:05:36.250. The end time of the AI ​​companion voice at 10:05:33.620 and the start time of the human intercom voice at 10:05: Comparing the two times, the time difference is calculated as 10:05:34.250 - 10:05:33.620 = 0.630 seconds. The threshold for the speech window interval is 0.50 seconds. When the calculated time difference exceeds this threshold, i.e., 0.630 > 0.50, a speech window interval signal is identified, with a start time of 10:05:33.620, an end time of 10:05:34.250, and a duration of 0.630 seconds. The result of the window interval identification is "A window interval exists, start time 10:05:33.620, end time 10:05:34.250".

[0031] The signal segment replacement submodule calls the empty window interval recognition results, matches the semantic chain structure, identifies the signal segment replacement logic, fills in the missing content, and obtains the preliminary reconstructed call content; The system retrieves the "empty window interval recognition result" which states that "an empty window interval exists, with a start time of 10:05:33.620 and an end time of 10:05:34.250". First, it analyzes the dialogue content before and after the empty window interval to match the semantic chain structure. For example, the AI ​​voice content before the empty window interval is "Hello, how can I help you?" and the human voice content after the empty window interval is "I want to reach the 20th floor". The system identifies that this semantic chain is missing a middle content item in the question-answer pair. Then, it identifies the signal segment replacement logic. Scenario 1: Simple Waiting. Based on the specific scenario of "AI waiting for a response during a question-and-answer session," the intelligent guidance logic is triggered. Since the system is configured with AI possessing high-precision speech recognition capabilities, there will be no "unclear" situations. Therefore, the silence is interpreted as the passenger thinking or hesitating. The AI's goal is to break the silence and guide the conversation, rather than simply requesting repetition. This logic pre-sets a series of context-based guiding or prompting phrases, such as "Can you tell me which floor you want to go to?", "Would you like me to introduce the layout of the floors in this building?", or "Don't rush, take your time." In the current scenario, the system determines that the most likely intention is to go to a certain floor, therefore selecting "Can you tell me which floor you want to go to?" as the most direct and effective supplementary content. This guiding statement is inserted at an appropriate time point within the silence (e.g., after the start time 10:05:33.620), supplementing the missing content, and the reconstructed call content is: "AI: Hello, how can I help you? (Passenger briefly silent) AI: Can you tell me which floor you want to go to? Passenger: I want to go to the 20th floor." Scenario 2: Intelligent Interruption and Emergency Screening (Two-way Interaction). While the AI ​​is outputting voice commands or waiting, the system monitors passenger input in real time. If the passenger's voice conflicts with the AI's current task or contains a higher-priority intent, the AI ​​should intelligently interrupt itself. For example, if the AI ​​is playing soothing music or reporting the weather, and at this time recognizes a passenger's voice saying, "Wait, I'm not going to the 20th floor, I'm going to the 10th floor," or "Someone has fainted, please help!" a. Intent and Priority Determination: The system will recognize the passenger's voice content in real time and classify and prioritize the intent (e.g., emergency assistance > target correction > routine inquiry > invalid noise). b. Response Logic Switching: If the response is determined to be an "emergency request for help" or a "goal correction," the AI ​​immediately interrupts its current task (e.g., stops broadcasting, abandons follow-up questions) and switches to the corresponding response process. For example, in response to "someone has fainted," the AI ​​should immediately execute the emergency plan, prioritize connecting to human customer service and reporting the situation, while simultaneously playing reassurance instructions. In response to "go to the 10th floor," the AI ​​should abandon its previous understanding of "go to the 20th floor" and confirm the new instruction. c. Call content reconstruction: The reconstruction result will reflect this interruption process. For example: "AI: Okay, heading to the 20th floor for you, estimated... Passenger: (interrupting) Wait, I'm not going to the 20th floor, it's the 10th floor. AI: Okay, I've corrected your destination floor to the 10th floor."

[0032] The semantic continuity verification submodule verifies the semantic continuity of the call content based on the initially reconstructed call content, judges the reconstruction effect, optimizes the semantic chain matching logic, and obtains the call content reconstruction result. Based on the initially reconstructed call content "AI: Hello, how can I help you? AI: What did you just say? Passenger: I want to go to the 20th floor," the semantic continuity of the reconstructed call content is first verified. This verification process is achieved by calculating a contextual coherence score. For example, "What did you just say?" and "I want to go to the 20th floor" are input into the language model, and the cosine similarity of their semantic vectors is calculated. With a semantic continuity verification threshold of 0.75, the current similarity is calculated to be 0.82, which is above the threshold, indicating successful reconstruction. If it is below the threshold, the effect is considered poor, and the semantic chain matching logic is optimized. If the verification finds a logical error in the AI's response to the user's command (e.g., the AI ​​is still confirming the original floor after the user corrects it), the system will mark this dialogue as "interaction failed" and trigger manual intervention or use a backup strategy. After completing the above verification, the module will further determine in real-time whether the AI ​​companionship task should end based on DialogueStateTracking. Termination conditions include: 1) the task is clearly completed (e.g., the passenger reaches the designated floor and makes no new requests); 2) the passenger issues a clear termination instruction (e.g., "Okay, thank you"); 3) after completing the task, the passenger has no effective interaction for an extended period; 4) a human customer service representative has intervened in the call. When any termination condition is met, the system will generate a corresponding termination interaction instruction (e.g., the AI ​​announces "We've reached the 20th floor, have a pleasant day"), ensuring that the AI ​​can exit appropriately and politely after completing the task, avoiding unnecessary disturbances, and ultimately outputting a complete reconstruction of the call content.

[0033] Specifically, such as Figure 2 , 6 As shown, the admission freeze module includes: The state lock value calculation submodule calculates the semantic task stack depth and remaining execution cycle based on the call content reconstruction result, identifies the state lock value, determines whether the task is in the critical execution segment, and obtains the critical execution segment judgment result. The semantic task stack depth is calculated using the following formula: ; Where D represents the semantic task stack depth. This represents the expected processing time for the i-th task. This represents the remaining execution time of the i-th task. This represents the execution cycle of the i-th task. The control parameters represent the i-th task, and n is the total number of tasks; Based on the call content reconstruction results, the semantic task stack depth D is calculated. This depth reflects the complexity and potential risks of the current system when handling multiple tasks. The formula is as follows: Where D represents the semantic task stack depth, indicating the overall complexity of the system's task processing; the larger the value, the more complex the task processing or the deviation from expectations. Representing the The expected processing time for each task, in seconds, represents the time required to complete the task under ideal conditions. Representing the The remaining execution time for each task, in seconds, indicates how much more time is needed for the task to complete. Representing the The execution cycle of a task, expressed in seconds, represents the periodic time interval from the start to the completion of the task. Representing the The control parameters for each task range from [0,1], representing the task's sensitivity to resource scheduling. A larger value indicates a more sensitive task to resource fluctuations. The total number of tasks, i.e., the number of tasks currently being processed or waiting to be processed, is given in the formula. The absolute difference between the expected processing time and the remaining execution time was calculated, reflecting the degree to which the task execution deviates from expectations. The larger the difference, the higher the unpredictability of the task execution. The stability factor representing task execution, where This ensures the weight of periodic tasks, while By taking the square root and absolute value of the control parameters, the impact of resource-sensitive tasks on the denominator is made relatively smooth, avoiding [the negative effects of excessive square roots]. An excessively large denominator leads to a sharp increase, while an excessively small denominator leads to an excessively small denominator. This balances the deviation from and stability of the task overall, resulting in the final summation sign. The individual evaluation values ​​of all tasks are summed to form a macro-level task stack depth indicator; Now assume that the current system has Two tasks are currently being processed: Task 1 is AI-powered voice companionship, and Task 2 is elevator destination response. Parameter settings are as follows: Task 1 (AI-powered voice companionship): Expected processing time. seconds, remaining execution time Seconds, execution cycle Seconds, control parameters Task 2 (Elevator Destination Response): Expected Processing Time seconds, remaining execution time Seconds, execution cycle Seconds, control parameters Among them, the expected processing time This is obtained by statistically analyzing the average execution time of similar historical tasks. For example, by monitoring the execution time of the past 1000 AI companion voice tasks and taking the average value. seconds, remaining execution time The progress is obtained by subtracting the executed time from the total time since the task started, through real-time monitoring of the task progress. For example, the AI ​​companion voice task has been completed. seconds, then Seconds, execution cycle The task scheduling period is set in the configuration, for example, the AI ​​companion task every [time period]. If the status is checked once per second, then Seconds, control parameters Experiments have shown that when a task's resource consumption fluctuates significantly, such as a sudden increase in CPU utilization from 10% to 80% during a computationally intensive phase, The value will be set relatively high, for example Conversely, set it to a lower value, for example... Currently, stress tests show that when the CPU usage of the AI ​​companionship task fluctuates by 30%, Set as When the network latency fluctuation for the elevator destination response task reaches 10%, Set as ; Substitute the parameters into the formula to calculate: For Task 1: ; For Task 2: ; Then the semantic task stack depth The system identifies state lock values ​​and sets a threshold of 0.15. When the D value is higher than the threshold, it determines that the current task stack depth is high and state locking is required; otherwise, locking is not necessary. The currently calculated... If the value is above the threshold of 0.15, the task is determined to be in a critical execution segment. For example, if the voice announcement before the elevator stops is in progress, it needs to be completed with high priority. Therefore, the result of the critical execution segment judgment is "yes".

[0034] The channel admission control submodule calls the judgment result of the key execution section, freezes the edge channel admission, puts the request into the queuing buffer, performs AI task release locking, and obtains the admission freeze mark; When the "critical execution segment judgment result is 'yes'" message is invoked, edge channel access is immediately frozen. New incoming voice interaction requests will be temporarily rejected from directly entering the processing queue. For example, when the judgment result is "yes", the channel access policy switches from "open" to "frozen". All access requests are no longer processed but are directed to the queuing buffer. The request is placed in the queuing buffer. For example, if a voice request arrives at 10:05:40.000, access is frozen at this time. The request is assigned to the buffer queue ID "Buf-001" and will be processed after the critical segment task is completed. At the same time, the AI ​​task release lock is performed, indicating that during the critical execution segment, the currently executing AI task (such as elevator floor broadcast) will not be interrupted or its resources will be released to other tasks to ensure the integrity of the execution. For example, when the judgment result is "yes", the release flag of the AI ​​task "floor broadcast" is changed from "releaseable" to "locked" until the critical execution segment ends. Through this operation, the access freeze mark is obtained as "frozen". The scheduling optimization submodule optimizes the voice channel scheduling logic based on the admission freeze flag, analyzes task independence and signal scheduling efficiency, and obtains an intelligent control scheme for elevator voice interaction. Based on the "frozen" access control flag, the voice channel scheduling logic is first optimized. The scheduler adjusts its scheduling strategy to "prioritize key segment tasks" according to the "frozen" flag. For example, the original polling scheduling strategy is paused, and a priority scheduling mechanism is adopted to prioritize tasks in the current key execution segment (such as elevator safety announcements) to ensure uninterrupted execution. Task independence and signal scheduling efficiency are analyzed. For example, it is identified that the key task "safety announcement" and the "background music playback" task in the queue buffer are independent, do not share core resources, and do not conflict. Simultaneously, the signal scheduling efficiency of key tasks under the current scheduling strategy is evaluated. For example, by monitoring the end-to-end latency of key tasks, it is found that the current latency is 50 milliseconds, lower than the efficiency benchmark of 100 milliseconds, indicating that scheduling efficiency is achieved. Through this analysis, the smooth execution of key tasks is ensured, and the efficiency of other tasks is maximized when conditions permit. An elevator voice interaction control scheme is obtained, which includes "high priority for key tasks, queuing and buffering new requests, and resuming normal scheduling after the key task is completed."

[0035] Specifically, such as Figure 2 , 7 As shown, the resource isolation module includes: The spatial overlap analysis submodule is based on the elevator voice interaction intelligent control scheme. It extracts the physical space overlap information of the multi-car system, analyzes the impact of the overlap area on the voice task, optimizes the resource distribution logic, and obtains the physical space overlap information. Based on the elevator voice interaction control scheme, spatial overlap information of multiple elevator platforms is extracted. For example, in a platform with two adjacent elevator cars (car A and car B), by reading the real-time position sensor data of the elevators, it is determined that car A is currently stopped on the 25th floor and car B is stopped on the 26th floor. Through physical space overlap rules, such as "spatial overlap is determined when the distance between adjacent cars is less than 0.5 meters, or when voice interaction occurs on the same floor", it is analyzed that there is voice signal interference between car A and car B in the area between the 25th and 26th floors, i.e., spatial overlap. The impact of the overlap area on voice tasks is analyzed as follows: it leads to voice privacy leakage and reduced voice recognition accuracy. To solve this impact, the resource distribution logic is optimized. For example, in the overlap area, different communication frequency bands or independent edge processing units are allocated to the voice tasks of car A and car B. The physical space overlap information is obtained as "car A and car B have spatial overlap between the 25th and 26th floors, causing voice interference and privacy risks".

[0036] The backup channel activation submodule calls the physical space overlap information, adjusts the backup channel activation logic of the edge gateway, optimizes the independence of voice tasks, and obtains the backup channel activation record. The system retrieves the physical space overlap information: "Car A and Car B overlap on floors 25-26, causing voice interference and privacy risks." Immediately, it adjusts the edge gateway's backup channel activation logic. This logic, pre-set to detect spatial overlap, selects different backup channel strategies based on the degree of overlap and task type. For example, when "voice interference and privacy risks" are detected, it activates a backup channel with high isolation encryption. It determines that voice interaction tasks for both Car A and Car B require independent states; therefore, it activates the backup channel "Ch_A_Backup_Enc" for Car A and "Ch_B_Backup_Enc" for Car B, optimizing the independent state of voice tasks. This ensures that even in physically overlapping areas, voice communication between the two cars does not interfere with each other, protecting privacy. For example, the backup channels use independent frequency bands and encryption protocols, completely separating data streams. The backup channel activation record is obtained as "Car A: Ch_A_Backup_Enc activated, Car B: Ch_B_Backup_Enc activated."

[0037] The resource distribution optimization submodule optimizes the distribution of voice task resources based on the backup channel activation record, adjusts the signal scheduling efficiency in multi-task concurrent scenarios, and obtains the voice task resource isolation result. Based on the backup channel activation records "Car A: Ch_A_Backup_Enc activated, Car B: Ch_B_Backup_Enc activated", the distribution of voice task resources is optimized. This involves optimizing the dedicated configuration for allocating computing and network resources to the newly activated backup channels to ensure independent operation. For example, "Ch_A_Backup_Enc" is allocated 3 CPU cores and 50MB of memory on the edge gateway, while "Ch_B_Backup_Enc" is allocated 4 CPU cores and 60MB of memory. In multi-task concurrent scenarios, signal scheduling efficiency is adjusted. When a passenger in car A issues the command "play news", the command is routed to CPU core 3, where "Ch_A_Backup_Enc" resides, and is given a higher scheduling priority than the normal type of voice task (for example, the scheduling priority is increased from 5 to 3). This ensures resource isolation without affecting response speed. Through this resource allocation and scheduling adjustment, the voice task resource isolation result is "the voice task in car A is isolated to CPU core 3 and 50MB of memory, and the voice task in car B is isolated to CPU core 4 and 60MB of memory, with the scheduling priority increased, achieving isolation and efficient concurrency".

[0038] In emergency communication or prolonged companionship scenarios, the system provides clear hang-up and exit paths. When the AI ​​makes an emergency call or provides reassurance and companionship, passengers can terminate the service at any time via voice commands such as "hang up," "end call," or "exit companionship" if they wish to end the call. Alternatively, passengers can operate via physical buttons, such as pressing and holding the "open door" or "close door" button for 3 seconds; the system will recognize this as a forced hang-up command and immediately terminate the communication. The end of the companionship mode follows a preset logic: after completing the passenger's requested task (such as reaching the target floor or playing specified content), the AI ​​will proactively ask if continued companionship is needed; if there is no response, the service will end. During a call, if the system asks the passenger three times consecutively (each time with a 10-second interval) and receives no voice or button response, the system will determine it to be in an unresponsive state. At this point, the system will make a final confirmation: "If you do not need help, I will hang up in 10 seconds." If there is still no response after the countdown ends, the system will automatically hang up to free up resources and avoid unnecessary disturbances, while recording a complete log of this interaction for subsequent analysis.

[0039] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An elevator intercom system based on AI real-time recognition, characterized in that, The system includes: The address index module generates three-dimensional location labels based on the triplet of building number, unit number and floor number. Based on the preset logical partitioning rules, it determines whether the labels meet the requirements. By adjusting the label structure, it realizes the effective mapping of the voice task in the spatial domain and associates it with the voice task to obtain the voice task binding number. Based on the voice task binding number, the channel layering module associates and extracts voice command channel information, attaches the user-triggered AI voice companionship task to the corresponding logical partition, generates command queue node number, and analyzes the task concurrency within the partition according to the number to obtain the voice channel layering control result. Based on the hierarchical control results of the voice channel, the silent segment reconstruction module identifies the intersection and conflict intervals between AI companion voice and human intercom voice during the call, extracts the semantic chain structure matching potential empty window intervals, and obtains the call content reconstruction results. Based on the reconstructed call content, the access freeze module calculates the semantic task stack depth and remaining execution cycle, forms a state lock value, determines whether the task is in a critical execution segment, and obtains an intelligent control scheme for elevator voice interaction.

2. The elevator intercom system based on AI real-time recognition according to claim 1, characterized in that: The voice task binding number includes a triplet label of building number, unit number, and floor number, logical partition mapping result, and adjusted label structure. The voice channel layer control result includes instruction queue node number, partition task concurrency record, and channel layer status. The call content reconstruction result includes signal segment replacement logic, semantic chain matching result, and empty window interval identification data. The elevator voice interaction intelligent control scheme includes state lock value generation logic, key execution section judgment result, and access freeze mark.

3. The elevator intercom system based on AI real-time recognition according to claim 1, characterized in that, The address index module includes: The label generation submodule extracts the main control board number and floor parameters based on the triplet of building number, unit number and floor number, identifies the initial three-dimensional position label, determines whether the label conforms to the logical partitioning rules, and obtains the label structure adjustment requirements. The tag optimization submodule calls the tag structure adjustment requirements, analyzes the redundant parts of the tags, optimizes the number of tags and their arrangement order, and generates three-dimensional location tags that conform to the partitioning rules. The logical partition mapping submodule extracts partition mapping rules based on the three-dimensional location labels that conform to the partitioning rules, compares the matching of labels with the rules, determines the validity of the labels, and obtains the voice task binding number.

4. The elevator intercom system based on AI real-time recognition according to claim 3, characterized in that, The channel layering module includes: The instruction mounting submodule extracts the voice instruction channel information based on the voice task binding number, mounts the user-triggered AI voice companion task to the corresponding logical partition, and generates an instruction queue node number. The partition concurrency analysis submodule calls the instruction queue node number to analyze the task concurrency within the partition, determine whether there is a risk of resource preemption, and obtain the partition task concurrency record; The channel layering control submodule extracts channel layering rules and analyzes channel layering status based on the concurrent records of the partitioned tasks to obtain the voice channel layering control results.

5. The elevator intercom system based on AI real-time recognition according to claim 4, characterized in that: The voice command channel information refers to the process of performing multi-dimensional feature analysis on the input voice signal when a user-triggered AI voice companion task is received, based on the recognition confidence threshold of the voice input signal not being lower than a preset threshold. According to the analyzed semantic matching degree parameter, when the semantic matching degree parameter is greater than the preset threshold, the voice task identifier and the voice task binding number are associated and stored. The concurrent task status within the partition refers to the recording of resource usage information for the corresponding task when the resource load rate within the partition exceeds a preset threshold, based on the partition resource load rate parameter when the instruction queue node number is used.

6. The elevator intercom system based on AI real-time recognition according to claim 4, characterized in that, The silent segment reconstruction module includes: Based on the hierarchical control results of the voice channel, the empty window interval identification submodule extracts the output time series and response time slice series of AI companion voice and human intercom voice, identifies potential voice empty window intervals, and obtains the empty window interval identification results. The signal segment replacement submodule calls the empty window interval identification result, matches the semantic chain structure, identifies the signal segment replacement logic, fills in the missing content, and obtains the preliminary reconstructed call content; The semantic continuity verification submodule verifies the semantic continuity of the call content based on the initially reconstructed call content, judges the reconstruction effect, optimizes the semantic chain matching logic, and obtains the call content reconstruction result.

7. The elevator intercom system based on AI real-time recognition according to claim 6, characterized in that, The access freeze module includes: Based on the reconstructed call content, the state lock value calculation submodule calculates the semantic task stack depth and remaining execution cycle, identifies the state lock value, determines whether the task is in a critical execution segment, and obtains the critical execution segment judgment result. The channel admission control submodule calls the judgment result of the key execution section, freezes the edge channel admission, puts the request into the queuing buffer, performs AI task release locking, and obtains the admission freeze mark; Based on the admission freeze flag, the scheduling optimization submodule optimizes the voice channel scheduling logic, analyzes task independence and signal scheduling efficiency, and obtains an intelligent control scheme for elevator voice interaction.

8. The elevator intercom system based on AI real-time recognition according to claim 1, characterized in that, The system also includes a resource isolation module: Based on the aforementioned intelligent control scheme for elevator voice interaction, the resource isolation module extracts the physical space overlap information of the multi-car system, optimizes the distribution of voice task resources, adjusts the activation logic of the backup channel of the edge gateway, and obtains the voice task resource isolation result.

9. The elevator intercom system based on AI real-time recognition according to claim 8, characterized in that: The voice task resource isolation results include physical space overlap information, backup channel activation records, and resource distribution optimization values.

10. The elevator intercom system based on AI real-time recognition according to claim 8, characterized in that, The resource isolation module includes: The spatial overlap analysis submodule, based on the aforementioned elevator voice interaction intelligent control scheme, extracts the physical space overlap information of the multi-car system, analyzes the impact of the overlap area on the voice task, optimizes the resource distribution logic, and obtains the physical space overlap information. The backup channel activation submodule calls the physical space overlap information, adjusts the backup channel activation logic of the edge gateway, optimizes the independence of voice tasks, and obtains the backup channel activation record. The resource distribution optimization submodule optimizes the distribution of voice task resources based on the backup channel activation record, adjusts the signal scheduling efficiency in multi-task concurrent scenarios, and obtains the voice task resource isolation result.