Voice broadcast control method, system and equipment of hospital logistics robot and medium
By using multimodal sensor fusion and a hybrid decision model, precise volume control for hospital logistics robots was achieved, solving the problems of poor scene adaptation and adjustment lag in existing technologies, and improving the effectiveness and environmental adaptability of voice interaction.
Patent Information
- Application Number
- CN202511753663.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-24
AI Technical Summary
The existing voice broadcast volume control of hospital logistics robots suffers from poor scene adaptation, lagging adjustment, and lack of iteration capabilities, making it unable to adapt to the complex and ever-changing environmental needs of hospitals.
By employing multimodal sensor fusion technology, acoustic, spatiotemporal, and dynamic scene data are collected collaboratively through microphone arrays, navigation systems, LiDAR, and depth cameras to construct a multi-dimensional perception system. A hybrid decision model is used to make volume decisions, and combined with refined calculations of noise, distance, and interaction intent, dynamic adaptive volume adjustment is achieved.
It achieves precise volume control in different hospital scenarios, reducing the false adjustment rate by 65% and the volume adaptation accuracy rate to 92%, ensuring the effectiveness of voice interaction and avoiding noise pollution.
Smart Images

Figure CN121568016A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of robotics, and in particular relates to a voice broadcast control method, system, device and medium for a hospital logistics robot. Background Technology
[0002] With the accelerated implementation of smart hospital construction, logistics and delivery robots have become a core support force for the transportation of goods within hospitals, undertaking key automated tasks such as drug delivery, specimen transfer, and medical equipment replenishment. Voice broadcasting, as the core interaction method between robots and medical staff, patients, and visitors, covers key scenarios such as obstacle avoidance prompts, arrival notifications, and task status synchronization, directly impacting transportation efficiency and the patient experience.
[0003] The current mainstream methods for controlling the volume of voice broadcasts in hospital logistics robots are basically fixed volume control models or simple voice feedback adjustments. Such control methods are no longer suitable for the complex and ever-changing environmental needs of hospitals. Therefore, there is an urgent need for a more intelligent, more accurate, and scene-understanding adaptive volume control solution. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide a voice broadcast control method, system, device and medium for hospital logistics robots, so as to solve the problems of "poor scene adaptation, lag in adjustment and lack of iterative capability" in the existing voice broadcast volume control of hospital logistics robots.
[0005] To address the aforementioned technical problems, this application provides the following technical solution: Firstly, this application provides a voice broadcast control method for a hospital logistics robot, including: Acquire raw multimodal environmental data in a hospital environment, wherein the raw multimodal environmental data includes: raw acoustic data, raw spatiotemporal data, and raw dynamic scenario data; Feature extraction processing is performed on the original multimodal environment data, and a unified environmental feature vector is constructed based on the results of the feature extraction processing; The environmental feature vector is input into a preset volume decision model to calculate the robot's target playback volume.
[0006] Furthermore, the method also includes: Obtain the volume playback feedback signal, and update the parameters of the volume decision model based on the feedback signal.
[0007] Furthermore, the acoustic raw data includes: the raw audio stream signal of ambient sound collected by the microphone array; the spatiotemporal raw data includes: real-time geographic location coordinates and the current timestamp output by the system clock; the dynamic scene raw data includes: environmental point cloud data and depth image.
[0008] Furthermore, the step of performing feature extraction processing on the original multimodal environment data, and constructing a unified environment feature vector based on the results of the feature extraction processing, includes: Acoustic feature extraction: Processing the raw audio stream signal, calculating the average sound pressure level, analyzing the spectral energy distribution, and identifying the type of sound event; Spatiotemporal context extraction: Match real-time geographic location coordinates with a pre-loaded high-precision semantic map to determine the robot's current semantic region; determine the current time period based on the timestamp and preset rules; Dynamic scene feature extraction: Analyze point cloud data and depth images to calculate the density of surrounding pedestrians, the distance to the nearest pedestrian, and identify whether pedestrians have interactive intentions; Based on the results of acoustic feature extraction, spatiotemporal context extraction, and dynamic scene feature extraction, a unified environmental characteristic vector is constructed.
[0009] Furthermore, the preset volume decision model uses the following calculation formula: V_target = [V_norm + ΔV_noise + ΔV_dist + ΔV_int] × k Where V_target is the target volume output; V_norm is the base volume; ΔV_noise is the noise compensation amount; ΔV_dist is the distance compensation amount; ΔV_int is the interaction intent compensation amount; and k is the dynamic coefficient.
[0010] Furthermore, the calculation rules for the noise compensation amount are as follows: when the average ambient sound pressure level is higher than the ambient threshold corresponding to the reference volume, compensation is increased by 3-5 dB for every 5 dB exceeding the threshold; the calculation rules for the distance compensation amount are as follows: when the pedestrian distance is less than 1.5 m, the volume is decreased by 5-8 dB based on the compensated volume; the calculation rules for the interaction intent compensation amount are as follows: when a pedestrian interaction intent is detected, the volume is increased by an additional 2-3 dB based on the compensated volume.
[0011] Further, the step of acquiring the volume playback feedback signal and updating the parameters of the volume decision model based on the feedback signal includes: The environmental feature vector, the broadcast volume, and the collected feedback signal for each broadcast are used as training samples and stored in the historical database. Periodically, or when the amount of data accumulates to a certain threshold, historical data from the historical database is used to retrain or fine-tune the volume decision model.
[0012] Secondly, this application also provides a voice broadcast control system for a hospital logistics robot, comprising: The data acquisition module is used to acquire raw multimodal environmental data in the hospital environment, wherein the raw multimodal environmental data includes: raw acoustic data, raw spatiotemporal data, and raw dynamic scenario data; The feature extraction module is used to perform feature extraction processing on the original multimodal environment data, and to construct a unified environment feature vector based on the result of the feature extraction processing; The playback control module is used to input the environmental feature vector into a preset volume decision model to calculate the robot's target playback volume.
[0013] Thirdly, this application also provides a computer electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the voice broadcast control method for a hospital logistics robot as described in any of the above.
[0014] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the voice broadcast control method for the hospital logistics robot described in any one of the above-mentioned methods.
[0015] The voice broadcast control method, system, equipment, and medium for a hospital logistics robot provided in this application have the following advantages: I. Overcoming the limitations of single-modal perception to achieve comprehensive and accurate understanding of hospital scenarios. Existing technologies often rely on acoustic data from a single microphone for volume adjustment, failing to fully capture the multidimensional information of complex hospital scenarios, leading to biased adjustment decisions. This invention utilizes a multi-sensor collaboration involving microphone arrays, navigation systems, LiDAR, and depth cameras to simultaneously collect three types of raw data: acoustic, spatiotemporal, and dynamic context data, constructing a multi-dimensional perception system encompassing "sound-location-time-dynamic target." On one hand, spectral analysis and event recognition of acoustic data can accurately distinguish between different noise types, such as "crowd noise" and "equipment alarms," avoiding volume interference caused by semantic misjudgment. On the other hand, matching spatiotemporal data with high-precision semantic maps, combined with pedestrian density, distance, and interaction intent recognition in dynamic contexts, enables precise positioning of specific scenarios such as "daytime in the outpatient hall" and "late night in the ICU." This comprehensive perception capability provides complete contextual semantic support for volume decisions, solving the "information bias" problem of traditional methods from the data source, making volume adjustment more aligned with the actual needs of hospital scenarios.
[0016] II. Constructing a hybrid decision-making model to achieve dynamic adaptive volume control and precise output. Traditional technologies with fixed volume modes or simple linear voice control cannot cope with the dynamic changes in hospital settings, often resulting in problems such as "inaudible in noisy areas and excessively loud in quiet areas." This invention employs a hybrid decision-making model combining a "rule engine + lightweight neural network," along with a standardized target volume calculation formula, to achieve collaborative optimization of multi-dimensional parameters: the rule engine provides a scenario-based basic volume through a "region-time period volume benchmark table," combined with a special event interception mechanism (such as forced volume reduction when equipment alarms) to ensure a standardized medical environment; the dynamic compensation coefficient k output by the lightweight neural network incorporates implicit correlations of multiple features, making adjustments more closely aligned with actual interaction needs; and noise, distance, and interaction intention... Figure 3 The refined calculation of compensation amounts further enables dynamic adaptation for "noise exceeding limits, volume reduction for pedestrians at close range, and active interactive volume supplementation." Actual testing shows that the volume adaptation accuracy of this invention reaches over 92% in different hospital scenarios. Compared to traditional linear adjustment methods, the misadjustment rate is reduced by 65%, effectively solving the "volume polarization" problem. This ensures the effectiveness of voice interaction while avoiding noise pollution. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a voice broadcast control method for a hospital logistics robot according to an embodiment of this application. Figure 2 This is a schematic diagram of the structure of a voice broadcast control system for a hospital logistics robot according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer electronic device according to an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] It should be noted that when an element is said to be "fixed" to another element, it can be directly on the other element or there may be an intervening element. When an element is said to be "connected" to another element, it can be directly connected to the other element or there may be an intervening element. Conversely, when an element is said to be "directly" on another element, there is no intervening element. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0021] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0022] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0023] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application are also intended to include the plural forms unless the context clearly indicates otherwise.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the template description is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0025] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this application, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "when".
[0026] Currently, the voice broadcast volume control method for hospital logistics robots has the following limitations: 1. Fixed volume mode: lack of environmental adaptability In this mode, the robot's broadcast volume remains fixed after factory calibration or initialization, and cannot be dynamically adjusted according to the acoustic characteristics of different areas in the hospital. In a hospital setting, the tolerance thresholds for environmental sound vary greatly among areas such as the outpatient hall (crowded and noisy), general wards (requiring basic quiet), and ICU corridors and wards (requiring absolute silence). A fixed volume inevitably leads to a "polarization" problem—insufficient volume in noisy areas, making it impossible to effectively convey interactive information and delaying task execution; excessive volume in quiet areas, creating noise pollution, interfering with patients' rest and medical staff's focused work.
[0027] II. Simple voice-controlled feedback: Limited adjustment strategy While some robots have incorporated adaptive adjustment based on environmental noise, their core logic relies solely on real-time noise decibel readings collected by the microphone. They adjust the broadcast volume using a "linear gain" method, failing to consider the specific characteristics of hospital settings, and thus exhibiting three major flaws: 1. Delayed response mechanism: It is a "passive adjustment" mechanism that can only adjust the volume after the noise is generated and collected. It cannot predict sudden noise (such as sudden crowds or temporary equipment operation), resulting in the "missed" or "mistransmitted" key interactive information.
[0028] 2. Lack of semantic context: Using only decibel values as the basis for adjustment fails to distinguish the semantic differences between noise types. For example, treating "low-frequency noise from normal conversation" and "high-frequency sharp noise from medical equipment alarms" the same way may lead to a sudden increase in volume in quiet areas such as ICUs due to temporary equipment alarms, which seriously violates medical environment regulations.
[0029] 3. Spatial and temporal dimension gap: Ignoring the difference in noise baseline in the same area at different times, such as the difference of more than 30 decibels between the baseline noise value of the ward area during the day (medical staff rounds, family visits) and late at night (patients rest). The problem of "too high" volume at night or "too low" volume during the day persists due to the single adjustment logic.
[0030] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes will not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0031] Please see Figure 1 The voice broadcast control method for a hospital logistics robot provided in this application embodiment includes at least the following steps: S10. Obtain multimodal environment raw data in the hospital environment, wherein the multimodal environment raw data includes: acoustic raw data, spatiotemporal raw data and dynamic scenario raw data.
[0032] In one embodiment of this application, the acoustic raw data includes: the raw audio stream signal of ambient sound collected by the microphone array; the spatiotemporal raw data includes: real-time geographic location coordinates and the current timestamp output by the system clock; the dynamic scene raw data includes: environmental point cloud data and depth image.
[0033] Specifically, the core objective of this step is to comprehensively and in real-time capture various raw data related to voice interaction in the hospital environment through the robot's multi-sensor fusion system, providing a complete data source for subsequent feature analysis. The specific implementation method is as follows: Acoustic raw data acquisition: A 4-microphone circular array is used as the acquisition unit to acquire the raw audio stream signal of ambient sound in real time with a sampling rate of 16kHz and quantization accuracy of 16bit. The array has spatial filtering capabilities, which can effectively suppress ambient echo and directional noise within a 360° range, ensuring that the acquired audio signal can truly reflect the acoustic state within a 10-meter range around the robot. The data format is continuous PCM (pulse code modulation) audio frames.
[0034] Spatiotemporal raw data acquisition: Spatiotemporal information is acquired using a dual-dimensional synchronous acquisition mechanism of "position-time". Position data is acquired through the robot's built-in multi-source navigation system—relying on GPS / BeiDou dual-mode positioning (positioning accuracy ±1m) in outdoor and open areas, and switching to laser SLAM navigation (positioning accuracy ±0.05m) in enclosed areas such as indoor wards and corridors. The output data is three-dimensional coordinates (X,Y,Z) based on the hospital's electronic map coordinate system. Time data is acquired from the robot's system real-time clock module, with timestamp accuracy down to the millisecond level, ensuring a synchronization error of ≤50ms with the position data.
[0035] Dynamic Scene Raw Data Acquisition: A complementary acquisition link is formed by integrating LiDAR and a depth camera. The LiDAR outputs environmental point cloud data at a frequency of 10Hz, capable of detecting obstacle outlines and motion states within a 50-meter range. The depth camera employs TOF (Time-of-Flight) technology, outputting 640×480 resolution depth images at a frame rate of 30fps, simultaneously capturing color information. The two types of data work together to achieve accurate perception of dynamic targets in the surrounding area, such as pedestrians, medical equipment, and mobile carts.
[0036] S20. Perform feature extraction processing on the original data of the multimodal environment, and construct a unified environmental feature vector based on the result of the feature extraction processing.
[0037] In one embodiment of this application, step S20 includes: S201. Acoustic Feature Extraction: Process the original audio stream signal, calculate the average sound pressure level, analyze the spectral energy distribution, and identify the type of sound event.
[0038] S202, Spatiotemporal Context Extraction: Match the real-time geographic location coordinates with the pre-loaded high-precision semantic map to determine the robot's current semantic region; determine the current time period based on the timestamp and preset rules.
[0039] S203. Dynamic Scene Feature Extraction: Analyze point cloud data and depth images to calculate the density of surrounding pedestrians, the distance to the nearest pedestrian, and identify whether pedestrians have interactive intentions.
[0040] S204. Based on the results of acoustic feature extraction, spatiotemporal context extraction, and dynamic scene feature extraction, a unified environmental characteristic vector is constructed.
[0041] Specifically, this step involves hierarchical processing of the multimodal raw data to extract features with clear physical meaning and scene semantics, and then integrating them into a standardized environmental feature vector to provide structured input for volume decision-making. The specific processing flow is as follows: Multi-dimensional feature extraction: Acoustic feature extraction: After pre-emphasis, framing (frame length 20ms, frame shift 10ms), and windowing (Hanning window) of PCM audio frames, the average sound pressure level (dB), spectral flatness, 13-dimensional Mel frequency cepstral coefficients (MFCC), and first-order difference coefficients are calculated. At the same time, through a pre-trained CNN sound event recognition model, the system classifies and recognizes four typical sound events, namely "human voice, medical equipment alarm sound, trolley sound, and background noise", and outputs classification labels (0-3).
[0042] Spatiotemporal feature extraction: The three-dimensional coordinates output by the navigation system are matched with the pre-loaded high-precision semantic map of the hospital (including boundary information of 12 functional areas such as the outpatient hall, ICU, and inpatient ward) to determine the semantic region label where the robot is currently located; combined with the timestamp and referring to the hospital scene rules (6:00-18:00 is daytime, 18:00-22:00 is nighttime, and 22:00-6:00 is late night), time period labels are output.
[0043] Dynamic scene feature extraction: Cluster analysis is performed on the LiDAR point cloud data to calculate the surrounding pedestrian density (unit: people / ㎡) and the straight-line distance between the nearest pedestrian and the robot (unit: m); Human skeleton key point detection is performed on the depth image. If a pedestrian is identified to be facing the robot for ≥1s, or if a preset interactive gesture such as "waving or nodding" is displayed, it is marked as "interactive intention exists" (marked as 1), otherwise it is marked as 0.
[0044] Environmental feature vector construction: The above features are integrated using a "numericalization + standardization" strategy. Semantic region labels and time period labels are converted into 8-dimensional and 3-dimensional numerical features through one-hot encoding; continuous features such as average sound pressure level, pedestrian density, and nearest pedestrian distance are normalized (mapped to the [0,1] interval); sound event labels and interaction intent identifiers are directly used as discrete numerical features. Finally, a unified environmental feature vector with a dimension of 20 is constructed, with the following format example: [0,1,0,...,0, 0,0,1, 0.65, 0.32, 0.48, 1, 2].
[0045] S30. Input the environmental feature vector into the preset volume decision model to calculate the target playback volume of the robot.
[0046] In one embodiment of this application, the preset volume decision model uses the following calculation formula: V_target = [V_norm + ΔV_noise + ΔV_dist + ΔV_int] × k Wherein, V_target is the target volume output; V_norm is the base volume correction volume; ΔV_noise is the noise compensation amount; ΔV_dist is the distance compensation amount; ΔV_int is the interaction intent compensation amount; and k is the dynamic coefficient.
[0047] In a specific embodiment of this application, the calculation rule for the noise compensation amount is as follows: when the average ambient sound pressure level is higher than the ambient threshold corresponding to the reference volume, compensation is increased by 3-5 dB for every 5 dB exceeding the threshold; the calculation rule for the distance compensation amount is as follows: when the pedestrian distance is less than 1.5 m, the volume is decreased by 5-8 dB based on the compensated volume; the calculation rule for the interaction intent compensation amount is as follows: when a pedestrian interaction intent is detected, the volume is increased by an additional 2-3 dB based on the compensated volume.
[0048] Specifically, this step employs a hybrid decision-making model combining a "rule engine" and a "lightweight neural network" to transform environmental feature vectors into target playback volume adapted to the current scenario, ensuring real-time and accurate volume adjustment. The specific process is as follows: Model Structure and Input Processing: The pre-defined volume decision model comprises two processing units—a rule engine as the basic decision layer, responsible for matching the scene's baseline volume and intercepting special events; and a lightweight neural network as the dynamic optimization layer, responsible for fusing multi-feature correlations to output compensation coefficients. The input environmental feature vector, after data validation (removing outliers), is simultaneously fed into the two processing units.
[0049] Core parameter calculation and model output: The rule engine first retrieves the baseline volume V_base from the preset "region-time period volume baseline table" based on the "semantic region-time period" combination in the feature vector (e.g., "outpatient hall - daytime" V_base=70dB, "ICU - late night" V_base=32dB). If the sound event label "medical equipment alarm sound" is detected, a special rule is immediately triggered to correct the baseline volume to V_norm=V_base×0.65, while limiting the broadcast duration to ≤1s. The lightweight neural network (3-layer fully connected structure, output activation function is Sigmoid) receives the complete feature vector and outputs a dynamic compensation coefficient k (value range 0.8-1.5), which is continuously optimized based on user feedback during robot operation.
[0050] Target playback volume determination: Combining the baseline volume, noise compensation, distance compensation, and interaction intent compensation, the data is input into a preset volume decision model to produce the final target volume. It should be noted that the final calculation result must satisfy 25dB (minimum audible threshold) ≤ V_target ≤ 85dB (safe volume upper limit). If the value exceeds this range, the boundary value is automatically used as the final output. For example, in the "Inpatient Ward - Night" scenario, if V_norm = 45dB, ΔV_noise = 4dB, ΔV_dist = -2dB, ΔV_int = 3dB, and k = 1.1, then V_target = (45 + 4 - 2 + 3) × 1.1 = 55dB, which meets the volume requirements for this scenario.
[0051] Through the above steps, the robot can achieve full-process volume control of "scene prediction - dynamic adjustment - precise output", which not only ensures the effectiveness of voice interaction, but also avoids interference with the quiet environment of the hospital.
[0052] In one embodiment of this application, the method further includes: S40. Obtain the volume playback feedback signal, and update the parameters of the volume decision model based on the feedback signal.
[0053] In one specific embodiment of this application, step S40 includes: S401. The environmental feature vector, the broadcast volume, and the collected feedback signal at each broadcast are used as a training sample and stored in the historical database. S402. Periodically or when the amount of data accumulates to a certain threshold, use historical data from the historical database to retrain or fine-tune the volume decision model.
[0054] Specifically, this step focuses on "user experience," simultaneously acquiring user feedback signals after voice broadcasting and associating them with environmental characteristics and volume during broadcasting to construct a complete training sample encompassing "input-output-evaluation." This sample is then periodically used to update the parameters of the volume decision model. Its core value lies in breaking the limitations of "static presets" in models, enabling the robot to learn volume preferences for different departments (such as pediatrics and geriatrics) and different times of day. For example, given the frequent family visits in pediatric wards, it automatically optimizes the compensation coefficient for daytime broadcast volume, improving the long-term user experience. The entire process requires no human intervention, achieving autonomous model evolution.
[0055] The voice broadcast control method for a hospital logistics robot provided in this application has the following advantages: I. Overcoming the limitations of single-modal perception to achieve comprehensive and accurate understanding of hospital scenarios. Existing technologies often rely on acoustic data from a single microphone for volume adjustment, failing to fully capture the multidimensional information of complex hospital scenarios, leading to biased adjustment decisions. This invention utilizes a multi-sensor collaboration involving microphone arrays, navigation systems, LiDAR, and depth cameras to simultaneously collect three types of raw data: acoustic, spatiotemporal, and dynamic context data, constructing a multi-dimensional perception system encompassing "sound-location-time-dynamic target." On one hand, spectral analysis and event recognition of acoustic data can accurately distinguish between different noise types, such as "crowd noise" and "equipment alarms," avoiding volume interference caused by semantic misjudgment. On the other hand, matching spatiotemporal data with high-precision semantic maps, combined with pedestrian density, distance, and interaction intent recognition in dynamic contexts, enables precise positioning of specific scenarios such as "daytime in the outpatient hall" and "late night in the ICU." This comprehensive perception capability provides complete contextual semantic support for volume decisions, solving the "information bias" problem of traditional methods from the data source, making volume adjustment more aligned with the actual needs of hospital scenarios.
[0056] II. Constructing a hybrid decision-making model to achieve dynamic adaptive volume control and precise output. Traditional technologies with fixed volume modes or simple linear voice control cannot cope with the dynamic changes in hospital settings, often resulting in problems such as "inaudible in noisy areas and excessively loud in quiet areas." This invention employs a hybrid decision-making model combining a "rule engine + lightweight neural network," along with a standardized target volume calculation formula, to achieve collaborative optimization of multi-dimensional parameters: the rule engine provides a scenario-based basic volume through a "region-time period volume benchmark table," combined with a special event interception mechanism (such as forced volume reduction when equipment alarms) to ensure a standardized medical environment; the dynamic compensation coefficient k output by the lightweight neural network incorporates implicit correlations of multiple features, making adjustments more closely aligned with actual interaction needs; and noise, distance, and interaction intention... Figure 3 The refined calculation of compensation amounts further enables dynamic adaptation for "noise exceeding limits, volume reduction for pedestrians at close range, and active interactive volume supplementation." Actual testing shows that the volume adaptation accuracy of this invention reaches over 92% in different hospital scenarios. Compared to traditional linear adjustment methods, the misadjustment rate is reduced by 65%, effectively solving the "volume polarization" problem. This ensures the effectiveness of voice interaction while avoiding noise pollution.
[0057] Please see Figure 2 This application also provides a voice broadcast control system 200 for a hospital logistics robot, comprising: The data acquisition module 201 is used to acquire multimodal environment raw data in the hospital environment, wherein the multimodal environment raw data includes: acoustic raw data, spatiotemporal raw data and dynamic scenario raw data; The feature extraction module 202 is used to perform feature extraction processing on the original multimodal environment data, and construct a unified environment feature vector based on the result of the feature extraction processing; The playback control module 203 is used to input the environmental feature vector into a preset volume decision model to calculate the target playback volume of the robot.
[0058] Please see Figure 3 This application also provides a computer electronic device 300, including a memory 303 and a processor 302. The memory 303 stores a computer program, and when the processor executes the computer program, it implements the steps of the voice broadcast control method for the hospital logistics robot described in any of the above claims.
[0059] Specifically, the electronic device 300 includes a transceiver 301, a bus interface, and a processor 302. The processor 302 is used to acquire multimodal environmental raw data in the hospital environment, wherein the multimodal environmental raw data includes acoustic raw data, spatiotemporal raw data, and dynamic scenario raw data; to perform feature extraction processing on the multimodal environmental raw data, and to construct a unified environmental feature vector based on the result of the feature extraction processing; and to input the environmental feature vector into a preset volume decision model to calculate the target playback volume of the robot.
[0060] In this embodiment of the application, the electronic device 300 further includes a memory 303. Figure 3 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 302) and memory (memory 303). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 301 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. The processor 302 is responsible for managing the bus architecture and general processing, and the memory 303 can store data used by the processor 302 during operation.
[0061] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the voice broadcast control method for the hospital logistics robot described above.
[0062] In this embodiment, the computer-readable storage medium can be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0063] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not as limitations; therefore, other examples of exemplary embodiments may have different values.
[0064] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0065] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0066] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0067] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0068] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A voice broadcast control method for a hospital logistics robot, characterized in that, include: Acquire raw multimodal environmental data in a hospital environment, wherein the raw multimodal environmental data includes: raw acoustic data, raw spatiotemporal data, and raw dynamic scenario data; Feature extraction processing is performed on the original multimodal environment data, and a unified environmental feature vector is constructed based on the results of the feature extraction processing; The environmental feature vector is input into a preset volume decision model to calculate the robot's target playback volume.
2. The voice broadcast control method according to claim 1, characterized in that, The method further includes: Obtain the volume playback feedback signal, and update the parameters of the volume decision model based on the feedback signal.
3. The voice broadcast control method according to claim 1, characterized in that, The acoustic raw data includes: the raw audio stream signal of ambient sound collected by the microphone array; the spatiotemporal raw data includes: real-time geographic location coordinates and the current timestamp output by the system clock; the dynamic scene raw data includes: environmental point cloud data and depth image.
4. The voice broadcast control method according to claim 1, characterized in that, The process involves feature extraction from the original multimodal environment data, and based on the results of the feature extraction, a unified environmental feature vector is constructed. include: Acoustic feature extraction: Processing the raw audio stream signal, calculating the average sound pressure level, analyzing the spectral energy distribution, and identifying the type of sound event; Spatiotemporal context extraction: Match real-time geographic location coordinates with a pre-loaded high-precision semantic map to determine the robot's current semantic region; determine the current time period based on the timestamp and preset rules; Dynamic scene feature extraction: Analyze point cloud data and depth images to calculate the density of surrounding pedestrians, the distance to the nearest pedestrian, and identify whether pedestrians have interactive intentions; Based on the results of acoustic feature extraction, spatiotemporal context extraction, and dynamic scene feature extraction, a unified environmental characteristic vector is constructed.
5. The voice broadcast control method according to claim 1, characterized in that, The preset volume decision model uses the following calculation formula: V_target = [V_norm + ΔV_noise + ΔV_dist + ΔV_int] × k Where V_target is the target volume output; V_norm is the base volume; ΔV_noise is the noise compensation amount; ΔV_dist is the distance compensation amount; ΔV_int is the interaction intent compensation amount; and k is the dynamic coefficient.
6. The voice broadcast control method according to claim 5, characterized in that, The calculation rule for the noise compensation amount is as follows: when the average ambient sound pressure level is higher than the ambient threshold corresponding to the reference volume, the compensation is increased by 3-5 dB for every 5 dB exceeding the threshold. The calculation rule for the distance compensation amount is as follows: when the pedestrian distance is less than 1.5 m, the volume is decreased by 5-8 dB based on the compensated volume. The calculation rule for the interaction intent compensation amount is as follows: when a pedestrian interaction intent is detected, the volume is increased by an additional 2-3dB based on the compensated volume.
7. The voice broadcast control method according to claim 2, characterized in that, The step of acquiring the volume playback feedback signal and updating the parameters of the volume decision model based on the feedback signal includes: The environmental feature vector, the broadcast volume, and the collected feedback signal for each broadcast are used as training samples and stored in the historical database. Periodically, or when the amount of data accumulates to a certain threshold, historical data from the historical database is used to retrain or fine-tune the volume decision model.
8. A voice broadcast control system for a hospital logistics robot, characterized in that, include: The data acquisition module is used to acquire raw multimodal environmental data in the hospital environment, wherein the raw multimodal environmental data includes: raw acoustic data, raw spatiotemporal data, and raw dynamic scenario data; The feature extraction module is used to perform feature extraction processing on the original multimodal environment data, and to construct a unified environment feature vector based on the result of the feature extraction processing; The playback control module is used to input the environmental feature vector into a preset volume decision model to calculate the robot's target playback volume.
9. A computer electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the voice broadcast control method for the hospital logistics robot according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the voice broadcast control method for the hospital logistics robot according to any one of claims 1-7.