Public space aging-adaptive interaction system based on multi-mode perception
By constructing a three-dimensional dynamic spatial model and behavior prediction through a multimodal perception system, the problem of lack of unified alignment and fusion modeling of multi-source perception data in age-friendly interactive systems in public spaces is solved. This enables the prediction of target object behavior and the calculation of the availability of interactive devices, ensuring the measurability and stability of interaction triggers and reducing the risk of device conflicts and repeated prompts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-28
AI Technical Summary
In existing public space age-friendly guidance and interaction systems, multi-source sensing data lacks unified alignment and fusion modeling, making it difficult to predict the behavioral trends of target objects. The timing and intensity of interaction triggers lack measurable basis, and environmental factors are not included in the constraints on the availability of output modalities, resulting in a mismatch between prompting methods and on-site conditions. Furthermore, the lack of concurrent budget and slot arrangement mechanisms easily leads to conflicts and duplicate prompts.
A multimodal sensing system is adopted to construct a three-dimensional dynamic spatial model through multi-source fusion sensing data, perform behavior prediction and calculate the availability of interactive device channels, determine the necessity and controllability of concurrency, generate concurrency budget and slot arrangement scheme, introduce environmental indicators to constrain output modes, and implement degradation and closed-loop write-back mechanisms to ensure the measurability and stability of interaction triggering.
An age-friendly interactive system for public spaces based on multimodal perception has been implemented. It can reliably trigger interactions in time-updated spatial and crowd dynamic representations, reduce device conflicts and execution uncertainties, and improve the reproducibility and stability of the system.
Smart Images

Figure CN121937263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of age-friendly guided interaction technology, and more specifically, to an age-friendly interactive system for public spaces based on multimodal perception. Background Technology
[0002] Age-friendly guidance and interaction in public spaces typically rely on fixed signs, broadcast prompts, directional screens, or manual service desks. Some systems introduce cameras, environmental sensors, and other equipment to count people or monitor the status of areas, and trigger prompts according to preset rules to help the elderly find their way, queue, or obtain services at entrances, service desks, and accessible pathways.
[0003] Existing technologies have the following shortcomings: multi-source sensing data lacks unified alignment and fusion modeling, making it difficult to form a dynamic representation of space and crowds that can be updated over time; there is a lack of prediction and uncertainty characterization of the behavioral trends of target objects, and the timing and intensity of interaction triggers lack measurable basis; environmental factors such as noise, glare, occlusion, and crowding are not included in the unified constraints on the availability of different output modes, and the prompting method is prone to mismatch with the on-site conditions; there is a lack of concurrent budget and slot arrangement mechanisms, which easily leads to conflicts between multiple devices in the same slot, uncontrollable prompting rhythm, or repeated prompts that are difficult to explain; and there is a lack of degradation and closed-loop write-back mechanisms for execution failures, resulting in insufficient reproducibility and stability.
[0004] To address the above problems, this invention proposes a solution. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an age-friendly interactive system for public spaces based on multimodal perception, in order to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The public space age-friendly interactive system based on multimodal perception includes the following modules: a perception prediction parameter generation module, used to construct a three-dimensional dynamic spatial model of the public space based on multi-source fusion perception data, and to predict the behavior of target elderly users, obtaining information including prediction uncertainty, peak intervention demand, and peak time; the urgency level is obtained by continuously monotonically decaying mapping based on the remaining time at the peak time and a preset urgency time constant; and the channel availability is calculated for each interactive device in the interactive output unit set, wherein the elements of the interactive output unit set are interactive devices with at least one output mode. The concurrent budget parameter calculation module is used to assess the single-channel failure risk based on the channel availability of the most reliable equipment, and to obtain the concurrent necessity by taking the smaller value between the urgency and the single-channel failure risk; the concurrent controllability is determined based on the prediction uncertainty and the peak of intervention demand; and the upper limit of the main channel concurrency within the intervention time window is determined by comparing the concurrent necessity and concurrent controllability with their corresponding preset thresholds. The device selection slot orchestration module is used to select the set of main role devices from the set of interactive output units based on the main channel concurrency limit, and generate a strong prompt token for the selected device within the intervention time window based on the main channel concurrency limit and bind it to the slot to generate an orchestration scheme; The instruction execution module is used to issue strong prompt instructions to token slots according to the orchestration scheme, and to issue weak prompt hold instructions to non-token slots.
[0007] In a preferred embodiment, the perception prediction parameter generation module pre-deploys a set of multi-source sensors and interactive output units in the public space, uses a LIO-SAM fusion positioning algorithm based on a factor graph optimization framework to register the lidar point cloud and IMU data and output a high-precision pose sequence. The pose sequence is then combined with the output of the semantic segmentation network MaskR-CNN to construct a three-dimensional dynamic space model. The three-dimensional dynamic space model includes a static layer containing structural boundaries and facility objects, and a dynamic layer that records crowd density, occlusion relationships, and temporary obstacles using time indexes.
[0008] In a preferred embodiment, the specific method by which the perception prediction parameter generation module performs behavior prediction includes: using a behavior prediction model based on spatiotemporal graph convolution and attention mechanism, constructing a spatiotemporal graph of the target user and surrounding pedestrians, extracting motion features and modeling social dependencies, and outputting the probability distribution of the target user's future location at each prediction step within the prediction window, wherein the probability distribution is given in the form of a position mean sequence and a position covariance sequence.
[0009] In a preferred embodiment, the budget parameter calculation module determines the main channel concurrency limit by: calculating the excess amount of concurrency necessity exceeding a preset concurrency necessity threshold, and the excess amount of concurrency controllability exceeding a preset concurrency controllability threshold; taking the smaller of the two excess amounts as the concurrency margin; comparing the concurrency margin with a preset tiered threshold sequence, and determining the concurrency level based on the interval it falls into; finally, based on the concurrency level and the system's preset main channel capacity limit, comprehensively determining the main channel concurrency limit; wherein the concurrency necessity is based on the prediction uncertainty and the peak value of intervention demand, taking the larger of the two, and using this larger value to characterize the overall uncertainty of system control concurrency, thereby deriving the corresponding concurrency necessity.
[0010] In a preferred embodiment, the concurrency budget parameter calculation module is also used to determine the minimum strong alert interval. The method is as follows: while determining the upper limit of the main channel concurrency, the concurrency budget parameter calculation module obtains the minimum strong alert interval by multiplying the basic interval by the rhythm conservative factor based on the preset basic interval and rhythm conservative factor. The intervention time window is discretized by the minimum strong alert interval, and the intervention time window is divided into several slots. The upper limit of the strong alert density is determined by the ratio of the upper limit of the main channel concurrency to the length of the intervention time window.
[0011] In a preferred embodiment, after receiving the main channel concurrency limit, strong cue density limit, and modal repetition budget, the device selection slot orchestration module first sorts the interactive devices in the interactive output unit set in descending order according to channel availability, and truncates them according to the preset participation device limit to obtain a candidate device set; iteratively selects the main role device set from the candidate device set, and adopts different device selection penalty logic according to the modal repetition budget: when the modal repetition budget is to suppress repetition, a penalty that increases with the number of repetitions is applied to the output modal that appears repeatedly in the candidate set; when the modal repetition budget is to allow one controlled repetition, the penalty is minimized when a modal repetition occurs once in the candidate set.
[0012] In a preferred embodiment, when constructing the main role device set, the device selection slot orchestration module calculates at least the following evaluation indicators for each candidate device after it is added: rhythm reachability violation number, modal repetition deviation, and minimum channel availability of the set. It then compares the candidate devices using a unified lexicographical selection criterion and selects the device with the best evaluation result to add to the main role device set until the main channel concurrency limit is met or no device can be added.
[0013] In a preferred embodiment, after the primary role device set is determined, the device selection slot orchestration module selects auxiliary role devices from the remaining part of the candidate device set in descending order of channel availability. Within the intervention time window, it generates a slot sequence according to the minimum strong cue interval, binds the maximum strong cue density tokens to the slot sequence at the most even intervals, and then allocates token slots among the primary role device sets using a cyclic allocation method to generate an orchestration scheme that includes the correspondence between slots and devices.
[0014] In a preferred embodiment, when executing the orchestration scheme, the instruction execution module issues strong prompt instructions only to the primary role device bound to the slot for each token slot, and issues weak prompt holding instructions or status indication instructions only to the primary role device or auxiliary role device for other slots within the intervention time window, and ensures that multiple devices do not execute strong prompts simultaneously in any slot.
[0015] The technical effects and advantages of this invention's multimodal perception-based age-friendly interactive system for public spaces are as follows: This invention forms a three-dimensional dynamic spatial model of a public space through time alignment, anomaly removal, and fusion modeling of multi-source sensor data, and outputs the pose and trajectory fragments of the target object, providing definite input for subsequent interactive decision-making. Based on the probability distribution of position within the prediction window, it constructs prediction uncertainty and peak intervention demand, and combines urgency and risk thresholds to determine the intervention intensity level, making the interaction triggering measurable. It introduces environmental indicators and channel availability calculation mechanisms to constrain the availability of output modes under conditions such as noise, glare, obstruction, and congestion, ensuring that equipment selection and prompting methods are consistent with on-site conditions. Under the constraints of concurrency budget and slot orchestration, it generates role mapping and instruction timing, and with the support of single-slot concurrency suppression, controlled mode repetition, and failure degradation rules, it reduces scheduling mismatch caused by slot conflicts and execution uncertainty. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the structure of the age-friendly interactive system for public spaces based on multimodal perception, as described in this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example Please see Figure 1 As shown, this invention discloses an age-friendly interactive system for public spaces based on multimodal perception, comprising the following modules: The perception prediction parameter generation module is used to construct a three-dimensional dynamic spatial model of the public space based on multi-source fusion perception data, and to predict the behavior of target elderly users, obtaining a result including prediction uncertainty, peak intervention demand, and peak time; the urgency level is obtained by continuously monotonically decaying mapping based on the remaining time at the peak time and a preset urgency time constant; and the channel availability is calculated for each interactive device in the interactive output unit set, wherein the elements of the interactive output unit set are interactive devices with at least one output mode. In this embodiment, a set of multi-source sensors and interactive output units are pre-deployed in the public space. The set of interactive output units is denoted as DevSet, and its elements are interactive devices with at least one output mode, such as visible light display, indicator light array, directional speaker, tactile or vibration, ground projection, etc. The system performs unified timestamp alignment and anomaly removal on the point cloud, image, inertial navigation, acoustic, and environmental data collected by the multi-source sensors to form a multi-source fusion data stream. In this embodiment, the multi-source fusion adopts the LIO-SAM fusion positioning algorithm based on the factor graph optimization framework: the lidar feature points are matched with the map to form lidar factors, and the IMU factors are formed based on the pre-integration of the IMU at adjacent times. In each iteration, the point cloud registration error and inertial error are minimized and a high-precision pose sequence is output to achieve accurate perception of public space environment and personnel movement information. The three-dimensional dynamic spatial model is constructed by inputting multi-source fused data into a semantic segmentation network: Mask R-CNN is used to extract features and predict masks from the fused data to obtain semantic segmentation results, and the spatial occupancy boundaries, passable areas, key facilities, and dynamic pedestrian instances are constructed by combining point cloud three-dimensional coordinates; the key facilities include, but are not limited to, service counters, entrances and exits, accessible entrances, stairs or escalators, etc.; and crowd density, occlusion relationships, and temporary obstacles are recorded with time index; at the same time, the target user is continuously tracked, and short-term trajectory and posture information are output; the three-dimensional dynamic spatial model contains a static layer and a dynamic layer, the static layer contains structural boundaries and facility objects, and the dynamic layer records crowd density, occlusion relationships, and temporary obstacles with time index; It should be noted that the determination and tracking of the target user mentioned in this invention as the target elderly user can be achieved using existing human detection, identity differentiation and multi-target tracking methods. This invention does not limit the specific algorithm, but at least outputs the target user identifier, current pose and historical trajectory fragments as the determination input for subsequent behavior prediction. In this invention, no further details will be provided. Furthermore, the system constructs a behavior prediction model based on a three-dimensional dynamic spatial model to form the user trend distribution within the prediction window. In this embodiment, a Social-STGCNN-Transformer network is used to construct a spatiotemporal graph of the target user and surrounding pedestrians. Local motion features are extracted through spatiotemporal graph convolution, and then long-distance social dependencies are modeled using Transformer multi-head attention. The system outputs the movement direction distribution, dwell probability, and interaction possibility within the prediction window, and derives the dwell probability accordingly. With directional divergence This serves as the input for subsequent target ambiguity determination; simultaneously, within the prediction window, for each prediction step k, the probability distribution of the target user's future location is output. In this embodiment, the probability distribution is a sequence of location mean values. With location covariance sequence The parameters are given in the form of, where This indicates the predicted position at the k-th prediction step. This indicates the degree of uncertainty diffusion at the predicted location; the prediction window parameters are given by the preset parameter library Π, including the upper and lower bounds of the prediction duration. With step size The prediction window is from the current time. From beginning to end The time interval, in which Limited by the upper and lower bounds of the prediction duration And the prediction step index k corresponds to the prediction time. ; It should be noted that the detailed construction methods and specific implementation steps of the three-dimensional dynamic spatial model and behavior prediction model can be directly completed by those skilled in the art based on existing mature technologies. This invention does not limit its internal processing flow, but only requires that it can output the aforementioned three-dimensional dynamic spatial model used to characterize the structure of public space and the dynamic distribution of crowds, as well as the behavior prediction results used to give the probability distribution of target user location and related characteristics within the prediction window. These will not be elaborated on in this invention. After obtaining the three-dimensional dynamic spatial model and behavior prediction results, a public space environment index vector is constructed. Where P represents the number of dimensions of the environmental indicator. This represents the value of the p-th environmental indicator. The public space environmental indicators include at least one of noise, glare, obstruction, and congestion, and each indicator has a clearly defined data source and unit. Simultaneously, the parameter library Π associates each output mode m with its associated indicators. Preset two-level thresholds and And store channel availability endpoints These are the preset full channel availability value and failed channel availability value, respectively. Example values can be used initially during deployment, for example... Its initial value is determined by the sample distribution collected during the deployment phase: for indicators that are worse the larger they are, the high quantile of the available samples is taken as the initial value. The low quantile of the sample cannot be used as The system sets upper and lower bounds to ensure that subsequent write-backs are monotonous and bounded; the system then processes each interactive output unit. DevSet calculates channel availability. Its construction logic is as follows: first, for each related indicator... in accordance with Perform piecewise monotonic mapping to obtain single-index channel availability Then, aggregate according to the weakest link principle, so that if any critical environmental factor makes the whole system unavailable, the entire system will be unavailable; the piecewise monotonic mapping can be directly implemented using if-else: if but ;like but Otherwise, Monotonically changing in a linear manner, for example, let... The aggregate expression for channel availability is: ,in For the set of indices associated with mode m(u); Furthermore, in this embodiment, the prediction uncertainty U is constructed as a normalized quantity jointly determined by the prediction variance statistic and the missing rate statistic. The system then uses the location covariance sequence... Constructing variance statistics In this embodiment, variance statistics Take the trace of the covariance matrix of each prediction step within the prediction window. The time average is used to characterize the overall level of prediction uncertainty; simultaneously, the in-window missing rate is calculated based on the missing markers. Variance scaling parameters are read from the preset parameter library Π respectively. Compared with the missing rate scaling parameter And normalize according to the defined rules: variance statistic With variance scaling parameter Find the ratio ; For the predicted in-window missing rate Compared with the missing rate scaling parameter Find the ratio The two ratios are then cropped to the range [0,1], i.e., 0 is taken when the ratio is less than 0, 1 is taken when the ratio is greater than 1, and the original values are retained for the rest, thus obtaining the prediction variance uncertainty. With the uncertainty of the missing rate The larger of the two values is ultimately taken as the prediction uncertainty. This ensures that unreliability from any source can trigger conservative budget inferences. Furthermore, within the prediction window, for each prediction step k, the prediction model outputs a discrete probability distribution of the movement direction. The number of discrete sectors in the direction The value is a fixed integer given by the preset parameter library Π; the system calculates the predicted velocity amplitude from the average position of adjacent prediction steps. And read the velocity scale parameters from the parameter library Π Calculate the original value of the tendency to stay. The data is then clipped to limit it to [0,1], thus obtaining the dwell tendency. Discrete probability distribution of the direction of movement Calculate information entropy and use Normalization yields the directional divergence. Finally, take As the intervention demand score at step k, and search for the peak step across all prediction steps. Thus, the peak intervention time is obtained. And at the moment of intervention peak Then, the remaining time is calculated. And read the urgent time constant from the preset parameter library Π The degree of urgency is determined by using a continuously monotonically decaying mapping. , that is to say And write it into the status packet G; where the urgent time constant is... The initial deployment values were obtained by statistically analyzing the response delay distribution of elderly users in the calibration sample set to observable responses to strong prompts, and the conservative quantile of this distribution was taken. ,For example The delay value corresponding to ≥0.8 is used as the tight time constant. ; Further intervention in peak demand Compared with the preset risk threshold in the preset parameter library Π A comparative assessment is performed to determine the level of intervention intensity, the peak intervention demand. For peak step Corresponding intervention requirement score: When satisfied The time-output intervention intensity level is When satisfied The time-output intervention intensity level is In other cases, the output intervention intensity level is Finally, the system intervenes at its peak time. Centered on and calling the preparation time constant from the parameter library Π With confirmation duration constant Generate intervention time window Among them, the preset risk threshold and The initial deployment values were determined statistically from the results of manual verification and labeling of intervention intensity in the calibration sample set. The high quantiles of samples where low intensity was acceptable and the low quantiles of samples requiring high intensity were selected, ensuring... ; The system encapsulates and outputs the key parameters from the three-dimensional dynamic space model and behavior prediction model, which are used for subsequent guidance decisions, into a structured dataset called the prediction parameter set. This prediction parameter set includes at least: the prediction uncertainty U, which characterizes the reliability of the prediction results, and the peak intervention demand, which characterizes the user's highest guidance demand within the prediction window. and according to The intervention intensity level is determined by a preset risk threshold; furthermore, depending on the actual implementation, this set may also include a calculated degree of urgency. Intervention peak time The prediction parameter set, including prediction window parameters and short-term user trajectory and attitude information, serves as a deterministic output of the perception prediction stage and is provided to the concurrent budget parameter calculation module for subsequent concurrent budgeting and resource planning.
[0019] The concurrent budget parameter calculation module is used to assess the single-channel failure risk based on the channel availability of the most reliable equipment, and to obtain the concurrent necessity by taking the smaller value between the urgency and the single-channel failure risk; the concurrent controllability is determined based on the prediction uncertainty and the peak of intervention demand; and the upper limit of the main channel concurrency within the intervention time window is determined by comparing the concurrent necessity and concurrent controllability with their corresponding preset thresholds. Based on the obtained channel availability The channel availability of all interactive devices is obtained by summarizing the data. The single-channel failure risk is assessed based on the channel availability of the most reliable equipment, denoted as... ; and combined with the obtained urgency The smaller of the two values is defined as the concurrency necessity. If and only if the degree of urgency Risk of single-channel failure When both are high, the necessity of concurrency This leads to an increase, thus avoiding the introduction of concurrency in situations where there is no urgent need or where a single channel is reliable enough; The system simultaneously reads the prediction uncertainty U and the peak value of the intervention demand. and will Take as the peak value of intervention demand Using peak ambiguity to characterize the peak moment, the formula for calculating concurrency controllability is as follows: ; The necessity of concurrency was obtained. With concurrency controllability Then, determine the maximum concurrent connection limit for the main channel. The system reads the necessity threshold from the parameter library. Controllability threshold With the main channel capacity limit It is a positive integer, reflecting the maximum number of main passageways that the public space venues and equipment architecture can support; Among them, the upper limit of the main channel capacity This represents the maximum number of concurrent main channels that the system can maintain within the same intervention time window, provided that the execution constraint of allowing only one strong cue output per slot is not violated; main channel capacity limit. The initial deployment value is determined by the number of output modal links that can be driven independently and simultaneously on the public space venue side, and the parallelism of its controller; specifically, the smaller of the two values is taken; necessity threshold. Controllability threshold The necessity of concurrent processing for samples that do not conflict within the calibration sample set. With concurrency controllability The distribution separation effect is determined, and the necessity of taking conflicting samples concurrently is determined. Conservative quantiles and controllability of non-conflicting sample concurrency The conservative quantiles are used as initial values; The system uses the smaller of the margins exceeding the corresponding thresholds for necessity and controllability as the concurrency margin, and then compares this concurrency margin with the preset hierarchical threshold sequence in the parameter library Π. Compare to obtain concurrency levels When the concurrency margin is negative, the concurrency level is used. When the concurrency margin falls into each of the different tiers, the concurrency level is selected sequentially. The final output is the maximum concurrency of the main channel. for: ; Output main channel concurrency limit Subsequently, the collaborative budget is further refined under the concurrency budget constraint to generate a modal repetition budget, specifically: based on the main channel concurrency limit. Calculate repeatable capacity ;in This represents the repeating modal capacity that the main channel set can form under size constraints; the dominant margin for system construction ambiguity. The determination of whether a repetitive modality needs to be introduced to improve understandability, given that concurrency is already allowed, is as follows: ; and the ambiguous dominant margin With preset dominant threshold Comparison of repeatable capacity The modal repetition budget is determined based on the relationship with zero, when the following conditions are met. and At that time, output modal repetition budget Otherwise, output modal repetition budget The preset dominant threshold is mentioned above. By statistically analyzing the dominant ambiguity in the calibrated sample set The distribution of the data is analyzed, and the conservative quantile that can distinguish between those requiring repeated confirmation and those not requiring repeated confirmation is selected as the preset dominant threshold. Initial value; Preferably, the system uses a preset basic interval. and upper and lower boundaries Construct the main channel occupancy ratio With rhythm conservative factor Combined with the necessity of concurrency Output minimum strong cue interval, i.e., based on the main channel occupancy ratio. Within [0, 1], the size will be a rhythmic conservative factor. The minimum strong cue interval is obtained by monotonically adjusting the interval. The formula for obtaining the minimum strong cue interval is as follows: ); among them, the necessity of concurrency The larger the value, the faster the pace tends to be, and the higher the concurrency limit of the main channel. The larger or The lower the value, the more conservative the pace; the system calculates the window length for the intervention time window W. The number of slots is determined by dividing the window length by the minimum strong prompt interval and rounding up. And write it into the budget package; the system further uses the total slot resources Given the maximum concurrency of the main channel, Integer amortization: Each main channel can receive a maximum of Each strong cue slot will be used as the upper limit for the strong cue density. And ensure strong indication of the upper limit of density. At least 1; and each main channel corresponds to at least one executable master role device. It should be noted that, Determined by the device's minimum controllable refresh cycle. The upper bound is determined by avoiding the possibility of missing the intervention window due to overly sparse prompts.
[0020] The device selection slot orchestration module is used to select the set of main role devices from the set of interactive output units based on the main channel concurrency limit, and generate a strong prompt token for the selected device within the intervention time window based on the main channel concurrency limit and bind it to the slot to generate an orchestration scheme; Based on the obtained channel availability According to channel availability The interactive output unit set DevSet is sorted in descending order and truncated to obtain a candidate set Cand; the size of the candidate set Cand must not exceed the preset upper limit of participating devices, which is set according to the specific needs of the user; based on the minimum strong prompt interval. Discretize the intervention time window as Each slot generates a set of strong hint tokens. Make its quantity equal to the upper limit of the strong hint density. This ensures that the number of strong prompts does not exceed the budget and that the interval between any two strong prompts meets the minimum strong prompt interval. ; Then, the main role set PR was determined and satisfied. This embodiment will perform modal repetition budgeting. The function acts as a piecewise function for the repetition bias penalty term; let the modal counts in set S be... The system defines the number of repeated modes as: The penalty for repetitive deviation is defined as follows: ; Therefore, when The penalty increases monotonically with the number of repetitions, and the main channels naturally tend to become complementary; when The time penalty is minimized when the number of repetitions is close to 1, and the main channel naturally tends to form an interpretable repetitive modal redundancy. During the iterative process of constructing the main role set PR, for each candidate device The system calculates the set after it is added. Three defining indicators: the first is the rate of change that can be achieved. That is, within the statistical S, the corresponding velocity scale parameter is greater than The number of candidate devices is determined to ensure that the main channel can execute strong prompts within the budget schedule; the second is the penalty for repeated deviations. Thirdly, the minimum channel availability of the set. The system uses a unified lexicographical selection criterion to determine the next primary role: prioritizing roles that ensure rhythmic reachability violates the rule. The smallest candidate device; when parallel, select one that penalizes repeatability deviation. The smallest candidate device; when grouped together again, select the one that minimizes the channel availability of the set. The largest candidate device; if still tied, the device with the greater channel availability is used as the deciding factor; the system repeats the above process until the set of primary roles is satisfied. ; Once the main character set is determined, the system selects from the remaining candidates... Complete the auxiliary role set RR in descending order, so that The auxiliary role set RR is used for weak cue persistence or status indication control commands and does not consume strong cue tokens; the system will Each token is bound to a slot sequence at even intervals. For example, starting from the first slot, several slots can be selected in fixed steps until all tokens are used up. Only one primary role is assigned to each token slot to trigger a strong prompt. Token slots are allocated among the primary roles in a cyclical order, ensuring that concurrent primary channels trigger in a determined rhythm within the intervention window and avoiding conflicts between multiple devices in the same slot. The system ultimately outputs a role mapping. The slot scheduling includes slot-to-device assignments and token consumption flags.
[0021] The instruction execution module is used to issue strong prompt instructions to token slots according to the orchestration scheme, and to issue weak prompt hold instructions to non-token slots. The instruction execution module receives the arrangement scheme output by the device selection slot arrangement module. Unified scheduling of slot times in the slot orchestration schedule: The minimum strong cue interval output by the concurrency budget parameter calculation module is used for calculation. As the slot step size, all slots within the intervention time window W are mapped to an absolute time sequence. The system generates a timestamped instruction packet Cmd(s) for each slot, the instruction packet containing at least: slot number, target device identifier, output mode identifier, strong or weak cue flag, and output level parameter; the system then uses role mapping... Mapping device roles to execution permissions: For the primary role set (PR), strong prompting commands are allowed to be triggered in the token slot; for the secondary role set (RR), only weak prompting hold or status indication commands are allowed to be triggered, thus ensuring that the number of strong prompts is limited. Constraints, strong cues, rhythm The constrained budget is implemented on the execution side; When slot s reaches time Upon arrival, the system reads the device assignment result for that slot from the slot scheduling schedule and issues control commands; for the assigned primary role device... PR, the system issues a strong warning command. The specific output content of the strong prompt instruction is not the subject of this analysis, but it at least includes three types of fields: output on, duration, and intensity level, and requires the device to return a confirmation receipt for process control; for auxiliary role devices in the same slot RR, the system issues a weak cues hold instruction without consuming the strong cues token. This is used to maintain the continuity of spatial prompts or to provide low-intensity supplementation to the main prompt; the system records the issuance time, receipt time, and effective time after each issuance, and in... The system updates the execution success flag and activation delay sample for this device to ensure that subsequent device slot selection and orchestration modules have a consistent source for using the delay field; to avoid cognitive load conflicts caused by multiple devices outputting simultaneously, the system implements a mechanism based on... Concurrency suppression: The system allows only one primary role device to execute strong hints in any slot; when the slot scheduling is abnormal, such as multiple strong hint assignments in the same slot or slot overlap caused by device acknowledgment delay, the system executes the decision according to the primary role priority order determined by the device selection slot scheduling module, retaining only the device with higher priority to output strong hints, and downgrading the remaining devices to weak hints or canceling their output in the slot, thereby ensuring that the slot-level single strong hint constraint is determined and implemented on the execution side; The system further calculates the modal repetition budget output by the concurrent budget parameter calculation module. Implement interpretable repeatability control: when During the execution phase, the system enables same-modality suppression. Specifically, if the previous strong hint modality is the same as the current planned strong hint modality within consecutive adjacent token slots, the system prioritizes calling devices with different modalities from the auxiliary role set RR to supplement with weak hints, thereby maintaining complementary output. At this time, the system allows controlled modal repetition without violating the single-slot single strong cue rule, and manages the repetition count by decrementing: each time a strong cue output for a repetitive modality occurs, the system will decrement the remaining repetition count. Decrease by 1 until it reaches 0, then automatically switch to same-mode suppression operation; Regarding failure degradation, the system maintains a continuous failure count for each device: when a device fails to return a confirmation receipt in a designated slot or the receipt indicates an execution failure, the system increments the failure count for that device within the current intervention window by 1; when the failure count reaches the failure limit in Π... Upon this, the system immediately revokes the strong prompt permission for the device in the remaining slots, selects the next priority device from the PR to take over the subsequent strong prompt slots, and reduces the channel availability of the device in a determined direction and writes it back. This ensures that the next round of concurrent budget parameter calculation module calculations... The candidate sorting in the equipment selection slot arrangement module can deterministically reflect the failure fact; After execution, the system summarizes the number of token consumptions, slot cancellations, average activation delay of strong prompts, and actual occurrences of repeated modes within the intervention time window into an execution summary and writes it into the process field of the status packet G. This summary is used for budget inference and orchestration stability analysis in subsequent rounds. If necessary, the system only applies the basic interval in Π. With controllability threshold Perform minimum step writeback: When multiple slot cancellations or slot conflict decisions occur, the base interval will be adjusted. Moving towards a conservative approach to advance the next round When consecutive failures and degradation lead to insufficient main channels, the controllability threshold will be adjusted. We move towards a more conservative approach to suppress concurrent entry conditions, ensuring that the closed loop maintains reproducible and stable convergence without introducing additional redundant thresholds.
[0022] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0023] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0024] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0025] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0026] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0027] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A public space age-friendly interactive system based on multimodal perception, characterized in that, It includes the following modules: a perception prediction parameter generation module, used to construct a three-dimensional dynamic spatial model of the public space based on multi-source fusion perception data, and to predict the behavior of target elderly users, obtaining the prediction uncertainty, peak intervention demand, and peak time; the urgency level is obtained by continuously monotonically decaying mapping based on the remaining time at the peak time and a preset urgency time constant; and the channel availability is calculated for each interactive device in the interactive output unit set, wherein the elements of the interactive output unit set are interactive devices with at least one output mode. The concurrent budget parameter calculation module is used to assess the single-channel failure risk based on the channel availability of the most reliable equipment, and to obtain the concurrent necessity by taking the smaller value between the urgency and the single-channel failure risk; the concurrent controllability is determined based on the prediction uncertainty and the peak of intervention demand; and the upper limit of the main channel concurrency within the intervention time window is determined by comparing the concurrent necessity and concurrent controllability with their corresponding preset thresholds. The device selection slot orchestration module is used to select the set of main role devices from the set of interactive output units based on the main channel concurrency limit, and generate a strong prompt token for the selected device within the intervention time window based on the main channel concurrency limit and bind it to the slot to generate an orchestration scheme; The instruction execution module is used to issue strong prompt instructions to token slots according to the orchestration scheme, and to issue weak prompt hold instructions to non-token slots.
2. The age-friendly interactive system for public spaces based on multimodal perception according to claim 1, characterized in that, The perception prediction parameter generation module pre-deploys a set of multi-source sensors and interactive output units in the public space, and uses a fusion positioning algorithm based on the factor graph optimization framework to register the lidar point cloud and IMU data and output the pose sequence. The pose sequence is then combined with the output of the semantic segmentation network to construct a three-dimensional dynamic space model. The three-dimensional dynamic space model includes a static layer containing structural boundaries and facility objects, and a dynamic layer that records crowd density, occlusion relationships and temporary obstacles with time index.
3. The age-friendly interactive system for public spaces based on multimodal perception according to claim 2, characterized in that, The specific methods for behavior prediction by the perception prediction parameter generation module include: using a behavior prediction model based on spatiotemporal graph convolution and attention mechanism, constructing a spatiotemporal graph of the target user and surrounding pedestrians, extracting motion features and modeling social dependencies, and outputting the probability distribution of the target user's future position at each prediction step within the prediction window. The probability distribution is given in the form of a position mean sequence and a position covariance sequence.
4. The age-friendly interactive system for public spaces based on multimodal perception according to claim 3, characterized in that, The budget parameter calculation module determines the main channel concurrency limit by: calculating the excess amount of concurrency necessity exceeding the preset concurrency necessity threshold, and the excess amount of concurrency controllability exceeding the preset concurrency controllability threshold; taking the smaller of the two excess amounts as the concurrency margin; comparing the concurrency margin with the preset graded threshold sequence, and determining the concurrency level based on the interval it falls into; finally, based on the concurrency level and the system's preset main channel capacity limit, comprehensively determining the main channel concurrency limit; wherein the concurrency necessity is based on the prediction uncertainty and the peak value of intervention demand, taking the larger of the two, and using this larger value to characterize the overall uncertainty of system control concurrency, thereby deriving the corresponding concurrency necessity.
5. The age-friendly interactive system for public spaces based on multimodal perception according to claim 1, characterized in that, The concurrent budget parameter calculation module is also used to determine the minimum strong alert interval. The method is as follows: while determining the upper limit of the main channel concurrency, the concurrent budget parameter calculation module obtains the minimum strong alert interval by multiplying the basic interval by the rhythm conservative factor based on the preset basic interval and rhythm conservative factor. The intervention time window is then discretized using the minimum strong alert interval, dividing the intervention time window into several slots. The upper limit of the strong alert density is determined by the ratio of the upper limit of the main channel concurrency to the length of the intervention time window.
6. The age-friendly interactive system for public spaces based on multimodal perception according to claim 5, characterized in that, The concurrent budget parameter calculation module is also used to determine the modal repetition budget. The method is as follows: the judgment is based on the main channel concurrency limit and the ambiguous dominant margin calculated according to the peak intervention demand. When the ambiguous dominant margin is greater than the dominant threshold and the repeatable modal capacity is greater than zero, the modal repetition budget is set to allow one controlled modal repetition. Otherwise, it is set to suppress modal repetition.
7. The age-friendly interactive system for public spaces based on multimodal perception according to claim 6, characterized in that, After receiving the main channel concurrency limit, strong cue density limit, and modal repetition budget, the device selection slot orchestration module first sorts the interactive devices in the interactive output unit set in descending order according to channel availability, and then truncates them according to the preset participation device limit to obtain a candidate device set. Iteratively selects the main role device set from the candidate device set, and adopts different device selection penalty logic according to the modal repetition budget: when the modal repetition budget is to suppress repetition, a penalty that increases with the number of repetitions is applied to the output modal that appears repeatedly in the candidate set; when the modal repetition budget is to allow one controlled repetition, the penalty is minimized when a modal repetition occurs once in the candidate set.
8. The age-friendly interactive system for public spaces based on multimodal perception according to claim 7, characterized in that, When constructing the main role device set, the device selection slot arrangement module calculates evaluation indicators for each candidate device after its addition, including at least the rhythm reachability violation number, modal repetition deviation, and minimum channel availability of the set. It then uses a unified lexicographical selection criterion to compare the candidate devices and selects the device with the best evaluation result to add to the main role device set until the main channel concurrency limit is met or no device can be added.
9. The age-friendly interactive system for public spaces based on multimodal perception according to claim 8, characterized in that, After the primary role device set is determined, the device selection slot arrangement module selects auxiliary role devices from the remaining part of the candidate device set in descending order of channel availability. Within the intervention time window, it generates a slot sequence according to the minimum strong cue interval, binds the maximum strong cue density tokens to the slot sequence at the most even intervals, and then uses a cyclic allocation method to allocate token slots among the primary role device sets, generating an arrangement scheme that includes the correspondence between slots and devices.
10. The age-friendly interactive system for public spaces based on multimodal perception according to claim 1, characterized in that, When executing the orchestration scheme, the instruction execution module issues strong prompt instructions only to the primary role device bound to the slot for each token slot, and issues weak prompt holding instructions or status indication instructions only to the primary role device or auxiliary role device for other slots within the intervention time window, and ensures that multiple devices do not execute strong prompts simultaneously in any slot.