Elevator advertisement audience behavior analysis system and method based on multi-modal data acquisition

By employing a multimodal data acquisition system and a dynamic fusion strategy, the problem of balancing privacy protection and perception accuracy in elevator advertising audience analysis was solved, enabling fine-grained audience behavior analysis and improving the accuracy of advertising effectiveness evaluation and privacy protection capabilities.

CN120987157APending Publication Date: 2025-11-21林家君
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511094394.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies for elevator advertising audience analysis suffer from several problems, including difficulty in balancing privacy protection and perception accuracy, poor adaptability of multimodal data fusion, and insufficient granularity of behavioral analysis. These issues fail to meet the comprehensive needs of advertising for privacy compliance, accurate perception, and in-depth insights.

Method used

A multimodal data acquisition system is adopted, including a millimeter-wave radar module, a low-resolution thermal imaging module, an ambient light and distance sensing module, and an elevator status interface module. By combining spatiotemporal alignment, dynamic weight adjustment, and conflict handling, non-intrusive multimodal data acquisition and fine-grained analysis are achieved, and anonymization is performed through a privacy protection module.

Benefits of technology

It achieves stable perception in the complex environment of elevators, provides fine-grained audience behavior analysis, improves the accuracy of advertising effectiveness evaluation and privacy protection, and complies with privacy regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120987157A_ABST
    Figure CN120987157A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal data processing, and discloses an elevator advertisement audience behavior analysis system and method based on multi-modal data acquisition, and the system comprises a data acquisition layer which is used for collecting multi-modal sensing information in an elevator space, the data acquisition layer comprises a millimeter wave radar module, a low-resolution thermal imaging module, an ambient light and distance sensing module, a directional microphone array module and an elevator state interface module; and the data processing and fusion layer is connected with the data acquisition layer and used for preprocessing and dynamically fusing the multi-modal sensing information, and the data processing and fusion layer comprises a space-time alignment unit, a dynamic weight adjustment unit and a conflict processing unit. According to the method, a complete audience behavior analysis scheme for the elevator closed space is formed through non-intrusive multi-modal data acquisition, a dynamic fusion strategy adaptive to the elevator scene and fine-grained behavior modeling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-modal data processing, and more particularly to an elevator advertisement audience behavior analysis system and method based on multi-modal data collection. BACKGROUND

[0002] As a closed and high-frequency touch public space, the elevator is an important scene for advertisement delivery, and its advertisement effect evaluation and audience behavior analysis depend on accurate perception of audience interaction state. In the prior art, elevator advertisement audience analysis mainly adopts the following schemes:

[0003] Single visual sensor scheme: image is collected by a camera, and whether the audience watches the advertisement is determined based on face recognition, gaze tracking and other technologies. This scheme has significant privacy risks: the camera directly collects biological feature information such as face and body shape, which easily triggers user privacy anxiety and does not comply with strict restrictions on the collection of personal sensitive information in privacy regulations such as GDPR and CCPA; at the same time, the lighting in the elevator is complex (such as light that is bright and dark at random, direct external light), and people are frequently blocked (such as crowded by multiple people), resulting in visual data being easily disturbed and the recognition accuracy being greatly reduced.

[0004] Simple presence sensor scheme: single sensors such as infrared sensors and ultrasonic sensors are used to only determine "presence / absence" or roughly estimate the length of stay, and cannot obtain key information such as audience orientation, attention allocation, and behavior patterns, so the analysis dimension is limited to "contact / no contact", and cannot support the depth of insight required for advertisement content optimization (such as "which time period attracts more attention" and "what is the audience interest tendency").

[0005] Primary multi-modal fusion scheme: some schemes attempt to combine cameras and infrared sensors, but do not design a fusion strategy for elevator scene characteristics: the elevator space is small (usually 1-3 m2), and the state changes dynamically (door opening and closing, rapid lifting, and instantaneous influx and outflow of people), resulting in natural contradictions in time synchronization (such as differences in sensor sampling rates) and spatial correlation (such as different sensor coordinate systems) of different modal data; and no dynamic weight adjustment mechanism based on scene state (such as vision being easily blocked when the door is open) is established, so when sensor data conflicts (such as the camera showing "facing the screen" but actually misjudging due to blocking), reliable fusion cannot be achieved, and the analysis result has low credibility.

[0006] Coarse granularity of behavior analysis: existing technologies model audience behavior in a binary judgment of "watched / not watched", without further exploring fine-grained features such as attention intensity (e.g. the difference between "intensive watching" and "glancing"), micro-behavior patterns (e.g. the difference between "actively approaching the screen" and "passively staying"), interest fluctuation trends (e.g. the rising / falling nodes of attention during ad playback), etc., resulting in advertisers being unable to obtain core decision-making criteria such as "which part of the content is effective" and "what type of information is the audience more sensitive to".

[0007] In summary, in the analysis of elevator advertising audiences, the existing technologies have the problems of difficulty in balancing privacy protection and perception accuracy, poor adaptability of multi-modal data fusion, and insufficient granularity of behavior analysis, which cannot meet the comprehensive needs of "privacy compliance, accurate perception, and deep insight" for ad delivery. Therefore, there is an urgent need for an elevator advertising audience behavior analysis system and method based on multi-modal data collection to solve the above problems. SUMMARY

[0008] In order to overcome the above-mentioned defects of the prior art, the present application provides an elevator advertising audience behavior analysis system and method based on multi-modal data collection to solve the problems in the above background art.

[0009] The present application provides the following technical solution: an elevator advertising audience behavior analysis system based on multi-modal data collection, comprising:

[0010] A data collection layer for collecting multi-modal perception information in an elevator space, the data collection layer comprising a millimeter wave radar module, a low-resolution thermal imaging module, an ambient light and distance sensor module, a directional microphone array module, and an elevator state interface module;

[0011] A data processing and fusion layer connected to the data collection layer for pre-processing and dynamic fusion of multi-modal perception information, the data processing and fusion layer comprising a space-time alignment unit, a dynamic weight adjustment unit, and a conflict processing unit;

[0012] A behavior analysis and modeling layer connected to the data processing and fusion layer for constructing a fine-grained audience behavior model based on the fused information, the behavior analysis and modeling layer comprising a multi-modal attention estimation unit, a micro-behavior recognition unit, and an interest curve modeling unit;

[0013] A privacy protection processing module connected to the data collection layer, the data processing and fusion layer, and the behavior analysis and modeling layer, respectively, for anonymizing the data of each layer.

[0014] Through the multi-module collaborative architecture, the integration of non-intrusive multi-modal data acquisition, dynamic fusion and fine-grained analysis is realized for the first time in the elevator scene, which not only solves the problem of insufficient information of a single sensor, but also balances the perception accuracy and privacy security through the privacy protection module from the system level.

[0015] As a further scheme of the application, the millimeter wave radar module is configured to detect the number, position coordinates, moving trajectory, stay duration and posture features of the audience, the posture features including standing, looking down or hugging arms, and no face or personal identity information is collected.

[0016] The millimeter wave radar completely avoids the collection of identity information such as face while providing accurate behavior features (such as posture and trajectory), eliminates privacy risks from the source compared with traditional visual sensors, and is not affected by light and occlusion, significantly improving the perception stability in complex elevator environments.

[0017] As a further scheme of the application, the dynamic weight adjustment unit is configured to dynamically adjust the fusion weight of each modal data based on the elevator door opening and closing state and running state obtained by the elevator state interface module; wherein when the elevator door is in an open state, the fusion weight of the data collected by the millimeter wave radar module is increased.

[0018] By dynamically adjusting the weight in combination with the real-time state of the elevator, the multi-modal fusion can adapt to changes in scenes such as door opening and closing (such as avoiding the influence of visual occlusion when the door is open), solving the problem that sensor data in a small space is easily affected by scene interference, and improving the reliability of the fusion result.

[0019] As a further scheme of the application, the multi-modal attention estimation unit is configured to combine the body orientation detected by the millimeter wave radar module, the approximate orientation detected by the low-resolution thermal imaging module, the distance parameter detected by the ambient light and distance sensing module, and the non-speech acoustic features detected by the directional microphone array module, and calculate the attention intensity of the audience through a preset scoring rule, the attention intensity being divided into 0-4 levels.

[0020] Breaking the traditional binary judgment of "watching or not watching", the multi-modal feature joint modeling realizes the quantitative grading of attention intensity, which can more accurately reflect the audience's attention degree to the advertisement, and provides a more fine-grained quantitative basis for advertisement effect evaluation.

[0021] As a further scheme of the application, an elevator advertisement audience behavior analysis method based on multi-modal data acquisition is proposed, comprising the following steps:

[0022] S1. Multi-modal information collection: Collect multi-modal perception information and elevator operation state in the elevator through millimeter wave radar, low-resolution thermal imaging, ambient light and distance sensor, directional microphone array and elevator state interface synchronously;

[0023] S2. Preprocessing and space-time alignment: Time synchronization and space coordinate unification are performed on the multi-modal perception information, the time synchronization is realized based on dynamic time warping algorithm, and the space coordinate unification is realized by mapping to a preset elevator space coordinate system;

[0024] S3. Multi-modal dynamic fusion: The fusion weight of each modal data is dynamically adjusted based on the elevator operation state, and the multi-modal data with conflict is fused by confidence weighting;

[0025] S4. Behavior state analysis and modeling: Based on the fused information, multi-modal attention estimation, micro-behavior recognition and interest curve modeling are performed to generate fine-grained audience behavior analysis results;

[0026] S5. Result output: The anonymized analysis results after privacy protection processing are output.

[0027] The method forms a closed loop of multi-modal data collection, alignment, fusion, analysis and privacy protection through standardized step process, especially optimizes the synchronization and fusion strategy for the space-time characteristics of the elevator scene, ensures efficient output of reliable anonymized analysis results in a dynamic small space, and has strong implementability.

[0028] As a further scheme of the application, in step S3, the rule of dynamically adjusting the fusion weight includes: when the elevator door is in an open state or the ambient light intensity is less than 50 lux, the fusion weight of millimeter wave radar data is increased to 0.6, and the fusion weight of visual related sensor data is reduced.

[0029] By explicitly defining the scene-based weight adjustment rule, the multi-modal fusion can specifically deal with typical problems such as elevator door opening (easy to be blocked) and dim light (visual failure), avoid invalid data interference, and improve the robustness of the fusion result in complex environment.

[0030] As a further scheme of the application, in step S4, the process of micro-behavior recognition includes: based on the moving track and posture features detected by millimeter wave radar and the relative activity level detected by low-resolution thermal imaging, the micro-behavior pattern library is matched, the micro-behavior pattern library includes "waiting to leave" mode and "interest close" mode.

[0031] Through the micro-behavior pattern library, the audience behavior can be classified in detail, different interaction types such as "active attention" and "passive stay" can be distinguished, and compared with traditional presence detection, the real interest tendency of the audience can be reflected, which provides a directional guide for advertisement content optimization.

[0032] As a further scheme of the application: in step S5, the privacy protection processing includes: only retaining the anonymized group behavior characteristics, not associating any individual identification, and the local storage duration of the original data is not more than 24 hours.

[0033] Through group anonymization and storage time limit control, privacy protection is strengthened at the end of data processing, ensuring that the analysis results only reflect the group rules and do not leak individual information, meeting the requirements of privacy regulations on data minimization and de-identification.

[0034] Technical effects and advantages of the application:

[0035] The application forms a complete audience behavior analysis scheme for the elevator closed space through non-invasive multi-modal data acquisition, dynamic fusion strategy suitable for elevator scene and fine-grained behavior modeling, and the technical effects and advantages are as follows:

[0036] The application uses non-visual sensors such as millimeter wave radar and low-resolution thermal imaging to avoid obtaining biological characteristics such as face and clear body shape that can identify individual identity from the source of acquisition; combined with the anonymization processing of the privacy protection processing module (only retaining group behavior characteristics, limiting local data storage duration), the privacy leakage risk of traditional visual solutions is eliminated from the technical architecture, naturally adapting to the restrictions of privacy regulations such as GDPR and CCPA on the collection of personal sensitive information, and solving the core contradiction between "precise perception" and "privacy protection" in the elevator scene.

[0037] In view of the characteristics of small elevator space, variable lighting (light switching, external light interference), frequent personnel flow (instant influx / exit when the door is opened / closed) and common shielding, the application breaks through the performance limitations of single sensor in complex environment through multi-sensor redundancy design (such as radar not affected by light, thermal imaging resistant to local shielding), and combined with the dynamic weight adjustment mechanism based on elevator state (door opening / closing, running / stop), such as increasing the weight of radar data to avoid visual shielding when the door is opened), ensuring that effective sensing information can still be stably output when the elevator state changes dramatically, overcoming the problem that existing solutions are prone to failure in dynamic scenes.

[0038] The application discards the traditional binary judgment mode of "watching / not watching", realizes the extraction of fine-grained features such as attention intensity grading, micro-behavior pattern (such as "interest close" and "wait to leave") recognition and interest fluctuation trend (high light moment / loss point) analysis through multi-modal data joint modeling. This deep analysis capability surpasses the information dimension of simple presence detection, and can capture the audience's attention allocation rules and potential interest tendency for advertising, providing more valuable insight dimensions for advertising content optimization and delivery strategy adjustment.

[0039] The innovative design of the application is based on dynamic fusion rules of elevator states, multi-sensor time synchronization is realized through dynamic time warping, spatial correlation is realized through a unified coordinate system, and conflict processing is performed in combination with environmental parameters (such as light) and sensor characteristics, so that multi-modal data can be efficiently coordinated when the elevator door is opened and closed, and the running / stop state is switched. This fusion strategy that adapts to the dynamic characteristics of the small space of the elevator solves the problem of data asynchronization and conflict difficult to adjust in the existing multi-modal scheme in a complex scene, and provides a reliable data basis for fine-grained analysis. BRIEF DESCRIPTION OF DRAWINGS

[0040] The application will be further described below with reference to the drawings.

[0041] Figure 1 is a system block diagram of an elevator advertisement audience behavior analysis system based on multi-modal data acquisition according to the application;

[0042] Figure 2 is a flowchart of an elevator advertisement audience behavior analysis method based on multi-modal data acquisition according to the application. DETAILED DESCRIPTION

[0043] The technical solutions of the application will be described in detail below with reference to the drawings and specific application scenarios. Those skilled in the art should understand that the specific embodiments described herein are only used to explain the application, and are not used to limit the protection scope of the application.

[0044] I. Overall architecture of the system

[0045] The elevator advertisement audience behavior analysis system based on multi-modal data acquisition according to the application is deployed in the elevator car, and the core processing unit is integrated in the edge computing module (using ARMCortex-A53 quad-core processor, main frequency 1.8GHz, memory 2GB) built in the elevator advertisement screen. Each module is connected through an internal bus to form a closed-loop architecture of "collection-processing-analysis-output". The system logic architecture is shown in Figure 1 As shown, it includes a data collection layer, a data processing and fusion layer, a behavior analysis and modeling layer, and a privacy protection processing module, and each layer cooperates to realize fine-grained analysis of audience behavior.

[0046] II. Specific configuration and working principle of each module

[0047] (I) Data collection layer

[0048] The data collection layer is deployed on the inner wall of the elevator car and the integrated area of the advertisement screen, and the specific configuration of each module is as follows:

[0049] Millimeter wave radar module

[0050] Installation position: top left of the advertisement screen (2.2m from the ground), horizontal depression angle 30°, covering the range of 0.5-3m inside the elevator car.

[0051] Hardware parameters: 77GHz FMCW radar, 8x8MIMO antenna array, sampling rate 10Hz, distance resolution 0.1m, angle resolution 5°, output point cloud data (including target x / y / z coordinates, velocity, and reflective cross-sectional area).

[0052] Working principle:

[0053] People counting: clustering human point cloud (reflective cross-sectional area -10~ -20dBsm) by DBSCAN clustering algorithm, the number of clusters is the number of people (supporting 1-8 people);

[0054] Posture recognition: analyzing the upper limb reflection point clustering degree (clustering degree >60% is determined as "hugging arms") by point cloud skeleton extraction algorithm, and analyzing the upper body point cloud pitch angle (<30° is determined as "bending over");

[0055] Trajectory tracking: fitting the coordinates of each cluster center in consecutive frames by Kalman filtering algorithm to generate a moving trajectory (e.g., if the coordinates change from (2.0, 0.5) to (1.0, 0.5), it is determined as "approaching the advertisement screen").

[0056] Low-resolution thermal imaging module

[0057] Installation position: top right of the advertisement screen (symmetrical to the radar), horizontal field of view 60°.

[0058] Hardware parameters: 32x32 resolution infrared focal plane array, temperature measurement range 20-40℃, output non-textured temperature matrix (does not include facial details).

[0059] Working principle:

[0060] Presence area detection: adaptive threshold method (areas higher than the ambient temperature by 5℃ are determined as human areas);

[0061] Approximate direction estimation: calculating the principal axis direction of the minimum bounding rectangle of the thermal area, and the angle with the normal of the advertisement screen is <30°, which is determined as "facing the screen area";

[0062] Relative activity level: calculating the shape change rate of the thermal area in consecutive 3 frames (change rate >50% is determined as "high activity state", reflecting anxiety or frequent movement).

[0063] Ambient light and distance sensing module

[0064] Installation position: center of the top frame of the advertisement screen.

[0065] Hardware parameters: Ambient light sensor (0-1000 lux, accuracy ±10 lux); Infrared distance sensor (0.5-3 m, sampling rate 5 Hz, accuracy ±0.1 m).

[0066] Working principle:

[0067] Distance parameter output: Real-time return of the vertical distance between the audience and the screen and the distance change rate (e.g. -0.2 m / s indicates "continuous approach");

[0068] Ambient light intensity is used to assist in determining the reliability of the visual sensor (marked as "low light environment" when <50 lux).

[0069] Directional microphone array module

[0070] Installation location: Lower frame of the advertising screen, 4-microphone linear array, directivity 120° (covering the area in front of the screen).

[0071] Hardware parameters: Sampling rate 16 kHz, frequency band analysis range 200-2000 Hz, no original audio stream is stored.

[0072] Working principle:

[0073] Extract the frequency band energy through short-time Fourier transform, and determine "low-frequency sound event" (such as laughter) when the burst energy in the 200-500 Hz frequency band is enhanced (> threshold 20 dB);

[0074] Determine "conversation state" when the sustained energy in the 1000-2000 Hz frequency band is > 3s.

[0075] Elevator state interface module

[0076] Hardware parameters: Communicate with the elevator controller through the RS485 interface to obtain the door opening signal (high level "door open", low level "door closed"), running state (up / down / stop), and current floor (update frequency 1 Hz).

[0077] (II) Data processing and fusion layer

[0078] This layer runs on the edge computing module, and the core algorithm is implemented through C++ language. The working principles of each unit are as follows:

[0079] Space-time alignment unit

[0080] Time synchronization: Use the dynamic time warping (DTW) algorithm to align the data of different sampling rates such as radar (10 Hz) and thermal imaging (5 Hz) with the timestamp (1 Hz) of the elevator state signal as the reference, ensuring that the time deviation is <100 ms;

[0081] Let the radar data sequence be:

[0082] R = [r1, r2,..., r m ], (sampling time t R = [t R1 , t R2 ,..., t Rm ]), thermal imaging data sequence T = [t1, t2,..., t n ](sampling time t T = [t T1 , t T2 ,..., t Tn ]), DTW is aligned by calculating the minimum cumulative distance D(R, T) : D(R, T) = min∑ i,j w i,j ·d(r i ,t j ), where w i,j is the path weight (satisfying the boundary condition w 1,1 = 1, w m,n = 1), d(r i ,t j ) is the Euclidean distance, and after alignment, it is unified to the target time sequence t = [0, 0.1, 0.2,..., 14.9] (15s advertisement length, step 0.1s).

[0083] Spatial unification: a two-dimensional coordinate system (x-axis parallel to the horizontal direction of the screen, y-axis perpendicular to the screen) is established with the center of the advertisement screen as the origin, and the radar point cloud coordinates and thermal imaging hot area coordinates are converted to this coordinate system through a preset mapping matrix M (based on installation position calibration) : [x', y'] T = M·[x, y] T , where [x, y] is the local coordinate of the sensor, and [x', y'] is the unified coordinate system coordinate, realizing spatial correlation.

[0084] Dynamic weight adjustment unit

[0085] The fusion weights of each modal data are dynamically adjusted based on the state of the elevator (the total weight is 1), and the rules are as follows:

[0086] Elevator "door open + stop" state: radar (0.6), thermal imaging (0.3), distance (0.1), microphone and ambient light do not participate (because the external noise interference is large when the door is open) ;

[0087] Elevator "door closed + running" state: radar (0.4), thermal imaging (0.2), distance (0.2), microphone (0.1), ambient light (0.1, auxiliary correction) ;

[0088] Low light environment (<50 lux) : thermal imaging weight is increased by 0.1, and other visual related weights are reduced by 0.1.

[0089] Conflict processing unit

[0090] When multimodal data conflict on the same behavior judgment (e.g. thermal imaging determines "facing screen" but radar shows "body orientation deviation"), weighted decision based on sensor historical accuracy library:

[0091] Establish accuracy matrix of each sensor in different scenarios (e.g. radar orientation judgment accuracy is 85% when the door is open, thermal imaging is 60%);

[0092] Calculate confidence score: sensor accuracy a k x current weight w k The final decision result D is the highest score: (e.g. radar score 0.85 x

[0093] 0.6 = 0.51, thermal imaging 0.6 x 0.3 = 0.18, accept radar result).

[0094] (Three) Behavior analysis and modeling layer

[0095] Multimodal attention estimation unit

[0096] Based on the fused multimodal features, the attention intensity (0-4 levels) is calculated through the preset scoring rules:

[0097] Let the feature variables be: body orientation matching degree A (0 or 1), thermal imaging orientation matching degree B (0 or 1), close distance factor C (0 or 1, distance < 1.5m is 1), close factor D (0 or 1, distance change rate < -0.1m / s is 1), acoustic event factor E (0 or 1, peak segment with low frequency event is 1), then the attention intensity S is: S = A + B + C + D + E.

[0098] For example:

[0099] Body orientation (radar): angle with screen normal < 15° (+1 point);

[0100] Approximate orientation (thermal imaging): facing screen area (+1 point);

[0101] Distance parameter: distance < 1.5m (+1 point), distance change rate < -0.1m / s (+1 point);

[0102] Acoustic features: low frequency sound event occurs in the peak segment of the advertisement (+1 point);

[0103] Total score corresponds to intensity: 0 points (no attention), 1-2 points (weak attention), 3-4 points (strong attention).

[0104] Micro-behavior recognition unit

[0105] Pre-set micro-behavior pattern library, identify through multi-modal feature matching:

[0106] "Interest approach" mode: rate of distance change < -0.1 m / s (for 2 s) + body orientation angle < 30° + no talking acoustic features;

[0107] "Wait to leave" mode: radar trajectory concentrated on the elevator door side (x > 1.5 m) + thermal imaging activity level > 50% + distance > 2 m;

[0108] "Passive stay" mode: rate of distance change < 0.05 m / s (basically static) + body orientation angle > 60° + no specific acoustic features.

[0109] Interest curve modeling unit

[0110] Take the advertisement playing time (0-15s) as the horizontal axis, calculate the average attention intensity every second, and generate the interest curve:

[0111] High light moment: 2s continuous slope > 0.5 (intensity rising), marked as "content attractiveness peak";

[0112] Dropout point: 2s continuous slope < -0.5 (intensity falling), marked as "content attractiveness trough".

[0113] (Four) Privacy protection processing module

[0114] Hardware carrier: built-in edge computing module in elevator advertising screen (no cloud upload channel).

[0115] Core processing rules:

[0116] Raw data localization: raw data such as radar point cloud and thermal imaging matrix are only processed in local memory and are not written into persistent storage;

[0117] Anonymization: analysis results only retain group features (such as "3 people average attention intensity 2.5" "1 time interest approach mode"), do not associate with any individual identifier;

[0118] Storage time limit control: local cache fusion feature data (non-raw data) retention time ≤24 hours, automatically cleared every morning.

[0119] II. Method execution process (see Figure 2 )

[0120] Take a certain brand beverage elevator advertisement (15s length) playing scene as an example, the method steps are as follows:

[0121] S1. Multi-modal information collection

[0122] Elevator stops at 3rd floor, door opens (elevator state interface output "door open, stop at 3rd floor"), 2 subjects enter the car:

[0123] Millimeter wave radar: 2 persons detected, position coordinates (1.2, 0.5), (0.8, 0.6), body orientation angles 20°, 45° respectively, no arm-holding / looking-down posture;

[0124] Thermal imaging: 2 thermal zones, roughly oriented at 30°, -15°, activity levels <30% for both;

[0125] Distance sensor: distances 1.8m, 2.2m respectively, rate of change 0m / s (stationary);

[0126] Microphone array: no low-frequency sound event detected.

[0127] S2. Preprocessing and spatio-temporal alignment

[0128] Temporal synchronization: DTW interpolation on radar (10Hz) and thermal imaging (5Hz) data with respect to elevator state timestamp (t=0s), unified to 10Hz time series;

[0129] Spatial unification: mapping radar coordinates and thermal imaging thermal zones to a unified coordinate system, confirming that both persons are located within 1-2m in front of the advertisement screen.

[0130] S3. Multimodal dynamic fusion

[0131] Due to "door open" state, dynamic weight adjustment: radar 0.6, thermal imaging 0.3, distance 0.1;

[0132] No conflict in orientation data for both persons (radar 20° + thermal imaging 30° both comply with "facing area"), fused as behavior feature sequence: 2 persons, average distance 2.0m, initial attention score 2 points (1 person gets 1 point and 3 points respectively).

[0133] S4. Behavior state analysis and modeling

[0134] Advertisement plays to 3s: 1 person distance rate of change -0.2m / s (approaching), attention score rises to 3 points;

[0135] Play to 8s: microphone detects 1 low-frequency sound event (laughter), this person's score adds 1 point (total score 4 points);

[0136] Play to 12s: elevator door closes (state switches to "door closed, going up"), the other person's body orientation angle drops to 25°, score rises to 2 points;

[0137] Interest curve shows 3-8s as the highlight moment (slope 0.6), no obvious drop-off point;

[0138] One "interest close" mode (consistent close + screen facing feature) is identified.

[0139] S5. Result output

[0140] Output the anonymized analysis result: "2-person audience, average attention intensity 2.5, 8s high-light moment, 1 interest close behavior detected", original data is destroyed in real time, and the result is cached locally (automatically deleted after 24 hours).

[0141] Three, the understanding of those skilled in the art

[0142] In the present application, the point cloud clustering of millimeter wave radar, the interpolation algorithm of dynamic time warping, and the scoring rules of attention intensity are all adjustable parameterized settings. Those skilled in the art can modify the parameter thresholds (such as distance determination threshold, angle threshold) to adapt to actual scenarios such as elevator car size (such as 1.5m x 2m or 2m x 2.5m) and advertisement screen installation position (side-mounted or top-mounted). Such adjustments are within the scope of protection of the present application.

Claims

1. An elevator advertisement audience behavior analysis system based on multi-modal data collection, characterized by, The application comprises: a data acquisition layer for collecting multi-modal perception information in an elevator space, the data acquisition layer comprising a millimeter wave radar module, a low-resolution thermal imaging module, an ambient light and distance sensor module, a directional microphone array module, and an elevator state interface module; a data processing and fusion layer connected to the data acquisition layer, for pre-processing and dynamic fusion of multi-modal perception information, the data processing and fusion layer comprising a space-time alignment unit, a dynamic weight adjustment unit, and a conflict processing unit; a behavior analysis and modeling layer connected to the data processing and fusion layer, for constructing a fine-grained audience behavior model based on the fused information, the behavior analysis and modeling layer comprising a multi-modal attention estimation unit, a micro-behavior recognition unit, and an interest curve modeling unit; a privacy protection processing module connected to the data acquisition layer, the data processing and fusion layer, and the behavior analysis and modeling layer, respectively, for anonymizing the data of each layer.

2. The elevator advertisement audience behavior analysis system based on multi-modal data collection according to claim 1, characterized in that: The millimeter wave radar module is configured to detect the number, position coordinates, movement trajectory, stay duration, and posture features of the audience, the posture features including standing, looking down, or arms crossed, and does not collect facial or personal identity information.

3. The elevator advertisement audience behavior analysis system based on multi-modal data collection of claim 1, wherein: The dynamic weight adjustment unit is configured to dynamically adjust the fusion weights of each modal data based on the elevator door opening and closing state and the running state obtained by the elevator state interface module; when the elevator door is open, the fusion weight of the millimeter wave radar module data is increased.

4. The elevator advertisement audience behavior analysis system based on multi-modal data collection of claim 1, wherein: The multi-modal attention estimation unit is configured to calculate the attention intensity of the audience by a preset scoring rule, combining the body orientation detected by the millimeter wave radar module, the general orientation detected by the low-resolution thermal imaging module, the distance parameter detected by the ambient light and distance sensor module, and the non-speech acoustic features detected by the directional microphone array module, the attention intensity being divided into 0-4 levels.

5. An elevator advertising audience behavior analysis method based on multi-modal data collection, characterized by, The application comprises the following steps: S1. Multi-modal information acquisition: synchronously collecting multi-modal perception information and elevator running state in the elevator through millimeter wave radar, low-resolution thermal imaging, ambient light and distance sensor, directional microphone array, and elevator state interface; S2. Pre-processing and space-time alignment: time synchronization and space coordinate unification of the multi-modal perception information, the time synchronization being realized based on dynamic time warping algorithm, and the space coordinate unification being realized by mapping to a preset elevator space coordinate system; S3. Multi-modal dynamic fusion: dynamically adjusting the fusion weights of each modal data based on the elevator running state, and performing confidence weighted fusion on multi-modal data with conflicts; S4. Behavior state analysis and modeling: generating fine-grained audience behavior analysis results through multi-modal attention estimation, micro-behavior recognition, and interest curve modeling based on the fused information; S5. Result output: outputting the anonymized analysis results after privacy protection processing.

6. The method of claim 5, wherein the method is based on multi-modal data collection. In step S3, the rules for dynamically adjusting the fusion weights include: when the elevator door is open or the ambient light intensity is less than 50 lux, the fusion weight of the millimeter wave radar data is increased to 0.6, and the fusion weight of the visual related sensor data is reduced.

7. The method of claim 5, wherein the method further comprises: In step S4, the micro-behavior identification process includes: matching the mobile trajectory and posture features detected by the millimeter wave radar, the relative activity level detected by the low-resolution thermal imaging, and the preset micro-behavior mode library, the micro-behavior mode library includes a "waiting to leave" mode and an "interest close" mode.

8. The method of claim 5, wherein the method is based on multi-modal data collection. In step S5, the privacy protection processing includes: only retaining the anonymized group behavior features, not associating any individual identifier, and the local storage duration of the original data is not more than 24 hours.

Citation Information

Cited By

  • Elevator multi-mode risk intelligent identification and disposal method

    CN121591076A