Multi-modal human body situation awareness method based on data fusion, terminal equipment and storage medium

By synchronously collecting infrared images and point cloud sequences, combining cloud collaborative verification and local neural network model, the problems of environmental interference and privacy exposure in traditional human monitoring technology are solved, and accurate identification of human body posture and privacy security protection are achieved.

CN120472542AActive Publication Date: 2025-08-12SHENZHEN LUJIANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510947681.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-12
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Traditional human body monitoring technology is in a dilemma due to its reliance on a single sensor. Thermal imaging solutions are disturbed by environmental heat sources, and visual monitoring requires continuous high-definition images to be taken, resulting in privacy data exposure.

Method used

By synchronously collecting infrared image sequences and point cloud sequences, the height change and projection area change characteristics are extracted, abnormal behavior is judged in real time, and infrared image verification is triggered when there is high suspicion, combining cloud collaborative verification and local neural network model for confirmation.

Benefits of technology

In complex environments, the accurate identification of human body posture and privacy security protection are achieved, reducing the frequency of impact of environmental heat source interference on system judgments, and avoiding the risk of privacy leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472542A_ABST
    Figure CN120472542A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of data processing, and discloses a multi-modal human body situation awareness method based on data fusion, terminal equipment and a storage medium. The multi-modal human body situation awareness method based on data fusion comprises the steps of performing synchronous acquisition of an infrared image sequence and a point cloud sequence on a living body in a target area; extracting height change features and projection area change features of the point cloud sequence under the preset visual angle; according to the height change characteristics and the projection area change characteristics, whether abnormal behaviors of the human body occur or not is judged in real time, and the abnormal behaviors include a falling behavior, a crawling behavior, a continuous falling behavior or a rapid getting-up behavior; and when the continuous falling behavior is detected, executing verification operation according to the infrared image sequence. According to the invention, under the condition of complex environment interference, accurate identification of the human body situation and protection of privacy security are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and in particular relates to a multimodal human situation awareness method based on data fusion, a terminal device and a storage medium. Background Art

[0002] Human monitoring technology refers to a collection of technologies that use sensors, algorithms, and systems to perceive, analyze, and provide feedback on a person's physiological state, behavioral activities, or location information in real time or in non-real time. This enables intelligent monitoring of an individual's health, safety, or behavioral patterns.

[0003] Traditional human monitoring technology faces a dilemma due to its reliance on a single sensor. Thermal imaging solutions are susceptible to interference from ambient heat sources, while visual monitoring requires continuous high-definition image capture, which can easily expose private data in sensitive areas such as bathrooms and bedrooms. A new technical approach is needed to address these issues. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a multimodal human situation awareness method, terminal device and storage medium based on data fusion, which can solve the problem in related technologies that easily leads to the exposure of private data.

[0005] A first aspect of the present invention provides a multimodal human situation awareness method based on data fusion, comprising: Synchronously collect infrared image sequences and point cloud sequences of living bodies in the target area; Extracting height change features and projection area change features of the point cloud sequence under a preset viewing angle; According to the height change characteristics and the projection area change characteristics, it is determined in real time whether the human body has abnormal behavior, wherein the abnormal behavior includes falling, crawling, continuously falling or quickly getting up; When the continuous falling behavior is detected, a verification operation is performed according to the infrared image sequence.

[0006] Optionally, in a first implementation of the first aspect of the present invention, when the continuous falling behavior is detected, the step of performing a verification operation according to the infrared image sequence includes: When the continuous falling behavior is detected, the infrared image sequence is sent to the cloud, wherein the cloud returns a verification result in response to the infrared image sequence; If the verification result is received, it is determined based on the verification result whether the abnormal behavior is the continuous falling behavior to complete the verification operation.

[0007] Optionally, in a second implementation of the first aspect of the present invention, after receiving the verification result and determining whether the abnormal behavior is the continuous falling behavior based on the verification result to complete the verification operation, the method further includes: When the verification operation indicates that the abnormal behavior is the continuous falling behavior, a manual review process is triggered according to the infrared image sequence and the point cloud sequence.

[0008] Optionally, in a third implementation of the first aspect of the present invention, the step of determining in real time whether abnormal behavior occurs in the human body based on the height change characteristics and the projection area change characteristics includes: If it is detected that the height change feature drops rapidly to a low position within a short time window and the projection area change feature expands, it is determined to be the fall behavior; If it is detected that the height change feature is at a low position and the projection area change feature continues to move over time, it is determined to be the crawling behavior; If it is detected that the height change feature remains at a low level for more than a preset time and the projection area change feature remains stable, it is determined to be the continuous falling behavior; If it is detected that the height change feature rises rapidly from a low position within a short time window, it is determined to be the rapid standing behavior.

[0009] Optionally, in a fourth implementation of the first aspect of the present invention, after the step of determining in real time whether abnormal behavior occurs in the human body based on the height change characteristics and the projection area change characteristics, the method further includes: If it is determined to be crawling or quickly getting up, the local voice prompt will be executed or silence will be maintained.

[0010] Optionally, in a fifth implementation of the first aspect of the present invention, when the continuous falling behavior is detected, the step of performing a verification operation according to the infrared image sequence includes: Inputting the infrared image sequence into a pre-trained neural network model to obtain a human posture classification result output by the neural network model; The verification operation is performed according to the human posture classification result.

[0011] Optionally, in a sixth implementation of the first aspect of the present invention, the step of synchronously acquiring a point cloud sequence and an infrared image sequence of a living body includes: Synchronous acquisition of infrared image sequences and original point cloud sequences for living bodies; The original point cloud sequence is subjected to data preprocessing to obtain the point cloud sequence, wherein the data preprocessing includes dynamic noise filtering and static interference removal.

[0012] Optionally, in a seventh implementation of the first aspect of the present invention, the step of extracting height change features and projection area change features from the point cloud sequence at a preset viewing angle includes: Decomposing the point cloud sequence into temporal spatial data units; Extracting, based on the temporal spatial data unit, a first original feature representing a vertical spatial distribution of a human body and a second original feature representing a horizontal projection distribution of a human body; The first original feature and the second original feature are subjected to a time domain dynamic evolution analysis to obtain a height change feature reflecting a continuous change in the vertical dimension and a projection area change feature reflecting a continuous evolution of the projection morphology.

[0013] In the second aspect, an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the above-mentioned multimodal human situation awareness method based on data fusion are implemented.

[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned multimodal human situation awareness method based on data fusion.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the above-mentioned multimodal human situation awareness method based on data fusion.

[0016] The beneficial effects of the embodiments of the present invention compared to the prior art are as follows: by synchronously acquiring infrared image sequences and point cloud sequences and extracting key features in a targeted manner, the limitations of traditional single sensor technology are overcome. On the one hand, the non-visual characteristics of point cloud data effectively circumvent the reliance of visual monitoring on high-definition facial or body images, eliminating the serious risk of privacy leakage caused by continuous high-definition shooting in private spaces such as bathrooms and bedrooms; on the other hand, under the premise of mainly real-time point cloud detection, infrared images are triggered for verification only when there is a high suspicion of a high-risk anomaly. This not only takes advantage of the advantages of infrared imaging that it does not rely on visible light, has strong penetration, and naturally blurs facial details, but also reduces the frequency of environmental heat source interference on system judgment by limiting its active use range. The present invention achieves accurate recognition of human posture and protection of privacy security under the condition of complex environmental interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 Schematic diagram of an embodiment of a multimodal human situation awareness method based on data fusion in an embodiment of the present invention; Figure 2 1 is a schematic diagram of a specific embodiment of step S104 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention; Figure 3 1 is a schematic diagram of a specific embodiment of step S101 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention; Figure 4 1 is a schematic diagram of a specific embodiment of step S102 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention; Figure 5 Schematic diagram of a terminal device in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are protected by the present invention.

[0020] It should be noted that the terms "include", "comprising" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, terminal, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices. In the claims, specification and drawings of the present invention, relational terms such as "first" and "second" are merely used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such real-time relationship or order between these entities / operations / objects.

[0021] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0022] Human monitoring technology refers to a collection of technologies that use sensors, algorithms, and systems to perceive, analyze, and provide feedback on a person's physiological state, behavioral activities, or location information in real time or in non-real time. This enables intelligent monitoring of an individual's health, safety, or behavioral patterns.

[0023] Traditional human monitoring technology faces a dilemma due to its reliance on a single sensor. Thermal imaging solutions are susceptible to interference from ambient heat sources, while visual monitoring requires continuous high-definition image capture, which can easily expose private data in sensitive areas such as bathrooms and bedrooms. A new technical approach is needed to address these issues.

[0024] In view of this, the embodiments of the present invention provide a multimodal human posture perception method, terminal device and storage medium based on data fusion. By synchronously collecting infrared image sequences and point cloud sequences and extracting key features in a targeted manner, the limitations of traditional single sensor technology are overcome. On the one hand, the non-visual characteristics of point cloud data effectively avoid the dependence of visual monitoring on high-definition face or body images, eliminating the serious privacy leakage risks caused by continuous high-definition shooting in private spaces such as bathrooms and bedrooms. On the other hand, under the premise of real-time point cloud detection, infrared images are triggered for verification only when there is a high suspicion of high-risk anomalies. This not only takes advantage of the advantages of infrared imaging that it does not rely on visible light, has strong penetration and naturally blurs facial details, but also reduces the frequency of environmental heat source interference on system judgment by limiting its active use range. The present invention achieves accurate recognition of human posture and protection of privacy security under complex environmental interference.

[0025] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0026] Figure 1 The following is a schematic diagram of a multimodal human situation awareness method based on data fusion, provided by an embodiment of the present invention. The method can be applied to a terminal device, such as a mobile phone, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), or netbook.

[0027] Specifically, the multimodal human situation awareness method based on data fusion may include the following steps S101 to S103.

[0028] Step S101 : synchronously collecting infrared image sequences and point cloud sequences of a living body in a target area.

[0029] In an embodiment of the present invention, an infrared imaging end continuously generates an infrared image sequence reflecting the thermal radiation distribution of the human body; and a point cloud acquisition end (such as a millimeter-wave radar) synchronously outputs a point cloud sequence representing the three-dimensional spatial coordinates of the human body.

[0030] Align the timestamp of each frame of infrared image with the corresponding point cloud data.

[0031] Step S102 : extracting height variation features and projection area variation features from the point cloud sequence at a preset viewing angle.

[0032] In an embodiment of the present invention, the spatial distribution extreme values of the human body point cloud in the vertical direction are calculated in real time to generate a time series characteristic curve reflecting the changes in the height of the human body posture; the three-dimensional point cloud is projected onto a horizontal reference plane to quantify the continuous change trend of the human body's ground coverage area; the dynamic evolution relationship between the coupling height and the projected area is generated to generate a composite characteristic vector representing the spatial motion pattern of the human body.

[0033] Step S103: judging in real time whether the human body has abnormal behavior based on the height change characteristics and the projection area change characteristics, wherein the abnormal behavior includes falling, crawling, continuously falling or quickly getting up.

[0034] In an embodiment of the present invention, if it is detected that the height change feature drops rapidly to a low position within a short time window and the projection area change feature expands, it is determined as the falling behavior; if it is detected that the height change feature is at a low position and the projection area change feature continues to move over time, it is determined as the crawling behavior; if it is detected that the height change feature continues to be at a low position for more than a preset time and the projection area change feature remains stable, it is determined as the continuous falling behavior; if it is detected that the height change feature rises rapidly from a low position within a short time window, it is determined as the quick getting up behavior.

[0035] Step S104 : when the continuous falling behavior is detected, performing a verification operation according to the infrared image sequence.

[0036] In an embodiment of the present invention, a sequence of infrared images collected synchronously during the current period is extracted; and the posture of the human body is analyzed by infrared heat distribution characteristics to eliminate the possibility of misjudgment.

[0037] The beneficial effects of the embodiments of the present invention compared to the prior art are as follows: by synchronously acquiring infrared image sequences and point cloud sequences and extracting key features in a targeted manner, the limitations of traditional single sensor technology are overcome. On the one hand, the non-visual characteristics of point cloud data effectively circumvent the reliance of visual monitoring on high-definition facial or body images, eliminating the serious risk of privacy leakage caused by continuous high-definition shooting in private spaces such as bathrooms and bedrooms; on the other hand, under the premise of mainly real-time point cloud detection, infrared images are triggered for verification only when there is a high suspicion of a high-risk anomaly. This not only takes advantage of the advantages of infrared imaging that it does not rely on visible light, has strong penetration, and naturally blurs facial details, but also reduces the frequency of environmental heat source interference on system judgment by limiting its active use range. The present invention achieves accurate recognition of human posture and protection of privacy security under the condition of complex environmental interference.

[0038] Traditional human monitoring technology faces a fundamental conflict between on-device computing power and recognition accuracy. Local devices are unable to run complex behavior recognition models, resulting in a high rate of misjudgment of critical behaviors such as falls. Based on this, the present invention proposes an optional embodiment.

[0039] Reference Figure 2 , Figure 2 4 is a schematic diagram of a specific embodiment of step S104 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention. Step S104 also includes the following specific implementation methods.

[0040] Step S1041 : When the continuous falling behavior is detected, the infrared image sequence is sent to the cloud, wherein the cloud returns a verification result in response to the infrared image sequence.

[0041] In an embodiment of the present invention, when the local point cloud feature analysis module determines that a continuous fall has occurred (the height is continuously low and the projection area is stable), the cloud collaboration mechanism is activated, and the infrared image sequence collected synchronously during the current abnormal time period is bound to the timestamp and location information; the encapsulated data packet is uploaded to the cloud analysis platform in real time through an encrypted channel.

[0042] After responding to data requests, the cloud platform performs multi-stage verification, analyzes infrared image sequences, and restores the thermal distribution evolution in the time dimension; executes highly complex algorithms such as spatiotemporal thermal map analysis and posture key point tracking through cloud computing resources; and finally outputs binary verification results.

[0043] Step S1042: If the verification result is received, determine whether the abnormal behavior is the continuous falling behavior based on the verification result to complete the verification operation.

[0044] In an embodiment of the present invention, after receiving feedback from the cloud, the terminal device maps the verification result returned from the cloud to a local abnormal behavior classifier.

[0045] If the verification result is confirmed to be a continuous fall, the original abnormal judgment is maintained; if it is excluded, the current alarm process is terminated.

[0046] In the embodiment of the present invention, by introducing a cloud-based collaborative verification mechanism, computationally intensive infrared image analysis can be migrated to the cloud, freeing up local terminal computing power limitations.

[0047] Optionally, if the behavior is determined to be crawling or rapid standing up, a local voice prompt is performed or silence is maintained. Specifically, when outputting the result of crawling or rapid standing up, the current behavior attribute is marked as non-emergency abnormality; and context parameters such as time information and area type are simultaneously obtained. The feedback mechanism is dynamically selected based on the environmental policy library. If the device is configured for active reminder mode, a pre-recorded customized voice message will be played. If the device is in a privacy-sensitive area or at night, only the behavior log will be updated without triggering an audible or visual alarm.

[0048] Traditional fall monitoring systems are plagued by a predicament of widespread false alarms and a lack of accountability. Single sensors frequently trigger false alarms, causing caregivers to overlook real dangers. Based on this, the present invention proposes an alternative embodiment.

[0049] The following specific implementation is also included after step S1042.

[0050] Step S1043 : When the verification operation indicates that the abnormal behavior is the continuous falling behavior, a manual review process is triggered according to the infrared image sequence and the point cloud sequence.

[0051] In an embodiment of the present invention, when the cloud-based verification result confirms that the behavior is a continuous fall, the infrared image sequence and point cloud sequence corresponding to the abnormal period are extracted to generate a spatiotemporally aligned behavior evidence package; key information such as the time of occurrence, location coordinates, and risk level of the incident are marked to obtain a structured audit task.

[0052] Priorities are assigned based on the duration of the fall and changes in the intensity of human body heat radiation; and audit tasks are distributed to the medical duty system, security console or mobile terminal through the API interface.

[0053] In the embodiment of the present invention, a balance is achieved between life safety and false alarm tolerance by introducing a manual review and final determination mechanism.

[0054] The present invention proposes an alternative embodiment.

[0055] Step S104 also includes the following specific implementation methods.

[0056] Step S1044: input the infrared image sequence into a pre-trained neural network model to obtain a human posture classification result output by the neural network model.

[0057] In an embodiment of the present invention, when the local point cloud analysis module determines that the falling behavior is continuous, the local verification engine is started: Retrieve the infrared image sequence with matching timestamps and activate the lightweight neural network pre-installed on the terminal.

[0058] Step S1045: performing the verification operation according to the human posture classification result.

[0059] In an embodiment of the present invention, an infrared sequence is input into a neural network to extract the spatiotemporal evolution characteristics of the body's thermal radiation distribution; the model outputs a structured classification result; If the model outputs "falling on the back or curled up on the side", the continuous falling behavior is confirmed and a subsequent alarm is triggered; if the output is a non-falling state such as "sitting or kneeling", the current alarm process is terminated; Optionally, after the original infrared image processing is completed, it is fragmented and cleared, and only the classification log is retained.

[0060] In the embodiment of the present invention, the efficiency of behavior determination can be improved by deploying a lightweight neural network on the terminal side.

[0061] Traditional point cloud monitoring solutions suffer from the dilemma of noise drowning and target confusion. Based on this, the present invention proposes an optional embodiment.

[0062] Reference Figure 3 , Figure 3 4 is a schematic diagram of a specific embodiment of step S101 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention. Step S101 also includes the following specific implementation methods.

[0063] Step S1011 , synchronously collecting infrared image sequences and original point cloud sequences of the living body.

[0064] In an embodiment of the present invention, raw data alignment is achieved through hardware collaborative control. Specifically, the thermal imaging sensor outputs a low-resolution thermal radiation sequence at a fixed frame rate, while the millimeter-wave radar simultaneously collects a three-dimensional point cloud containing environmental noise.

[0065] Step S1012 : performing data preprocessing on the original point cloud sequence to obtain the point cloud sequence. The data preprocessing includes dynamic noise filtering and static interference removal.

[0066] In an embodiment of the present invention, instantaneous interference points are eliminated by time-domain bandpass filtering; fixed object point clouds are identified and removed based on a historical position database; and dynamic point cloud clusters that meet a human body size threshold are retained.

[0067] In the embodiment of the present invention, interference rejection of dynamic and static scenes can be achieved through point cloud preprocessing.

[0068] The height and area of a point cloud at a single moment are easily affected by limb occlusion, clothing material, or ground clutter, which can lead to the failure of instantaneous features. Based on this, the present invention proposes an optional embodiment.

[0069] Reference Figure 4 , Figure 4 4 is a schematic diagram of a specific embodiment of step S102 of the multimodal human situation awareness method based on data fusion in an embodiment of the present invention. Step S102 also includes the following specific implementation methods.

[0070] Step S1021: decompose the point cloud sequence into temporal spatial data units.

[0071] In an embodiment of the present invention, the point cloud sequence is segmented according to a fixed time window to form spatial data units with time stamps.

[0072] Step S1022: extracting a first original feature representing the vertical spatial distribution of the human body and a second original feature representing the horizontal projection distribution of the human body according to the temporal spatial data unit.

[0073] In an embodiment of the present invention, for each spatial data unit, the extreme value distribution of the point cloud in the Z-axis direction is calculated to generate a first original feature representing the degree of uprightness of the human body; the point cloud is orthogonally projected onto the XY plane, the area of the outer envelope rectangle is calculated, and a second original feature representing the area occupied by the human body is generated.

[0074] Step S1023 , performing a time domain dynamic evolution analysis on the first original feature and the second original feature to obtain a height change feature reflecting a continuous change in the vertical dimension and a projection area change feature reflecting a continuous evolution of the projection morphology.

[0075] In an embodiment of the present invention, polynomial fitting is performed on n consecutive first original features, and the derivative function value reflecting the vertical movement trend is extracted to obtain the height change feature; the sliding standard deviation calculation is performed on the second original feature to quantify the projection morphological stability and obtain the projection area change feature.

[0076] In an embodiment of the present invention, through time series decomposition and dynamic evolution analysis, the continuous motion trajectories of height and projection features can be separated, and the physical quantification decoupling and collaborative verification of human posture changes can be achieved, which significantly reduces the misjudgment rate of a single sensor.

[0077] like Figure 5FIG2 is a schematic diagram of a terminal device according to an embodiment of the present invention. The terminal device 500 may include a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as a multimodal human situation awareness program based on data fusion. When the processor 501 executes the computer program 503, the steps described in the aforementioned multimodal human situation awareness embodiments based on data fusion are implemented.

[0078] The computer program can be divided into one or more modules / units, which are stored in the memory 502 and executed by the processor 501 to implement the present invention. One or more modules / units can be a series of computer program instruction segments that can perform specific functions. These instruction segments are used to describe the execution process of the computer program in the terminal device.

[0079] The terminal device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art will appreciate that Figure 5 It is only an example of a terminal device and does not constitute a limitation of the terminal device. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device may also include input and output devices, network access devices, buses, etc.

[0080] The processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0081] Memory 502 can be an internal storage unit of the terminal device, such as the terminal device's hard drive or memory. Memory 502 can also be an external storage device of the terminal device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 502 can include both the terminal device's internal storage unit and an external storage device. Memory 502 is used to store computer programs and other programs and data required by the terminal device. Memory 502 can also be used to temporarily store data that has been output or is about to be output.

[0082] It should be noted that, for the convenience and brevity of description, the structure of the above-mentioned terminal device can also refer to the specific description of the structure in the method embodiment, which will not be repeated here.

[0083] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the multimodal human situation awareness method based on data fusion can be implemented.

[0084] An embodiment of the present invention provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the multimodal human situation awareness method based on data fusion when executing the computer program product.

[0085] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0086] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0087] In the embodiments provided herein, it should be understood that the disclosed terminal devices and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection via some interface, device, or unit, which may be electrical, mechanical, or other means.

[0088] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0089] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0090] If the integrated module / unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0091] The above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may be modified or some of the technical features thereof may be replaced with equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention and are therefore intended to be included within the scope of protection of the present invention.

Claims

1. A multimodal human situation awareness method based on data fusion, characterized in that: include: Synchronously collect infrared image sequences and point cloud sequences of living bodies in the target area; Extracting height change features and projection area change features of the point cloud sequence under a preset viewing angle; According to the height change characteristics and the projection area change characteristics, it is determined in real time whether the human body has abnormal behavior, wherein the abnormal behavior includes falling, crawling, continuously falling or quickly getting up; When the continuous falling behavior is detected, a verification operation is performed according to the infrared image sequence.

2. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: When the continuous falling behavior is detected, the step of performing a verification operation according to the infrared image sequence includes: When the continuous falling behavior is detected, the infrared image sequence is sent to the cloud, wherein the cloud returns a verification result in response to the infrared image sequence; If the verification result is received, it is determined based on the verification result whether the abnormal behavior is the continuous falling behavior to complete the verification operation.

3. The multimodal human situation awareness method based on data fusion according to claim 2, characterized in that: If the verification result is received, determining whether the abnormal behavior is the continuous falling behavior based on the verification result to complete the verification operation, the method further includes: When the verification operation indicates that the abnormal behavior is the continuous falling behavior, a manual review process is triggered according to the infrared image sequence and the point cloud sequence.

4. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: The step of determining in real time whether abnormal behavior of the human body occurs based on the height change characteristics and the projection area change characteristics comprises: If it is detected that the height change feature drops rapidly to a low position within a short time window and the projection area change feature expands, it is determined to be the fall behavior; If it is detected that the height change feature is at a low position and the projection area change feature continues to move over time, it is determined to be the crawling behavior; If it is detected that the height change feature remains at a low level for more than a preset time and the projection area change feature remains stable, it is determined to be the continuous falling behavior; If it is detected that the height change feature rises rapidly from a low position within a short time window, it is determined to be the rapid standing behavior.

5. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: After the step of determining in real time whether abnormal behavior occurs in the human body based on the height change characteristics and the projection area change characteristics, the method further includes: If it is determined to be crawling or quickly getting up, the local voice prompt will be executed or silence will be maintained.

6. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: When the continuous falling behavior is detected, the step of performing a verification operation according to the infrared image sequence includes: Inputting the infrared image sequence into a pre-trained neural network model to obtain a human posture classification result output by the neural network model; The verification operation is performed according to the human posture classification result.

7. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: The step of synchronously collecting a point cloud sequence and an infrared image sequence for a living body comprises: Synchronous acquisition of infrared image sequences and original point cloud sequences for living bodies; The original point cloud sequence is subjected to data preprocessing to obtain the point cloud sequence, wherein the data preprocessing includes dynamic noise filtering and static interference removal.

8. The multimodal human situation awareness method based on data fusion according to claim 1, characterized in that: The step of extracting height change features and projection area change features from the point cloud sequence at a preset viewing angle includes: Decomposing the point cloud sequence into temporal spatial data units; Extracting, based on the temporal spatial data unit, a first original feature representing a vertical spatial distribution of a human body and a second original feature representing a horizontal projection distribution of a human body; The first original feature and the second original feature are subjected to a time domain dynamic evolution analysis to obtain a height change feature reflecting a continuous change in the vertical dimension and a projection area change feature reflecting a continuous evolution of the projection morphology.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the multimodal human situation awareness method based on data fusion as claimed in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multimodal human situation awareness method based on data fusion as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Non-contact indoor personnel falling identification method based on Internet of Things platform

    CN112435440A

  • Pedestrian tumble detection method based on laser radar

    CN115331302A

  • Behavior recognition method, device, equipment, storage medium and computer program product

    CN119760475A

  • Multi-mode tumble detection method, device and equipment

    CN120220230A

  • Self-adaptive human body posture detection method based on privacy protection camera and radar

    CN120257179A