Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

138 results about "Visual modality" patented technology

Visual Modality - A Visual Learner. Learns by seeing and by watching demonstrations Likes visual stimuli such as pictures, slides, graphs, demonstrations, etc. Conjures up the image of a form by seeing it in the “mind’s eye” Often has a vivid imagination.

Double-branch diffusion three-dimensional scene generation method based on multi-modal semantic graph

The invention belongs to the technical field of three-dimensional scene modeling, and discloses a dual-branch diffusion three-dimensional scene generation method based on a multi-modal semantic graph, which comprises the following steps of: firstly, receiving multi-modal data such as sketches, texts, automatic completion instructions and scene general knowledge, extracting features and fusing the features into a unified multi-modal semantic graph; utilizing a graph neural network and an attention mechanism to enhance semantic graph features, and optimizing physical constraints through a physical engine; complementing the missing visual modality and graph structure relationship; performing quality scoring on the scene based on semantics, spatial relationships and physical constraints; and finally, respectively generating a spatial layout and a geometric shape through a double-branch diffusion model, and ensuring the coordination of the layout and the shape. The method has the advantages of multi-modal information fusion, physical rationality guarantee, high structure complementation capability, high-quality score optimization and efficient generation process, and is suitable for three-dimensional scene modeling requirements in the fields of virtual reality, augmented reality, robots and the like.
Owner:CHINA JILIANG UNIV

Multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning

The invention relates to the technical field of natural language processing, in particular to a multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning. According to the method, learnable residual scaling and pre-normalization are introduced through a cross-modal encoder, deep interaction of texts, voices and visual modals is stabilized, and gradient explosion is inhibited; a dialogue graph fusing semantic similarity and time proximity is constructed online through a semantic-time sequence graph enhancement module, and a graph attention network is used for explicitly modeling long-distance round dependence and cross-speaker emotion transmission; through an adaptive comparison and alignment module, a dynamic scheduling comparison loss and index moving average updated mode-emotion prototype library is adopted to realize data distribution adaptive cross-mode alignment; through cooperative work of the modules, the problems that in the prior art, cross-modal fusion is unstable, long-distance and cross-speaker dependence modeling is insufficient, and cross-dataset alignment capacity is weak are solved.
Owner:CHONGQING TELECOMM PLAN & DESIGN INST

Method and system for evaluating infusion operation of intern medical staff

The invention discloses a method and system for evaluating infusion operation of intern medical staff, and belongs to the field of infusion, and the method comprises the steps that the operation of the intern medical staff is collected in multiple modes, and the multiple modes comprise a visual mode, a touch mode and a physiological signal mode; inputting the collected operation of the intern medical staff into the multi-modal infusion evaluation model, performing feature extraction by the multi-modal infusion evaluation model in a modal manner, and performing cross-modal attention fusion on feature vectors of different modals; and finally, comparing the calculated total feature tensor Hfusion with a standard total feature tensor Hfusion ', and outputting a final score of the practice medical personnel according to a preset scoring standard. According to the scheme, multi-dimensional data such as visual, tactile and physiological signals are synchronously collected through a multi-modal fusion method, operation details are comprehensively covered, comprehensive evaluation is carried out, and evaluation objectivity is improved.
Owner:SICHUAN HEALTH REHABILITATION VOCATIONAL COLLEGE +1

Textile cloth defect real-time detection method and system based on multi-modal feature fusion

The invention provides a textile cloth defect real-time detection method and system based on multi-modal feature fusion, and the method comprises the steps: collecting visual image data, infrared thermal imaging data and ultrasonic acoustic data of textile cloth through a multi-sensor array, and forming multi-modal input; performing time sequence alignment and noise filtering preprocessing on the multi-modal data to eliminate motion artifacts and environmental interference; a parallel feature extraction module is used for extracting texture features from the visual data, extracting temperature distribution features from the thermal imaging data and extracting acoustic impedance features from the acoustic data. By deeply fusing complementary information of three modes of vision, thermal imaging and ultrasonic wave, the detection capability is improved, the vision mode captures surface texture details, the thermal imaging mode reveals thermodynamic anomalies related to friction and materials, the ultrasonic wave mode perceives subcutaneous structure defects, hidden flaws which cannot be recognized by a single mode can be found, and the detection efficiency is improved. Therefore, the omission ratio is greatly reduced, and flaw types are distinguished more accurately.
Owner:NANCHANG ZHONGTUO KNITWEAR CORP LTD

Multi-modal fusion man-machine interaction control method, system and equipment and storage medium

The embodiment of the invention provides a multi-mode fusion man-machine interaction control method, system and device and a storage medium, and relates to the technical field of intelligent driving, and the method comprises the steps: obtaining vehicle driving data and driver state data; based on the vehicle driving data and the driver state data, calculating fusion weights of a visual mode, an auditory mode and a tactile mode to obtain a multi-mode fusion weight matrix; generating a multi-modal interaction signal according to the multi-modal fusion weight matrix, wherein the multi-modal interaction signal comprises a visual signal, an auditory signal and a tactile signal; and triggering corresponding visual warning, auditory warning and tactile warning according to the multi-mode interaction signal. In this way, multi-mode fusion is conducted on vision, hearing and touch, the output intensity of visual warning, hearing warning and touch warning is adjusted according to the weight of each mode of vision, hearing and touch, a driver is helped to make the most urgent operation at present according to the warning of each mode, and the driving safety of man-machine interaction in the driving process is improved.
Owner:FAW HAIMA AUTOMOBILE CO LTD +1

Video segmentation method, server, storage medium, and program product

The present application provides a video segmentation method, a server, a storage medium, and a program product. In the method of the present application, video data to be segmented is segmented into multiple data segments, unimodal features of the data segments, including text features of a text modality and visual features of a visual modality, are respectively extracted by means of a video topics segmentation model, and then the text features and visual features of the data segments are fused, so that the fusion of multimodal information can be performed at the intermediate representation level, the relationship and interaction between different modalities can be better captured, and higher-quality multimodal fusion features of the data segments are obtained. Furthermore, on the basis of the multimodal fusion features of the data segments, whether the data segments are topic boundaries is predicted, so that the topic boundaries of the video data can be accurately predicted, improving the accuracy of topic boundary recognition, thereby improving the accuracy and quality of video topics segmentation results.
Owner:ALIBABA (CHINA) CO LTD

Graph convolution multi-modal dialogue emotion recognition method based on emotion dimension compensation

The invention relates to a graph convolution multi-modal dialogue emotion recognition method based on emotion dimension compensation. The method belongs to the field of multi-modal emotion recognition. Comprising the following steps: respectively extracting original features of texts, audios and visual modalities by utilizing a pre-training model; detecting a mode missing condition, and generating a pseudo feature by using an available mode feature and context information; the multi-modal features are mapped to a VAD three-dimensional space, and prediction and enhancement are carried out; the global level is based on VAD similarity to connect cross-utterance nodes, and emotional consistency is quantified to capture long-distance interaction. In the local level, information spreading and node updating are carried out by calculating the similarity between different modal features in the same utterance; neighbor information is accumulated through multiple layers of image volumes, node weights are dynamically adjusted to suppress noise propagation, and finally utterance-level emotional representation is generated by fusing multi-modal features through mean pooling operation. According to the method, the emotion recognition performance in a modal missing scene is remarkably improved.
Owner:KUNMING UNIV OF SCI & TECH

Multi-modal sentiment analysis model based on multi-granularity features and adaptive fusion

The invention discloses a multi-modal sentiment analysis model based on hierarchical adaptive cross-modal fusion, belongs to the field of natural language processing, and is used for solving the problems that in the prior art, multi-granularity sentiment feature extraction is insufficient, a cross-modal fusion mechanism is rigid, and the distribution difference between different-source modals is large. The method comprises the following steps: firstly, extracting features of texts, audios and visual modalities from original video data, and coding the features into advanced semantic features; secondly, multi-granularity information is fused through a hierarchical feature extractor to generate enhanced single-mode features; then, a self-adaptive cross-modal fusion network with a text as a core is adopted to realize bidirectional interaction and dynamic weighted fusion between modals; further, a dynamic contrast learning mechanism is introduced to align modal distribution in a unified potential space; and finally, inputting the optimized multi-modal features into a classifier and outputting an emotion analysis result. According to the model, through collaborative optimization of multi-granularity feature extraction, adaptive fusion and comparative learning, the accuracy of sentiment analysis and the robustness of the model are remarkably improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Pet behavior prediction method and device, equipment and storage medium

The invention provides a pet behavior prediction method and apparatus, a device and a storage medium. The method comprises the steps of obtaining a pet type, image data, audio data and text data describing pet behaviors; performing feature extraction and multi-modal feature fusion on the image data, the audio data and the text data to obtain fusion features of the pet; and based on the type of the pet and the fusion feature, predicting the behavior of the pet by adopting a pre-constructed pet behavior prediction large model to obtain a behavior prediction result of the pet. When pet behaviors are predicted, visual modal data are introduced, auditory features of an audio modal and situational semantics described by an owner text are synchronously fused, behavior ambiguity of single visual data is eliminated through a cross-modal dynamic weighting mechanism, emotion clues in sound signals and recessive states described by the text are captured, and the pet behaviors are predicted. Therefore, a multi-dimensional cognitive model of pet behaviors is comprehensively constructed, and the prediction accuracy is remarkably improved.
Owner:HANGZHOU QIANWAN TECH CO LTD

Multimodal personality perception method and device based on multiple correlation features and graph relationship attention

The present invention relates to the field of computer vision, and more particularly to a multimodal personality perception method and apparatus that utilizes multiple association features and graph relationship attention. The method comprises: obtaining an input video for personality perception; inputting the input video into a data preprocessing module to obtain visual modality input, audio modality input, and text modality input; inputting the visual modality input, audio modality input, and text modality input into a modal feature extraction network module to obtain scene-audio association features, scene-description word association features, audio-description word association features, and text modality features; inputting the scene-audio association features, scene-description word association features, audio-description word association features, and text modality features into a feature fusion module to obtain multimodal fusion features; and inputting the multimodal fusion features into a perception prediction module to obtain personality perception results. The present invention proposes a multimodal attention fusion framework for personality perception.
Owner:UNIV OF SCI & TECH BEIJING

Label-assisted report generation method and device

The invention relates to a method and a device for generating a report under the assistance of a label, and the method comprises the following steps: 1) extracting a structured label set from a text report of a sample based on a large language model, wherein the structured label set comprises a multi-classification group consisting of dichotomous labels and mutual exclusion options; 2) aggregating the labels in batches, after a threshold value is reached, merging and de-weighting, performing specification and mutual exclusion group merging on synonymous, near-synonymous and redundant labels, and converging into a unified label library; 3) based on the text report and the tag library, outputting a tag subset of each sample through a large language model; 4) multi-modal multi-label classification model training: extracting each visual modal feature, performing weighted aggregation and splicing, and outputting each label group logits through a classification head to perform weighted group loss optimization; (5) carrying out joint training on the multi-modal large language model by using samples of'only images-reports' and'images + labels-reports', and (6) carrying out label prediction and screening on the images by using the classification model, and inputting'images + prediction labels' into the multi-modal large language model to obtain a final report.
Owner:ZHEJIANG UNIV

Parkinson's disease early recognition system and method based on multi-task learning

ActiveCN120913884AMedical data miningBiological modelsData setSymptom perception
The invention provides a Parkinson's disease early recognition system and method based on multi-task learning, and belongs to the technical field of machine learning. The method comprises the following steps: firstly, constructing a Parkinson's disease discrimination data set covering actions, languages and visual modalities, and providing comprehensive symptom data and identities for subsequent analysis; then, in combination with a medical priori rule, constructing a symptom recognition module integrated with a plurality of traditional machine learning models, and performing preliminary judgment on action and language data; a short-term recognition model is constructed to enhance the multi-modal symptom perception ability and improve the discrimination effect of Parkinson's disease short-term recognition; and finally, carrying out mathematical modeling analysis based on the recognition result sequence of the multiple time periods, fusing statistical trend and group difference, and realizing robust comprehensive judgment on the Parkinson's disease. According to the invention, through cooperative training and knowledge sharing among tasks, the comprehensive discrimination capability of the model on the multi-modal pathological signals of the Parkinson's disease is enhanced, so that more accurate and more sensitive early recognition is realized.
Owner:YANTAI LIAN BIOTECHNOLOGY CO LTD

Adversarial camouflage and decoy generation across visual and non-visual modalities

Disclosed is a system to generate an adversarial pattern. The system receives an input specifying data describing a target object, an objective indicating whether to reduce detectability of the target object or to induce a false detection of a target type, and context parameters. The system generates candidate adversarial patterns for the target object using an AI-based generative algorithm. Each candidate adversarial pattern represents a potential solution for achieving the objective under context parameters. The system simulates a representation of the target object with candidate adversarial pattern in a virtual environment that models a plurality of sensor modalities. The simulation may use machine learning object detection models. The system analyzes the target object and generates, for each target object, a performance metric for ranking the candidate adversarial pattern based on an effectiveness of the simulated representation. The candidate adversarial pattern with highest rank is generated as an optimized adversarial pattern.
Owner:THE ADVERSARIAL CO INC

System

A system is provided.SOLUTION: A system comprising: means for storing information; means for extracting key claims from the information using generative artificial intelligence to analyze the information; means for quantifying the key claims; and means for visually displaying the quantified key claims.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

A hand pose reconstruction method based on multi-stage joint enhancement

The present invention discloses a method for reconstructing hand posture based on multi-stage joint point enhancement. The method of the present invention proposes a multi-stage joint enhancement network, the first stage of which consists of a network that only trains the fingertip joint positions and the base of palm joint positions; the second stage combines the input image features obtained from the feature extractor and the 12 joint points output by the first stage, and performs joint reasoning through the Transformer module and the features extracted from the image to enhance the accuracy of the prediction of the remaining 30 joint points. The present invention combines visual modalities and geometric modalities, and improves the model's prediction ability for the remaining 30 joint points through multimodal fusion, thereby enhancing the accuracy of joint point prediction.
Owner:HANGZHOU DIANZI UNIV

Electronic vaping device

An electronic vaping device may be designed to enhance or facilitate its use. For example, the electronic vaping device may: allow a capability of the electronic vaping device to provide vapor to be altered (e.g., disabled, reduced, enabled, or increased) in some situations (e.g., to prevent unauthorized vaping by a child, teenager or other individual); be able to communicate with an external communication device (e.g., a smartphone, a computer, etc.) to convey a notification of potential unauthorized use of the electronic vaping device (e.g., by a child, teenager or other unauthorized user); implement a physical deterrent to its unauthorized use; be able to visually convey information (e.g., advertisements, notifications, etc.); and / or be able to capture images and / or sounds (e.g., record pictures and / or video, speech, music, etc.).
Owner:RAI STRATEGIC HOLDINGS INC

Systems, devices and methods for line guidance of product application

Systems, Devices, and Methods for Line-by-Line Guidance of Product Application: The cosmetic deposition device includes an applicator component, a position sensor, and a reservoir for cosmetic styling compositions. A display of the cosmetic deposition device represents the skin area as a guide segment and the applicator component as a visual indicator relative to the guide segment, based on the position sensor. The visual indicator responds to changes in the applicator component's position relative to the skin area to visually guide the user to accurately apply the cosmetic style. A feedback device alerts the user if the application deviates from the skin area. These approaches allow the user to correctly apply the composition in real time. Figure for abstract: None
Owner:LOREAL SA

Augmented reality-based smart socket visualization

Methods and systems for visualizing and controlling smart electrical sockets are disclosed. A mobile device with a camera captures an image of a smart electrical socket and displays the image on its screen. The mobile device determines a socket's unique identifier and receives electrical parameters associated with power delivered through the socket's plug receptacle. These parameters are displayed concurrently with the socket image, visually associated such as through superimposition or connecting lines. The smart socket includes sensors for monitoring electrical parameters, a controller for power switching, and a wireless interface. The mobile device can receive parameters directly from the socket, via a gateway device or from a server connected to the gateway device. The system enables pairing between smart sockets and gateway devices using visual codes. Control of the smart socket is achieved through user gestures on the mobile device interface, allowing power switching and adjustment of power levels.
Owner:HONEYWELL INTERNATIONAL INC

Information presenting device, information presenting method, and program

When diminishing an area Er to be diminished or an object Et to be diminished in an image (video) Prm captured by an imaging unit 11 and displayed to a display unit 16 by applying diminishing processing thereto, the present invention visually presents attention attraction information at for causing an inattentional overlook phenomenon to a worker (user) HW by showing the same in an attention attraction information display area Dr on the Prm, or (and) auditorily presents the same by outputting from a sound output unit 17. Then, after the attention attraction information at is presented, the object Er to be diminished or the object Et to be diminished in the image (video) Prm is diminished by applying diminishing processing (Fe).
Owner:NT T INC

Method for neurofeedback training to output brain-region reality

A method for neurofeedback training to output brain-area reality is disclosed. The method includes transmitting a physical and mental parameter related to a subject as a neurophysiological signal; performing signal processing, feature extraction and pattern determination on the neurophysiological signal; providing a neurophysiological feedback parameter and conducting a brain region network activity; and converting a brain-area reality through a brain-computer interface to present an interactive scene and an interactive element to the subject for brain / brain-area (an Electroencephalography (EGG) brain waves and / or brain network) training. In this way, the subject's brain area training status can be known in real time and the subject can understand the state of his own brain area through visual means, so as to facilitate communication between subjects (or their family members or related persons) and professionals (such as doctors).
Owner:EXEBRAIN CO LTD

A multi-modal entity linking method based on double encoders and hybrid expert mechanism

A multimodal entity linking method based on dual encoders and a hybrid expert mechanism is proposed. This invention relates to multimodal entity linking technology at the intersection of natural language processing and computer vision. Addressing the problems of low inference efficiency, insufficient cross-modal interaction, and shallow modal fusion in existing methods, this invention proposes a multimodal entity linking method based on dual encoders and a hybrid expert mechanism. A dual-tower architecture is used to independently encode mentions and entities. Entity embeddings can be pre-computed offline and indexed, and linking is completed during inference through fast vector retrieval. A hybrid expert mechanism is introduced to achieve adaptive feature transformation of samples, and a gating network dynamically selects expert combinations. Bidirectional cross-modal attention is used to establish fine-grained alignment at the word-image block granularity. A channel attention mechanism dynamically balances the contributions of textual and visual modalities. The model is jointly optimized by multiple constraints, including load balancing loss. This invention achieves efficient retrieval while maintaining deep inference capabilities, simplifies inference time complexity, and is suitable for scenarios such as knowledge graph construction and intelligent question answering systems.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

A parkinson's disease early identification system and method based on multi-task learning

The application provides a Parkinson disease early identification system and method based on multi-task learning, and belongs to the technical field of machine learning. First, a Parkinson disease discrimination data set covering action, language and visual modalities is constructed to provide comprehensive symptom data and identity for subsequent analysis. Then, in combination with medical prior rules, a symptom recognition module integrating multiple traditional machine learning models is constructed to preliminarily discriminate the action and language data. A short-term recognition model is constructed to enhance the multi-modal disease perception ability and improve the discrimination effect of short-term Parkinson disease recognition. Finally, based on the recognition result sequence of multiple time periods, mathematical modeling analysis is carried out to fuse statistical trends and group differences, and to realize stable comprehensive judgment of Parkinson disease. Through the collaborative training and knowledge sharing among tasks, the model can enhance the comprehensive discrimination ability of multi-modal pathological signals of Parkinson disease, so as to realize more accurate and sensitive early identification.
Owner:YANTAI LIAN BIOTECHNOLOGY CO LTD

A multi-modal sentiment analysis method and system based on semantic guidance denoising

PendingCN122758328AData ingestionThresholding
The application provides a multi-modal sentiment analysis method and system based on semantic guidance denoising, and belongs to the field of multi-modal intelligent perception and sentiment analysis. The method comprises the following steps: acquiring multi-modal video data, and extracting initial feature representations of text, speech and visual modalities; constructing semantic guidance information based on the initial feature representation of the text modality, performing multi-scale time sequence modeling on the initial feature representations of the speech and visual modalities, and performing adaptive residual enhancement and fusion based on the semantic guidance information to obtain enhanced non-text modality representations; generating a dynamic denoising threshold by combining the semantic guidance information and the enhanced non-text modality representations, and performing soft threshold denoising on the enhanced non-text modality representations; mapping the denoised non-text modality representations and the initial feature representations of the text modality to perform multi-modal fusion; and inputting the multi-modal fusion representations into a regression prediction layer to output a sentiment prediction result. The application effectively improves the accuracy and robustness of multi-modal sentiment analysis.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

A multi-modal interactive sentiment analysis method based on semi-supervised learning

The application discloses a multimodal interactive sentiment analysis method based on semi-supervised learning. The method first extracts features of voice, text and visual modalities, then fuses the features through two interactive branch networks of intra-modal and inter-modal, and finally adaptively extracts fusion information of the two interactive branches based on a gating mechanism. The intra-modal and inter-modal interactive networks are based on the encoder structure of the transformer, except that the multi-head self-attention layer introduces the proposed attention mask mechanism. In terms of training method, the method generates pseudo-labels to assist model training in a semi-supervised learning manner, thereby reducing the dependence on labeled data. The application solves the problems of high artificial labeling cost and how to mine effective interactive information in the field of multimodal sentiment analysis, and improves the sentiment recognition accuracy.
Owner:SOUTH CHINA UNIV OF TECH

Video synthesis via multimodal conditioning

A multimodal video generation framework (MMVID) that benefits from text and images provided jointly or separately as input. Quantized representations of videos are utilized with a bidirectional transformer with multiple modalities as inputs to predict a discrete video representation. A new video token trained with self-learning and an improved mask-prediction algorithm for sampling video tokens is used to improve video quality and consistency. Text augmentation is utilized to improve the robustness of the textual representation and diversity of generated videos. The framework incorporates various visual modalities, such as segmentation masks, drawings, and partially occluded images. In addition, the MMVID extracts visual information as suggested by a textual prompt.
Owner:SNAP INC

system

We provide the system. [Solution] A means of analyzing electronic communications and determining priorities, Means for generating notifications based on the aforementioned priority, A means of automatically generating replies according to criteria specified by the user, A means for sending the automatically generated reply, An information display device provides means for providing notifications by sound and visual means, A system that includes this.
Owner:SOFTBANK GROUP CORP

Video generation method for generating video showing air permeability of constituent element of wearable article

Provided is a video generation method for generating a video showing the air permeability of a constituent element of a wearable article. The excellent air permeability of the components of the wearable article is visually shown so that the wearable article can be easily recognized by a user. The method comprises the following steps: a mounting step, a part of a constituent element of a wearable article (for example, a waist strip element having improved air permeability when stretched) is attached to a water vapor discharge device for discharging predetermined visually recognizable water vapor in a stretched state (for example, in a range from 60% of the maximum elongation to the maximum elongation) in which the actual use state is reproduced; and an imaging step for imaging, in a specific discharge direction (for example, + / -20 degrees in the vertical direction or 45 degrees in the horizontal-oblique upper direction), a case in which the water vapor discharged from the water vapor discharge device passes through the constituent element (for example, a state in which the water vapor rises linearly or rises upward by 20 cm or more).
Owner:UNI CHARM CORP

Zero-shot video classification method based on test time visual proxy tuning

The application provides a zero-shot video classification method based on test time visual proxy tuning, which realizes zero-shot action recognition of video actions by constructing a visual proxy using a support set and simultaneously fine-tuning visual and text prompts, and avoids the semantic gap problem between the two modalities of video and text in the video classification task. The visual proxy construction module and the dual-modal prompt collaborative tuning module constitute a zero-shot learning framework TPC. In the visual proxy construction module, a pre-trained video encoder is used to extract support set video features to construct a visual proxy, and a learnable visual prompt is added to the sampled support set video to make the visual proxy adjustable. In the dual-modal prompt collaborative tuning module, the learnable visual and text prompts are fine-tuned by minimizing the KL divergence between the prediction probability distribution of the visual proxy and the text proxy, the information of the text and visual modalities is used to optimize the visual proxy, and the zero-shot classification performance of the visual proxy is improved.
Owner:NANJING UNIV OF SCI & TECH

College student behavior anomaly early warning method and system based on multi-source data fusion

PendingCN122332841AImage manipulationNoise
This invention discloses a method and system for early warning of abnormal behavior among college students based on multi-source data fusion, relating to the technical field of image processing. The method involves real-time acquisition of video streams of individuals within a target area, extraction of behavioral features from the video streams for classification, and a determination list based on these behavioral features. Multi-source data is collected, and anomaly coefficients are calculated based on the multi-source data. If the anomaly coefficient exceeds a threshold, the target individual is identified as exhibiting abnormal behavior. Anomaly warnings are then issued based on the behavioral features corresponding to the abnormal behavior. By calculating the anomaly coefficient through multi-source data fusion, a preliminary screening of hidden risks is achieved, reducing the detection delay and computational redundancy caused by traditional full-process monitoring analysis. Furthermore, real-time video stream acquisition and cloud-based behavioral feature extraction are initiated for high-risk individuals, avoiding the limitations of audio noise interference and single visual modalities, and improving the robustness of identification and the real-time performance of warnings in complex monitoring scenarios.
Owner:ZHEJIANG GUGEL INTELLIGENT TECHNOLOGY CO LTD