Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

115 results about "Temporal context" patented technology

What is Temporal Context. 1. Temporal information which may impact interpretation of a study. At different points in time, different browsing environments and activities emerge and become part of users’ experiences. Temporal factors which can be reported include the date of the study and duration of the study.

Multi-modal news recommendation method and system in combination with clock interests

The invention discloses a multi-modal news recommendation method and system in combination with clock interests, and relates to the technical field of data mining and recommendation methods.The method comprises the steps that a user behavior sequence, a candidate news set and a related timestamp set are extracted; obtaining the multi-modal coding representation of the historical interaction news through the time perception multi-modal feature coding; performing clock interest modeling based on multi-modal coding representation of historical interactive news, and fusing long and short-term interests through Gaussian weighted aggregation and long-term interest enhancement operation to obtain a user interest vector; and calculating the matching degree of the candidate news and the current user interest vector, applying time period sensitive suppression through time sequence gating, performing click probability prediction on the output of the time sequence gating by adopting a linear network, and performing a dynamic recommendation decision. According to the method, by means of the feature fusion technology of time context perception, the interest continuity of the hour granularity is captured, and the recommended content is more accurately matched with the requirements of the user in different time periods.
Owner:SOUTH CHINA AGRICULTURAL UNIVERSITY

Weakly supervised group behavior identification method based on dynamic prompt tuning

The invention discloses a weak supervision group behavior recognition method based on dynamic prompt tuning, and belongs to the field of video understanding. According to the method, a pre-trained vision-language model is adaptively expanded to a group behavior recognition task for challenges such as complex group behavior semantics, strong spatio-temporal context dependency and lack of individual labels in a video. According to the technical scheme, a dynamic prompt generation technology with visual conditions is provided, instance-level text prompts can be automatically generated according to input video content, the accuracy of vision-text semantic alignment is enhanced, meanwhile, a time sequence fusion module is integrated in a model, and the purpose is to fuse key time information in multiple frames through modeling inter-frame time sequence dependence. The model calculates a similarity score between visual and text features and takes the similarity score as a classification prediction basis, and finally, a cross entropy loss function is adopted to carry out end-to-end training on the network. The validity of the method is verified on a volleyball data set and an NBA data set.
Owner:BEIJING UNIV OF TECH

Sequence recommendation method based on time-aware hierarchical attention network

The invention discloses a sequence recommendation method based on a time-aware hierarchical attention network, and the method comprises the steps: obtaining a user interaction sequence of a target user, and obtaining a time interval sequence and a time context sequence of the interaction of the target user based on the user interaction sequence; inputting the user interaction sequence, the time interval sequence and the time context sequence into a trained sequence recommendation model to obtain a user interest representation based on a time interval and a user interest representation fused with time context information; and determining a correlation score with the user interaction sequence according to the user interest representation, obtaining a final correlation score through weighted fusion, determining a recommended item list according to the final correlation score, and recommending the recommended item list to a target user. According to the method, the problem of insufficient noise interference and context fusion in time information modeling of a traditional method is solved.
Owner:CHONGQING UNIV OF TECH

Temporal context from eye tracking for generative ai

Examples relate to systems and methods for enhancing generative AI outputs using eye tracking data. An eye tracking system accesses eye gaze information associated with a field of view of a head-wearable apparatus and generates contextual information associated with the field of view of the head-wearable apparatus based on the eye gaze information. The eye tracking system processes, by a generative machine learning model, the contextual information and at least one image of the field of view of the head-wearable apparatus to generate an output and presents on a display of the head-wearable apparatus the output generated by the generative machine learning model.
Owner:SNAP INC

Retrieval enhancement generation method and data set generation method for time-sensitive problems

The invention discloses a retrieval enhancement generation method for a time-sensitive problem, which comprises the following steps of: mixed time perception retrieval: enhancing document retrieval by adding time constraint on the basis of semantic relevance, and guiding by a time card to ensure that the retrieved document not only conforms to the meaning of query, but also conforms to the semantic relevance; the time context is met; the progressive multi-step reflection comprises the following steps of: firstly, acquiring and evaluating an initial document set by applying mixed time perception retrieval; if a document is retrieved, generating a final answer by using a large language model; otherwise, entering a reflection stage, and summarizing useful time information in the retrieved document into a context; and merging document sets accumulated in all iterations to generate a final answer. According to the method, a new framework integrating dynamic knowledge updating and time reasoning into the retrieval and generation process is provided, and accurate and timely response can be made to time-related problems.
Owner:NAT UNIV OF DEFENSE TECH

Multi-modal agricultural technology question and answer method and system

The invention provides a multi-modal agricultural technology question and answer method and system, and belongs to the technical field of artificial intelligence, and the method comprises the steps: extracting a text feature vector, an image feature vector, a time sequence feature vector and a spatio-temporal context feature vector according to a query text and a crop image; calculating an initial fusion feature according to the spatio-temporal context feature vector, the image feature vector and the time sequence feature vector; when the query text contains a semantic entity of a preset type, calculating to obtain a final fusion feature according to the text feature vector and the initial fusion feature; splicing the final fusion feature and the text feature vector, and mapping the spliced vector to a space-time knowledge graph for reasoning to obtain a causal reasoning path; and generating a question and answer result according to the causal reasoning path. According to the method, deep alignment of time and space and semantics is carried out on the multi-modal agricultural data, and causal reasoning is carried out in combination with the knowledge graph with time and space constraints, so that the accuracy of a question and answer result is remarkably improved.
Owner:BEIJING ACADEMY OF AGRICULTURE & FORESTRY SCIENCES

Multi-mode weak supervision video anomaly detection method and device, equipment and medium

The invention belongs to the field of computer science and technology, and particularly relates to a multi-mode weak supervision video anomaly detection method and device, equipment and a medium. The method comprises the following steps: firstly, carrying out dynamic modeling on inter-fragment information by utilizing an external attention mechanism to obtain cross-fragment global context information; the temporal context aggregation module and the multi-scale time network are used to capture global and local information of the visual information and the text information within the fragment, and a feature representation containing local context details and the global information is generated. In addition, a multi-modal adaptive fusion mode is adopted, key modal features are focused in combination with target weights, further processing is performed through a multi-scale convolution attention module, and feature representation with higher discrimination is extracted. According to the method, the dependence on accurate annotation data in traditional video anomaly detection is effectively reduced; through hierarchical context modeling and an adaptive attention mechanism, the feature expression ability and the key information capture efficiency are enhanced, and a reliable anomaly detection solution is provided for an intelligent monitoring system.
Owner:TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Depression state prediction method and system based on voice multi-scale time domain perception

The invention discloses a depression state prediction method and system based on voice multi-scale time domain perception, and the method comprises the steps: collecting voice signals of a participant reading a unified standardized text, and carrying out the preprocessing of the voice signals, and generating a corresponding Mel spectrogram; extracting time context features by using a spectrum-time domain feature extraction algorithm to obtain joint feature representation; carrying out multi-scale division on the combined feature along the time dimension, and fusing the features of each scale into a global feature; and obtaining depression state discrimination results of the participants based on the global features. By combining the spectrum-time domain feature extraction algorithm and the frame-level time attention, the time domain locality features of depression speech such as inter-sentence pause, hesitant pause, inter-word transition and formant blurring can be explicitly analyzed and accurately positioned on the Mel spectrum, so that the recognition accuracy and the result stability are remarkably improved.
Owner:NANJING MEDICAL UNIV

Drought monitoring method and system based on multi-source data and spatio-temporal context embedding

The invention belongs to the technical field of drought monitoring, and particularly discloses a drought monitoring method and system based on multi-source data and spatio-temporal context embedding, and the method comprises the following steps: obtaining drought index data, multi-source dynamic remote sensing factor data and static geographic data of a meteorological station in a research area, and carrying out the preprocessing, obtaining original remote sensing features; extracting attribute values from the static geographic data according to the longitude and latitude of a site, extracting a space context embedding vector and a time context embedding vector based on a multi-layer perceptron and a long-short-term memory network, generating a modulated space-time context embedding vector, and splicing the modulated space-time context embedding vector with the original remote sensing features to obtain a final feature vector; and inputting the information enhanced final feature vector and the monthly standardized rainfall evapotranspiration index SPEI-3 of each station into a GXGBDM model, and outputting a drought monitoring result. By adopting the technical scheme, the drought monitoring result can be efficiently and accurately generated by fusing the multi-source data and the spatio-temporal context embedding information.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Flight guarantee time prediction method and device, medium and equipment

The invention relates to the technical field of data prediction, in particular to a flight guarantee time prediction method and device, a medium and equipment, by constructing a multi-dimensional initial feature set and fusing a channel attention mechanism, and by capturing a long-time sequence dependency relationship, the capturing capability of a preorder guarantee node delay conduction effect is remarkably improved; through dual-module adaptive selection based on node attributes, prediction demands of different types of guarantee nodes are accurately matched, and the overall precision of whole-process node prediction is improved; the prediction value of the ith guarantee node and the time sequence context feature are spliced to form the dynamic enhancement feature of the subsequent guarantee node, so that the transmission and fusion of the preorder prediction information to the subsequent node are realized, and the prediction of the subsequent guarantee node can adapt to the running state change of the preorder guarantee node in real time; the subsequent prediction deviation caused by the delay of the preorder guarantee node is effectively reduced, and the dynamic adaptability and timeliness of the prediction result are improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Recursive-temporal models for autonomous or semi-autonomous perception systems and applications

In various examples, machine learning models that benefit from temporal context while being computationally efficient to train and use are described herein. For instance, the disclosed systems and methods may apply a temporal series of images to a model and use intermediate features output from one or more backbone layers of the model as training data. In some examples, one or more recursive layers and / or one or more head layers of the model—or another model—may be trained using the training data by applying the intermediate features to the recursive layer(s). The recursive layer(s) may output a state representative of a temporal combination of the intermediate features, and the state may be applied to the head layer(s) to make one or more predictions. During inference, the recursive layer(s) may, in some examples, continuously update the state based on previous states of the recursive layer(s).
Owner:NVIDIA CORP

Path perception using temporal modeling for autonomous systems and applications

In various examples, to improve path perception in machine learning implementations, a temporal model includes a backbone model trained to predict one or more path perception outputs, such as, path geometry, path class, path uncertainty and / or other path attributes, for a current input frame. To create temporal context, the temporal model enables the backbone model to separately operate (in parallel or otherwise) on a set of frames that are temporally related to the current input frame. The outputs of the separate executions of the backbone model are then concatenated and processed via one or more convolution operations to generate a set of features that will be fed to the final output layer of the pipeline that encapsulates one or more path perception outputs that are generated based on temporal context.
Owner:NVIDIA CORP

Bridge construction crack dynamic detection method and system based on AI visual identification

The invention relates to the technical field of bridge construction monitoring, and discloses a bridge construction crack dynamic detection method and system based on AI visual identification. According to the method, pixel-level time domain and space domain feature separation and fusion are carried out by analyzing an obtained multi-frame continuous image sequence, a target structure feature map fused with spatio-temporal context information is generated, and a potential crack region is positioned according to the target structure feature map. And constructing a region evolution model by analyzing a pixel intensity multi-direction discrete change rule of the candidate region, and outputting quantitative description of a crack growth trend by using a crack growth prediction model. And a structural integrity dynamic score is generated through comprehensive growth trend and multi-scale texture analysis, and real-time detection and evaluation of the crack are achieved. According to the method, the accuracy and robustness of crack detection in a complex construction environment are improved, and the dynamic development trend of the crack can be effectively predicted.
Owner:ANSHAN URBAN & RURAL PLANNING & DESIGN INST CO LTD

Space-time trajectory prediction method and device based on multi-modal condition potential diffusion generation model

The invention discloses a spatio-temporal trajectory prediction method and device based on a multi-modal condition potential diffusion generation model, electronic equipment and a storage medium, and the method comprises the steps: cleaning and aligning multi-source data such as trajectory data, environment semantics and a road network, and constructing standardized input; extracting spatio-temporal context representation through multi-modal attention, mining an environment topological structure in combination with an iterative graph network, and generating structured node embedding; adopting dual-channel cross-modal attention fusion trajectory dynamic features and graph structure information to form unified semantic representation; the conditional variation auto-encoder fuses feature codes to a low-dimensional potential space, conditional reverse denoising generation is executed by using a potential diffusion model, and a future trajectory sequence is recovered step by step; and light weight of the model is realized through diffusion consistency distillation, and reverse sampling is compressed. Through the multimode topology-diffusion distillation integrated architecture, the precision, continuity and reasoning efficiency of trajectory prediction under sparse noise data are improved, and the method is suitable for scenes such as ecological monitoring, intelligent traffic and navigation.
Owner:BEIJING FORESTRY UNIVERSITY

Depression state prediction method and system based on voice multi-scale time domain perception

The application discloses a depression state prediction method and system based on voice multiscale time domain perception, and the method comprises the following steps: collecting voice signals of participants reading unified standardized texts, and generating corresponding Mel spectrum graphs after preprocessing; extracting time context features by using a spectrum-time domain feature extraction algorithm to obtain joint feature representation; performing multiscale division on the joint features along the time dimension, and then fusing the features of each scale into global features; and obtaining a depression state discrimination result of the participants based on the global features. By combining the spectrum-time domain feature extraction algorithm and frame-level time attention, the application can explicitly analyze and accurately locate the time domain local features of depression voice, such as inter-sentence pause, hesitation pause, inter-word transition and resonance peak fuzziness, so that the accuracy of recognition and the stability of results are significantly improved.
Owner:NANJING MEDICAL UNIV

Bird's eye view generation method based on multi-scale feature transformation and temporal context

The application discloses an aerial view generation method based on multi-scale feature transformation and time sequence context, and relates to the technical field of map generation. The method fully utilizes the complementarity of image and laser radar data by fusing the image and the laser radar data, improves the accuracy and robustness of aerial view generation, and can still remain stable under bad weather; a multi-scale space conversion module extracts different scale features, enhances the feature expression capability, and makes the aerial view clearer and more accurate; time sequence information is introduced, past time features are used to enhance current features, and dynamic perception capability is improved; an advanced backbone network and a feature alignment and fusion module are adopted, so that high efficiency and flexibility are ensured; and specific network structures such as Swin-T, PointPillars and random inactivation layers are applied, so that the accuracy and generalization capability are further improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Intelligent lamp control method based on multi-dimensional spatio-temporal context snapshot and edge side instruction fine-tuning enlarging model

The invention relates to an intelligent lamp control method based on a multi-dimensional spatio-temporal context snapshot and an edge side instruction fine-tuning enlarging model, and the method comprises the steps: receiving a user adjustment operation, and obtaining a light state before adjustment and a corresponding target light state; generating adjustment feature data according to the difference value; collecting multi-source context data corresponding to the adjustment operation, performing time alignment with the adjustment feature data, packaging the multi-source context data into spatio-temporal context snapshot data, and storing the spatio-temporal context snapshot data; combining continuous adjustments in a preset time window into an adjustment session as statistical sample data; performing conditional statistical aggregation on the statistical sample data to obtain statistical index data, generating cue words, inputting an edge side instruction to finely tune the large language model for reasoning, and outputting structured user portrait data under the constraint of JSON Schema; and generating or updating an intelligent adjustment strategy to control the intelligent lamp. According to the method, user adjustment and context information identification preferences are combined on the edge side, and an executable strategy is formed, so that the intelligence, stability and personalized adaptation capability of automatic adjustment are improved.
Owner:YUEYING INNOVATION TECH (GUANGDONG) CO LTD

Information processing apparatus and method, and program

The present technique relates to an information processing apparatus, an information processing method, and a program that enable a contextual relationship of data to be proved even when troubleshooting of equipment has been carried out.The information processing apparatus is an information processing apparatus that generates output data including object data to be a verification object of a temporal contextual relationship, the information processing apparatus including: a control unit configured to generate the output data including a hash value calculated on the basis of a part of or all of output data that temporally immediately-precedes the output data, the object data, and a plurality of mutually-different types of ID information related to the object data. The present technique can be applied to cameras.
Owner:SONY SEMICON SOLUTIONS CORP

A video salient object detection method and system based on spatio-temporal context scene relationship propagation

The application provides a video salient object detection method based on spatio-temporal context scene relationship propagation, and relates to the technical field of video salient object detection. First, scene analysis is performed on each frame of video in a video frame sequence to obtain an instance-level object corresponding to each frame of video. Then, global instance-level features, local instance-level features and intra-frame low-level features of each frame of video and the corresponding instance-level object are extracted. Then, a matching frame of any frame of video is randomly sampled from the same video frame sequence, global instance-level features of the matching frame are extracted, the global instance-level features are integrated into global instance-level features of the matching frame, and time features between video frames are obtained. The dense spatial attention mechanism is used to integrate the local instance-level features into the global instance-level features to obtain spatial features. The time features and the spatial features are spliced to obtain spatio-temporal features. The spatio-temporal features and the global instance-level features are input into a convolutional neural network based on a gated recurrent unit for updating to obtain high-level spatio-temporal features. Finally, the intra-frame low-level features and the high-level spatio-temporal features are fused and decoded to generate a video salient object mask detection result. The application utilizes rich inter-frame and intra-frame scene relationship information in the video, and improves the accuracy of video salient object detection in complex scenes.
Owner:GUANGDONG UNIV OF TECH

A state diagnosis method, system, device and medium based on unmanned aerial vehicle operation data

This application discloses a state diagnosis method, system, device, and medium based on UAV operational data, mainly relating to the field of state diagnosis technology. It addresses the problems of single-channel statistical features failing to capture cross-channel gradient correlations and dynamic coupling relationships between sensor signals, and unidirectional LSTMs only utilizing historical information and neglecting the inverse constraints of future temporal context on the current state. The method includes: inputting training samples into a forward LSTM and a backward LSTM respectively to obtain forward and backward outputs; concatenating the forward and backward outputs; compressing the concatenated data into a fixed-length semantic vector; obtaining the labeled state of the semantic vector; and obtaining a trained classification function based on the semantic vector and the labeled state; when new multi-channel UAV operational data is obtained, obtaining the corresponding new semantic vector; and inputting the new semantic vector into the trained classification function to obtain the predicted state.
Owner:SHANDONG ZHENGCHEN TECH CO LTD

A method and system for real-time monitoring of network device operating status based on multi-source log fusion

PendingCN122316929APathPingTemporal context
This invention relates to the field of equipment monitoring technology and discloses a method and system for real-time monitoring of network device operating status based on multi-source log fusion. The method acquires multi-source log data from network devices and constructs a fusion window sample. Based on device state transition graph constraints, it performs latent state soft allocation and temporal smoothing correction. It then performs differentiated feature enhancement according to feature physical semantics and generates a global temporal context signal. Furthermore, it constructs a temporal decomposition convolutional neural network, extracts trend and impulse information through trend paths and impulse paths respectively, fuses the dual-path information using state transition memory gating vectors, and finally outputs the device status identification result through state probability-guided attention pooling and classification. This significantly improves the accuracy of real-time monitoring of network device operating status.
Owner:SICHUAN YONGFENG TECH CO LTD

Knowledge recommendation method and device of online education platform, computer device and medium

The application discloses a kind of knowledge recommendation method, device, computer equipment and medium of online education platform, comprising: constructing the heterogeneous information network of online education platform, and dynamically generating the semantic embedding of knowledge concept, obtain basic embedding information;The enhancement optimization of knowledge concept is carried out to basic embedding information, and enhanced embedding information is obtained;Based on enhanced embedding information, construct the multi-type user interaction graph that integrates time context information;Message transmission is carried out in multi-type user interaction graph, and the node embedding of user and knowledge concept is updated, and target embedding information is obtained;Different sources of target embedding information are weighted and fused by semantic attention module, and fusion embedding information is obtained;The preference of user to knowledge concept is predicted and recommendation information is generated using extended matrix decomposition model and fusion embedding information, and the accuracy of knowledge recommendation of online education platform is improved using the application.
Owner:ANHUI POLYTECHNIC UNIV

Time action positioning method and system based on time sequence context maximum pooling

The invention discloses a time action positioning method and system based on time sequence context maximum pooling. The method comprises the following steps: acquiring a to-be-identified action video, and extracting and acquiring a feature coding sequence of the to-be-identified action video; presetting a time action positioning model, inputting the feature coding sequence into the time action positioning model, and obtaining an action classification result; by introducing the time sequence context maximum pooling operation and the multi-scale time feature pyramid structure, the calculation complexity is effectively reduced, and the reasoning speed and efficiency are improved. Meanwhile, by designing a time action positioning model comprising an encoder, a long-term time context module and a decoder, the model can simultaneously capture short-time and long-time dependency, so that the accuracy of time action positioning is improved. Due to the advantages, the method has important application value in the field of time action positioning.
Owner:GUIZHOU POWER GRID CO LTD

Chain prompt and multi-modal large model-based sarcopenia detection method and device

The application discloses a sarcopenia detection method and device based on a chain prompt and a multi-modal large model, and belongs to the technical field of video processing. The method comprises the following steps: collecting a patient action video, and dividing each frame into a plurality of small blocks to be mapped into a feature vector sequence; information fusion is performed between the same frame and different time points, and a global visual feature vector with spatial and temporal context is output; a continuous feature sequence is cut into a plurality of semantic coherent action stages; corresponding segmented prompt text is retrieved from a pre-defined mapping table based on the test type; a natural language description containing quantitative indicators and action features is generated at one time; a sarcopenia diagnosis and its detailed reasons are output; a probability value is taken as the confidence of the diagnosis; and a complete diagnosis report containing quantitative and qualitative analysis is output by one key. The application can greatly improve the universality and robustness across scenes and devices, and significantly improve the explainability of model decision.
Owner:XIYUAN HOSPITAL OF CHINA ACAD OF CHINESE MEDICAL SCI

Video editing model training method and device, computing equipment and storage medium

The invention relates to a video editing model training method and device, computing equipment and a storage medium, and belongs to the technical field of videos. In the method, since the difference between the to-be-edited video and the target video only lies in different mouth shapes of the dubbing object, the to-be-edited video provides complete spatio-temporal context information of the dubbing object, so that when a video model is trained, the dubbing object can be quickly edited; the target video is taken as expected output, the target voice of the dubbing object in the target video is taken as a constraint condition, the video editing model is based on the mouth shape of the dubbing object in the target video, only the mouth shape of the dubbing object in the to-be-edited video is edited, and parts except the mouth part of the dubbing object do not need to be edited. In this way, the mouth shape of only the dubbing object in the edited video is changed, the environment where the dubbing object is located and the posture and identity of the dubbing object are kept unchanged, and therefore the quality of the visual dubbing result of the video editing model can be improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Anti-interference industrial vision counting method and device based on spatio-temporal context perception

ActiveCN122199550BVisual CountTemporal context
The application provides an anti-interference industrial visual counting method and device based on space-time context perception, which comprises the following steps: acquiring real-time video stream of a monitoring area and recognizing a detection target; calculating intersection proportion of the detection target and a target area; triggering video recording to acquire a target video segment when the intersection proportion is greater than or equal to an intersection threshold; extracting instantaneous position and instantaneous category of a counting target in each video frame of the target video segment; generating a position sequence based on the instantaneous position and calculating a decision displacement direction; statistically analyzing the instantaneous category and determining a decision category; comparing the decision displacement direction and the decision category with preset effective direction and effective category to generate a detection result; and counting the number of effective results in a preset detection time to obtain a counting result. The application realizes classified and directional industrial visual counting by analyzing the motion state of the counting target in continuous time.
Owner:SHANGHAI DONGFANG HOPE SOFTWARE TECH CO LTD

User interaction intention recognition method and system based on context awareness

The invention discloses a user interaction intention recognition method and system based on context awareness, and belongs to the technical field of artificial intelligence intention recognition, and the method comprises the steps: collecting user multi-modal interaction input data and corresponding spatio-temporal context data, and combining the entity association strength and the spatio-temporal association weight coefficient of the multi-modal interaction input data, generating initial intention confidence, determining a memory layer priority factor according to a weighted value of a session continuous interaction round proportion and a historical interaction recall frequency proportion, generating a dynamic space-time weight to update a space-time association weight coefficient based on the memory layer weighted confidence and the memory layer priority factor, and if the memory layer weighted confidence is lower than a preset threshold, determining that the time-space association weight coefficient is greater than the preset threshold. According to the method, through deep fusion of space-time perception, memory layer priority and confidence coefficient closed-loop feedback, a dynamic self-adaptive intention recognition mechanism is formed.
Owner:CHENGDU MINGTU TECH CO LTD

Station long-time abnormal behavior identification method based on multi-granularity spatio-temporal context fusion

The invention discloses a station long-time abnormal behavior identification method based on multi-granularity spatio-temporal context fusion, and the method comprises the steps: collecting multi-source sensing data, carrying out the spatio-temporal alignment, and generating a multi-modal data flow of a unified coordinate system; constructing an atomic event code based on the attitude sequence, generating a dynamic scene graph according to a target-environment relationship, and forming a dual-channel feature primitive; injecting the feature elements into a short-term memory layer STM, abstracting a middle-term behavior pattern, storing the abstracted middle-term behavior pattern into a middle-term memory layer MTM, and fusing a cross-camera scene graph to construct a long-term memory layer LTM to form a layered space-time memory library; performing cross-level retrieval on the associated memory in the STM / MTM / LTM through a deformable space-time attention lens, and outputting context features of multi-granularity fusion; and calculating short / medium / long-term abnormal scores based on the fusion features, and dynamically adjusting a threshold value in combination with the crowd density to realize collaborative judgment. According to the method, the problem of fragmentation of long-time behavior understanding is effectively solved, and instantaneous anomaly and long-time mode anomaly can be accurately identified at the same time.
Owner:NANJING MODERN MULTIMODAL TRANSPORTATION LABORATORY

A transformer-based video decoding filtering method, system, terminal and storage medium

This invention discloses a transformer-based video decoding and filtering method, system, terminal, and storage medium. The method includes: acquiring a reference frame and a current frame; passing the reference frame and the current frame through a motion estimation network to obtain motion vectors; performing motion encoding and decoding on the motion vectors to obtain decoded motion vectors; performing motion compensation on the decoded motion vectors to obtain multi-scale context information; inputting the current frame and the multi-scale context information into a residual coding network to obtain a feature representation of the current frame; inputting the feature representation and reference features of the reference frame into a preprocessing module of a temporal context filter for preprocessing to obtain key vectors, value vectors, and query vectors constituting an attention mechanism; inputting the key vectors, value vectors, and query vectors into the filtering main module of the temporal context filter for filtering optimization; outputting a final reconstructed frame; and participating the final reconstructed frame in subsequent video decoding and filtering processes. This invention significantly improves the quality of reconstructed video.
Owner:PENG CHENG LAB