Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

89 results about "Temporal context" patented technology

What is Temporal Context. 1. Temporal information which may impact interpretation of a study. At different points in time, different browsing environments and activities emerge and become part of users’ experiences. Temporal factors which can be reported include the date of the study and duration of the study.

Weakly supervised group behavior identification method based on dynamic prompt tuning

The invention discloses a weak supervision group behavior recognition method based on dynamic prompt tuning, and belongs to the field of video understanding. According to the method, a pre-trained vision-language model is adaptively expanded to a group behavior recognition task for challenges such as complex group behavior semantics, strong spatio-temporal context dependency and lack of individual labels in a video. According to the technical scheme, a dynamic prompt generation technology with visual conditions is provided, instance-level text prompts can be automatically generated according to input video content, the accuracy of vision-text semantic alignment is enhanced, meanwhile, a time sequence fusion module is integrated in a model, and the purpose is to fuse key time information in multiple frames through modeling inter-frame time sequence dependence. The model calculates a similarity score between visual and text features and takes the similarity score as a classification prediction basis, and finally, a cross entropy loss function is adopted to carry out end-to-end training on the network. The validity of the method is verified on a volleyball data set and an NBA data set.
Owner:BEIJING UNIV OF TECH

Temporal context from eye tracking for generative ai

Examples relate to systems and methods for enhancing generative AI outputs using eye tracking data. An eye tracking system accesses eye gaze information associated with a field of view of a head-wearable apparatus and generates contextual information associated with the field of view of the head-wearable apparatus based on the eye gaze information. The eye tracking system processes, by a generative machine learning model, the contextual information and at least one image of the field of view of the head-wearable apparatus to generate an output and presents on a display of the head-wearable apparatus the output generated by the generative machine learning model.
Owner:SNAP INC

Multi-modal agricultural technology question and answer method and system

The invention provides a multi-modal agricultural technology question and answer method and system, and belongs to the technical field of artificial intelligence, and the method comprises the steps: extracting a text feature vector, an image feature vector, a time sequence feature vector and a spatio-temporal context feature vector according to a query text and a crop image; calculating an initial fusion feature according to the spatio-temporal context feature vector, the image feature vector and the time sequence feature vector; when the query text contains a semantic entity of a preset type, calculating to obtain a final fusion feature according to the text feature vector and the initial fusion feature; splicing the final fusion feature and the text feature vector, and mapping the spliced vector to a space-time knowledge graph for reasoning to obtain a causal reasoning path; and generating a question and answer result according to the causal reasoning path. According to the method, deep alignment of time and space and semantics is carried out on the multi-modal agricultural data, and causal reasoning is carried out in combination with the knowledge graph with time and space constraints, so that the accuracy of a question and answer result is remarkably improved.
Owner:BEIJING ACADEMY OF AGRICULTURE & FORESTRY SCIENCES

Drought monitoring method and system based on multi-source data and spatio-temporal context embedding

The invention belongs to the technical field of drought monitoring, and particularly discloses a drought monitoring method and system based on multi-source data and spatio-temporal context embedding, and the method comprises the following steps: obtaining drought index data, multi-source dynamic remote sensing factor data and static geographic data of a meteorological station in a research area, and carrying out the preprocessing, obtaining original remote sensing features; extracting attribute values from the static geographic data according to the longitude and latitude of a site, extracting a space context embedding vector and a time context embedding vector based on a multi-layer perceptron and a long-short-term memory network, generating a modulated space-time context embedding vector, and splicing the modulated space-time context embedding vector with the original remote sensing features to obtain a final feature vector; and inputting the information enhanced final feature vector and the monthly standardized rainfall evapotranspiration index SPEI-3 of each station into a GXGBDM model, and outputting a drought monitoring result. By adopting the technical scheme, the drought monitoring result can be efficiently and accurately generated by fusing the multi-source data and the spatio-temporal context embedding information.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Flight guarantee time prediction method and device, medium and equipment

The invention relates to the technical field of data prediction, in particular to a flight guarantee time prediction method and device, a medium and equipment, by constructing a multi-dimensional initial feature set and fusing a channel attention mechanism, and by capturing a long-time sequence dependency relationship, the capturing capability of a preorder guarantee node delay conduction effect is remarkably improved; through dual-module adaptive selection based on node attributes, prediction demands of different types of guarantee nodes are accurately matched, and the overall precision of whole-process node prediction is improved; the prediction value of the ith guarantee node and the time sequence context feature are spliced to form the dynamic enhancement feature of the subsequent guarantee node, so that the transmission and fusion of the preorder prediction information to the subsequent node are realized, and the prediction of the subsequent guarantee node can adapt to the running state change of the preorder guarantee node in real time; the subsequent prediction deviation caused by the delay of the preorder guarantee node is effectively reduced, and the dynamic adaptability and timeliness of the prediction result are improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Recursive-temporal models for autonomous or semi-autonomous perception systems and applications

In various examples, machine learning models that benefit from temporal context while being computationally efficient to train and use are described herein. For instance, the disclosed systems and methods may apply a temporal series of images to a model and use intermediate features output from one or more backbone layers of the model as training data. In some examples, one or more recursive layers and / or one or more head layers of the model—or another model—may be trained using the training data by applying the intermediate features to the recursive layer(s). The recursive layer(s) may output a state representative of a temporal combination of the intermediate features, and the state may be applied to the head layer(s) to make one or more predictions. During inference, the recursive layer(s) may, in some examples, continuously update the state based on previous states of the recursive layer(s).
Owner:NVIDIA CORP

Bridge construction crack dynamic detection method and system based on AI visual identification

The invention relates to the technical field of bridge construction monitoring, and discloses a bridge construction crack dynamic detection method and system based on AI visual identification. According to the method, pixel-level time domain and space domain feature separation and fusion are carried out by analyzing an obtained multi-frame continuous image sequence, a target structure feature map fused with spatio-temporal context information is generated, and a potential crack region is positioned according to the target structure feature map. And constructing a region evolution model by analyzing a pixel intensity multi-direction discrete change rule of the candidate region, and outputting quantitative description of a crack growth trend by using a crack growth prediction model. And a structural integrity dynamic score is generated through comprehensive growth trend and multi-scale texture analysis, and real-time detection and evaluation of the crack are achieved. According to the method, the accuracy and robustness of crack detection in a complex construction environment are improved, and the dynamic development trend of the crack can be effectively predicted.
Owner:ANSHAN URBAN & RURAL PLANNING & DESIGN INST CO LTD

Space-time trajectory prediction method and device based on multi-modal condition potential diffusion generation model

The invention discloses a spatio-temporal trajectory prediction method and device based on a multi-modal condition potential diffusion generation model, electronic equipment and a storage medium, and the method comprises the steps: cleaning and aligning multi-source data such as trajectory data, environment semantics and a road network, and constructing standardized input; extracting spatio-temporal context representation through multi-modal attention, mining an environment topological structure in combination with an iterative graph network, and generating structured node embedding; adopting dual-channel cross-modal attention fusion trajectory dynamic features and graph structure information to form unified semantic representation; the conditional variation auto-encoder fuses feature codes to a low-dimensional potential space, conditional reverse denoising generation is executed by using a potential diffusion model, and a future trajectory sequence is recovered step by step; and light weight of the model is realized through diffusion consistency distillation, and reverse sampling is compressed. Through the multimode topology-diffusion distillation integrated architecture, the precision, continuity and reasoning efficiency of trajectory prediction under sparse noise data are improved, and the method is suitable for scenes such as ecological monitoring, intelligent traffic and navigation.
Owner:BEIJING FORESTRY UNIVERSITY

Intelligent lamp control method based on multi-dimensional spatio-temporal context snapshot and edge side instruction fine-tuning enlarging model

The invention relates to an intelligent lamp control method based on a multi-dimensional spatio-temporal context snapshot and an edge side instruction fine-tuning enlarging model, and the method comprises the steps: receiving a user adjustment operation, and obtaining a light state before adjustment and a corresponding target light state; generating adjustment feature data according to the difference value; collecting multi-source context data corresponding to the adjustment operation, performing time alignment with the adjustment feature data, packaging the multi-source context data into spatio-temporal context snapshot data, and storing the spatio-temporal context snapshot data; combining continuous adjustments in a preset time window into an adjustment session as statistical sample data; performing conditional statistical aggregation on the statistical sample data to obtain statistical index data, generating cue words, inputting an edge side instruction to finely tune the large language model for reasoning, and outputting structured user portrait data under the constraint of JSON Schema; and generating or updating an intelligent adjustment strategy to control the intelligent lamp. According to the method, user adjustment and context information identification preferences are combined on the edge side, and an executable strategy is formed, so that the intelligence, stability and personalized adaptation capability of automatic adjustment are improved.
Owner:YUEYING INNOVATION TECH (GUANGDONG) CO LTD

Information processing apparatus and method, and program

The present technique relates to an information processing apparatus, an information processing method, and a program that enable a contextual relationship of data to be proved even when troubleshooting of equipment has been carried out.The information processing apparatus is an information processing apparatus that generates output data including object data to be a verification object of a temporal contextual relationship, the information processing apparatus including: a control unit configured to generate the output data including a hash value calculated on the basis of a part of or all of output data that temporally immediately-precedes the output data, the object data, and a plurality of mutually-different types of ID information related to the object data. The present technique can be applied to cameras.
Owner:SONY SEMICON SOLUTIONS CORP

A state diagnosis method, system, device and medium based on unmanned aerial vehicle operation data

This application discloses a state diagnosis method, system, device, and medium based on UAV operational data, mainly relating to the field of state diagnosis technology. It addresses the problems of single-channel statistical features failing to capture cross-channel gradient correlations and dynamic coupling relationships between sensor signals, and unidirectional LSTMs only utilizing historical information and neglecting the inverse constraints of future temporal context on the current state. The method includes: inputting training samples into a forward LSTM and a backward LSTM respectively to obtain forward and backward outputs; concatenating the forward and backward outputs; compressing the concatenated data into a fixed-length semantic vector; obtaining the labeled state of the semantic vector; and obtaining a trained classification function based on the semantic vector and the labeled state; when new multi-channel UAV operational data is obtained, obtaining the corresponding new semantic vector; and inputting the new semantic vector into the trained classification function to obtain the predicted state.
Owner:SHANDONG ZHENGCHEN TECH CO LTD

A method and system for real-time monitoring of network device operating status based on multi-source log fusion

PendingCN122316929APathPingTemporal context
This invention relates to the field of equipment monitoring technology and discloses a method and system for real-time monitoring of network device operating status based on multi-source log fusion. The method acquires multi-source log data from network devices and constructs a fusion window sample. Based on device state transition graph constraints, it performs latent state soft allocation and temporal smoothing correction. It then performs differentiated feature enhancement according to feature physical semantics and generates a global temporal context signal. Furthermore, it constructs a temporal decomposition convolutional neural network, extracts trend and impulse information through trend paths and impulse paths respectively, fuses the dual-path information using state transition memory gating vectors, and finally outputs the device status identification result through state probability-guided attention pooling and classification. This significantly improves the accuracy of real-time monitoring of network device operating status.
Owner:SICHUAN YONGFENG TECH CO LTD

Time action positioning method and system based on time sequence context maximum pooling

The invention discloses a time action positioning method and system based on time sequence context maximum pooling. The method comprises the following steps: acquiring a to-be-identified action video, and extracting and acquiring a feature coding sequence of the to-be-identified action video; presetting a time action positioning model, inputting the feature coding sequence into the time action positioning model, and obtaining an action classification result; by introducing the time sequence context maximum pooling operation and the multi-scale time feature pyramid structure, the calculation complexity is effectively reduced, and the reasoning speed and efficiency are improved. Meanwhile, by designing a time action positioning model comprising an encoder, a long-term time context module and a decoder, the model can simultaneously capture short-time and long-time dependency, so that the accuracy of time action positioning is improved. Due to the advantages, the method has important application value in the field of time action positioning.
Owner:GUIZHOU POWER GRID CO LTD

Chain prompt and multi-modal large model-based sarcopenia detection method and device

The application discloses a sarcopenia detection method and device based on a chain prompt and a multi-modal large model, and belongs to the technical field of video processing. The method comprises the following steps: collecting a patient action video, and dividing each frame into a plurality of small blocks to be mapped into a feature vector sequence; information fusion is performed between the same frame and different time points, and a global visual feature vector with spatial and temporal context is output; a continuous feature sequence is cut into a plurality of semantic coherent action stages; corresponding segmented prompt text is retrieved from a pre-defined mapping table based on the test type; a natural language description containing quantitative indicators and action features is generated at one time; a sarcopenia diagnosis and its detailed reasons are output; a probability value is taken as the confidence of the diagnosis; and a complete diagnosis report containing quantitative and qualitative analysis is output by one key. The application can greatly improve the universality and robustness across scenes and devices, and significantly improve the explainability of model decision.
Owner:XIYUAN HOSPITAL OF CHINA ACAD OF CHINESE MEDICAL SCI

Video editing model training method and device, computing equipment and storage medium

The invention relates to a video editing model training method and device, computing equipment and a storage medium, and belongs to the technical field of videos. In the method, since the difference between the to-be-edited video and the target video only lies in different mouth shapes of the dubbing object, the to-be-edited video provides complete spatio-temporal context information of the dubbing object, so that when a video model is trained, the dubbing object can be quickly edited; the target video is taken as expected output, the target voice of the dubbing object in the target video is taken as a constraint condition, the video editing model is based on the mouth shape of the dubbing object in the target video, only the mouth shape of the dubbing object in the to-be-edited video is edited, and parts except the mouth part of the dubbing object do not need to be edited. In this way, the mouth shape of only the dubbing object in the edited video is changed, the environment where the dubbing object is located and the posture and identity of the dubbing object are kept unchanged, and therefore the quality of the visual dubbing result of the video editing model can be improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Anti-interference industrial vision counting method and device based on spatio-temporal context perception

ActiveCN122199550BVisual CountTemporal context
The application provides an anti-interference industrial visual counting method and device based on space-time context perception, which comprises the following steps: acquiring real-time video stream of a monitoring area and recognizing a detection target; calculating intersection proportion of the detection target and a target area; triggering video recording to acquire a target video segment when the intersection proportion is greater than or equal to an intersection threshold; extracting instantaneous position and instantaneous category of a counting target in each video frame of the target video segment; generating a position sequence based on the instantaneous position and calculating a decision displacement direction; statistically analyzing the instantaneous category and determining a decision category; comparing the decision displacement direction and the decision category with preset effective direction and effective category to generate a detection result; and counting the number of effective results in a preset detection time to obtain a counting result. The application realizes classified and directional industrial visual counting by analyzing the motion state of the counting target in continuous time.
Owner:SHANGHAI DONGFANG HOPE SOFTWARE TECH CO LTD

User interaction intention recognition method and system based on context awareness

The invention discloses a user interaction intention recognition method and system based on context awareness, and belongs to the technical field of artificial intelligence intention recognition, and the method comprises the steps: collecting user multi-modal interaction input data and corresponding spatio-temporal context data, and combining the entity association strength and the spatio-temporal association weight coefficient of the multi-modal interaction input data, generating initial intention confidence, determining a memory layer priority factor according to a weighted value of a session continuous interaction round proportion and a historical interaction recall frequency proportion, generating a dynamic space-time weight to update a space-time association weight coefficient based on the memory layer weighted confidence and the memory layer priority factor, and if the memory layer weighted confidence is lower than a preset threshold, determining that the time-space association weight coefficient is greater than the preset threshold. According to the method, through deep fusion of space-time perception, memory layer priority and confidence coefficient closed-loop feedback, a dynamic self-adaptive intention recognition mechanism is formed.
Owner:CHENGDU MINGTU TECH CO LTD

Station long-time abnormal behavior identification method based on multi-granularity spatio-temporal context fusion

The invention discloses a station long-time abnormal behavior identification method based on multi-granularity spatio-temporal context fusion, and the method comprises the steps: collecting multi-source sensing data, carrying out the spatio-temporal alignment, and generating a multi-modal data flow of a unified coordinate system; constructing an atomic event code based on the attitude sequence, generating a dynamic scene graph according to a target-environment relationship, and forming a dual-channel feature primitive; injecting the feature elements into a short-term memory layer STM, abstracting a middle-term behavior pattern, storing the abstracted middle-term behavior pattern into a middle-term memory layer MTM, and fusing a cross-camera scene graph to construct a long-term memory layer LTM to form a layered space-time memory library; performing cross-level retrieval on the associated memory in the STM / MTM / LTM through a deformable space-time attention lens, and outputting context features of multi-granularity fusion; and calculating short / medium / long-term abnormal scores based on the fusion features, and dynamically adjusting a threshold value in combination with the crowd density to realize collaborative judgment. According to the method, the problem of fragmentation of long-time behavior understanding is effectively solved, and instantaneous anomaly and long-time mode anomaly can be accurately identified at the same time.
Owner:NANJING MODERN MULTIMODAL TRANSPORTATION LABORATORY

A transformer-based video decoding filtering method, system, terminal and storage medium

This invention discloses a transformer-based video decoding and filtering method, system, terminal, and storage medium. The method includes: acquiring a reference frame and a current frame; passing the reference frame and the current frame through a motion estimation network to obtain motion vectors; performing motion encoding and decoding on the motion vectors to obtain decoded motion vectors; performing motion compensation on the decoded motion vectors to obtain multi-scale context information; inputting the current frame and the multi-scale context information into a residual coding network to obtain a feature representation of the current frame; inputting the feature representation and reference features of the reference frame into a preprocessing module of a temporal context filter for preprocessing to obtain key vectors, value vectors, and query vectors constituting an attention mechanism; inputting the key vectors, value vectors, and query vectors into the filtering main module of the temporal context filter for filtering optimization; outputting a final reconstructed frame; and participating the final reconstructed frame in subsequent video decoding and filtering processes. This invention significantly improves the quality of reconstructed video.
Owner:PENG CHENG LAB

End-to-end video compression method based on temporal context mining and frame enhancement reconstruction

This invention relates to the field of video coding technology, specifically to an end-to-end video compression method based on temporal context mining and reconstructed frame enhancement. The method includes: extracting the decoded frame of the previous video frame; inputting the current video frame and the decoded frame into a motion estimation module to output a motion vector; compressing and reconstructing the motion vector in a lossy manner using a motion vector encoder and a motion vector decoder to obtain a reconstructed motion vector; inputting the reconstructed motion vector and features of the previous video frame into an enhanced context mining module to obtain a context; inputting the current video frame and its context into a conditional encoder and a conditional decoder to obtain a reconstructed frame; and inputting the reconstructed motion vector, features, and reconstructed frame into a reconstructed frame enhancement module to obtain a decoded frame. This invention improves the overall quality of the final output frame.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Dynamic spatio-temporal graph learning microservice anomaly detection method for multi-modal data

The application discloses a kind of dynamic spatio-temporal graph learning microservice exception detection methods for multi-modal data, which comprises: collecting data information from the microservice system to be detected and inputting into computer system and carrying out data preprocessing, obtaining multi-modal data;The multi-modal data is processed based on the sliding window mechanism, and a dynamic dependency graph sequence is constructed;Based on dynamic dependency graph sequence, spatial relationship modeling and time series dynamic evolution modeling are performed, a representation integrating time series information is generated, and a graph-level representation vector containing complete spatio-temporal context is generated;The graph-level representation vector is input into a deep vector data description anomaly detection model based on contrast learning, and the detection result is output.The microservice exception detection based on multi-modal dynamic spatio-temporal graph learning and contrast enhancement SVDD of the application explicitly captures the dynamic evolution characteristics of microservice topology structure, and realizes the deep fusion of spatio-temporal coupling characteristics, while suppressing the decision boundary ambiguity problem in unsupervised detection.
Owner:GUIZHOU UNIV

End-to-end video compression method based on enhanced spatio-temporal context mining and enhanced reconstructed frames

The present application belongs to the technical field of scalable video coding, and particularly relates to an end-to-end video compression method based on enhanced spatio-temporal context mining and enhanced reconstructed frame, which is realized through an end-to-end deep learning model, wherein the model efficiently utilizes motion vector time series information through an improved spatio-temporal context mining module at the encoding end, and utilizes the motion vector generated by an optical flow network and the motion vector compensated to the current time to jointly generate the final required motion vector at the current time, so as to significantly improve the accuracy of the motion vector. The scheme can effectively utilize the redundant information between different levels, reduce unnecessary data transmission and improve the video compression efficiency through inter-layer prediction, fusion of motion vector information of different layers, fusion of texture information of different layers and the like.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Multi-source sensing data fusion and event reasoning method based on edge computing and knowledge graph

PendingCN122310431ATimestampEvent level
This invention discloses a multi-source sensing data fusion and event reasoning method based on edge computing and knowledge graphs, relating to the field of remote monitoring technology. The method includes the following steps: Step 1, synchronous acquisition of multi-source sensing data; Step 2, location semantic mapping and appliance operating status identification; Step 3, the edge gateway constructs a set of current observation facts based on the current functional area number, current area dwell time, appliance status register content, and the time interval to which the current timestamp belongs, determining the current event type and current event level; Step 4, the cloud server performs incremental updates to the home scene event knowledge graph based on event feedback records and distributes the updates to the edge gateway. This invention achieves three-dimensional fusion sensing of location information, electricity consumption behavior, and temporal context, possessing advantages such as strong privacy protection, good interpretability, and a response measure that matches the level of risk.
Owner:SHANDONG VOCATIONAL COLLEGE OF ECONOMICS & TRADE

A viewpoint extraction and evolution path tracing system and method

PendingCN122309993ATemporal contextEngineering
This invention relates to a system and method for opinion extraction and evolution path tracking, belonging to the field of film review analysis technology. It includes a data input module, a data processing module, a temporal alignment module, a temporal enhancement module, a deep extraction module, an analysis module, and a visualization module. The data input module acquires long film review text data, temporal metadata, and external event data. The text processing module outputs text feature vectors, temporal decay factors, and external event feature vectors. The temporal alignment module outputs aligned temporal context. The temporal enhancement module performs temporal enhancement attention calculation. The deep extraction module extracts deep temporal features. The analysis module outputs sentiment curves and topic drift detection results. The visualization module generates a dynamic reputation evolution map and a visualization analysis report. This application effectively demonstrates the evolution of opinions in long film reviews.
Owner:NANTONG UNIV

Ripple-like image segmentation optimization method driven by edge enhancement and region features

This invention proposes a ripple-like image segmentation optimization method driven by edge enhancement and region features. The proposed "ripple mechanism" refers to a modeling strategy that uses edges and regions as structural centers, progressively spreading semantic features through direction perception and structural interaction. This mechanism mainly consists of two parts: a scan enhancement module and a nested state weight key-value network (NSKV). The scan enhancement module includes edge enhancement, region enhancement, and interactive dual scan modules for multi-directional, multi-level modeling of local and structural features. The state weight key-value network is a sequence modeling structure with state preservation capabilities, integrating the memory and attention mechanisms of recursive modeling. This invention can model long-range dependencies and temporal context information while maintaining high computational efficiency, thereby improving overall structure perception and segmentation accuracy.
Owner:CHINA UNIV OF MINING & TECH

Contextualized filtering of large language model content

In some implementations, there is provided a computer-implemented method including receiving a query to grant user access to content generated by a large-language model, the query including a user identifier; verifying, based on the user identifier and using a first filter of a filter pipeline, a clearance level associated with the user identifier; granting, based on the verifying, the user access to at least a subset of the content generated by the large-language model; verifying, using at least a second filter of the filter pipeline, a temporal context of the query and a spatial context of the query, the temporal context comprising a time at which the query is received and the spatial context comprising a location from which the query is received; and providing, based on the verifying of the clearance level, the temporal context, and the spatial context, content generated by the large-language model.
Owner:SAP SE

A multi-modal and time-aware based multi-task point of interest recommendation method

The present application belongs to the technical field of intelligent recommendation system and location service, and in particular to a multi-task point-of-interest recommendation method based on multi-modal and time perception. The method comprises the following steps: S1: data collection; S2: feature extraction; S3: construction of MAST-POI model. The present application introduces a large language model text encoder to improve semantic expression and cross-modal interaction modeling capability; introduces similar historical check-in trajectory embedding in the modeling process to effectively improve the recommendation accuracy in the data sparse scene; adopts a rotating position encoder to process the check-in time difference to enhance the model's ability to capture the time context; uses a gating weighting module and a hybrid expert feedforward network to realize dynamic adjustment of feature weights and expert selection, improve the model's personalized adaptability and generalization ability; can improve the accuracy, real-time performance and personalization level of the sequence POI recommendation system, and is suitable for large-scale user and multi-task real-time prediction scenarios.
Owner:JILIN UNIVERSITY

A robot vision control method and system based on spatio-temporal context perception

The application relates to a robot vision control method and system based on spatiotemporal context perception and belongs to the field of vision control. The method comprises the following steps: acquiring an original image sequence, performing spatiotemporal context feature coding, and outputting a spatiotemporal context feature map; mapping the spatiotemporal context feature map into an initial attention map through a lightweight convolution network, introducing an attention diffusion process to output an attention potential field; obtaining a resource allocation matrix based on the attention potential field and performing calculation power allocation; obtaining a historical attention potential field based on the attention potential field, performing attention dynamics prediction according to the historical attention potential field, and obtaining a predicted attention potential field; and realizing prospective vision control based on the attention potential field and the predicted attention potential field. Through the introduction of spatiotemporal context perception coding, gated fusion, attention diffusion, nonlinear resource allocation and the like, the intelligent level of robot vision control is comprehensively improved.
Owner:SHANGHAI YINNI TECHNOLOGY CO LTD

A context-based audio adaptive entropy encoding method and processing terminal

This invention discloses a context-based adaptive entropy coding method for audio, comprising the following steps: First, the quantized codebook index sequence Q and quantization side information S of the audio signal are sequentially fed into a causal convolutional context model, a cross-band attention network, and a cascaded codebook dependency network for processing, to obtain temporal context feature vectors, frequency context feature vectors, and quantization-level context feature vectors for the audio signal in the time domain, respectively. Second, the temporal context feature vectors, frequency context feature vectors, and quantization-level feature vectors are input into a neural probability prediction model for processing, estimating the conditional probability distribution of each codebook index. Third, the predicted conditional probability distributions are input into an adaptive arithmetic encoder to generate the encoded bitstream, and necessary header information is added to the encoded bitstream to form a complete bitstream. This invention can further improve the compression efficiency of audio coding.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD