Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

82 results about "Visual attentiveness" patented technology

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Robot automatic grabbing path planning method based on visual identification

The invention discloses a robot automatic grabbing path planning method based on visual identification, and relates to the technical field of intelligent grabbing. The method comprises the following steps: acquiring multi-view visual data of a target scene, and generating a scene three-dimensional compact reconstruction model through a multi-modal image fusion algorithm; performing target detection and feature extraction on the model, and screening an optimal capture point in combination with a visual attention mechanism; constructing a dynamic environment obstacle probability map, updating an obstacle state through time sequence visual tracking, and quantifying an interference weight; an initial grabbing path is planned based on an improved fast expansion random tree algorithm, path smoothness constraints and robot joint movement limit parameters are introduced, and path nodes are optimized through a Bezier curve; visual servo feedback and path deviation prediction are fused, path parameters are corrected in real time, and a continuous movement track is generated. The method effectively adapts to the dynamic environment, gives consideration to path safety, smoothness and mechanical arm motion characteristics, and remarkably improves the grabbing success rate and operation reliability.
Owner:TIANJIN UNIV OF SCI & TECH

Multi-modal large model illusion detection and suppression method based on attention time sequence difference

A multi-modal large model illusion detection and inhibition method based on attention time sequence difference comprises the following steps: inputting text lexical elements of cue words and visual lexical elements of images into a multi-modal large model, and obtaining an internal attention graph of the decoding stage of the multi-modal large model; then calculating the attention proportion of the visual lexical units at the current generation moment, and making a difference between the attention proportion and the proportion at the previous moment; and if the difference value exceeds a set threshold value, determining that the lexical elements are visual related lexical elements. When the visual related lexical elements are recognized, performing secondary forward propagation of primary visual enhancement to obtain more accurate output; and if not, directly entering the next step of generation. According to the method, visual related lexical elements in text generation are recognized and refined through attention time sequence difference, two-time forward propagation is adopted, the visual attention of second-time forward propagation is enhanced based on a visual attention graph of first-time forward propagation, and illusion can be recognized and corrected on the premise that the language expression ability is not reduced.
Owner:HANGZHOU DIANZI UNIV

Intelligent enhancement method and system for brightness of LED light-emitting module

The invention relates to the technical field of intelligent illumination control, and discloses an intelligent enhancement method and system for the brightness of an LED light-emitting module, and the method comprises the steps: collecting environment illumination data and LED module operation state data, and carrying out the preprocessing; acquiring cultural relic material information, performing exhibit change detection, determining an illumination constraint threshold, calculating effective illumination and tracking accumulated exposure; determining an illumination safety boundary and dynamically modulating an LED spectrum; performing hierarchical coordination by adopting a three-layer game decision framework; a three-dimensional light field representation framework is constructed through audience behavior perception and visual attention prediction, and space illumination optimization is carried out; environment sudden changes and audience behaviors are detected, and quick response adjustment is executed; analyzing lighting space-time distribution characteristics, performing dynamic power distribution and generating an intelligent dimming curve; according to the invention, environmental perception, cultural relic protection constraint, light attenuation compensation, multi-module cooperation and intelligent decision can be comprehensively considered.
Owner:HUBEI XIEJIN SEMICON TECH CO LTD

Exhibition and display streamline optimization method and system based on viewpoint thermodynamic diagram and spatial syntax

The invention belongs to the field of spatial design optimization, and particularly relates to an exhibition streamline optimization method and system based on a viewpoint thermodynamic diagram and a spatial syntax, and the method comprises the steps: firstly constructing an initial three-dimensional layout model, carrying out the dynamic discretization of the initial three-dimensional layout model into visual or path-connected spatial units, and calculating the integration degree, understandability and other indexes of the units in combination with the spatial syntax. Generating a theoretical streamline sequence; collecting viewpoint data of the test group along a preset path through eye movement tracking, and generating a viewpoint thermodynamic value distribution diagram in combination with a thermodynamic diagram algorithm, group interest distribution and visual attention characteristics; based on the theoretical streamline and the thermodynamic diagram, the coupling coordination degree of the space unit is calculated through a coupling algorithm, a preset optimization strategy is called after a defect unit is recognized, and an optimized exhibition streamline scheme is generated by combining the initial model and simulation algorithm adjustment. And the exhibition viewing experience and the space use efficiency are effectively improved.
Owner:NANJING TECH UNIV

Evaluation method and device for model generation image and storage medium

The invention discloses a model generation image evaluation method and device and a storage medium, and relates to the technical field of image processing. According to the method, the image generated based on the text cue word is obtained, wherein the image comprises the image element corresponding to the text cue word; filtering the image to obtain a filtering response value of each image pixel in the image, and generating a visual attention map of the image based on the filtering response values, the visual attention map representing frequency domain energy distribution of the image; and determining a visual saliency value of the image element according to the visual attention map, and determining a layout score of the image according to the visual saliency value of the image element. According to the method, spatial relations such as distances, alignment modes and hierarchical structures among elements are quantitatively evaluated through visual saliency values, and the defect that a previous model cannot accurately evaluate and adjust image layout is overcome.
Owner:SHENZHEN DONSON CLOUD TECHNOLOGY CO LTD

Automobile wire harness quality inspection method based on visual inspection

The invention discloses an automobile wire harness quality inspection method based on visual inspection, and belongs to the field of visual inspection and automobile wire harness detection.The method comprises the steps that a normal sample image set is constructed based on visual images of qualified automobile wire harnesses; self-adaptive preprocessing is carried out; inputting a pre-trained visual attention network, extracting key visual features of qualified wire harnesses, and constructing a normal feature library; collecting a visual image of a to-be-detected wire harness, performing adaptive preprocessing, and extracting key visual features of the to-be-detected image; calculating the feature deviation degree between the key visual features of the to-be-detected image and the normal feature library, and judging whether defects exist or not through the feature deviation degree; dynamically updating a normal feature library and a judgment threshold value through an online self-calibration module; and for the to-be-detected image which is judged to have the defect, positioning a defect area through a visual attention thermodynamic diagram, matching a preset defect feature template library, determining a defect type and outputting a quality inspection conclusion. According to the invention, appearance and assembly precision defects can be covered.
Owner:ZHUHAI QINCHUANG ELECTRONIC TECH CO LTD

Image feature enhancement method and system based on learnable unary function gating

PendingCN122367778ARadiologyImaging Feature
This invention relates to the field of image feature enhancement technology, providing an image feature enhancement method and system based on learnable unary function gating. The method includes: dividing the input image into image patches and mapping them to visual tokens; in a multi-head attention layer, calculating a visual attention aggregation score based on an attention weight matrix to quantify the degree of abnormal attention received by the image patch; inputting the normalized visual token, aggregation score, and two-dimensional position code into a gating module composed of a learnable unary function to generate a gating matrix; after each attention head completes SDPA output and before multi-head stitching, performing element-wise gating modulation using the gating matrix, and stitching and projecting to obtain the enhanced image features. This invention, through the synergy of aggregation score and learnable unary function, suppresses abnormal attention propagation from background noise, enhances the expression of key features of small targets, and improves the feature discriminativeness and task adaptability of the visual Transformer.
Owner:CCTEG BEIJING HUAYU ENG

Video large language model illusion relieving method and system

The invention belongs to the technical field of artificial intelligence, and particularly relates to a video large language model illusion relieving method and system. The method comprises the following steps: inputting a video and a text into a video big language model, and starting generation of a plurality of candidate replies in parallel; in the process of generating each candidate reply, monitoring a visual attention degeneration point; upon detecting the visual attention recession point, stopping generation of the corresponding candidate reply, and calculating a timing attention collapse value of the corresponding candidate reply; and selecting the candidate reply with the maximum time sequence attention collapse value to continuously generate until the candidate reply is finished. According to the method, the video large language model can give more attention to the global content of the video, perception illusion is reduced, so that correct reply is made, additional training of the model is not needed, and the calculation cost is low.
Owner:FUDAN UNIVERSITY

Precise transport control method and system for remediation agents in in-situ groundwater remediation

This invention provides a method and system for precise transport control of remediation agents in in-situ groundwater remediation, relating to the field of groundwater pollution remediation technology. The method includes: acquiring images and sensor data from borehole cores at contaminated sites; extracting features through visual attention networks and temporal attention networks; fusing these features into a conditional probability diffusion model to generate hydrogeological feature vectors; constructing a high-precision hydrogeological feature model by combining spatial attention mapping and a generative diffusion model; inputting this model into a multi-regional collaborative simulation environment for optimization calculation; fusing monitoring data through edge computing nodes; and training a lightweight control model to execute the transport control of remediation agents.
Owner:JIANGSU ZHONGWU ENVIRONMENTAL PROTECTION IND DEV CO LTD

A research method for visual attention mechanisms based on EEG microstates

This invention discloses a method for studying visual attention mechanisms based on EEG microstates, comprising: collecting EEG signals from subjects while watching videos and preprocessing them; extracting saliency maps from the videos, and extracting statistical features from the perspectives of spatial saliency information in the local temporal domain and spatiotemporal saliency information changes in the global time series, respectively, to obtain sIQR features and tsIQR features; converting the preprocessed EEG signals into EEG topology map sequences and performing spatial clustering to extract EEG microstate templates, determining the number of templates and selecting the final microstate templates, and then backfitting them to the EEG signals to obtain microstate sequences; extracting microstate features and depth features from the microstate sequences, examining the statistical differences between microstate features and sIQR and tsIQR features, constructing decoding models based on microstate features or depth features, and using the decoding models to decode segments and videos respectively to obtain segment labels and video labels.
Owner:SHENZHEN UNIV

Personalized recommendation method for virtual digital humans in enterprise publicity

The invention relates to the technical field of enterprise propaganda, and discloses a personalized recommendation method for virtual digital humans in enterprise propaganda. According to the method, multi-modal interaction data such as visual attention data and voice feedback data in the interaction process of a user and a virtual digital human are collected in real time, and an original interaction flow is generated; performing multi-dimensional fusion analysis on the original interaction flow, analyzing a dependency relationship among different dimensions, identifying a hidden association between a user interest mode and a virtual digital human performance feature, labeling an analysis result as an initial interest index, and generating an interest labeling data set with confidence; optimizing model adaptability and dynamically updating a model state by utilizing a fusion analysis result; matching the real-time interaction data with the dynamic interest evolution model, and identifying recommendation candidates and opportunities to obtain recommendation contexts; and mining a resource library based on a matching result, identifying a deep recommendation strategy and an adjustment signal, and tracing the strategy to identify a core factor and an optimization direction, thereby realizing accurate and adaptive personalized recommendation.
Owner:ANHUI RUIXUAN SUPPLY CHAIN TECH CO LTD

Ultra low friction gestural interface for artificial reality

Aspects of the present disclosure are directed to gesture-based user interfaces (UIs) for artificial reality (XR) messaging applications. By supplementing or replacing “gaze to tap” user interfaces with “ultra low friction” (ULF) gestures, a user is not required to repeatedly remove his focus from the real world to gaze at menu options in order to select them. The ULF gestures can include, for example, a single pinch motion to tell a messaging application to start recording a voice message. Releasing the pinch stops the recording and allows for editing, while a snap (or tug right) can stop the recording and send the message immediately. A tug left can delete the message unsent. Adding these ULF gestures to the messaging application's UI allows the user to fully engage with the application while maintaining visual focus on the real world, thus encouraging the user to remain connected through the XR system.
Owner:META PLATFORMS TECHNOLOGIES LLC

Course video key frame intelligent identification method based on AI visual attention mechanism

The application relates to the technical field of video recognition, and discloses a course video key frame intelligent recognition method based on an AI visual attention mechanism, which comprises the following steps: acquiring a visual saliency feature map of a video frame and calculating a global attention gravity center coordinate, constructing a spatial second moment tensor by using the visual saliency feature map to determine an anisotropy coefficient, then performing nonlinear weighted processing on the trajectory distribution density in a space-time trajectory space, calculating a second acceleration residual based on the processed trajectory evolution process, and determining a key frame in combination with the trajectory distribution density and the second acceleration residual. The application uses an anisotropy regulation mechanism to suppress non-content dynamic interference, checks the integrity of teaching content generation through the second acceleration residual, solves the problem of lag in semantic turning point capture under a dynamic background, and enhances the semantic density of extracted sequences and the recognition stability.
Owner:HUNAN YUNPAN NETWORK TECH CO LTD

A steel plate surface defect recognition method fusing visual attention mechanism

The application provides a steel plate surface defect recognition method fusing visual attention mechanism, and belongs to the technical field of defect detection based on computer vision; multi-source images of a steel plate, production line process and quality detection data are synchronously collected to construct a standardized tensor benchmark with aligned physical attributes and unified data modalities. A visual attention feature coding network fusing physical information constraints is built, and PINN constraint correction is used to remove the interference of ambient light brightness gain, so that the intrinsic reflection attribute features of defects are accurately extracted. Through spatiotemporal collaborative alignment and correlation modeling of defects and process heterogeneous data, the contribution weight of process parameters to defects is quantified by adopting double-track machine learning, and a process target evaluation baseline is built. Finally, the defect topology reasoning and multi-task prediction are completed by relying on the graph attention network GAT, the local visual features and global process constraints are deeply fused, and the defect position mask and category recognition result are output; the application significantly improves the steel plate quality inspection accuracy and system decision reliability.
Owner:RIZHAO YULAN NEW MATERIAL CO LTD

Visual attention tracking using gaze and visual content analysis

A method for detecting content of interest to a user includes obtaining a first data stream indicative of eye movement and / or gaze direction of the user as the user is viewing a scene in a field of view of the user, obtaining a second data stream indicative of visual content in the field of view of the user, determining, based on the first data stream and the second data stream, that content of interest to the user is present in the scene in the field of view of the user, and, in response to determining that content of interest to the user is present in the scene in the field of view of the user, triggering, with the processor, an operation to be performed with respect to the scene in the field of view of the user.
Owner:THE RGT UNIV OF MICHIGAN

Children visual attention abnormal screening method, involves extracting eye movement, facial expression and head movement multi-mode characteristic of children based on multimodal data learning

The present invention relates to a method for screening mobile terminal visual attention abnormalities in children based on multimodal data learning. A calibration video and a testing video are set up, and a head-face video of children while watching the calibration video and the testing video on smartphones is recorded, respectively. An eye-tracking estimation model is constructed to predict the fixation point location from the head-face video corresponding to the testing video frame by frame and to extract the eye-tracking features. Facial expression features and head posture features are extracted. A Long Short-Term Memory (LSTM) network is used to fuse different modal features and realize the mapping from multimodal features to category labels. In the testing stage, the head-face video of children to be classified while watching the videos on smartphones is recorded, and the features are extracted and input into the post-training model to determine whether they are abnormal.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Deep reinforcement learning target positioning method and device, equipment and storage medium

The invention discloses a deep reinforcement learning target positioning method, and the method comprises the steps: carrying out the multi-scale feature extraction and attention weighted fusion of an original image containing rain and fog through an FFA-Net model, outputting a denoised clear image, calculating the weight of particles through the edge intensity based on the clear image through a particle filtering algorithm, and carrying out the weighted average, thereby obtaining a target positioning result. Screening a high-weight pixel region in combination with a historical state and a visual attention mechanism, completing target focusing, carrying out region coarse adjustment under the dominance of a deep Q network until conditions are met, outputting a coarse positioning region, repairing missing features of the region, inputting the repaired region features into a regression network to predict a coordinate offset, and completing target focusing. And outputting a final high-precision positioning result. According to the deep reinforcement learning target positioning method provided by the embodiment of the invention, the problems of imaging blur, feature deficiency, positioning deviation and the like caused by temperature can be solved.
Owner:HUBEI SANJIANG AEROSPACE WANFENG TECH DEV

Revision apparatus, recording medium, and design revision method

A design revision apparatus includes a hardware processor that: applies, to an input image, attentional property evaluation that evaluates a portion easily attracting visual attention; performs impression evaluation that estimates an impression given by the input image to an observer; and combines a result of the attentional property evaluation with a result of the impression evaluation and presents a design revision proposal about an attentional property or an impression of the input image.
Owner:KONICA MINOLTA INC

Multi-modal large language model visual attention guiding method

The invention relates to a multi-mode large language model visual attention guiding method. The method comprises the following steps: screening attention heads responsible for capturing effective visual information in a multi-mode large language model; establishing a mapping relationship between the model output answer and the original image region corresponding to the effective attention head; calculating an attention ratio of the general problem to the specific problem, and screening a visual attention region strongly related to the model input problem to obtain a relative attention graph; performing comprehensive scoring on each middle layer according to the focusing degree and the certainty of the relative attention maps, and fusing the relative attention maps generated by a set number of middle layers with the highest scores; and constructing a binary mask based on the fused attention map to carry out binary segmentation on the model input image to obtain a processed image with a reserved high attention area, and inputting the processed image into the multi-modal large language model for reasoning. According to the method, the model can be guided to more accurately pay attention to the image region related to the problem, so that the visual alignment capability and the output accuracy are remarkably improved.
Owner:TONGJI UNIV

Implantable brain-ear collaborative interface microsystem, implantable component and hearing enhancement method

The invention provides an implantable brain-ear collaborative interface microsystem, an implantable component and a hearing enhancement method, the system comprises the implantable component and an in-vitro component, the implantable component collects a neural signal indicating a visual attention direction in a brain and receives an auditory stimulation instruction from the in-vitro component, and the auditory stimulation instruction is sent to the in-vitro component. The pulse electrical stimulation signal is converted into a pulse electrical stimulation signal to stimulate a cochlea; the in-vitro assembly collects a plurality of sound source signals in an environment where a user is located, receives a neural signal sent by the implantable assembly, determines a target visual direction according to the neural signal, enhances a sound source in a corresponding direction, encodes an enhanced audio into an auditory stimulation instruction, and returns the auditory stimulation instruction to the implantable assembly. According to the method and the device, a technical path for guiding auditory signal processing according to visual attention related neural signals is realized, so that the generation of auditory stimulation is associated with a spatial direction concerned by a user.
Owner:SHANGHAI JIAOTONG UNIV

Collision failure detection method and related device

The invention discloses a collision failure detection method and a related device. The method comprises the following steps: extracting visual features to be detected of each frame of to-be-detected image in multiple frames of to-be-detected images of a to-be-detected collision video; calculating feature differences among the plurality of to-be-detected visual features corresponding to the plurality of frames of to-be-detected images to obtain to-be-detected inter-frame differences corresponding to the plurality of frames of to-be-detected images; inputting the inter-frame difference to be detected and the plurality of visual features to be detected into a visual attention layer in a collision detection model to carry out feature transformation based on visual attention, and outputting a plurality of attention features to be detected; inputting the plurality of attention features to be detected into a collision detection layer in a collision detection model for collision detection, and outputting a plurality of collision detection results; and when the plurality of collision detection categories represented by the plurality of collision detection results comprise a collision failure category, determining that the to-be-detected collision video comprises a collision failure image. According to the method, collision failure detection is more automatic and convenient, and detection errors are reduced, so that collision failure detection is efficiently and accurately realized.
Owner:SHENZHEN WANGYU COMPUTER NETWORK CO LTD

An autism child attention prediction method based on a Mamba-Unet structure

The application discloses an autism children attention prediction method based on a Mamba-Unet structure, and comprises the following steps: firstly, a multi-modal data set for autism children attention prediction is constructed; secondly, an autism children attention prediction model based on the Mamba-Unet structure is constructed and trained; and finally, a non-typical visual attention prediction is performed on a to-be-detected image by using the trained model to obtain a prediction result. The model comprises a pretreatment module, an encoder part, a bottleneck part, a decoder part and an output part which are connected in sequence, wherein an efficient adaptive visual state space block, a frequency domain attention module, a conditional visual state space block and a conditional attention fusion module are arranged; and the output of the encoder part is further transmitted to the decoder part through a jump connection. The method can significantly improve the accuracy of autism spectrum disorder non-typical visual attention prediction.
Owner:JIANGXI NORMAL UNIV

An AI emotional interaction guiding system for children's picture book reading

The application provides an AI emotional interaction guiding system for children's picture book reading, and relates to the technical field of data processing, which comprises: a part for generating the basic visual composition and frame layout of a picture book page according to an AI picture book strategy parameter set; a part for determining the position, angle and form adjustment rules of each component of the basic visual composition based on the emotional guiding parameters in the AI picture book strategy parameter set; a part for generating the picture book visual content by moving the position, rotating the angle and adapting the form of the basic visual composition and frame layout according to the form adjustment rules; a part for collecting the dynamic visual attention data stream of the reader in real time during the presentation process of the picture book visual content, converting the dynamic visual attention data stream into updated cognitive distribution characteristic values, feeding back the updated cognitive distribution characteristic values to an emotional state analysis model, dynamically adjusting the generation of subsequent picture book visual content, and realizing emotional interaction guiding. The application effectively improves the reading concentration and reading experience of children.
Owner:XIAMEN SANDU EDUCATION TECH CO LTD

Method for hallucination detection and suppression of multi-modal large model based on attention time difference

The application discloses a hallucination detection and suppression method based on attention timing difference of a multimodal large model, which comprises the following steps: inputting text word elements of a prompt word and visual word elements of an image into a multimodal large model to obtain an internal attention graph in a decoding stage of the multimodal large model; then, calculating an attention proportion of the visual word elements in a current generation moment and performing difference with the proportion in a previous moment; if the difference exceeds a set threshold, it is determined that the visual word elements are related; when the visual related word elements are identified, a second forward propagation with visual enhancement is performed to obtain more accurate output; if not, the next generation is directly entered. According to the application, the visual related word elements in text generation are identified and refined through attention timing difference, twice forward propagation is adopted, the visual attention of the second forward propagation is enhanced based on the visual attention graph of the first forward propagation, and hallucination can be identified and corrected without reducing the language expression ability.
Owner:HANGZHOU DIANZI UNIV

Method, system, terminal and medium for constructing visual attention prediction model

The application provides a construction method of a visual attention prediction model, and faces an autism population, and comprises the following steps: constructing a visual attention prediction model based on atypical salient region enhancement; pre-training the visual attention prediction model based on atypical salient region enhancement by using a known eye movement data set, and correcting the model by using an eye movement data set of an autism population, so as to complete end-to-end training of the visual attention prediction model based on atypical salient region enhancement; testing the trained visual attention prediction model based on atypical salient region enhancement by using test images in the known eye movement data set, and constructing a final visual attention prediction model. Meanwhile, a corresponding construction system, an application method, a terminal and a medium are provided. The application is started from the special visual preference of autism patients, has the characteristics of high prediction efficiency, low cost, easy implementation, and very flexible deployment, and the like.
Owner:SHANGHAI UNIV

A traffic car customer service marketing method and system based on a large language model

ActiveCN120851956Baccurate perceptionaccurate quantitative analysisInput/output for user-computer interactionBiological modelsPersonalizationData set
The application relates to the technical field of intelligent interaction systems and automobile customer service marketing, and discloses a traffic automobile customer service marketing method and system based on a large language model, wherein the traffic automobile customer service marketing method based on the large language model comprises the following steps: collecting driver fixation point data by using an eye movement tracking algorithm to generate eye movement trajectory data sets; constructing a visual attention heat map to form a user attention distribution model; calculating multi-medium characteristic parameters to establish a medium characteristic model; using the large language model to generate and adapt personalized marketing content and output multi-medium compatible marketing information; realizing intelligent attention guidance according to the user attention distribution model, the medium characteristic model and the multi-medium compatible marketing information; and realizing accurate perception and quantitative analysis of the visual attention of the driver through the eye movement tracking technology, so that the marketing system can master the user focus points in real time, and the technical problem that a traditional interface cannot perceive actual user focus points is solved.
Owner:CHINACHEM PUHUI (CHANGCHUN) DATA SERVICE CO LTD

Course video key frame intelligent identification method based on AI visual attention mechanism

The invention relates to the technical field of video recognition, and discloses a course video key frame intelligent recognition method based on an AI visual attention mechanism, which comprises the following steps: acquiring a visual saliency feature map of a video frame, calculating a global attention barycentric coordinate, constructing a spatial second moment tensor by using the visual saliency feature map to determine an anisotropy coefficient, and calculating an anisotropy coefficient of the video frame; the method comprises the following steps of: performing nonlinear weighting processing on a trajectory distribution density in a space-time trajectory space, calculating a second-order acceleration residual error based on a processed trajectory evolution process, and determining a key frame by combining the trajectory distribution density and the second-order acceleration residual error. The integrity of teaching content generation is verified through the second-order acceleration residual error, the problem of semantic turning capture lagging under the dynamic background is solved, and the semantic density of the extracted sequence and the recognition stability of the extracted sequence are enhanced.
Owner:HUNAN YUNPAN NETWORK TECH CO LTD

A deep learning-based stereo matching method and system

The application discloses a kind of stereomatching method and system based on deep learning, comprising: build manufacturing scene data acquisition platform, synchronously collect stereoscopic image pair and real depth information, construct the training dataset containing left view, right view and real parallax graph;The stereomatching neural network model is constructed, which includes unary feature extraction, cost volume construction, context aggregation and parallax prediction module;Image is processed in turn through each module, extracts features, constructs cost volume, aggregates context information, calculates predicted parallax;Model is trained using training dataset and multi-task joint loss function, and is deployed on acquisition platform after reaching standard, processes real-time image and outputs parallax graph.The method introduces the visual attention module VABlock architecture of stacking, improves feature expression capability, combined with double-path aggregation and other modules, enhances the matching accuracy and robustness in complex scene, can realize manufacturing scene real-time accurate three-dimensional perception, provides data basis for subsequent tasks.
Owner:HUNAN UNIV

An interface determination method and device, a storage medium and an electronic device

PendingCN122285167AImprove adaptation efficiencyImaging processingSimulation
This application discloses an interface determination method, device, storage medium, and electronic device, relating to the technical fields of image processing, artificial intelligence, and mobile application interface adaptation. It simulates the user's natural visual scanning process in a right-to-left reading mode by constructing a visual attention energy field. During this process, a directional cognitive flow model is used to process the directional cognitive flow, obtaining the reading path and operation decision path formed under the current interface. Based on the modeling of a preset interface structure model, spatial relationships, and potential paths, the current interface is matched and determined to be consistent with preset interface habits. This allows for the determination of the actual user experience of the current interface in real-world use. Furthermore, this solution does not rely on testers with language and cultural backgrounds for manual experience and repeated verification during the matching and determination of preset language interface habits, thereby improving the efficiency of verifying Arabic interface adaptation.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD