Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

114 results about "Visual space" patented technology

Visual space is the experience of space by an aware observer. It is the subjective counterpart of the space of physical objects. There is a long history in philosophy, and later psychology of writings describing visual space, and its relationship to the space of physical objects. A partial list would include René Descartes, Immanuel Kant, Hermann von Helmholtz, William James, to name just a few.

Zero sample anomaly detection method and device and electronic equipment

The invention provides a zero sample anomaly detection method and device and electronic equipment, and relates to the technical field of image anomaly detection.The method comprises the steps that a to-be-detected image is acquired, and visual features are extracted through a CLIP model; performing visual enhancement on the visual features to obtain enhanced visual features; injecting the enhanced visual features into a learnable text prompt template to generate an adaptive text prompt; injecting the adaptive text prompt into a text encoder for encoding to obtain a text embedding representation; mapping the adaptive text prompt to a visual space to obtain a visual prompt, and inputting the visual prompt into the local visual features to obtain scale visual features; and based on the text embedding representation and the scale visual features, separately calculating an anomaly score and an anomaly positioning map of a preset scale, and fusing to obtain a detection result. According to the zero sample anomaly detection method and device and the electronic equipment provided by the invention, the generalization ability of the whole detection process is effectively improved, and the actual application requirements are further met.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Human motion posture recognition method and system based on multi-modal data fusion

The invention discloses a human motion posture recognition method and system based on multi-modal data fusion, and relates to the technical field of posture recognition, and the method comprises the steps: firstly, synchronously obtaining a human motion posture video frame and an IMU segment, then carrying out the feature extraction of the two kinds of heterogeneous data in a shunt parallel mode, and carrying out the feature extraction of the two kinds of heterogeneous data; and respectively capturing visual space attitude information and dynamic inertial characteristics of the IMU. Afterwards, specific features of the modals are uniformly expressed through a mixed Token and embedding mechanism, and a space-inertia collaborative attention fusion mechanism is further introduced to realize dynamic association and deep fusion of cross-modal information; and finally, classifying the multi-modal fusion representation vector obtained by fusion so as to realize accurate recognition of the human motion posture. In this way, the defect that traditional fusion is insufficient in capturing subtle action differences can be overcome, and the accuracy and stability of action recognition in a complex scene are improved.
Owner:ZHEJIANG FUBAO INTELLIGENT TECH CO LTD

Defect image enhancement method integrating reasoning and generation

The invention belongs to the technical field of electrical equipment detection, and discloses a defect image enhancement method fusing reasoning and generation, which integrates visible light, infrared and laser radar data through a multi-modal feature fusion network, breaks through the limitation that a contrast file CN114281093A only depends on a visible light image, and improves the detection accuracy. The dynamic attention mechanism can flexibly deploy visual, spatial and semantic feature weights according to defect types, key features can still be captured in complex environments such as strong light and shielding, and meanwhile, the spatial form of the defects is analyzed by means of three-dimensional point cloud; by means of the design, missing detection caused by insufficient characteristics of tiny parts such as hardware fittings and pins is effectively avoided. Aiming at the problem of distortion of a sample generated by a traditional data enhancement method in a comparison file, the sample quality is guaranteed through double mechanisms of reasoning constraint and physical verification, defect features output by a reasoning model directly constrain feature distribution of the generated sample, and meanwhile, a material mechanics rule is introduced to verify the physical rationality of the generated sample.
Owner:STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST

Robot control method and device, control equipment and storage medium

The invention provides a robot control method and device, control equipment and a storage medium, and relates to the technical field of intelligent control, and the method comprises the steps: obtaining an input task command for a target robot, and obtaining a static scene image of a task scene collected by the target robot based on the task command; according to the static scene image and the task command, adopting a preset visual space reasoning model to generate a task action sequence corresponding to the task command and object space state information corresponding to each action in the task action sequence; according to the task action sequence and the object space state information corresponding to each action, adopting a preset task generator to generate an action sequence instruction corresponding to the task command; and according to each action instruction in the action sequence instruction, sequentially controlling the target robot to execute corresponding actions. The control efficiency and the control accuracy of the robot are improved, so that the control requirement in a complex scene is met.
Owner:BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

PDF drawing data extraction method and system based on intelligent identification

The invention relates to the field of drawing recognition, in particular to a PDF drawing data extraction method and system based on intelligent recognition. Comprising the following steps: reading an internal structure of a PDF engineering drawing to obtain a native text stream, a vector path and a grating image; identifying the native text flow through a shunt preprocessing framework to form structured text data; rendering the vector path and the grating image to obtain a background image; analyzing the structured text data by utilizing the intelligent recognition model through the character recognition and extraction sub-model, obtaining drawing metadata and recording the position, and obtaining a character recognition result; analyzing the background image through a graphic element recognition and classification sub-model, recognizing and classifying component elements, and obtaining a graphic recognition result; and performing fusion according to the visual space corresponding relation to form a drawing analysis result. According to the method, the adaptive capacity of engineering drawings with various sources and different qualities is improved through the shunting preprocessing framework and the intelligent identification model.
Owner:TAIZHOU HUAWEI INFORMATION TECH CO LTD

Visual-semantic collaborative awareness method and system for automatic driving vehicle

The invention relates to a vision-semantic collaborative perception method and system for an automatic driving vehicle. Comprising a visual-semantic feature coding module and a visual-semantic feature fusion module. The visual-semantic feature encoding module comprises a text semantic encoder and a visual space-time encoder; a visual space-time encoder extracts single-frame semantic and structural features based on a 2D visual encoding model, and a 3D visual encoding model is adopted for modeling to form 3D frame-level fusion visual perception features; and the visual-semantic feature fusion module comprises a time converter, a cross converter and a context enrichment converter, and is used for adaptively adjusting a fusion weight based on a dynamic weight fusion strategy to generate required visual-semantic feature fusion information. According to the invention, the multi-modal weight is automatically adjusted, the environment perception error and decision delay are reduced, the adaptability and response precision of the automatic driving vehicle to the trunk line logistics complex traffic scene are improved, and the safe and stable operation of the intelligent driving system in the complex road environment is ensured.
Owner:NANJING UNIV OF SCI & TECH

Multi-modal human body activity identification method and device based on WIFI and video

The invention relates to a multi-mode human body activity identification method based on WIFI and videos, and the method comprises the steps: S1A, processing WIFI data, and obtaining WIFI space features; s1B, extracting human skeleton information in the video image to obtain visual features; s2A, performing spatial processing on the visual spatial features and the WIFI spatial features to obtain fused spatial features; s2B, performing time processing on the visual spatial features and the WIFI spatial features to obtain fusion time features; and S3, analyzing the fusion space feature and the fusion time feature to obtain a human body activity type. The multi-mode human body activity identification method based on the WIFI and the video has the advantages of being real-time and low in power consumption.
Owner:INNER MONGOLIA UNIV OF SCI & TECH

Mathematical formula identification coding method

The invention discloses a mathematical formula identification coding method, and particularly relates to the technical field of formula coding. The method comprises the following steps: acquiring mathematical formula image data to be identified, and performing symbol boundary extraction and preprocessing to generate symbol feature expression data; and performing visual spatial layout analysis and symbol type semantic classification based on the symbol feature expression data to generate formula layout structure data and symbol semantic classification data. And through graph structure analysis based on a topological relation, determining an inter-symbol topological relation of the mathematical formula, and generating symbol topological relation data. And deriving a dimension constraint relationship and an operator dependency relationship between symbols through a mathematical meta-knowledge mining technology, and generating mathematical meta-knowledge constraint data. Initial mathematical formula structure expression data is generated through decoding of structure and semantic constraint fusion, and a formula coding sequence is generated through formula consistency verification and semantic constraint reconstruction. According to the invention, the accuracy and efficiency of mathematical formula identification can be effectively improved.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

SAM-based mirror surface detection method

The invention discloses an SAM-based mirror surface detection method, which constructs a MirrorSAM based on segmentation of all models and is specially used for mirror surface detection. Specifically, due to the fact that reflections generated by a mirror at different positions are different, and a complex visual space can interfere with positioning, HMDE (Hierarchical Mixture of Direct Experts) is designed in a low-rank space, so that bias of SAM on an entity is reduced, and experts are dynamically adjusted according to input. Meanwhile, it is observed that a depth difference exists between a mirror surface and an adjacent area, depth mark calibration (DTC) is provided, learnable depth marks are introduced to generate a depth map and serve as error correction factors, in addition, selective pixel-prototype contrast (SPPC) loss is further formulated, and the error correction factors are used as error correction factors. Partially confusable samples are selected to facilitate decoupling of specular and non-specular representations. Wide experiments on four mirror reference data sets and three settings show that the method exceeds the existing most advanced method under the condition of a small number of trainable parameters and FLOPs.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Character interaction detection method based on spatial fine-grained context interaction feature fusion

The invention relates to the technical field of computer vision, and discloses a figure interaction detection method based on spatial fine-grained context interaction feature fusion, which comprises the following steps of: firstly, performing target detection on an input image to obtain a target detection result set and a person-object pairing feature; performing gridding projection on the image to obtain image global features, and inputting the image global features into a spatial fine-grained feature learning module to obtain spatial fine-grained features; inputting the spatial fine-grained features and the human-object pairing features into a spatial context interaction feature fusion module to obtain spatial context interaction features, and inputting the spatial context interaction features and the spatial fine-grained features into a visual encoder to obtain enhanced human-object pairing features; obtaining an interaction category score according to the feature and a text embedding feature of a character interaction category; and iteratively optimizing the character interaction detection model until convergence. According to the method, the description of local details of a visual space is enhanced, the context relationship between a person and an object is enhanced, and the accuracy of person interaction detection is improved.
Owner:HANGZHOU DIANZI UNIV

Urban landscape style assessment method and system based on three-dimensional reconstruction

The invention relates to the technical field of urban planning and spatial information, in particular to an urban landscape style assessment method and system based on three-dimensional reconstruction, and the method comprises the following steps: S1, generating a live-action three-dimensional model of a target urban area; s2, dividing the target urban area into a plurality of space grid units on a horizontal plane; s3, extracting three-dimensional visual feature parameters in each vision field range; s4, calculating a visual space oppression index according to the geometrical relationship of the buildings in the visual field range; s5, calculating a landscape style and feature comprehensive evaluation value of the space grid unit; and S6, mapping the comprehensive evaluation value of the landscape style to the geographic position of the landscape style, and generating an urban landscape style evaluation thermodynamic diagram. According to the method, through three-dimensional modeling, visual feature extraction and comprehensive quantitative evaluation, three-dimensional, standardized and visual evaluation of urban landscape styles and features is realized, and scientific support is provided for urban planning and styles and features management.
Owner:NANJING FORESTRY UNIV

Low-cost robot imitation learning method and system based on human video

The invention discloses a low-cost robot imitation learning method and system based on a human video. The method comprises the following steps: S1, data acquisition; s2, data extraction and physical alignment are carried out to eliminate man-machine physical differences; the step is divided into two parallel processing modules of action space alignment and visual space alignment; s3, data set construction: mixing the aligned human data with real robot teleoperation data, carrying out balanced sampling, and constructing a mixed data set Dmix; and S4, cooperative training: constructing a strategy network based on diffusion Transform for training. According to the method, data can be acquired only through the monocular RGB camera, expensive robot teleoperation data are replaced with cheap and easily available human videos, and the data acquisition threshold is greatly reduced. Through a visual alignment strategy of random color grid rendering, a network can learn neglect skin color textures and pay attention to geometric structures without a complex generative model, so that the robot can be seamlessly migrated to robots in different forms.
Owner:RENMIN UNIVERSITY OF CHINA

Vehicle-mounted audio augmented reality system and method based on virtual-real fusion space anchoring

PendingCN121957331AEliminate fragmentationShorten emergency response timeInput/output for user-computer interactionSound input/outputVisual spaceSound sources
The invention discloses a vehicle-mounted audio augmented reality system and method based on virtual-real fusion space anchoring, and relates to the technical field of intelligent cabin man-machine interaction, and the system comprises a controller, an augmented reality display device, a distributed loudspeaker array, and a sensor group. The controller receives the virtual image pixel coordinates, the eyeball position coordinates and the vehicle state semantic identifier, and stores a cabin three-dimensional digital model. The system converts a dynamic pixel coordinate into a three-dimensional virtual sound source coordinate by utilizing perspective inverse projection through vision-space mapping logic; the retrieval model anchors the semantic identifier to the physical component coordinates through the semantic-space mapping logic. The audio rendering module calculates a speaker drive gain based on the virtual sound source coordinates to synthesize a virtual sound source. By constructing a unified cabin three-dimensional digital model and a double-channel mapping mechanism, spatial alignment of an audio sound image, an AR visual track and a physical part of a vehicle body is realized, and audio-visual perception splitting is eliminated.
Owner:CHINA FAW CO LTD

Photovoltaic power station fire hazard monitoring system based on computer vision

The invention relates to the technical field of photovoltaic safety monitoring, and discloses a photovoltaic power station fire hazard monitoring system based on computer vision. The system comprises an image acquisition unit, a visual feature extraction unit, a feature fusion unit, a hidden danger weight distribution unit and a hidden danger tracing unit. The image acquisition unit acquires real-time visible light image data and infrared thermal imaging data; a visual feature extraction unit extracts a visual space feature vector and a heat distribution feature vector; the feature fusion unit performs multi-modal fusion on the two types of vectors to generate a fusion monitoring feature set; the hidden danger weight distribution unit calculates a hidden danger association degree score set through a pre-trained convolutional neural network model, wherein the set comprises association strength of abnormal heat distribution and visual abnormality of different regions; and the hidden danger tracing unit generates a fire hidden danger positioning result based on association intensity distribution joint positioning. The system can comprehensively capture the state of the photovoltaic power station, improves the accuracy and comprehensiveness of fire hazard monitoring, and assists the safe operation of the photovoltaic power station.
Owner:GUANGZHOU DEV ZONE YUEDIAN NEW ENERGY CO LTD

Laboratory safety assessment method based on machine vision

PendingCN121121639ACharacter and pattern recognitionVisual spaceVirtual lab
The invention discloses a laboratory safety assessment method based on machine vision, and relates to the technical field of laboratory safety monitoring, and the method comprises the steps: carrying out the standard processing of collected laboratory comprehensive data, obtaining a standard monitoring image, carrying out the feature element setting of the standard monitoring image, and obtaining a feature recognition matrix; constructing a virtual visual space, performing supplementary positioning on the virtual visual space to obtain a virtual laboratory model and an image acquisition node, and performing feature acquisition on the standard monitoring image through a feature recognition matrix based on the virtual visual space to obtain a monitoring feature picture; performing periodic comparison on the monitoring feature pictures to obtain a difference comparison result, and performing safety judgment on the laboratory according to the difference comparison result to obtain an anomaly evaluation result; laboratory safety is greatly improved, accident loss is reduced, and the safety management level of a laboratory is improved.
Owner:SHENZHEN TUOYOU SOFTWARE TECH CO LTD

Using affordance plans for robot control

Implementations for robot control are provided. A method involves, based on vision data depicting an environment of a robot and a natural language instruction for the robot, determining an affordance plan for performing a task. The affordance plan comprises a sequence of intermediate representations of the robot in visual space, such as end effector poses. An action input prompt is assembled with data indicative of the vision data, the natural language instruction, and the affordance plan. The action input prompt is processed using one or more generative models to generate action output indicative of one or more actions to be performed by the robot. Subsequently, a robot control signal is generated based on the one or more actions. This provides a spatially precise and dimensionally concise form of guidance for robot manipulation tasks, which can improve performance and generalization.
Owner:GDM HOLDING LLC

A double-branch zero-shot remote sensing scene classification method based on a knowledge graph

The application discloses a kind of double-branch zero sample remote sensing scene classification method based on knowledge graph, it is related to scene identification technical field, by analyzing the spatial distribution information of local area ground object composition, automatically constructs a kind of " scene-landscape-ground object " three-level knowledge graph;On this basis, the double-branch supervision zero sample remote sensing scene classification network DBSS designed from global and local two branches respectively to the mapping of semantic vector to visual space is supervised, to promote visual space can fully reflect the correlation structure contained in semantic space;Through a large number of experiments on UCM, AID and NWPU data set, it shows that the class average accuracy and overall accuracy of AKG-DBSS for the classification of invisible class scene can reach 98% and 59.56% respectively, the standard deviation is less than 6.91%, significantly better than other prior art methods, with good scalability.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Image processing method, apparatus, and storage medium

The embodiments of the present disclosure provide an image processing method, apparatus, device and a storage medium. The method includes: obtaining a target object detected in a current key frame and a rendered virtual object corresponding to the current key frame, wherein the key frame is a frame which triggers object detection; determining, based on the detected target object and a vision space queue, a virtual object to be newly added; determining, based on the rendered virtual object and the vision space queue, a virtual object to be deleted; and updating, based on the virtual object to be newly added and the virtual object to be deleted, a virtual object corresponding to the current key frame.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Intraoperative gauze counting and tracking method based on visual perception

The invention provides an intraoperative gauze counting and tracking method based on visual perception, and relates to the technical field of visual pattern recognition. A collaborative perception framework integrating multispectral physical fingerprint recognition, visual space-time tracking and adaptive entangled state filtering is constructed; the core problem of target identity confirmation and persistent tracking in the prior art is solved. High-robustness fluorescent fingerprints are introduced to serve as identity anchor points, deep coupling and mutual correction of physical identities and spatial-temporal trajectories are achieved through an adaptive fusion algorithm, and finally absolute identity confirmation and high-precision and high-robustness continuous tracking of each target are achieved. According to the method, the recognition accuracy in a seriously polluted and shielded environment is improved, error accumulation in a long-term tracking process is effectively inhibited, and the self-adaptive capability and reliability of the whole system in a dynamic complex scene are enhanced.
Owner:THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL

A multi-scale image matching method and system

The present invention relates to a multi-scale image matching method and system. The method comprises the following steps: S1, obtaining a data set for image matching training; S2, constructing an image matching model, which includes a visual space cue extraction module, a visual space fusion module and a visual space latent map transformation module. The visual space cue extraction module extracts visual cues and spatial cues from the input image pair. The visual space fusion module fuses the extracted visual cues and spatial cues into the same space. The visual space latent map transformation module extracts local features from the visual space fusion features by constructing a feature map and fuses the local features with global features, so as to effectively capture multi-scale context information corresponding to the visual space; training the image matching model through the data set; S3, performing image matching on the image pair to be matched through the trained image matching model. The method and system can improve the speed and accuracy of image matching.
Owner:XIAMEN UNIV OF TECH

Electric fire hazard dynamic monitoring and alarm system based on big data analysis

This invention discloses a dynamic monitoring and alarm system for electrical fire hazards based on big data analysis, belonging to the field of electrical safety monitoring and intelligent early warning technology. It includes modules for visual space modeling and interference prediction, temporal thermal anomaly extraction, structural contour recognition and consistency judgment, temperature disturbance assessment and credibility grading, composite credibility modeling and hotspot correction, and dynamic scoring and identification model update. The visual space modeling and interference prediction module acquires the field-of-view projection model of the thermal imaging monitoring area and establishes a high-risk pixel distribution model for reflection interference based on camera installation parameters and spatial geometric features. This invention constructs an intelligent identification closed-loop system from perception to decision-making, integrating spatial modeling, thermal behavior analysis, and dynamic credibility assessment to achieve accurate identification and continuous adjustment of false hotspots, effectively avoiding false alarms and missed alarms, and significantly improving the accuracy and system stability of electrical fire early warning.
Owner:杭州天卓网络有限公司

Spatial position instruction fine tuning method based on multi-modal large language model

The invention relates to a spatial position instruction fine tuning method based on a multi-modal large language model, and the method comprises the following steps: S1, converting a spatial position reasoning data set into a visual instruction format through employing a dialogue template, and obtaining a visual spatial position reasoning data set; s2, acquiring a large language model InternVL as a multi-modal large language model, performing pre-training on the general data set to obtain a pre-training model, reasoning the data set based on the visual spatial position, adjusting parameters of the pre-training model by adopting a low-rank adaptation method to obtain a trained large language model, and outputting a description corresponding to a spatial task by the large language model; and S3, introducing a text-based large language model, and optimizing the description corresponding to the space task based on the large language model. Compared with the prior art, the method has the advantages that the ability of the multi-modal large language model in understanding and generating context rich description is fully utilized, and the ability of the model in generating accurate and detailed description is enhanced.
Owner:SHANGHAI JIAOTONG UNIV

A belt deviation monitoring method based on visual space mapping

The application provides a belt deviation monitoring method based on visual space mapping, and belongs to the technical field of belt deviation monitoring, solves the problem that the existing belt deviation monitoring technology is complex in calculation and cannot directly reflect the specific deviation of the belt, the belt target detection module is used to convert a video stream into an image frame, and the belt in the image frame is detected to draw the contour of the belt, the coordinate mapping module and the center point calculation module are used to calculate the corresponding coordinates of the world coordinate system of the center of the belt in the image frame, the angle correction module is used to detect the change of the angle of the belt, the center coordinates of the world coordinate system are corrected according to the change angle of the belt, and the specific deviation of the belt is obtained by using the deviation calculation module. The mapping relationship between the image coordinates and the space coordinates is used, the specific deviation of the belt is obtained by using the matrix calculation mode, the fuzzy classification of the monitoring result is abandoned, and the deviation of the belt in the working process is monitored in real time.
Owner:JINCHUAN GROUP NICKEL COBALT CO LTD

Self-adaptive braking energy recovery control method based on multi-mode sensing fusion

The invention discloses a self-adaptive braking energy recovery control method based on multi-mode perception fusion, and belongs to the technical field of new energy automobile control. The method comprises the following steps: synchronously acquiring RGB image flow of a front-view camera of a vehicle and dynamics state data of a CAN bus of a chassis, and performing space-time alignment by adopting an improved linear interpolation method; respectively extracting visual spatial features and dynamic time sequence features through an improved ResNet-18 network and a time domain convolutional network; performing cross attention operation to generate fusion features by taking the dynamic features as Query and the visual features as Key / Value; inputting the fusion features into a multi-task prediction head, and synchronously outputting a braking condition category and a regenerative braking distribution coefficient lambda; and the vehicle control unit calculates target electro-hydraulic braking force according to the total braking torque of the driver and lambda, and issues and executes the target electro-hydraulic braking force after ABS activation and battery SOC and temperature safety rule correction. The braking intention recognition precision and the energy recovery efficiency are improved at the same time under the complex working condition.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Sign language word formation quality measurement method, device and equipment

The invention relates to the technical field of data processing, in particular to a sign language word formation quality measurement method, device and equipment, and the method comprises the steps: obtaining an RGB video set of sign language actions, including sign language videos and corresponding meaning tags; extracting visual features and semantic features of corresponding meaning labels; based on the visual features, calculating a first Grubrum matrix of the visual feature matrix; on the basis of the semantic features, calculating a second Gramer matrix of the semantic features; obtaining a visual matrix based on the first Grubrum matrix and the n-dimensional visual space; obtaining a semantic matrix based on the second Grubrum matrix and the n-dimensional semantic space; converting the visual matrix into an n-dimensional hypergeometric space to obtain a first representation of the visual features in the n-dimensional hypergeometric space; and converting the semantic matrix into an n-dimensional hypergeometric space to obtain a second representation of the semantic features in the n-dimensional hypergeometric space, and determining a sign language word formation quality measurement result based on the first representation and the second representation, thereby solving the problem that two heterogeneous features cannot be directly calculated.
Owner:LESHAN NORMAL UNIV

Defect image enhancement method fusing reasoning and generation

The application belongs to the technical field of power equipment detection, and discloses a defect image enhancement method fusing reasoning and generation, which integrates visible light, infrared and laser radar data through a multi-modal feature fusion network, breaks through the limitation of relying only on visible light images in the contrast file CN114281093A, and dynamically adjusts the weights of visual, spatial and semantic features according to the defect type, so that key features can still be captured in complex environments such as strong light and shielding, and the spatial form of the defect is analyzed with the help of three-dimensional point cloud; this design effectively avoids the missed detection of small components such as hardware pins due to insufficient features; in view of the problem that the generated samples are distorted in the traditional data enhancement method in the contrast file, the sample quality is guaranteed through a double mechanism of "reasoning constraint + physical verification", the defect features output by the reasoning model directly constrain the feature distribution of the generated samples, and the physical rationality of the generated samples is verified by introducing the law of material mechanics.
Owner:STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST

Multi-modal space-time alignment safe driving emotion recognition method based on sensitive word guidance

The invention discloses a multi-modal space-time alignment safe driving emotion recognition method based on sensitive word guidance. Sensitive word features are extracted through driving monitoring videos and voice data. A plurality of candidate time intervals in the video where emotions may be present are determined by analyzing the temporal attention of the sensitive word-speech text. And adaptively determining a key time frame with a driving emotion by means of sensitive word-visual time cross attention. In a key time frame, component features are extracted in a multi-space region of a face component, an alignment relationship between the component features and voice sensitive words is analyzed, and a key space region with driving emotion is adaptively determined by means of sensitive word-visual space cross attention. According to the method, the space-time asynchronous condition of the sensitive words in the multi-modal data is fully considered, the candidate time interval, the key time frame and the key space area are sequentially analyzed, the key clues of the driving emotion are gradually searched, and the accuracy of safe driving emotion recognition can be effectively improved.
Owner:HEFEI UNIV OF TECH

A method for implementing visual space ability evaluation based on non-immersive virtual reality

The application provides a method for evaluating visual space ability based on non-immersive virtual reality. The method is used for a computer device, and the method comprises the following steps: displaying a path learning interface through a display screen, so that a subject performs path learning; the path learning interface displays a two-dimensional plane map; the map has a specified path marked with a starting point and an ending point; a virtual three-dimensional space is displayed through the display screen, so that the subject performs a three-dimensional virtual space walking test; when the subject performs the three-dimensional virtual space walking test, the subject starts from the starting point and walks along the street to the ending point according to the specified path remembered by the subject, or returns to the starting point from the ending point; the eye movement track and the mouse action of the subject are used to calculate an evaluation index of the subject; and the visual space ability of the subject is evaluated according to the evaluation index. The detection paradigm of the application can evaluate the visual space ability of subjects of different ages or different cognitive function levels, and realizes a visual space test close to a real scene.
Owner:CHIMEDICAL UNIVERSITY

Video fusion method and system based on shadow map, and program product

The invention relates to the technical field of video fusion, and discloses a video fusion method and system based on a shadow map and a program product, and the method comprises the following steps: creating a virtual perspective camera in a three-dimensional scene, and generating the shadow map according to a visual cone of the virtual perspective camera; pixels in the three-dimensional scene are converted from a visual space coordinate system of a physical world camera to a virtual perspective camera coordinate system, NDC coordinates of the pixels are obtained, the NDC coordinates are aligned with texture coordinates of the shadow map, and sampling texture coordinates are obtained; judging whether the pixels in the visual range of the virtual perspective camera are in the visual area of the shadow map or not, and obtaining the color of the current pixel of the three-dimensional scene; and obtaining the color of the current pixel of the fused three-dimensional scene. According to the invention, the problems of shielding area processing errors and the like in the prior art are solved.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD