Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Visual information processing" patented technology

Intelligent agent visual language navigation method and system based on task completion prediction

The invention provides an agent visual language navigation method and system based on task completion prediction. The method comprises the step of constructing a dual-drive structure composed of a self-adaptive mixed pooling mechanism and a task completion analysis module. Firstly, in the visual information processing process, a dynamic weight distribution strategy is adopted to carry out multi-scale adaptive mixed pooling on panoramic features, so that the fusion effect of local and global semantic information is optimized, and the retention capability and semantic integrity of navigation historical information in a dynamic topological map are improved. And then, inspired by a human navigation cognitive behavior mechanism, a task completion analysis module is designed and introduced, and the task execution progress is dynamically estimated based on the recognition condition of a key landmark in a navigation path, so that an intelligent agent is driven to preferentially select a key path node and invalid exploration is reduced. And finally, realizing efficient understanding and execution of the natural language instruction by the intelligent agent through a multi-round cyclic cross-modal reasoning and action prediction mechanism.
Owner:FUZHOU UNIV

Communication robot, communication robot control method, and program

A communication robot includes an auditory information processing portion configured to recognize a volume of voice collected by a sound collection portion and generate an auditory attention map by projecting a sound position in a three-dimensional space onto a two-dimensional attention map in which the robot is located at a center, a visual information processing portion configured to generate a visual attention map using a face detection result obtained by detecting a face of a person from an image captured by an imaging portion and a motion detection result obtained by detecting a motion of the person, an attention map generation portion configured to generate an attention map by integrating the auditory attention map and the visual attention map, and a motion processing portion configured to control eyeball movements and motions of the communication robot using the attention map.
Owner:HONDA MOTOR CO LTD

Cargo stabilizing mechanism of automatic driving transport vehicle

The utility model discloses a cargo stabilizing mechanism of an automatic driving transport vehicle, and belongs to the field of cargo transportation. The cargo stabilizing mechanism comprises a vehicle body, a controller, driving wheels fixedly installed on the vehicle body and electrically connected with the controller, a power module fixedly installed on the vehicle body and electrically connected with the controller, and a visual module fixedly installed on the vehicle body and electrically connected with the controller. The visual information processing module is fixedly installed on the vehicle body and electrically connected with the controller, and the compartment is fixedly installed on the vehicle body. According to the goods stabilizing mechanism, one pressing plate is taken, the two screws are aligned, the first pressing plate is placed in the goods storage box, then goods are placed, finally one pressing plate is taken, the screws are aligned, the goods are placed on the uppermost layer of goods, then knobs are taken out to be arranged on the screws in a sleeving mode, the uppermost layer of pressing plate is pressed, and therefore all layers of goods are limited by the pressing plates; therefore, irregular collision displacement of the goods can be effectively avoided, and the effects of protecting and stabilizing the goods are achieved.
Owner:NONGZHENG QIMIN TECH (TIANJIN) CO LTD

Pick ball automatic scoring system based on multi-view high-speed camera shooting and AI identification

The invention relates to the technical field of sports event auxiliary decision, and discloses a multi-view high-speed camera shooting and AI identification-based Pick ball automatic scoring system, which comprises a data acquisition subsystem, a data processing subsystem, a system scheduling and fusion decision subsystem and a result output module. The core of the system is a dynamic refocusing mechanism triggered by an acoustic event: the system detects key acoustic events such as ball hitting and the like in real time through a distributed microphone, and accurately calculates the three-dimensional position and timestamp of the events; and the acoustic event information is used as a trigger signal to guide the system to generate a control instruction, so that the visual information processing module is switched from a low-power-consumption conventional monitoring mode to a high-precision fine analysis mode. The system fuses the information of the acoustic mode and the visual mode, and a decision is generated according to a preset competition rule base. According to the invention, the power consumption of the system can be reduced, and the multi-modal information is used for mutual evidence, so that the accuracy and robustness of judgment are improved, and the real-time performance of processing is ensured.
Owner:GUANGZHOU UNIVERSITY

Magnetic-Sync Engine-based Digital Therapeutic Software and Driving Method Thereof for Cognitive Path Restructuring and Neural Plasticity Induction

The present invention relates to a real-time interactive cognitive guidance system and method using a Magnetic-Sync Engine (MSE) that monitors a user's cognitive state in real time and pre-projects a target trajectory to compensate for the delay in the brain's visual information processing. The MSE of the present invention calculates a virtual gravitational force corresponding to the distance between the user's input point and the target trajectory to pull the user toward the target path like a magnet, and forms a cognitive highway by projecting a visual guide 0.1 to 0.2 seconds earlier than the expected point the user will reach through pre-pulling timing control. By fusing multimodal sensors such as eye tracking and tactile input, the cognitive load is precisely calculated, and accordingly, the size of the virtual gravitational force and the shape of the trajectory can be varied in real time. Upon successful synchronization, neurofeedback is provided through solfège frequency audio and blooming animation to maximize learning efficiency, and the invention can be applied to various fields such as medical rehabilitation, cognitive training, and precision work education.
Owner:김선경

Virtual riding exercise system based on VR technology

The application discloses a virtual riding exercise system based on VR technology and relates to the technical field of virtual reality.The system comprises a riding assembly, a VR head-mounted device, a motion control card, a four-degree-of-freedom motion platform, a stepless fan and a visual information processing device, wherein the riding assembly is located on the four-degree-of-freedom motion platform, the stepless fan is located in front of the riding assembly and faces the rider, and through the communication connection relationship and the function design of the riding assembly, the stepless fan and the visual information processing device, the system can not only provide a riding audio-visual experience through the VR head-mounted device, but also can realize the purpose that the posture of the riding vehicle is adaptively adjusted along with the change of the virtual scene road condition through the motion control card and the four-degree-of-freedom motion platform, and the stepless fan can provide the rider with the wind feeling caused by the riding, so that the riding experience can be greatly enriched, the riding experience immersion can be effectively improved, and the system is convenient for actual application and popularization.
Owner:LEBAO SPORTS INTERNET (WUHAN) CO LTD

Mask art multi-dimensional visual state display method and mask art multi-dimensional visual state display system

The invention relates to the technical field of visual information processing, and particularly discloses a multi-dimensional visual state display method and system for mask art, and the method comprises the steps: carrying out the quantification of quality data, visual optimization data and terminal presentation data, and obtaining a source quality characteristic coefficient, a rendering optimization coefficient and a terminal presentation coefficient; acquiring a real-time rendering frame rate and an interaction response delay, and processing the real-time rendering frame rate and the interaction response delay based on the rendering optimization coefficient and the terminal presentation coefficient to obtain a real-time collaboration coefficient; determining the target precision of the three-dimensional model based on the real-time collaboration coefficient and the source quality characteristic coefficient; according to the method, by constructing a plurality of quantitative models such as source quality characteristics, rendering optimization, terminal presentation and fluency-real-time collaboration, accurate measurement and fusion analysis of display of full-link multi-dimensional influence factors are realized, and stable and efficient operation of the system is ensured while mask art details are fully mined and presented.
Owner:ANHUI UNIVERSITY OF ARCHITECTURE

Pesticide risk prevention and control system for rice and shrimp co-cropping in cold region

The invention relates to the technical field of agricultural pest control, in particular to a cold region rice-shrimp co-cropping pesticide risk control system, which comprises a monitoring unit for acquiring environmental parameters such as temperature, humidity, illumination, water level, soil acidity and alkalinity and the like of a rice field, shooting images of the rice field and acquiring visual information of diseases, weeds and pests. And the AI processing unit is used for receiving the environment parameters and the image information transmitted by the monitoring unit, identifying and classifying diseases, weeds and weeds through a trained deep learning model, and judging the types, the occurrence degrees and the stages of the diseases, the weeds and the weeds. And the decision-making unit is used for consulting a database according to the identification result of the AI processing unit and formulating a targeted pesticide application scheme. According to the execution unit, a communication module can send an instruction needing manual operation to related personnel, and pesticide spraying or other prevention and control operations are automatically carried out by pesticide spraying equipment according to a pesticide spraying scheme; precise prediction, identification and scientific prevention and control of diseases, weeds and weeds in the rice field are achieved, the pesticide risk is reduced, and safe production of rice and crayfish is guaranteed.
Owner:HEILONGJIANG RIVER FISHERY RES INST CHINESE ACADEMY OF FISHERIES SCI

A large multi-modal model guided adaptive image fog removal method

The application discloses a large multi-modal model guided adaptive image fog removal method, belongs to the technical field of visual information processing, and is suitable for foggy scene image processing. The application realizes the method as follows: 1, obtaining degradation priori by using a pre-trained large multi-modal model; 2, inputting a foggy image and the degradation priori into a MoE-SSM model to perform training, so as to dynamically adjust model parameters and then perform dynamic perception defogging; 3, fusing shallow feature F s and optimized features to reconstruct a clean and fog-free image; 4, training a neural network by using a clean and fog-free image and a true value clear image by using a Charbonnier loss function shown in formula (8), so as to punish the deviation of the recovered image from the true value clear image and encourage consistent image gradients; compared with the prior art, the application solves the compromise processing problem of receptive field and calculation efficiency in image processing of the prior defogging model, simultaneously reduces the requirement for a large amount of labeled data, and then improves the image processing accuracy and stability.
Owner:BEIJING INST OF TECH

Visual information generation method and system based on document semantic retrieval

The invention is suitable for the field of visual information processing, and provides a visual information generation method and system based on document semantic retrieval. According to the system, a document is obtained through a receiving module, a text ambiguity analysis module identifies a visual subject and description keywords thereof, semantic vectors are generated after ambiguity is detected and classified, a main semantic instruction set and adjustable intermediate state data are constructed through weight fusion and selection, and a first version of visual result is generated through a main vision generation module. When a user modifies a specific subject through the editing interface module, the interactive visual optimization module quickly generates and presents a local visual alternative scheme based on the adjustable intermediate state data, and a final result is exported by the output module after the user confirms the local visual alternative scheme. According to the method, semantic conflicts and fuzziness in the document are actively analyzed and managed before generation, and unadopted semantic paths are structured and archived, so that the consistency of a visual result and the intention of the document and the control force of a user on the generation process are effectively improved.
Owner:CHENGDU ENHANCED VIEW TECH CO LTD

Multi-angle crack detection device for water conservancy detection

The utility model belongs to the technical field of crack detection, and particularly relates to a multi-angle crack detection device for water conservancy detection, which comprises a base, the top of the base is fixedly connected with a visual information processor, the top of the base is fixedly connected with a supporting plate, and the left side wall of the supporting plate is provided with a sliding groove. A motor is slidably connected outside the sliding groove, the driving end of the motor is fixedly connected with a rotating rod, the outer side wall of the rotating rod is slidably sleeved with a rotating cylinder, the rotating cylinder is provided with an arc-shaped opening, the outer side wall of the rotating rod is fixedly connected with a protruding rod penetrating through the arc-shaped opening, and the outer side wall of the protruding rod makes contact with the inner side of the arc-shaped opening. The motor drives the rotating rod to move, when the rotating rod moves, the rotating rod drives the rotating cylinder to rotate through the protruding rod, the rotating cylinder rotates due to the arc-shaped opening and horizontally moves along the rotating rod, and when the rotating cylinder moves, the detection camera is placed in the water pipe, so that personnel can conveniently detect the inner wall of the water pipe.
Owner:山东省调水工程运行维护中心昌邑管理站

Document image processing method, electronic equipment and storage medium

The invention discloses a document image processing method, electronic equipment and a storage medium. The method comprises the steps that a visual information extraction task is obtained, and the visual information extraction task is used for determining a to-be-processed document image; performing multi-modal feature modeling on the document image to obtain multi-modal feature embedding; performing content adaptation on the multi-modal feature embedding to obtain target text embedding; a target recognition result is generated based on the target text embedding, and the target recognition result is used for describing the text content displayed in the document image. The technical problems of high difficulty in model training and poor visual information identification accuracy of a scheme for performing visual information processing on a document image in related technologies are solved.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Mixed linear layer enhanced image generation method and system

The invention provides a mixed linear layer enhanced image generation method and system, and the method comprises the steps: a visual information processing step: carrying out the marking processing of the obtained visual information, extracting the original features of the visual information, and outputting a noise enhanced residual visual mark; a text information processing step: obtaining weight parameters of an INR encoder through a super network; a color feature acquisition step: performing mapping processing on the coordinate information of the visual information by using a mixed linear layer and an activation function of an INR encoder, and acquiring a color feature value of a corresponding space coordinate; and an image synthesis step: fusing and inputting the color features generated in the color feature acquisition step and the noise residual error marks obtained in the visual information processing step into a multi-layer perceptron network, and synthesizing an image with high continuity. The method supports any resolution output, breaks through the limitation of traditional pixel representation, and improves the expression capability and adaptability of the model.
Owner:SHENZHEN YIDAO DIGITAL TECHNOLOGY R&D CO LTD

Method for recognizing pattern skipping rope action and intelligent judging based on visual space-time feature matching

The application provides a pattern skipping rope action recognition and intelligent judging method based on visual space-time feature matching, and belongs to the technical field of visual information processing. The first step is multi-view video data acquisition and synchronization; the second step is video preprocessing and frame extraction; the third step is skeleton key point extraction based on improved YOLOv8-pose; the fourth step is skeleton key point sequence post-processing and kinematic feature extraction; the fifth step is rope body motion trajectory extraction; the sixth step is space-time feature modeling and action sequence recognition; the seventh step is action quality evaluation and intelligent judging score; the eighth step is system deployment and real-time inference. The application effectively enhances the recognition ability of the complex action of the pattern skipping rope by fusing YOLOv8-pose and CA attention mechanism and introducing a space-time feature matching module, and has higher precision in distinguishing similar actions such as cross jumping and side swing crossing.
Owner:INNER MONGOLIA UNIV OF SCI & TECH

Hydrogen leakage fluorescence detection system and method based on unmanned aerial vehicle

The invention discloses a hydrogen leakage fluorescence detection system and method based on an unmanned aerial vehicle. The system comprises an unmanned aerial vehicle unit, a remote control unit and a fluorescence color-changing adhesive tape arranged on the surface of a pipeline. An RFID tag is arranged on the fluorescent color-changing adhesive tape; pre-stored coordinates are pre-stored in the RFID tag; the unmanned aerial vehicle unit comprises an unmanned aerial vehicle, and a laser excitation assembly, a visual sensor and a laser range finder on the unmanned aerial vehicle; the unmanned aerial vehicle is provided with a Beidou positioning module, a wireless communication module, an RFID reader-writer, a visual information processing module and a control module; the visual information processing module is used for receiving the fluorescent color-changing adhesive tape image and comparing the fluorescent color-changing adhesive tape image with a fluorescent color-changing adhesive tape image pre-stored in the visual information processing module so as to judge whether the color of the fluorescent color-changing adhesive tape is changed or not; the remote control unit is connected with the control module through the wireless communication module so as to control the unmanned aerial vehicle, obtain color changing information of the fluorescent color changing adhesive tape and issue early warning information. According to the invention, the hydrogen leakage detection efficiency, the detection precision and the detection safety can be effectively improved.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Neuromorphic vision system

A retinomorphic array is used to convert visual information into electrical signals, and the neural network performs information processing on the input electrical signals to obtain the result of visual cognition; the perception and synchronous preprocessing of visual information is achieved through the retinomorphic array, avoiding the transmission of a large number of redundant visual information from the photoreceptor end to the image information processor, saving bandwidth resources, and improving the efficiency of visual information processing; the use of the crossbar array allows the configuration of a neural network with a more complex structure and more diverse functions, and the higher-level processing of visual information by the neural network realizes a novel neuromorphic vision system integrated therein with image recognition, dynamic tracking, and trajectory prediction.
Owner:NANJING UNIV