Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

322 results about "Visual task" patented technology

Visual Task Board is a graphic-rich environment. It transforms the navigation list & forms in an interactive way. It will allow users to view, update multiple tasks. An activity stream will display recent activity. so users can see the changes in tasks. Users can also add a task in it. It is also possible to edit and update the tasks directly.

Multimodal large model task processing method, device and equipment based on reinforcement learning

The invention provides a multi-modal large model task processing method, device and equipment based on reinforcement learning, and the method comprises the steps: employing a multi-modal large model, generating G groups of responses for multi-modal visual task data, carrying out the scoring of each group of responses, obtaining a reward value, carrying out the standardization of the reward value, obtaining a dominance score, building a strategy updating gradient through the dominance score, and carrying out the calculation of the strategy updating gradient. A final result is obtained after expectation reward maximization, model parameter adjustment and multi-round iteration, and an expectation reward objective function comprises an expectation calculation item and KL divergence constraint, so that a current strategy can be updated by comparing the performance of candidate strategies through group relative strategy optimization, and local optimum can be helped to be jumped out; the strategy updating amplitude is limited by the KL divergence constraint, and the model is prevented from violently changing in the optimization process, so that the training stability is improved; the dynamic strategy iteration allows the model to adjust the balance between exploration and utilization according to the learning progress on the basis of keeping the stability, thereby further ensuring the effectiveness and stability of strategy optimization.
Owner:STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2

Visual algorithm self-training method based on multi-agent collaborative optimization

The invention discloses a visual algorithm self-training method based on multi-agent collaborative optimization, and the method comprises the following steps: constructing a multi-agent system architecture which comprises a user interaction layer, an intelligent scheduling layer, an A2A protocol communication layer and a professional agent cluster layer; the user interaction layer analyzes a user task intention and generates an execution plan; the scheduling agent calls the professional agent to complete data processing, model construction, training, testing and deployment; a task process is coordinated through a standardized communication mechanism, and task execution is supported by combining an MCP tool set, a knowledge base module and a memory system; and when the task fails, automatically executing rescheduling operation, and finally outputting a self-training result. According to the method, the development efficiency, the self-adaptability and the intelligent level are remarkably improved, and the method is suitable for computer vision tasks such as industrial detection, intelligent security and protection and automatic driving.
Owner:ANHUI HEQING INTELLIGENT ROBOT CO LTD

Visual task generation method based on Token

The invention discloses a Token-based visual task generation method, and belongs to the technical field of intelligent task automation, and the method comprises the steps: S1, cross-modal alignment; s2, performing visual Token processing; s3, constructing a task description Token sequence: constructing the task description Token sequence based on a predefined visual task template library according to requirements of a user or a specific application scene; s4, checking task feasibility; s5, task priority scheduling; s6, training a task generation model; s7, dynamic task allocation; and S8, optimizing the model. According to the method, the long sequence processing capability is optimized through a hierarchical merging strategy, the compactness of feature expression is realized while space position information is reserved, linear projection and enhanced position coding are combined to form a visual Token sequence with strong representation capability, local detail features are contained, a global context relationship is kept, and the method is suitable for the visual Token sequence with high representation capability. High-information-density feature input is provided for subsequent task processing, and the processing precision of various visual algorithms is effectively improved.
Owner:BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD

Resource and task aware visual processing edge adaptive decision-making method

The invention belongs to the technical field of artificial intelligence and computer vision, particularly relates to a visual processing edge adaptive decision-making method for resource and task perception, and aims to solve the problem of scheduling mismatch caused by resource dynamic change and task demand diversity in visual task processing in an edge computing environment. The method comprises the following steps: collecting multi-dimensional resource state data of edge nodes in real time to form a resource state vector with high time resolution; analyzing the visual task request, and constructing a quantifiable task feature vector; and establishing a resource-task association mapping model based on a dynamic weight distribution mechanism. The method also supports cross-edge domain collaborative decision, and processes a pipeline dynamic reconstruction and security isolation mechanism. According to the technical scheme, the fluctuation of the resource utilization rate is reduced to 15% or below, the average task processing delay is reduced to 60%, the scheduling satisfaction degree is improved by 40% or above, and the self-adaptability and the service quality guarantee capability of the edge vision system are remarkably enhanced.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Space-time adaptive threshold-based spiking neural network image classification method and system

The invention discloses a pulse neural network image classification method and system based on a space-time adaptive threshold, mainly solving the problems of poor nonlinear expression and limited time sequence and space feature processing ability in the prior art, and the scheme comprises the following steps: obtaining an image data set, and dividing the image data set into a training set and a test set; a spiking neural network main body structure comprising an input layer, a hidden layer and an output layer is selected, an existing neuron model is improved by introducing a space-time joint threshold adjustment mechanism, and improved neurons are placed in each neuron layer in the hidden layer to form a spiking neural network based on a space-time adaptive threshold. The training set is used to carry out iterative training; and inputting the test set into the trained pulse neural network to obtain an image classification result. According to the method, a space-time adaptive threshold mechanism is introduced, the threshold can be dynamically adjusted to adapt to time and space features, the processing capacity of the network on time sequence data and complex features and the classification accuracy of images are remarkably improved, and the method can be widely applied to dynamic visual tasks and event-driven scenes.
Owner:XIDIAN UNIV

Multi-robot collaborative indoor scene semantic segmentation method and system based on multivariate interactive learning

The invention belongs to the field of computer vision image signal processing, and relates to a multi-robot collaborative indoor scene semantic segmentation method and system based on multivariate interactive learning. The method comprises the steps of adopting an image feature encoder based on a double attention integration module to realize cross-modal feature fusion on an RGB image and a depth image; a dimension anisotropy attention enhancement module is adopted, and feature optimization is carried out through a multi-dimensional direction attention mechanism; a cross-task collaborative interaction module is adopted, and information is shared among different tasks by utilizing an interaction mechanism based on a feature level; in a multi-task heterogeneous decoding stage, each robot uses a decoder with a specific task to execute a specific task, and semantic segmentation, instance segmentation and scene classification are realized; and a joint optimization loss function is adopted to perform end-to-end joint training on multiple tasks, so that overall optimization and collaborative awareness performance improvement are realized. According to the method, the precision and robustness can be improved when complex indoor scene semantic segmentation and other visual tasks are executed.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Visual large model Token adaptive optimization method, system and device based on differential evolution and medium

The invention discloses a visual large model Token adaptive optimization method, system and device based on differential evolution and a medium, and the method comprises the steps: carrying out the data processing of an image classification data set, an instance segmentation data set and a saliency target detection data set, and obtaining all Tokens corresponding to each image through a Patch Embedding and position coding method; obtaining a plurality of groups of Tokens corresponding to each image through a random selection mode, and performing data processing to output all Tokens corresponding to each image and the plurality of groups of Tokens selected from each image; constructing a Token adaptive selection module, a self-attention optimization module and a downstream task output module; a complete Token adaptive optimization visual large model is constructed; training a reconstruction model and a complete Token self-adaptive optimized visual large model; performing model reasoning to obtain an image classification result, an instance segmentation result image and a saliency target detection result image; the system, the equipment and the medium are used for implementing the method. The method can be widely applied to various visual tasks such as image classification, instance segmentation and saliency target detection.
Owner:XIDIAN UNIV +1

Continuous attention nerve feedback training method and system based on brain-computer interface

The invention discloses a continuous attention neural feedback training method and system based on a brain-computer interface, and relates to the technical field of neural feedback, and the method comprises the steps: collecting a multi-channel electroencephalogram signal of a user in visual task training in real time; extracting power spectral density characteristics of the multi-channel electroencephalogram signals in a beta frequency band, classifying the power spectral density characteristics by adopting a support vector machine algorithm, and outputting a judgment result of an alert or non-alert state; and according to a judgment result, dynamically adjusting an information fusion proportion alpha value in the visual task through a reward-punishment mechanism, updating image information feedback in the visual task in real time, and adjusting the attention state of the user through an image information feedback result. Neural feedback and a dynamic reward and punishment system are fused, real-time excitation feedback is obtained by autonomously adjusting electroencephalogram activity, the problem of insufficient training power caused by traditional static tasks or single positive feedback is solved, and the long-term training effect is enhanced.
Owner:XI AN JIAOTONG UNIV

Vision processing and model training method, device, storage medium and program product

The present disclosure provides a vision processing and model training method, device, storage medium and program product. A specific implementation solution is as follows: establishing an image classification network with the same backbone network as the vision model, performing a self-monitoring training on the image classification network by using an unlabeled first data set; initializing a weight of a backbone network of the vision model according to a weight of a backbone network of the trained image classification network to obtain a pre-training model, the structure of the pre-training model being consistent with that of the vision model, and optimize the weight of the backbone network by using real data set in a current computer vision task scenario, so as to be more suitable for the current computer vision task; then, training the pre-training model by using a labeled second data set to obtain a trained vision model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Multi-degraded image restoration method based on frequency domain decomposition

The invention discloses a multi-degraded image recovery method based on frequency domain decomposition, and aims to solve the problems that a single model is difficult to deal with various image degradation and recovery processes of different frequency domains are mutually coupled in the prior art. According to the method, a degraded image is decomposed into a high-frequency space and a low-frequency space through fast Fourier transform, and a double-branch network architecture is adopted for targeted processing: for the high-frequency part, a high-frequency feature adaptive processing module HFPM is designed, and detail texture features are effectively extracted and interference is suppressed through feature enhancement and cross-layer fusion technologies; and for the low-frequency part, constructing a low-frequency feature conversion enhancement module LTEM, and capturing global context information by using cyclic convolution to improve the integrity of the structure contour. According to the method, decoupling processing of frequency domain features is realized, and the image restoration performance of the model in various degradation scenes such as rain removal, noise removal and defogging is remarkably improved through the synergistic effect of high-frequency detail enhancement and low-frequency structure optimization. Experimental results show that the method has excellent recovery effect and robustness when a plurality of image degradation tasks are processed at the same time, and can be effectively applied to visual tasks such as traffic accidents with high image quality requirements.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Visual task processing method and device based on multi-modal large model of low-rank adaptive enhancement

The invention provides a visual task processing method and device based on a multi-modal large model of low-rank adaptive enhancement, and relates to the technical field of data processing. The method comprises the following steps: receiving a plurality of visual task requests sent by a visual application and visual input data corresponding to each visual task request in the plurality of visual task requests; dynamically scheduling the plurality of visual task requests into a plurality of batches of processing sequences through an orchestrator, and determining a reasoning mode of a low-rank adapter used by each batch of visual input data; and inputting the visual input data of each batch and the corresponding low-rank adapter into a target multi-modal large model, and performing dynamic fragmentation calculation on the visual input data and a decomposition matrix of the corresponding low-rank adapter through an adaptive fragmentation matrix multiplication operator ATMM to obtain a response result corresponding to each visual task request. Through the method provided by the invention, diversified visual task requirements are met.
Owner:TSINGHUA UNIVERSITY

Positive feedback self-supervised learning method based on characterization self-distribution and clustering attributes

The invention discloses a positive feedback self-supervised learning method based on characterization of self-distribution and clustering attributes, and belongs to the field of computer vision, and the method specifically comprises the steps: firstly, building a joint embedded architecture JEA composed of a data enhancement module, an encoder, a projection head and an optional prediction head; then, an input set D of given images randomly samples a group of images and inputs the images into a data enhancement module to obtain enhancement results, then the enhancement results respectively pass through an encoder to obtain corresponding encoding results, and then the enhancement results respectively pass through a projection head to obtain corresponding projection results; for the online network, based on an output result of the projection head, calculating an output result after passing through the prediction head; calculating and optimizing a cross entropy loss function on the basis of an output result of each module of the joint embedded architecture JEA so as to update parameters of the JEA, and finally performing performance evaluation on the trained model on a computer vision task; according to the method, inherent clustering attributes of model coding are utilized, and the model performance is further improved through positive feedback learning.
Owner:BEIHANG UNIV

YOLO model parameter efficient fine tuning method based on low-rank self-adaption

The invention discloses a YOLO model parameter efficient fine tuning method based on low-rank self-adaption. Aiming at the problems that a traditional full-parameter fine tuning method is high in calculation cost and large in storage pressure and an existing efficient parameter fine tuning method is insufficient in adaptation in a visual task, a low-rank adaptive layer is introduced into a convolutional layer of a YOLO model, a weight update quantity is decomposed into a low-rank matrix product, and a feature fusion process is optimized in combination with a lightweight adapter. In specific implementation, customized improvement is performed on a core module of the YOLO model, including multi-scale feature fusion optimization, attention weight dynamic adjustment and the like, and meanwhile, pre-training parameters are frozen and only a low-rank matrix is updated during training. The parameter efficiency and the detection performance are effectively balanced under the condition that only a small number of parameters are finely adjusted, and the method is suitable for edge device deployment and real-time detection scenes.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

CNN-Transform-based unsupervised low-illumination image multi-degradation problem recovery method

The invention is suitable for the technical field of computer vision, and provides an unsupervised low-illumination image multi-degradation problem recovery method based on CNN-Transform, which constructs a network comprising a low-illumination image enhancement and exposure suppression module and an image denoising module. Adjusting the brightness pixel by pixel and suppressing overexposure by using a low light enhancement and exposure suppression curve; and the latter realizes noise removal through noise addition processing and a C-T module. The design comprises seven unsupervised loss functions, and training can be carried out without pairwise labeling data. The advantages of CNN local feature extraction and Transform global dependence modeling are combined, and the lightweight and real-time performance of the model are ensured. According to the method, the image brightness is effectively enhanced, overexposure is inhibited, noise is removed, a high-quality image basis is provided for the fields of automatic driving, security and protection monitoring, medical images and the like, and meanwhile advanced visual tasks such as target detection and the like are assisted.
Owner:JILIN UNIVERSITY

Dynamic self-adaptive edge server visual task processing method and system

The invention discloses a dynamic self-adaptive edge server visual task processing method and system, and belongs to the technical field of visual task processing, and the method comprises the steps: constructing an edge load model and an energy risk model, outputting an edge load rate and an energy risk coefficient, and generating a local execution confidence coefficient based on the edge load rate and the energy risk coefficient; generating a cloud execution confidence coefficient in combination with the network index; quantizing the data quality and the environment state through the data evaluation model and the data acquisition environment model; fusing the data timeliness factors to construct a local-data adaptation degree and a cloud-data adaptation degree; local / cloud execution is decided according to the decision model, and the resolution is optimized based on the execution adaptation degree. According to the method, resource dynamic adaptation and task quality optimization are realized, and the edge visual processing efficiency is remarkably improved.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Image enhancement method, system and equipment for automatic driving vision task

The invention provides an image enhancement method, system and device for an automatic driving vision task, and belongs to the technical field of combination of computer vision and automatic driving, and the method comprises the steps: obtaining a multi-view image of a to-be-detected target, and splicing adjacent view images; applying a first size grid mask on the spliced image; removing a mask in an overlapping region of the spliced image to obtain an intermediate image; detecting a preset target area in the intermediate image, and applying a second size grid mask to the preset target area; and replacing the first size grid mask at the corresponding position with the second size grid mask to obtain an enhanced image. Based on the method, the invention further provides an image enhancement system for the automatic driving vision task. According to the method, a multi-view grid mask enhancement method is adopted, the first-size mask and the second-size mask are used for carrying out global subject covering and local small target covering respectively, and the robustness and generalization ability of a visual perception model in an automatic driving system in a complex environment are improved.
Owner:ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3

Domain generalization semantic segmentation enhancement method based on efficient multi-scale and simple attention module

The invention discloses a field generalization semantic segmentation enhancement method based on an efficient multi-scale and simple attention module, and belongs to the field of computer vision. According to the invention, the model is based on a TQDM (Text Query-Driven Mask Transfer) framework, and an EMA (Efficient Multi-Scale Attention Module) and a SimAM (Simple Parameter-free Attention Module) are fused, so that the adaptability and the robustness of a semantic segmentation task on a plurality of domains are improved. According to the invention, by introducing the EMA module, multi-scale feature aggregation and cross-space information interaction are realized, so that the understanding ability of the model to a complex scene is enhanced; meanwhile, in combination with a SimAM module, feature expression is optimized under the condition of not increasing extra parameters, and the precision of small target segmentation and target boundary detection is improved. The method can be widely applied to computer vision tasks related to cross-domain semantic segmentation, such as automatic driving, intelligent monitoring and medical image analysis.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Visual task generation method based on field Token

The invention relates to the technical field of artificial intelligence image generation, in particular to a visual task generation method based on domain Token. Comprising the steps of obtaining field feature information of a target field and generating or determining at least one field lexical element; initializing an initial representation of an image to be generated; in an autoregression generation process, according to an image lexical element sequence corresponding to a previously generated image part and the at least one field lexical element, utilizing an autoregression model to predict a next image lexical element of the current position; repeating the prediction step until an image lexical element sequence representing the complete image is generated; and generating a final pixel space image by decoding based on the image lexical element sequence representing the complete image. According to the method, by introducing and fusing the field lexical elements, the field specific attributes of the generated image can be more accurately controlled, the generation fidelity and consistency of the image in the specific field are remarkably improved, and the controllability of the field style, element and structure of the generated content is improved.
Owner:BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD

Asymmetric scattering myopia control lens based on visual behaviors of human eyes

The invention relates to an asymmetric scattering myopia control lens based on human eye visual behaviors, which is applied to the technical field of eye vision optics, and has the design core that an asymmetric visual behavior mode shown by human eyes in an actual visual task is deeply integrated into the optical design of the lens. Therefore, an asymmetric optical structure with scattering intensity distribution highly matched with individual visual habits is constructed. According to the lens, scattering characteristics can be accurately configured in a specific area of the lens according to a preset human eye visual behavior data model, especially in the lower half part of the lens, namely a visual field area corresponding to near-distance tasks such as near-distance reading, writing or electronic equipment use of a wearer; the scattering intensity is higher than that of the upper half part (corresponding to a far task view field area), so that the myopia control efficiency is remarkably improved, and meanwhile, the visual comfort and definition of a wearer are effectively considered.
Owner:南通诺瞳奕目医疗科技有限公司 +1

Method and system for obtaining an optical prescription based on determined values of a fatigue parameter

A method for obtaining an optical prescription. The method includes determining a value of a fatigue parameter for at least two intermediate optical prescriptions among a plurality of intermediate optical prescriptions, the fatigue parameter associated with a visual fatigue level of a subject carrying out a visual task involving any kind of visual content, wherein each determined value of the fatigue parameter is associated with a respective one of the plurality of intermediate optical prescriptions; comparing the determined values of the fatigue parameter; and obtaining the optical prescription, based on the comparison of the determined values of the fatigue parameter. A system for obtaining an optical prescription is also disclosed.
Owner:ESSILOR INTERNATIONAL(COMPAGNIE GENERALE D OPTIQUE)

Deep learning image compression method and system based on semantic discriminator

The invention discloses a deep learning image compression method and system based on a semantic discriminator, and mainly solves the problem that the visual task performance of a downstream machine is remarkably reduced due to serious semantic information loss under a high compression rate in the conventional image compression method. According to the implementation scheme, a group of images are selected from an existing image data set and are divided into a training set, a verification set and a test set, and the training set, the verification set and the test set are preprocessed respectively; constructing an image compression network comprising an image codec, a semantic extraction network and a semantic guide discriminator under a Pytorch framework; inputting the training set into an image compression network, carrying out two-stage iterative training, and verifying through a verification set; and inputting the test set into the trained image compression network, and only calling the image codec to output the compressed reconstructed image. According to the method, the structural integrity and semantic consistency of the reconstructed image are remarkably improved, higher visual quality and task accuracy can be kept at a low code rate, better image compression performance is embodied, and the method can be used for efficient transmission and storage of image data.
Owner:XIDIAN UNIV

Visual language model prompt fine tuning method based on meta prompt

The invention discloses a visual language model prompt fine tuning method based on meta prompts, which comprises the following steps of: acquiring a data set, constructing a trainable multi-modal prompt, and generating text data comprising a spliced text prompt and image data comprising a spliced image prompt; the method comprises the following steps: constructing a pre-defined element prompt fine tuning framework A, freezing an open source model CLIP of basic model parameters, and then inserting A into a specified layer of a CLIP visual encoder to construct a multi-modal visual language model B adapted to a downstream visual task; training B by using the data set, wherein the optimization targets are cross entropy loss and diversity loss; and outputting a prediction result of a given image by using the trained multi-modal visual language model B. The method is suitable for migrating a pre-trained large visual language model to a scene on a downstream visual task, the model can be promoted to extract more discriminative visual features by utilizing the meta prompts, and the generalization performance of the model is greatly improved.
Owner:ZHEJIANG UNIV

Multi-modal enhanced representation collaborative learning pneumonia image recognition method

The invention discloses a multi-modal enhanced representation collaborative learning pneumonia image recognition method, which comprises the following steps of: aiming at chest radiograph image data, respectively extracting visual modal features and text modal features, and generating rich feature representation fused with context semantics through a multi-modal feature coding strategy; a collaborative learning mechanism is adopted, complementarity of visual and text features is combined, the model is guided to carry out feature optimization and decision reasoning, and robustness and interpretability of the model are improved; in the classification reasoning stage, multi-modal auxiliary information is utilized to refine pathological region features, and fine-grained differences are effectively captured; and through an auxiliary information constraint mechanism, the recognition capability of the model on a tiny pathological mode is enhanced, and the pneumonia classification accuracy and generalization capability are improved. The method can be widely applied to computer vision tasks in the fields of medical image auxiliary diagnosis, disease detection and the like.
Owner:HANGZHOU VOCATIONAL & TECHN COLLEGE

Single scanning Mama feature extraction method and system for image global modeling

The invention provides an image global modeling-oriented single-scan Mama feature extraction method and system, and the method comprises the steps: obtaining an input image, and extracting a multi-channel feature map with a channel dimension and a spatial dimension through convolution operation; state space modeling of single-time forward calculation is carried out on the multi-channel feature map so as to model the global dependency relationship between channels, and the state space modeling is achieved by carrying out matrix operation on features obtained after space dimension flattening and a channel interaction matrix built based on state space model parameters; and outputting the feature map after global dependency modeling, and applying the feature map to a downstream vision task. According to the method, redundant calculation caused by multiple times of space scanning is eliminated, the model structure is simplified, and the calculation complexity and reasoning delay are remarkably reduced while the downstream task performance such as image classification is ensured.
Owner:HANGZHOU DIANZI UNIV

Visual model system, visual model construction method, visual task processing method, image generation method, video generation method, information processing method based on visual model, and cloud training platform

The embodiment of the invention provides a visual model system, a visual model construction method, a visual task processing method, an image generation method, a video generation method, an information processing method based on a visual model and a cloud training platform. The visual model comprises at least two feature processing modules and at least one tuning module, the feature processing module comprises a plurality of feature processing layers, the tuning module comprises a plurality of tuning layers, and the plurality of tuning layers are respectively connected with the feature processing layers of the corresponding levels in the at least two feature processing modules; and the tuning layer is used for receiving the feature data output by the connected feature processing layer and performing tuning processing on the feature data to obtain a tuning result, and the tuning result is used for generating visual data. According to the method, detail information on multiple levels is reserved, more accurate visual data is generated, a feature processing module and an adjustment and optimization module are decoupled, the universality and generation efficiency of a visual model system are improved, the training overhead is reduced, and the training efficiency is improved.
Owner:ALIBABA (CHINA) CO LTD

Illumination environment evaluation system for simulating underground multi-scene space

The invention discloses a lighting environment evaluation system for simulating an underground working space, and relates to the technical field of lighting environments. Comprising an immersive virtual reality environment experience system which is used for experiencing different space illumination scenes by people in an immersive manner; a human body physiological index acquisition system; the human body subjective perception acquisition system is used for acquiring subjective feelings of personnel through questionnaires preset in a virtual environment; the task performance system is determined by the task collection time length of the visual task preset in the virtual environment and the answer condition, and the work efficiency performance is evaluated through the inverse efficiency score; and the comprehensive data processing system is used for receiving, sorting and analyzing the data collected by each system. According to the method, the limitation of traditional single information feedback and qualitative method evaluation is broken through, diachronic and instantaneous perception dynamic changes are considered, a quantitative evaluation system integrating subjective perception, objective physiological indexes and working efficiency performance is established based on multi-source information feedback, and the method can be applied to long-term closed space place illumination environment evaluation.
Owner:BEIJING JIAOTONG UNIV

Pseudo-supervision diffusion type underwater image enhancement method based on cross-modal semantic guidance

The invention provides a pseudo-supervised diffusion type underwater image enhancement method based on cross-modal semantic guidance, and the method comprises the following steps: obtaining an original degraded image and an actual reference picture of a UIEB, inputting the original degraded image into a multi-modal large model LLaVA to generate a semantic prompt, and selecting a plurality of different existing underwater enhancement methods, the method comprises the following steps: respectively obtaining a plurality of groups of enhanced images, integrating the enhanced images and semantic prompts into image-text pairs, inputting the image-text pairs into a Zip-CLIP model, calculating the similarity of the image-text pairs to generate a plurality of groups of similarity heat maps, inputting the similarity heat maps and high-order semantic information into a special convolutional neural network architecture, generating a pseudo tag as a weak supervision signal for diffusion model training, and carrying out diffusion model training. And a final enhancement result is obtained. According to the method, through the characteristics of cross-modal semantic alignment and gradual reconstruction of the diffusion model, the image detail reduction and structure retention capability is remarkably improved, a more reliable visual basis is provided for an underwater visual task, and theoretical innovation and practical values are both achieved.
Owner:SICHUAN POLICE COLLEGE +1

Image recognition model training method and device, electronic equipment and storage medium

The invention discloses an image recognition model training method and device, electronic equipment and a storage medium, and relates to the technical field of feature learning, and the main technical scheme comprises the steps: obtaining training question and answer data; training a visual universal model according to the training question and answer data and the training image data to obtain a prediction answer, generated by the visual universal model, of the training question and answer data; and calculating a loss function of the visual universal model according to the standard answer and the predicted answer, and performing parameter adjustment on the visual universal model according to the loss function. According to the scheme that various tasks are unified into question-answer data pairs, various visual task data are uniformly trained through a language interface, so that a new visual universal model is obtained, a network has better visual-language space alignment capability, visual information of various levels can be better processed and captured, and the visual effect is improved. And the capability and the effect of a mainstream multi-modal large language model can be effectively improved.
Owner:TSINGHUA UNIVERSITY

Turbulence recovery method combining super-resolution and multi-scale network

The invention discloses a turbulence restoration method combining super-resolution and a multi-scale network, belongs to the field of image restoration, and is suitable for remote sensing image restoration under the influence of atmospheric turbulence. The method comprises the following steps: constructing a training set and a test set; a multi-scale attention module is introduced into the generator, and shallow image features and deep image features under different scales are fully extracted; performing super-resolution up-sampling on the deep features to recover the spatial resolution, and fusing the deep features with the shallow features of the corresponding scales to generate a restored image; a multi-scale discriminator is adopted to carry out quality evaluation on the generated image, the image passing the evaluation is directly output, and the image not passing the evaluation is fed back to a generator by the discriminator to carry out reconstruction optimization; and finally outputting a restored image and obtaining a trained network model. According to the method, the spatial resolution and the structure restoration quality of the image can be effectively improved under the influence of atmospheric turbulence, and the identifiability of the image and the robustness of a subsequent visual task are enhanced.
Owner:CHANGCHUN UNIV OF SCI & TECH

OpenVX framework system for NPU acceleration

The invention provides an OpenVX framework system for NPU acceleration, which comprises an application layer, an OpenVX framework layer, an NPU runtime layer and an NPU hardware layer, and is characterized in that the OpenVX framework layer comprises an NPU perception graph optimizer. According to the invention, a set of OpenVX framework for completing the visual task on the NPU acceleration chip is designed, the calculation graph of the visual task constructed by the standard OpenVX API can be optimized according to the hardware characteristics of the NPU, and the memory management of the hardware layer of the NPU is designed, so that the visual task can be completed on the NPU in an accelerated manner, and the calculation efficiency of the visual task is improved.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD