Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

238 results about "Visual task" patented technology

Visual Task Board is a graphic-rich environment. It transforms the navigation list & forms in an interactive way. It will allow users to view, update multiple tasks. An activity stream will display recent activity. so users can see the changes in tasks. Users can also add a task in it. It is also possible to edit and update the tasks directly.

Visual algorithm self-training method based on multi-agent collaborative optimization

The invention discloses a visual algorithm self-training method based on multi-agent collaborative optimization, and the method comprises the following steps: constructing a multi-agent system architecture which comprises a user interaction layer, an intelligent scheduling layer, an A2A protocol communication layer and a professional agent cluster layer; the user interaction layer analyzes a user task intention and generates an execution plan; the scheduling agent calls the professional agent to complete data processing, model construction, training, testing and deployment; a task process is coordinated through a standardized communication mechanism, and task execution is supported by combining an MCP tool set, a knowledge base module and a memory system; and when the task fails, automatically executing rescheduling operation, and finally outputting a self-training result. According to the method, the development efficiency, the self-adaptability and the intelligent level are remarkably improved, and the method is suitable for computer vision tasks such as industrial detection, intelligent security and protection and automatic driving.
Owner:ANHUI HEQING INTELLIGENT ROBOT CO LTD

Resource and task aware visual processing edge adaptive decision-making method

The invention belongs to the technical field of artificial intelligence and computer vision, particularly relates to a visual processing edge adaptive decision-making method for resource and task perception, and aims to solve the problem of scheduling mismatch caused by resource dynamic change and task demand diversity in visual task processing in an edge computing environment. The method comprises the following steps: collecting multi-dimensional resource state data of edge nodes in real time to form a resource state vector with high time resolution; analyzing the visual task request, and constructing a quantifiable task feature vector; and establishing a resource-task association mapping model based on a dynamic weight distribution mechanism. The method also supports cross-edge domain collaborative decision, and processes a pipeline dynamic reconstruction and security isolation mechanism. According to the technical scheme, the fluctuation of the resource utilization rate is reduced to 15% or below, the average task processing delay is reduced to 60%, the scheduling satisfaction degree is improved by 40% or above, and the self-adaptability and the service quality guarantee capability of the edge vision system are remarkably enhanced.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Visual large model Token adaptive optimization method, system and device based on differential evolution and medium

The invention discloses a visual large model Token adaptive optimization method, system and device based on differential evolution and a medium, and the method comprises the steps: carrying out the data processing of an image classification data set, an instance segmentation data set and a saliency target detection data set, and obtaining all Tokens corresponding to each image through a Patch Embedding and position coding method; obtaining a plurality of groups of Tokens corresponding to each image through a random selection mode, and performing data processing to output all Tokens corresponding to each image and the plurality of groups of Tokens selected from each image; constructing a Token adaptive selection module, a self-attention optimization module and a downstream task output module; a complete Token adaptive optimization visual large model is constructed; training a reconstruction model and a complete Token self-adaptive optimized visual large model; performing model reasoning to obtain an image classification result, an instance segmentation result image and a saliency target detection result image; the system, the equipment and the medium are used for implementing the method. The method can be widely applied to various visual tasks such as image classification, instance segmentation and saliency target detection.
Owner:XIDIAN UNIV +1

Continuous attention nerve feedback training method and system based on brain-computer interface

The invention discloses a continuous attention neural feedback training method and system based on a brain-computer interface, and relates to the technical field of neural feedback, and the method comprises the steps: collecting a multi-channel electroencephalogram signal of a user in visual task training in real time; extracting power spectral density characteristics of the multi-channel electroencephalogram signals in a beta frequency band, classifying the power spectral density characteristics by adopting a support vector machine algorithm, and outputting a judgment result of an alert or non-alert state; and according to a judgment result, dynamically adjusting an information fusion proportion alpha value in the visual task through a reward-punishment mechanism, updating image information feedback in the visual task in real time, and adjusting the attention state of the user through an image information feedback result. Neural feedback and a dynamic reward and punishment system are fused, real-time excitation feedback is obtained by autonomously adjusting electroencephalogram activity, the problem of insufficient training power caused by traditional static tasks or single positive feedback is solved, and the long-term training effect is enhanced.
Owner:XI AN JIAOTONG UNIV

Positive feedback self-supervised learning method based on characterization self-distribution and clustering attributes

The invention discloses a positive feedback self-supervised learning method based on characterization of self-distribution and clustering attributes, and belongs to the field of computer vision, and the method specifically comprises the steps: firstly, building a joint embedded architecture JEA composed of a data enhancement module, an encoder, a projection head and an optional prediction head; then, an input set D of given images randomly samples a group of images and inputs the images into a data enhancement module to obtain enhancement results, then the enhancement results respectively pass through an encoder to obtain corresponding encoding results, and then the enhancement results respectively pass through a projection head to obtain corresponding projection results; for the online network, based on an output result of the projection head, calculating an output result after passing through the prediction head; calculating and optimizing a cross entropy loss function on the basis of an output result of each module of the joint embedded architecture JEA so as to update parameters of the JEA, and finally performing performance evaluation on the trained model on a computer vision task; according to the method, inherent clustering attributes of model coding are utilized, and the model performance is further improved through positive feedback learning.
Owner:BEIHANG UNIV

YOLO model parameter efficient fine tuning method based on low-rank self-adaption

The invention discloses a YOLO model parameter efficient fine tuning method based on low-rank self-adaption. Aiming at the problems that a traditional full-parameter fine tuning method is high in calculation cost and large in storage pressure and an existing efficient parameter fine tuning method is insufficient in adaptation in a visual task, a low-rank adaptive layer is introduced into a convolutional layer of a YOLO model, a weight update quantity is decomposed into a low-rank matrix product, and a feature fusion process is optimized in combination with a lightweight adapter. In specific implementation, customized improvement is performed on a core module of the YOLO model, including multi-scale feature fusion optimization, attention weight dynamic adjustment and the like, and meanwhile, pre-training parameters are frozen and only a low-rank matrix is updated during training. The parameter efficiency and the detection performance are effectively balanced under the condition that only a small number of parameters are finely adjusted, and the method is suitable for edge device deployment and real-time detection scenes.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Dynamic self-adaptive edge server visual task processing method and system

The invention discloses a dynamic self-adaptive edge server visual task processing method and system, and belongs to the technical field of visual task processing, and the method comprises the steps: constructing an edge load model and an energy risk model, outputting an edge load rate and an energy risk coefficient, and generating a local execution confidence coefficient based on the edge load rate and the energy risk coefficient; generating a cloud execution confidence coefficient in combination with the network index; quantizing the data quality and the environment state through the data evaluation model and the data acquisition environment model; fusing the data timeliness factors to construct a local-data adaptation degree and a cloud-data adaptation degree; local / cloud execution is decided according to the decision model, and the resolution is optimized based on the execution adaptation degree. According to the method, resource dynamic adaptation and task quality optimization are realized, and the edge visual processing efficiency is remarkably improved.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Asymmetric scattering myopia control lens based on visual behaviors of human eyes

The invention relates to an asymmetric scattering myopia control lens based on human eye visual behaviors, which is applied to the technical field of eye vision optics, and has the design core that an asymmetric visual behavior mode shown by human eyes in an actual visual task is deeply integrated into the optical design of the lens. Therefore, an asymmetric optical structure with scattering intensity distribution highly matched with individual visual habits is constructed. According to the lens, scattering characteristics can be accurately configured in a specific area of the lens according to a preset human eye visual behavior data model, especially in the lower half part of the lens, namely a visual field area corresponding to near-distance tasks such as near-distance reading, writing or electronic equipment use of a wearer; the scattering intensity is higher than that of the upper half part (corresponding to a far task view field area), so that the myopia control efficiency is remarkably improved, and meanwhile, the visual comfort and definition of a wearer are effectively considered.
Owner:南通诺瞳奕目医疗科技有限公司 +1

Deep learning image compression method and system based on semantic discriminator

The invention discloses a deep learning image compression method and system based on a semantic discriminator, and mainly solves the problem that the visual task performance of a downstream machine is remarkably reduced due to serious semantic information loss under a high compression rate in the conventional image compression method. According to the implementation scheme, a group of images are selected from an existing image data set and are divided into a training set, a verification set and a test set, and the training set, the verification set and the test set are preprocessed respectively; constructing an image compression network comprising an image codec, a semantic extraction network and a semantic guide discriminator under a Pytorch framework; inputting the training set into an image compression network, carrying out two-stage iterative training, and verifying through a verification set; and inputting the test set into the trained image compression network, and only calling the image codec to output the compressed reconstructed image. According to the method, the structural integrity and semantic consistency of the reconstructed image are remarkably improved, higher visual quality and task accuracy can be kept at a low code rate, better image compression performance is embodied, and the method can be used for efficient transmission and storage of image data.
Owner:XIDIAN UNIV

Single scanning Mama feature extraction method and system for image global modeling

The invention provides an image global modeling-oriented single-scan Mama feature extraction method and system, and the method comprises the steps: obtaining an input image, and extracting a multi-channel feature map with a channel dimension and a spatial dimension through convolution operation; state space modeling of single-time forward calculation is carried out on the multi-channel feature map so as to model the global dependency relationship between channels, and the state space modeling is achieved by carrying out matrix operation on features obtained after space dimension flattening and a channel interaction matrix built based on state space model parameters; and outputting the feature map after global dependency modeling, and applying the feature map to a downstream vision task. According to the method, redundant calculation caused by multiple times of space scanning is eliminated, the model structure is simplified, and the calculation complexity and reasoning delay are remarkably reduced while the downstream task performance such as image classification is ensured.
Owner:HANGZHOU DIANZI UNIV

Pseudo-supervision diffusion type underwater image enhancement method based on cross-modal semantic guidance

The invention provides a pseudo-supervised diffusion type underwater image enhancement method based on cross-modal semantic guidance, and the method comprises the following steps: obtaining an original degraded image and an actual reference picture of a UIEB, inputting the original degraded image into a multi-modal large model LLaVA to generate a semantic prompt, and selecting a plurality of different existing underwater enhancement methods, the method comprises the following steps: respectively obtaining a plurality of groups of enhanced images, integrating the enhanced images and semantic prompts into image-text pairs, inputting the image-text pairs into a Zip-CLIP model, calculating the similarity of the image-text pairs to generate a plurality of groups of similarity heat maps, inputting the similarity heat maps and high-order semantic information into a special convolutional neural network architecture, generating a pseudo tag as a weak supervision signal for diffusion model training, and carrying out diffusion model training. And a final enhancement result is obtained. According to the method, through the characteristics of cross-modal semantic alignment and gradual reconstruction of the diffusion model, the image detail reduction and structure retention capability is remarkably improved, a more reliable visual basis is provided for an underwater visual task, and theoretical innovation and practical values are both achieved.
Owner:SICHUAN POLICE COLLEGE +1

Turbulence recovery method combining super-resolution and multi-scale network

The invention discloses a turbulence restoration method combining super-resolution and a multi-scale network, belongs to the field of image restoration, and is suitable for remote sensing image restoration under the influence of atmospheric turbulence. The method comprises the following steps: constructing a training set and a test set; a multi-scale attention module is introduced into the generator, and shallow image features and deep image features under different scales are fully extracted; performing super-resolution up-sampling on the deep features to recover the spatial resolution, and fusing the deep features with the shallow features of the corresponding scales to generate a restored image; a multi-scale discriminator is adopted to carry out quality evaluation on the generated image, the image passing the evaluation is directly output, and the image not passing the evaluation is fed back to a generator by the discriminator to carry out reconstruction optimization; and finally outputting a restored image and obtaining a trained network model. According to the method, the spatial resolution and the structure restoration quality of the image can be effectively improved under the influence of atmospheric turbulence, and the identifiability of the image and the robustness of a subsequent visual task are enhanced.
Owner:CHANGCHUN UNIV OF SCI & TECH

OpenVX framework system for NPU acceleration

The invention provides an OpenVX framework system for NPU acceleration, which comprises an application layer, an OpenVX framework layer, an NPU runtime layer and an NPU hardware layer, and is characterized in that the OpenVX framework layer comprises an NPU perception graph optimizer. According to the invention, a set of OpenVX framework for completing the visual task on the NPU acceleration chip is designed, the calculation graph of the visual task constructed by the standard OpenVX API can be optimized according to the hardware characteristics of the NPU, and the memory management of the hardware layer of the NPU is designed, so that the visual task can be completed on the NPU in an accelerated manner, and the calculation efficiency of the visual task is improved.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Video processing method, device, system, equipment, storage medium and program product

The invention discloses a video processing method, device, system and equipment, a storage medium and a program product. The method comprises the following steps: obtaining a reconstructed key frame and a second code stream based on a first code stream of a key frame of a target video; the first code stream is obtained by encoding a video encoding model which does not conform to a video encoding standard; converting a third code stream of a non-key frame of the target video into a fourth code stream; the third code stream is obtained by referring to the reconstructed key frame by an encoder following the video coding standard; fusing the first code stream and the fourth code stream to obtain a first target code stream; and fusing the second code stream and the third code stream to obtain a second target code stream. The first target code stream contains important features in the original frame, so that the requirement of a visual task for the features can be met, the performance of the visual task is enhanced, the second target code stream contains the performance of AI coding in the aspect of image details, and the compatibility and the stability of traditional coding are utilized, so that the visual experience of human eyes can be met, and the image quality is improved. And decoding and displaying are carried out on different hardware.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Systems and methods for assessing and mitigating accommodation spasm

A user's accommodative spasm can be evaluated and mitigated via a virtual reality (VR) system, which can include a VR headset in electronic communication with a computing device. The computing device causes virtual environments, which can include objects, optotypes, various lighting conditions, and various weather conditions, to be displayed on the VR headset. Using varying combinations of eye-tracking sensors, eye-tracking cameras, motion-tracking sensors, handheld devices, and microphones, the VR headset collects data about the user as she focuses on various objects placed at different distances away from her in the virtual environments. Optionally, advanced algorithms in the computing device dynamically alter the virtual environments and the visual tasks and analyze the user's responses to evaluate the user's accommodative spasm.
Owner:ZENNI OPTICAL

Self-adaptive brightness adjusting method and system for mobile phone backlight plate

The invention relates to the technical field of mobile equipment display, in particular to a self-adaptive brightness adjusting method and system for a mobile phone backlight plate. The method comprises the following steps: acquiring an environment light and shadow dynamic information set and user behavior intention information, and analyzing cross-modal fusion deduction information of light and shadow artistic features and user task intention based on the environment light and shadow dynamic information set in combination with the user behavior intention information to obtain a light and shadow-user situation information set; based on the light and shadow-user situation information set, analyzing a dynamic tension relationship among an art immersion demand, visual task efficiency and visual health, and balancing an active light and shadow fusion strategy, user instantaneous discomfort and long-term eye movement load to obtain a situation backlight fusion strategy set; and based on the contextualized backlight fusion strategy set, flexible intervention matched with the artistic atmosphere and the instantaneous demand of the user is executed on the mobile phone backlight. The smooth backlight change is ensured, the matching with the artistic atmosphere and the instantaneous demand of the user is realized, and the comfort and satisfaction of the user are improved.
Owner:JIANGXI LANHAOHONG TECHNOLOGY CO LTD

Task orchestration methods, robot control methods, and robot control systems

This application provides a task scheduling method, a robot control method, and a robot control system. Tasks are scheduled within a control interface that displays task tracks, and each robot is controlled to execute tasks based on these tracks. The visual task tracks provide a clear view of the task execution sequence and its correspondence with the robots, allowing operators to quickly predict spatiotemporal conflicts during task execution. This also prevents resource imbalances caused by mismatches between robot functionalities and tasks, ensuring efficient collaboration among the robots.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Dark light field remote sensing image enhancement method and device based on data driving and medium

The invention discloses a dim light field remote sensing image enhancement method and device based on data driving and a medium, and the method comprises the steps: S1, constructing an optical remote sensing image enhancement data set in a dim light field, and obtaining a collected to-be-enhanced remote sensing image; s2, constructing a decomposition network based on a Retinex theory, and training the decomposition network by using the enhanced data set to obtain an image decomposition model; inputting the to-be-enhanced remote sensing image into a trained image decomposition model, and performing decomposition to obtain a reflection feature map and an illumination feature map; s3, constructing a feature enhancement module based on frequency domain Fourier convolution, inputting the reflection feature map and the illumination feature map into the feature enhancement module, and performing frequency domain and space domain combined feature extraction and enhancement to obtain an enhanced feature map; and S4, outputting the enhanced remote sensing image. The method can improve the brightness, contrast and details of the image, effectively suppress the noise, keep the color natural, and improve the downstream visual task performance.
Owner:BEIJING INSTITUTE OF TECHNOLOGY (ZHUHAI)

Attitude motion and visual training combined motion visual evaluation system and method

PendingCN121943190ASensorsDiagnostic recording/measuringBiomechanicsVisual adaptation
The invention discloses a motion visual evaluation system and method combining attitude motion and visual training, and relates to the technical field of digital medical treatment and visual health, and the system comprises a biomechanics-attitude acquisition module which integrates a three-dimensional force sensor, a high-precision joint angle sensor, an inertial measurement unit sensor group and a depth camera; synchronously collecting joint torque, ground reaction force, muscle power output, joint angle and posture offset; through a mechanics-vision correlation model of the data processing module, how the biomechanical characteristics influence the visual response and the feedback effect of the visual task on the biomechanical state can be deeply excavated, the improvement of the posture and the visual ability is concerned, the training strategy can be optimized from the biomechanical level, and the training efficiency is improved. And a key basis is provided for the evaluation feedback module to output a multi-dimensional report containing collaborative rating, biomechanical optimization direction and visual adaptation suggestions.
Owner:TSINGHUA UNIVERSITY

A mirror highlight detection and removal method based on a double-flow convolutional neural network

ActiveCN115311157BImage enhancementImage analysisSpecular highlightComputer vision
The application discloses a mirror highlight detection and removal method based on a double-flow convolutional neural network, which calculates the gradient of a pixel point in an x direction and a y direction of an input original image with highlights, extracts a first highlight feature mapping, subtracts the first highlight feature mapping after being processed by a first convolution block attention module CBAM from a preprocessed image, and outputs a first-stage highlight-free image; the first highlight feature mapping is progressively down-sampled and reduced in size, each highlight feature mapping after being down-sampled and reduced in size is processed by a convolution block attention module CBAM at each stage, and is subtracted from a highlight-free image output by a previous stage, and finally a rough highlight-free image is obtained. Then, highlight extraction and refinement are performed on the rough highlight-free image by a highlight extraction module to obtain a final highlight-free image. The application can effectively solve the image information degradation problem caused by the mirror highlight, thereby reducing the interference of the highlight on visual tasks such as target detection.
Owner:ZHEJIANG UNIV OF TECH

A method and apparatus for generalized supervision representation learning

ActiveCN115272806BData packData set
The application discloses a kind of pan supervision representation learning method and device, the method includes: obtaining training data;Wherein, training data includes image data and the label information corresponding to image data;Image data corresponding first spatial feature is obtained by the feature extraction of visual network model to the training data input, and second spatial feature is obtained based on first spatial feature mapping, and third spatial feature is obtained based on second spatial feature mapping;First loss function value of first spatial feature and second spatial feature is calculated, and the second loss function value of third spatial feature is calculated according to label information;The parameter of the visual network model is updated based on first loss function value and second loss function value, to obtain the visual network model after training.The application can make that the representation learned not only can obtain superior performance on training data set, but also can obtain good transfer performance on other visual tasks, can realize image detection and segmentation.
Owner:TSINGHUA UNIVERSITY

Agricultural greenhouse automatic dimming image acquisition method based on visual task feedback

The invention discloses an agricultural greenhouse automatic dimming image acquisition method based on visual task feedback, and relates to the field of agricultural data acquisition, and the method comprises the steps: collecting a crop image through a depth camera, carrying out the target detection, extracting an ROI region of a crop, and carrying out the target detection based on an RGB image and a depth image in a detection frame of the crop; the method comprises the steps of calculating a comprehensive image quality index used for representing the overall texture richness and the overall detection reliability of each crop, and then dynamically adjusting the driving current of an LED lamp panel based on the comprehensive image quality index so as to dynamically adjust light. Even if the illumination conditions of different areas of the agricultural greenhouse change and the crops have the phenomena of strong reflection, local shielding and the like, closed-loop dimming can be carried out based on the local image quality of the crops, so that the quality of the collected crop images is improved, and the accuracy and reliability of image analysis are improved.
Owner:GUOCHUANG WISDOM (JIANGSU) AGRICULTURAL ROBOT CO LTD

Machine vision task processing method and apparatus, device, and storage medium

The application discloses a machine vision task processing method and device, equipment and a storage medium, and relates to the technical field of machine vision, and is used for at least timely outputting a machine vision task processing result, and avoiding the problem that processing speed fluctuation in a previous task processing process easily leads to the problem that a beat cannot be met. The method comprises the following steps: processing at least two machine vision tasks according to a preset beat, wherein the at least two machine vision tasks comprise a first task and a second task executed in sequence; wherein, in response to the fact that the execution time length of the first task reaches a preset time length, a task result of the first task is sent; and the preset time length is less than or equal to the interval time length of the preset beat.
Owner:HANGZHOU HIKROBOT TECH CO LTD

A visual token pruning method based on graph information propagation

The application provides a visual token pruning method based on graph information propagation, comprising the following steps: visual extraction is performed on an input image to obtain visual tokens; importance scores of the visual tokens are initialized; a graph structure about the visual tokens is constructed, each visual token is taken as a node, an adjacency matrix is calculated to construct connections between the visual tokens, and the graph structure is initialized; the adjacency matrix is updated through a preset similarity threshold to obtain visual token subgraph structures of different regions; each row of the adjacency matrix is normalized, node information is iteratively propagated, and final scores of each visual token are calculated; k visual tokens with the highest scores are selected according to the final scores of the visual tokens and are projected; the visual tokens obtained through the projection are spliced with text tokens, input into a large language model, and output results are obtained. The method can improve the calculation efficiency of the model, significantly reduces the calculation cost while maintaining the performance of the visual task.
Owner:XIAMEN UNIV

Image optimization method and device, equipment and medium

The invention relates to the technical field of image processing, and discloses an image optimization method and device, equipment and a medium, and the method comprises the steps: recognizing the degradation information of a target image, carrying out the image transformation processing of the target image according to the image degradation information and a preset image transformation process, and obtaining an optimized image, performing multi-type visual task detection on the optimized image, generating a quantized image quality index, judging whether the image reaches a quality standard by comparing a detection parameter with a preset threshold value, if not, calculating a reward signal of a reinforcement learning agent model according to a detection result, and updating an image transformation process by using the signal, so as to improve the quality of the image. And then returning to the optimization step to process the image again, if the current strategy reaches the standard, indicating that the current strategy is effective, directly processing the subsequent to-be-processed image by the agent by using the updated strategy, and finally outputting a high-quality target optimized image. According to the invention, the efficiency and precision of image processing are improved.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

Low-illumination image enhancement method based on wavelet driving

The invention discloses a low-illumination image enhancement method based on wavelet driving, and aims to solve the problems of insufficient local and global feature interaction, single frequency domain feature description view angle and lack of prior feature guidance in the existing low-illumination image enhancement technology. According to the method, a progressive frequency domain multi-view feature collaborative sensing network (WALLIE) is constructed, U-Net is taken as a basic framework, and the method comprises the following steps: decomposing input image features into multi-view frequency domain features through wavelet transform; a cooperative space-wavelet domain multi-view detail compensation (S2WD) module is used to realize spatial domain and wavelet domain feature cooperative compensation, and the local detail recovery capability is improved. According to the method, key indexes such as the PSNR and the SSIM on a plurality of public data sets are superior to those of an existing advanced method, the generalization ability is excellent, the advanced visual task performance such as downstream low-light semantic segmentation and target detection can be effectively improved, and the method has wide application prospects.
Owner:XIANGNAN UNIV

Depth completion method, system and terminal based on three-dimensional feature extraction and fusion

ActiveCN118115556BRadiologyColor map
This invention discloses a self-supervised depth completion method, system, and terminal based on 3D feature extraction and fusion. The method includes: acquiring a discrete depth map and a color map; preprocessing the discrete depth map to obtain a target discrete depth map; downsampling the color map to obtain a target color feature map; performing a preset number of downsampling and feature extraction operations on the target discrete depth map and the color map to obtain a first color feature map and a first target discrete depth feature map; performing channel concatenation and upsampling operations to obtain a fused image feature map; inputting the fused image feature map and the target discrete depth map into a cross-attention feature fusion module to output the fused feature map; and performing channel concatenation and upsampling on the target color feature map and the fused feature map to obtain a completed depth map. This invention can obtain a completed depth map after information completion, thereby enabling accurate processing of subsequent computer vision tasks.
Owner:SHENZHEN UNIV

Task processing method and related device therefor

Disclosed in the present application are a task processing method and a related device therefor, which can relatively accurately execute visual tasks. The method of the present application comprises: first, first data and second data of different modalities for a visual task can be input into a target model; next, the target model can perform feature extraction on the first data and the second data to obtain a first feature and a second feature, the first feature comprising N first sub-features, and the second feature comprising N second sub-features; then, the target model can use the i-th second sub-feature to process the i-th first sub-feature to obtain an i-th third sub-feature, such that a third feature comprising N third sub-features can be obtained; similarly, the target model can further use the i-th first sub-feature to process the i-th second sub-feature to obtain an i-th fourth sub-feature, such that a fourth feature comprising N fourth sub-features can be obtained; and finally, on the basis of the third feature and the fourth feature, the target model can acquire a processing result of the visual task.
Owner:HUAWEI TECH CO LTD

Visual task processing method, visual processing model training method and image generation method

PCT designated stageWO2026175084A1Pattern recognitionVision processing
Provided in the embodiments of the present description are a visual task processing method, a visual processing model training method and an image generation method. The visual task processing method comprises: acquiring an image to be processed and prompt text; inputting said image into a visual encoding layer of a visual processing model for encoding, so as to obtain visual input features, and inputting the prompt text into a text encoding layer of the visual processing model for encoding, so as to obtain text input features; inputting the visual input features and the text input features into a feature processing layer, and executing a target visual task, so as to obtain multi-modal output features; and inputting the multi-modal output features into a decoding layer, and on the basis of a plurality of text feature units, positioning a corresponding target image region in said image, so as to obtain a target image in which the target image region is highlighted. No additional task decoder is added for a model, a variety of visual tasks for fine-grained perception are completed, the complexity of the architecture of the model is not increased, the scalability of the model is retained, and the synergistic effect between different visual tasks is also promoted.
Owner:ALIBABA (CHINA) CO LTD

Remote sensing target detection feature extraction algorithm based on depth multi-scale

This invention proposes a feature extraction algorithm for remote sensing target detection based on deep multi-scale, comprising: S1, segmenting a downloaded public dataset; S2, constructing a deep multi-scale remote sensing target detection model; and S3, using the trained target detection algorithm model to detect the remote sensing image to be detected and generating the final detected image. Through convolutional operations of various sizes and channel attention mechanisms, features at different scales can be captured, and dynamic weighting enhances the model's focus on important features. This makes the network model more capable of handling complex visual tasks, especially when processing remote sensing images with detail and hierarchy. The proposed backbone network has been extensively tested on multiple remote sensing datasets, demonstrating excellent performance and effectively detecting various types of targets, proving its potential and feasibility in practical applications. This innovative network structure also brings new ideas to the field of remote sensing image processing.
Owner:CHINA THREE GORGES UNIV