Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3946 results about "Layers" patented technology

Layers are used in digital image editing to separate different elements of an image. A layer can be compared to a transparency on which imaging effects or images are applied and placed over or under an image. Today they are an integral feature of image editors.

Adaptive Real-Time Multi-Modal Compression System with Dynamic Resource Allocation

A system and method for adaptive real-time multi-modal compression with dynamic resource allocation provides intelligent compression optimization based on continuously monitored device conditions. The system monitors battery level, CPU utilization, and memory availability while classifying incoming multi-modal data streams comprising image, audio, text, and sensor data to determine processing priorities. Multi-objective optimization balances compression efficiency, reconstruction quality, and energy consumption using evolutionary algorithms that generate optimal parameters for an adaptive variational autoencoder. The autoencoder features dynamically selectable processing complexity, adjustable latent space dimensionality, and modality-specific processing layers. The system automatically switches between operational modes including emergency mode triggered by resource constraints, which applies maximum compression settings and intelligent data triage. Continuous learning adapts compression parameters based on observed performance outcomes, improving future optimization decisions. The system enables homomorphic operations on compressed data and provides enhanced compression performance under varying resource constraints across diverse edge computing applications.
Owner:ATOMBEAM TECH INC

Special equipment defect automatic identification method and system

The invention relates to the technical field of defect detection, in particular to a special equipment defect automatic identification method and system, and the method comprises the following steps: marking a continuous path based on an image gray scale gradient and adjacent differences, identifying the pixel connection intensity to obtain a contour image layer, extracting continuous pixels, analyzing gray scale and gradient features, and expanding a texture direction to generate a defect image block. And marking and connecting the small regions to obtain a closed layer, analyzing texture and edge differences, matching classification tags, and clustering and coding to generate a defect identification tag set. According to the method, boundary structure recognition is enhanced through gray gradient and direction continuity analysis, a texture aggregation area is expanded in combination with gray consistency and edge direction stability, brightness abrupt change points are removed, contour extraction precision is improved, micro crack continuity is recovered, and pseudo defects are eliminated; and multi-dimensional image attributes are fused to identify key region feature differences, so that the classification precision and the spatial mapping consistency are improved, and efficient and accurate defect identification and stable classification are realized.
Owner:SHUNDAAN TECHNOLOGY GROUP CO LTD

Unmanned aerial vehicle image-based small object detection method for target areas

The present invention relates to the technical field of deep learning and computer vision. Disclosed is an unmanned aerial vehicle image-based small object detection method for target areas. The present invention crops images of obvious small objects in certain target areas, and annotates the small objects of different categories to form a raw training and testing dataset, so as to ensure the accuracy of data required in the early stage of the algorithm and further ensure the scientificity of the algorithm; uses the computing capability of an improved YOLOv7 detection model to collect image features of different degrees in the dataset, the improved YOLOv7 detection model using YOLOv7 as a basic model and adding to a neck network an MS-CET module, which is constituted by an improved self-attention mechanism and convolution module SPPCSP, and a BHC-FB module, which is constituted by bidirectional mixed convolution modules NConv and RPConv connected in parallel; and finally fuses different feature layers as a final judgment basis of an unmanned aerial vehicle for small object detection in the target areas, to further check the accuracy of the algorithm and criteria for dataset selection, thereby improving recognition accuracy.
Owner:CHONGQING UNIV OF TECH

Intelligent planning method for space monitoring of unmanned aerial vehicle

The invention discloses an intelligent planning method for space monitoring of an unmanned aerial vehicle. The method comprises the steps that intelligent path allocation is realized by constructing a task demand priority matrix; the method comprises the following steps: firstly, collecting geographical, climate and environmental parameters of a monitoring area, quantifying regional complexity and color features of a monitoring target by combining high-resolution image data with a neural network model, and generating a priority matrix according to task importance, change frequency and risk level; a Dijkstra algorithm is adopted to plan an initial flight path giving consideration to priority and flight limitation, a path complexity index is calculated, and the index comprehensively considers a target priority weight, a task detouring coefficient and a path relaxation degree; and finally, dividing a monitoring area into height layers according to a path complexity index threshold value, performing height layer adjustment on the initial path, and generating a dynamic flight path containing height layer switching, thereby realizing efficient resource allocation and accurate risk prevention and control in a complex monitoring scene.
Owner:BEIJING JUNDE SPACETIME TECH CO LTD

Coal rock fracture intelligent extraction method based on improved U-Net

The invention discloses a coal rock fracture intelligent extraction method based on improved U-Net. The method comprises the following steps: S1, constructing a coal rock fracture CT image data set; s2, constructing an improved U-Net segmentation model, specifically comprising the following steps: S2.1, taking VGG16 as a backbone network, and introducing a depth separable convolution module; s2.2, a PPA attention module is added after each layer of depth separable convolution of the decoder, the PPA attention module is introduced after each up-sampling stage of the decoder, and the output of the PPA attention module is subjected to batch normalization and Dropout layer processing; s2.3, defining a composite loss function; s3, training and optimizing a segmentation model, wherein the specific steps comprise: S3.1, setting hyper-parameters; and S3.2, training the model by using the training set, adjusting hyper-parameters by using the verification set, and evaluating the performance by using the test set, wherein the evaluation indexes comprise MIoU, MAcc and FWIoU. According to the method, the problems of difficult identification of small fractures, large model calculation amount, poor multi-scale information fusion and class imbalance in the coal rock fracture image can be solved, and the robustness, segmentation precision and practicability of the model are improved.
Owner:CHINA UNIV OF MINING & TECH

Infrared small target detection method fusing local prior and multi-scale global background

The invention discloses an infrared small target detection method fusing local prior and a multi-scale global background, and the method comprises the steps: firstly obtaining image data containing an infrared image and a mask label corresponding to the infrared image, and carrying out the preprocessing; secondly, a target detection model of an encoder-decoder architecture is constructed, an encoder comprises a local detail prior mining branch and a multi-scale global background perception branch which are parallel, step-by-step feature extraction is performed on the preprocessed image data, and a decoder comprises a progressive feature fusion decoding branch; and inputting the features of each level of the encoder double branches into decoder branches for decoding step by step to obtain a detection result. And finally, a weighted depth supervision mechanism is introduced in training, auxiliary prediction output is set in a plurality of decoding layers, and weighting loss is calculated. According to the method, the problems of insufficient local detail modeling, insufficient multi-scale global background perception of Mamba, difficulty in global and local feature fusion and the like in the existing method are solved, and the detection precision of the infrared small target is improved.
Owner:HANGZHOU DIANZI UNIV

Style transfer using generative diffusion features

The present invention sets forth techniques for performing style transfer from multiple supplied style images to a supplied content image to generate novel images that include style elements from the multiple supplied style images and content elements from the supplied content image. The techniques include guiding one or more self-attention and cross-attention layers included in a machine learning model based on the multiple supplied style images, such that content elements and style elements included in the style images are not entangled when generating the novel images. The techniques also distill a small subset of representative attention map values from multiple style images, improving performance while reducing computational costs compared to processing all attention map values from the multiple style images.
Owner:DISNEY ENTERPRISES INC

Defect detection method and system based on honeycomb catalyst stacking

The invention belongs to the technical field of industrial detection, and discloses a defect detection method and system based on honeycomb catalyst stacking. Omnibearing image data of honeycomb catalyst stacking are obtained through a multi-angle polarization imaging technology, pixel-level polarization degree parameters are calculated to construct a global polarization feature map, and accurate distinguishing between an intrinsic porous structure and suspected defects is achieved. A blind area identification and virtual view angle reconstruction mechanism is introduced, so that the problem of a stacked edge detection blind area is solved; and a layered reflectivity compensation function is adopted, so that the optical interference of an interlayer overlapping region is eliminated. Texture features are extracted through multi-scale morphological filtering, multi-dimensional feature fusion is carried out in combination with polarization features, edge continuity indexes and correction reflection intensity, and a high-precision defect discrimination model is established. And for a low-confidence region, dynamically adjusting detection parameters and performing iterative optimization to form an adaptive detection closed loop. According to the invention, the detection precision and reliability are improved, and the defect position, type and severity can be accurately output.
Owner:TIANHE BAODING ENVIRONMENTAL ENG

Three-dimensional human body posture estimation method and system based on multi-view visual information fusion and storage medium

The invention provides a three-dimensional human body posture estimation method and system based on multi-view visual information fusion and a storage medium, and the method comprises the steps: 1, designing the front half part of a model into Ender Layers with the same layer number as a Transform decoder at a multi-view feature fusion layer, carrying out the data enhancement of an input multi-view original image, and carrying out the reconstruction of the Ender Layers in the multi-view feature fusion layer; inputting the CNN Backbone with the shared weight to extract an initial feature map; 2, introducing a micro-reprojection optimization mechanism, deeply fusing the multi-view geometric consistency constraint into a model training process, and guiding the model to predict a three-dimensional attitude end to end; and step 3, constructing a dynamic projection compensation module. The method has the beneficial effects that the method is particularly suitable for capturing human body posture information in a multi-person interaction scene, the shielding problem and depth estimation ambiguity in a single view angle can be effectively overcome, and the robustness, precision and efficiency of three-dimensional human body posture estimation are remarkably improved.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Mobile terminal streetscape image real-time segmentation method based on lightweight neural network

The invention discloses a mobile terminal streetscape image real-time segmentation method based on a lightweight neural network, and relates to the technical field of image segmentation. The method comprises the following steps: firstly, carrying out 320 * 320 adjustment, Z-score standardization, adaptive histogram equalization and 3 * 3 Gaussian filtering preprocessing on an input streetscape image; then, an improved MobileNetV3 backbone network is used, and a five-scale feature map is output in combination with DropBlock regularization through eight feature extraction stages including depth separable convolution and an SE attention module; multi-scale features are fused through a U-shaped structure, and a fusion feature map is generated through up-sampling, element-by-element addition of dimension reduction low-layer features and an attention gating module; and during reasoning, outputting a segmentation mask by using a convolutional layer, Softmax and a conditional random field, and finally performing knowledge distillation, weight pruning, 8-bit quantization and TensorRT optimization. According to the invention, high-precision real-time street view segmentation is realized, the robustness is high, and the method is suitable for different devices and scenes.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Self-isolation type image sensing structure, sensor and preparation method

The invention discloses a self-isolation type image sensing structure, a sensor and a preparation method, and belongs to the field of semiconductors, the self-isolation type image sensing structure is characterized in that a shallow trench isolation structure is exposed after the back surface of a substrate is thinned, then a first semiconductor layer is extended on the back surface of the substrate, and a first isolation structure is prepared in the first semiconductor layer; sequentially depositing a second semiconductor material layer and a metal grating material layer on the first semiconductor layer; sequentially etching the metal grating material layer and the second semiconductor material layer to form a plurality of metal grating structures distributed at intervals and corresponding self-isolation structures; extending a plurality of photosensitive layers in the photosensitive region groove between the adjacent self-isolation structures to form a photosensitive region; and then activating, preparing a filter layer and the like are carried out to obtain a complete back-illuminated image sensor. The self-isolation photodiode is obtained by changing the structural distribution of the photodiode, unexpectedly, the problem that the substrate is damaged by doping is solved, meanwhile, the crosstalk effect is greatly reduced, and the performance of the image sensor is improved.
Owner:NEXCHIP SEMICON CO LTD

Visual navigation method and system based on improved optical flow method

The invention provides a visual navigation method and system based on an improved optical flow method, and relates to the technical field of computer vision and inertial navigation. A collaborative optimization framework of an inertial navigation system and visual feature navigation is constructed, carrier motion state parameters are obtained through dead reckoning of an inertial measurement unit, variance parameters of pose variation are calculated, and the layering depth of a pyramid LK optical flow method is dynamically adjusted based on the variance parameters. When the inertial navigation output variance exceeds a preset threshold value, the number of layers of an image pyramid is increased to cope with large-range motion, and otherwise, calculation levels are reduced to improve real-time performance. A dual-period optical flow / feature matching tracking mechanism is designed, improved sub-pixel-level optical flow tracking is adopted in a short period, the calculation complexity is reduced while the positioning precision is guaranteed, feature matching is switched to in a long period, accumulative errors are eliminated, and a period parameter N is determined by inertial navigation precision. According to the method, the calculation complexity can be reduced, and a better effect can be achieved on a low-calculation-performance platform.
Owner:THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION

Using a Container-Aware Storage System to Deploy Virtual Machines on a Container Orchestration Platform

An illustrative storage management system is configured to provide persistent data storage to workloads managed by a container orchestration platform deployed in a cluster of nodes. The storage management system is configured to receive a request comprising identifiers indicative of a plurality of layers associated with a virtual machine to be run in the cluster, generate, based on the request, a virtual machine disk image comprising the plurality of layers, the virtual machine disk image configured to be used to run the virtual machine in the cluster, and provide the virtual machine disk image in response to the request. In some implementations, the virtual machine disk image is used by a virtual machine handler associated with the container orchestration platform to run the virtual machine in a virtual machine type pod that is managed by the container orchestration platform.
Owner:PURE STORAGE INC

Medical image classification method and system based on multi-scale spatial state modeling

The invention discloses a medical image classification method and system based on multi-scale spatial state modeling, and the method comprises the steps: firstly dividing an input medical image into a plurality of non-overlapping image blocks, and mapping the non-overlapping image blocks to a feature space through a learnable linear projection layer to obtain an initial feature map; then, multiple layers of stacked MS-SMamba blocks are used for carrying out layer-by-layer feature extraction, each MS-SMamba block comprises a main branch, an auxiliary branch, a dynamic gating fusion network, a residual error connection unit and a feedforward network, and long-range dependency relation capture and multi-scale feature fusion are achieved; and finally, processing the last-layer output feature map through a global feature aggregation and classification module, generating a global feature vector, and outputting a classification result. According to the method, the capturing capability of complex pathological features in the medical image is improved, the calculation efficiency and clinical applicability are improved, and the method is suitable for scenes such as disease screening and auxiliary decision making in medical image diagnosis.
Owner:XIANGJIANG LAB

Fan metal surface defect detection method and device based on lightweight YOLO11 and medium

The invention discloses a fan metal surface defect detection method and device based on lightweight YOLO11 and a medium, and relates to the technical field of computer vision and industrial detection. The method comprises the following steps: acquiring fan metal surface defect image data, and preprocessing to obtain a training data set; yOLO11n is used as a basic model, and a lightweight StarNet adopting a four-level layered architecture and a star operation feature fusion mechanism is used as a model backbone network; combining a bottleneck module with a multi-scale convolution block, and constructing a neck network by applying a global heterogeneous kernel selection mechanism and an efficient up-sampling module; a heavy parameterized detail enhanced convolution and group normalization GN strategy is used to construct a detail enhanced lightweight shared convolution detection head; and sequentially connecting the backbone network, the neck network and the output layer of the detection head to form the lightweight metal surface defect detection model. On the premise that the detection efficiency is guaranteed, high-precision identification of metal weld defects can be achieved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Steel bar binding point detection method based on improved YOLOv8

The invention provides a reinforcing steel bar binding point detection method based on improved YOLOv8, and the method comprises the following steps: obtaining a construction site image, carrying out the preprocessing, inputting a to-be-detected image into an improved target detection model, sequentially carrying out the local space attention calculation, depth separable convolution and channel attention weighting operation, and carrying out the detection of the to-be-detected image. Detail feature expression of reinforcing steel bar intersection points is enhanced; performing dynamic weight fusion on the top-down first feature path and the bottom-up second feature path, and embedding a lightweight residual unit in a fusion node; a classification task and a regression task are respectively processed by adopting a classification branch and a regression branch, the classification branch outputs category confidence, and the regression branch outputs coordinate positioning through depth separable convolution and a probability distribution prediction layer; and decoding output data of the target detection model, and generating a detection result containing the position and the category of the binding point. According to the method, the requirement of real-time detection can be met while the accuracy of steel bar binding point detection is improved.
Owner:HEBEI ZHUCHENG DIGITAL TECHNOLOGY CO LTD

Semantic segmentation method and device with enhanced depth estimation, equipment and medium

The invention provides a depth estimation enhanced semantic segmentation method and device, equipment and a medium. The method comprises the following steps: extracting a corresponding depth image from an acquired RGB image by using a depth estimation large model; constructing a scene perception model, wherein the scene perception model comprises a coding layer and a decoding layer; the coding layer comprises an RGB image coding branch and a depth image coding branch, and is used for correcting and fusing the characteristics of the RGB mode and the depth mode output by the RGB feature extraction layer and the depth feature extraction layer based on data fusion modules arranged between the coding layers respectively to obtain the first fusion characteristics of each layer; a first fusion feature obtained after fusion of the data fusion modules except the last layer of data fusion module is input to a multi-scale fusion decoding module, step-by-step recovery of a feature map is achieved through a multi-scale feature fusion module and a frequency perception feature fusion device of the multi-scale fusion decoding module, and finally a semantic segmentation result is obtained. According to the method, scene perception is accurate, and meanwhile, only the minimum cost is needed.
Owner:TRAFFIC CONTROL TECH CO LTD

Medical image segmentation method based on hierarchical pre-training model

The invention discloses a medical image segmentation method and system based on a hierarchical pre-training model, and the method comprises the steps: carrying out the size adjustment of an input image, and carrying out the data enhancement operation; extracting multi-scale features of the medical image by using an encoder based on a hierarchical pre-training visual model, and inserting an adapter in front of each encoding block; a decoder based on a dense connection structure is adopted to gradually recover the spatial resolution, a feature enhancement module is applied behind each decoder block for feature enhancement, and dense connection is achieved through layer-by-layer up-sampling and cross-layer connection; multi-scale supervision and prediction are realized through three output heads, a main output head generates a main segmentation result, and two auxiliary output heads respectively generate segmentation results from an intermediate decoding layer to provide a multi-scale supervision signal. Therefore, the problems of precision, accuracy and efficiency of an existing medical image segmentation technology in tiny focus recognition are solved, and the recognition capability of tiny boundaries and complex geometrical shapes is remarkably improved.
Owner:HUBEI UNIV OF TECH

Target detection method based on YOLO model, electronic equipment and storage medium

The invention discloses a target detection method based on a YOLO model, electronic equipment and a storage medium, and relates to the technical field of target detection. Comprising the following steps: inputting an aerial image of an unmanned aerial vehicle into a trained target YOLO model; the target YOLO model comprises a backbone network, a neck network and a head network, and a feature extraction module in the backbone network performs multi-scale feature extraction by adopting a double-branch architecture attention mechanism; performing multi-scale feature extraction on the aerial image by adopting a dual-branch architecture attention mechanism through a feature extraction module in the backbone network, and constructing to obtain a plurality of layers of first comprehensive image features of the aerial image; performing feature fusion on the first comprehensive image features of different levels through a neck network to obtain second comprehensive image features of multiple levels; and inputting the multiple levels of second comprehensive image features into a head network to obtain a target detection result of the aerial image. According to the invention, the accuracy of small target detection can be improved.
Owner:HUNAN UNIV OF TECH

Lightweight visible light ship target detection method based on edge feature guidance

The invention provides a lightweight visible light ship target detection method based on edge feature guidance, and relates to the technical field of ship detection image data processing, and the method comprises the steps: collecting remote sensing satellite images, and carrying out the random distribution of the images after screening and marking, and obtaining a training set and a verification set; the backbone network module comprises a plurality of Conv modules and C3k2 modules which are mutually stacked; the neck module comprises a detail-enhanced convolution module and a hierarchical pyramid module based on dynamic feature aggregation; in the head module, after the features of all detection layers are subjected to independent convolution processing, feature transformation is carried out through a multi-branch detail enhancement convolution module; performing data enhancement on the training set; and obtaining a trained ship target detection model through a back propagation algorithm and a gradient descent optimization method. According to the invention, the lightweight and precision improvement of the detection head are realized, the robustness of the model to the illumination change is enhanced, and the global semantic information and the local detail features are fused to balance the detection of the small target and the large target.
Owner:HARBIN INST OF TECH AT WEIHAI

Method for detecting quality of functional layer of outer wall of building by unmanned aerial vehicle

The invention discloses a method for detecting the quality of a functional layer of a building outer wall by an unmanned aerial vehicle, and particularly relates to the technical field of building outer wall detection.The method comprises the steps that firstly, an infrared image of a building outer wall facing object is collected through the unmanned aerial vehicle, pixel points and remaining pixel points of a serious hollowing area are determined, and the abnormal degree of the remaining pixel points is calculated; clustering is carried out by adopting a defect probability-based region growing method; based on seed point screening of morphological preprocessing, structural elements of different sizes are comprehensively determined to be used through weighted summation calculation according to the jitter frequency of the unmanned aerial vehicle and the image resolution, and noise is filtered step by step. Generating a multi-scale pyramid for the original image, and detecting candidate seed points at each level; sorting and preferentially selecting the candidate seed points based on the morphological closure degree and the shape regularity of each candidate seed point; pixel points in the neighborhood of the growth seed points are merged to obtain all hollow defect connected domains; and judging the quality condition of the building outer wall functional layer according to the area of the hollowing defect connected domain.
Owner:SHANXI ARCHITECTURE KEXUE RES YUAN

Neural network codec with hybrid entropy model and flexible quantization

Innovations in systems, methods, and software for features of a neural image or video codec are described herein. For example, a neural video encoder can receive a current video frame, encode the current video frame to produce encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation for the current video frame, and encode the current latent representation using an entropy model network that includes one or more convolutional layers. As part of the encoding the current latent representation, the encoder can estimate statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation for a previous video frame, and entropy code the quantized version of the current latent representation based at least in part on the estimated statistical characteristics.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Pipeline all-position multi-layer multi-channel TIG welding bead morphology measuring device and method

The invention discloses a pipeline all-position multi-layer multi-channel TIG welding bead morphology measuring device and method, and relates to the technical field of intelligent welding. A molten pool image collecting module is used for collecting morphology information of a molten pool, and the light emitting end of a laser transmitter is used for facing a welding bead; the active and passive vision module is used for collecting an image of light emitted by the laser emitter projected on a weld pool groove and identifying a weld bead corner opening position and a weld bead width, and the active and passive vision module and the weld pool image collecting module are both electrically connected with the control module; the control module can fuse image features collected by the active and passive vision module and the molten pool image collection module, the axes of the molten pool image collection module, the welding module and the active and passive vision module are located in the same plane, and the molten pool image collection module and the active and passive vision module are located on the two sides of the welding module respectively. According to the method, manual participation in the welding process can be reduced, the welding automation degree is improved, and meanwhile, the welding bead morphology measurement accuracy is improved.
Owner:TIANJIN UNIV

Lightweight rice leaf disease identification method based on YOLO target detection

The invention discloses a lightweight rice leaf disease identification method based on YOLO target detection. The method comprises the following steps: S1, inputting a rice leaf disease image; s2, carrying out super-resolution reconstruction on the image by using an improved Real-ESRGAN algorithm; s3, inputting the picture into an improved YOLOv8n target detection algorithm; s4, obtaining disease spot category and position information; and S5, visualizing the information on the image. The invention relates to the technical field of target detection, and has the beneficial effects that image super-resolution reconstruction is carried out aiming at the problems that a rice leaf scab target is relatively small and an image acquired in real time is relatively fuzzy, so that the resolution of the small target is improved, and the definition and texture features are improved. On the basis of a super-division model Real-ESRGAN, a group of residual dense modules containing five layers of cavity convolution layers are designed to help the network to acquire receptive fields and information of different scales.
Owner:JILIN UNIVERSITY

Sea temperature complementation method and system based on recursive double-current Mama

The invention belongs to the technical field of image processing, and particularly relates to a sea temperature complementation method and system based on recursive double-flow Mama, and the method comprises the following steps: splicing a damaged SST image and a weekly average SST image as input, and outputting a predicted weekly average SST image and a predicted abnormal SST image through N times of recursive iteration of N same recursive hierarchical Mama blocks, and the two output images are added to obtain a final complemented SST image. According to the method, a recursive hierarchical Mama block for a sea surface temperature completion task is set, two parallel layers, namely a stable information representation module and an abnormal information representation module, are integrated in each block, features related to stability and features related to anomalies are extracted respectively, long-range dependency relationship modeling under large-area deficiency is achieved, and completion accuracy is improved.
Owner:OCEAN UNIV OF CHINA

Vision-to-text pairwise model for content tagging using an interest graph

Techniques for automated tagging of visual content are described. A pairwise model is used to encode images and videos with a vision encoder, and encode keywords and phrases from an interest graph with a text encoder. Similarity layers compare these cross-modality embeddings by calculating distance in a shared embedding space. Scores indicate the association between visual features and text. A pairwise loss function brings together matched pairs while separating non-matches during training. Scores exceeding a threshold tag content with relevant keywords and phrases. The pairwise architecture relates images and text despite limited associated text. It leverages vision-to-text understanding for accurate tagging without per-class labels. Contrastive similarity techniques associate visual patterns with textual concepts. Automated tagging organizes user-generated content by topics using this scalable cross-modality approach.
Owner:SNAP INC

Object class inpainting in digital images utilizing class-specific inpainting neural networks

The present disclosure relates to systems, methods, and non-transitory computer readable media that generate inpainted digital images utilizing class-specific cascaded modulation inpainting neural network. For example, the disclosed systems utilize a class-specific cascaded modulation inpainting neural network that includes cascaded modulation decoder layers to generate replacement pixels portraying a particular target object class. To illustrate, in response to user selection of a replacement region and target object class, the disclosed systems utilize a class-specific cascaded modulation inpainting neural network corresponding to the target object class to generate an inpainted digital image that portrays an instance of the target object class within the replacement region. Moreover, in one or more embodiments the disclosed systems train class-specific cascaded modulation inpainting neural networks corresponding to a variety of target object classes, such as a sky object class, a water object class, a ground object class, or a human object class.
Owner:ADOBE INC

Video language model training method and human body interaction behavior recognition method

The invention provides a video language model training method and a human body interaction behavior recognition method, and relates to the technical field of computer vision recognition, and the method comprises the steps: obtaining a video sample and motion description text data for a human body interaction behavior in the video sample; determining a first video feature and a first object position feature corresponding to the video sample; determining visual joint features output by each layer of multi-head self-attention blocks in the L layers of multi-head self-attention blocks based on the first video features and the first object position features; and based on the action description text data and the visual joint features, determining visual representation, text representation and multi-modal representation output by the last layer of multi-modal refined learning module in the L layers of multi-modal refined learning modules, and based on the visual representation, the text representation and the multi-modal representation, updating model parameters of the video language model. And obtaining the trained target video language model. According to the invention, the accuracy of human body interaction behavior recognition can be improved.
Owner:NORTHEASTERN UNIV CHINA

Container Image Vulnerability Scanning Based on Vulnerability Signatures

Mechanisms are provided for scanning container images for vulnerabilities. With these mechanisms, a container image is received for inclusion in a container image registry and each layer of the container image is scanned to generate file signatures for each file referenced in each layer. Vulnerability signature(s) are applied to each layer, based on the file signatures of the layer, to determine if criteria of the vulnerability rule(s) of the vulnerability signature(s) are satisfied by at least one layer of the container image. Registration of the container image in the container image registry is accepted or denied based on results of the application of the vulnerability signature(s). Each vulnerability signature comprises one or more vulnerability rules generated from a scanning and indexing of layers of one or more other container images.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

AI image recognition and grading method for field crop leaf diseases and insect pests

The invention relates to the technical field of disease and insect pest image analysis, in particular to an AI image recognition and grading method for field crop leaf disease and insect pests, which comprises the following steps: under the irradiation of a field fixed light source, synchronously acquiring a plurality of polarized reflection images around crop leaves at preset angle intervals; extracting a pixel polarization degree matrix of a leaf area in each polarization reflection image; inputting the pixel polarization degree matrix into a polarization transmission model, outputting a cuticle anomaly coefficient graph, and marking an area exceeding a preset anomaly threshold in the cuticle anomaly coefficient graph as a highlight display area; matching an infection type template library according to a highlight display area distribution mode in the abnormal coefficient graph; and calculating an infection intensity value by combining the diffusion gradient of the highlight area, and outputting a pest grade. According to the method, the boundary of the optical mutation region of the focus region is depicted, so that the physical interpretation of disease detection is improved, and distinguishable feature spaces are provided for different infection mechanisms (such as fungal growth layers and insect pest piercing and sucking points).
Owner:BEIJING BANGWEIKE TECH CO LTD