Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

98 results about "Image complexity" patented technology

Artificial intelligence machine vision image acquisition system

The invention discloses an artificial intelligence machine vision image acquisition system, and the system comprises a multi-mode perception layer which integrates a self-adaptive optical module, inhibits metal reflection, captures a visible light to short wave infrared image, and captures a motion edge; the dynamic adaptive layer adopts an illumination compensation and motion compensation module to dynamically adjust camera parameters and micro displacement compensation, feeds back an illumination trend, outputs a motion vector to the cognitive layer, generates a confrontation sample through a GAN, simulates virtual defects in combination with a physical engine, and expands training data; the cognitive reasoning layer is used for deploying a dynamic routing network, distributing computing resources according to image complexity and optimizing feature extraction efficiency; reducing data deviation through anti-fact analysis, and generating a thermodynamic diagram to explain a detection basis; and the collaborative decision-making layer is used for rapidly screening samples by edge nodes, training a global model by cloud aggregated data, automatically triggering manual rechecking when the confidence coefficient of the model is insufficient, synchronously optimizing a training set and a causal reasoning module by a rechecking result, and improving the labeling efficiency through AR assistance.
Owner:南昌理工学院

System and Method for Low-Light Image Enhancement Using Hierarchical Adaptive Wavelet Decomposition with Cross-Scale Feature Fusion

A system and method are disclosed for low-light image enhancement using hierarchical adaptive wavelet decomposition with cross-scale feature fusion. The system analyzes a raw input image to determine image characteristics and preprocessing parameters. A hierarchical adaptive wavelet decomposition process creates a variable-depth decomposition tree comprising frequency domain nodes, with decomposition depth determined by local image complexity. Cross-scale feature fusion implements attention mechanisms between nodes at different decomposition levels, enabling bidirectional information flow across scales. A dynamic network pool allocates specialized neural networks to process nodes based on their frequency characteristics, with weight sharing between similar nodes for efficiency. An adaptive reconstruction engine traverses the decomposition tree using learned filters and multi-scale residual learning to produce an enhanced image. The hierarchical approach enables superior low-light image enhancement by allocating computational resources based on content complexity, achieving better quality than fixed decomposition methods while maintaining compatibility with existing image signal processing pipelines.
Owner:ATOMBEAM TECH INC

Identification and classification method for lesions in medical images

The invention relates to the technical field of medical image processing and analysis, and discloses a method for identifying and classifying lesions in medical images, which comprises the following steps: acquiring a plurality of medical image data, and constructing a multimode medical image data set containing CT, MRI and PET images; training the multi-mode medical image data set by using a hierarchical attention feature fusion network to generate a focus recognition and classification model; a target medical image is partitioned by a dynamic architecture partitioning algorithm based on image complexity, and analysis is performed by using the focus recognition and classification model to obtain focus features; and classifying the lesion features by using an iterative classification algorithm of Bayesian uncertainty estimation, calculating a classification threshold in combination with a medical expert knowledge base, and outputting a final classification result about the lesion. According to the invention, the accuracy and robustness of identification and analysis are improved, and the accuracy and reliability of a focus classification result are ensured.
Owner:NANJING MAITUO MEDICAL TECH CO LTD

Video encoder parameter dynamic adjustment method

The invention relates to a method for dynamically adjusting parameters of a video encoder, which comprises the following steps of: acquiring a visual characteristic index corresponding to a video frame image in real time, and calculating an image complexity score of the image according to the visual characteristic index; the visual feature indexes comprise motion intensity, edge density and information entropy; acquiring a network index of a video transmission network in real time; the network indexes comprise available bandwidth, delay, jitter and packet loss rate; obtaining a current encoder control parameter according to the current image complexity score and the network index; the control parameters of the encoder comprise GOP, GP, FPS and Buffer Size; and adjusting the encoder according to the obtained control parameters of the encoder. The method is superior to an existing fixed parameter coding system in the aspects of image quality, delay control, system robustness and adaptability.
Owner:BROAD VISION (XIAMEN) TECHNOLOGY CO LTD

Cloud gaming benchmark testing

The technology disclosed teaches a method of testing performance of a device-under-test during cloud gaming over a live cellular network. The method comprises instrumenting the device-under-test with at least one instrument app that interacts with a browser on the device-under-test and captures performance metrics from gaming network traffic. The browser and the instrument app can be invoked using a test controller separated from the device-under-test, causing the browser to connect to a gaming simulation over the live cellular network. A segmented gaming image stream is transmitted to the browser, with segmented playing at varying bit rates and image complexity while the instrument app causes the browser to transmit artificial gameplay events to a gaming simulation test server. Performance metrics from the gaming network traffic are captured, as well as gaming images rendered by the browser during the segmented gaming image stream.
Owner:SPIRENT COMM INC

High-density microalgae detection method oriented to perception enhancement and characteristic distillation

The invention relates to the technical field of artificial intelligence image processing and biological detection crossing, and discloses a perception enhancement and feature distillation-oriented high-density microalgae detection method, which comprises the following steps: acquiring a high-density microalgae image containing cell overlapping, boundary blur and cross-scale distribution, and inputting the image into a backbone network to extract a multi-scale feature map; optimizing the bounding box through potential consistency mapping processing; generating a microalgae density map through density sensing auxiliary processing, and calculating density loss; self-adaptive denoising is carried out based on image complexity; a student model is optimized by using a dual-feature distillation framework, and the small-scale microalgae detection capability is enhanced; and finally, fusing the results of the modules, and outputting the position, category and quantity of the microalgae. According to the method, the overall average accuracy of high-density microalgae detection can be improved, bounding box jitter and small-scale microalgae omission ratio are reduced, the average reasoning time is shortened, the real-time detection requirement is met, and reliable data support is provided for microalgae culture process control and optimization.
Owner:SOUTH CHINA NORMAL UNIV +1

Ultra-wide-angle traffic sign detection method and system based on adaptive dynamic pruning

The invention discloses an ultra-wide-angle traffic sign detection method and system based on adaptive dynamic pruning, and belongs to the field of computer vision and target detection, and the method comprises the steps: obtaining to-be-detected ultra-wide-angle traffic sign data, and carrying out the splicing and fusion; inputting the spliced and fused data into the trained traffic sign detection model to obtain a traffic sign detection result; wherein the training of the traffic sign detection model comprises the following steps: constructing an ultra-wide-angle traffic sign data set; a traffic sign detection model is established and comprises a human eye attention concentration position detection module, a detection module based on dynamic pruning and a detection result output module. And training the constructed traffic sign detection model by using the constructed ultra-wide-angle traffic sign data set to obtain a trained traffic sign detection model. According to the method, the image complexity is evaluated in real time, the calculation path of the model is flexibly adjusted, the allocation and use of calculation resources are optimized, and the sensing and processing capability on a large-view-angle target in a complex dynamic scene is improved.
Owner:WUHAN UNIV

Contour extraction method based on edge detection

The invention discloses a contour extraction method based on edge detection, and belongs to the technical field of image processing, and the method comprises the following steps: obtaining a plurality of images to be subjected to contour extraction, and real labeled contours corresponding to the images to be subjected to contour extraction; performing noise level quantization and edge ambiguity quantization analysis on each image to be subjected to contour extraction, and calculating to obtain an image complexity level of each image to be subjected to contour extraction; performing complexity level classification on each image to be subjected to contour extraction, and constructing a complexity classification image data set; constructing a contour extraction model based on edge detection; a trained contour extraction model is obtained; and obtaining a new image to be subjected to contour extraction and an image complexity level, inputting the image to be subjected to contour extraction into the trained contour extraction model for contour prediction, and obtaining a predicted image contour of the image as a contour extraction result. The problem that it is difficult to accurately extract the object contour in a complex image scene is solved.
Owner:CHENGDU RUIGAN TECH

Large model collaborative road disease data efficient intelligent labeling method

The invention discloses an efficient and intelligent road disease data marking method based on large model cooperation, and particularly relates to the technical field of road detection. Various core features are extracted from two dimensions of disease ontology and identifiability, the complexity of the disease and the complexity of the identification difficulty are comprehensively covered, the limitation of a single feature is avoided, the two types of complexity are quantified through a disease evaluation coefficient and a disturbance evaluation coefficient respectively, then the complexity is fused into a complex evaluation coefficient, and the complexity judgment is completed by grading according to a coefficient interval. The problems that in the prior art, a fixed model full-amount calling mode is mostly adopted, judgment on image complexity mostly depends on a single feature, and the cooperative influence of a disease ontology feature and an environment interference feature is ignored are solved.
Owner:HANGZHOU TOPWAY VIEW INFORMATION TECH CO LTD

Methods and apparatus for dynamic codec configuration

Systems, apparatus, and methods for dynamic encoder configuration. In one exemplary embodiment, a machine-learning model uses pixel features and encoding features from previous stages of an image processing pipeline (IPP) to dynamically adjust bitrate. The machine-learning model is trained to select bitrate adjustments for an encoder such that the expected image quality of a video stream remains at a selected quality level (e.g., SSIM, VMAF, VIF, HVS-PSNR, etc.). Conventional dynamic encoding solutions are focused on encode-once-deliver-often (best-effort) applications, the exemplary IPP is designed for real-time applications that may not have the benefit of actual subsequent encoding quality analysis; instead proxy data (pixel features and encoding features) that are representative approximations of image complexity are used.
Owner:GOPRO INC

Adaptive image generation method and system based on multi-modal routing

The invention discloses a self-adaptive image generation method and system based on multi-modal routing, and the method comprises the steps: extracting the multi-scale features of an input image, and generating three different tokens with progressive increase of continuous information keeping capability: respectively extracting visual and text modal information from the image and corresponding text description, and generating a multi-modal information abstract after fusion; inputting the multi-modal information abstract into a learnable soft router, and dynamically selecting a token modeling path based on an image complexity label; a three-stage training strategy optimization model is adopted; in the inference stage, a trained soft router dynamically selects a token path to complete image generation according to input text description and an image information abstract predicted by an autoregressive Transform. According to the method, a dynamic router and three quantification and modeling strategies with different complexities are fused, and dynamic modeling path selection is realized in a reasoning stage through a soft router module. According to the method, on the premise of ensuring the generation quality, the reasoning efficiency is effectively improved, and good efficiency is shown.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

Cloud particle data target detection method based on image complexity

The invention relates to a cloud particle data target detection method based on image complexity. The method comprises the following steps: firstly, calculating weighted information entropy, weighted texture complexity and weighted local standard deviation of data of each cloud particle image, and evaluating the complexity of the image data; and marking according to the complexity of the image data, further calculating the comprehensive complexity of the data group, and judging the uniformity degree of complexity distribution of the data group. On the basis, different splitting and combining strategies are adopted according to the overall complexity distribution uniformity and marks of the data sets, and it is ensured that the deep learning detection model can select the most suitable strategy according to the complexity. Through the method, the computing resources and the detection precision of the deep learning model can be optimized in the face of cloud particle image data with different complexities, the image processing efficiency is improved, and the method is suitable for large-scale cloud particle data sets with large complexity differences.
Owner:CHENGDU UNIV OF INFORMATION TECH

Image encryption method based on interpretable block model driving

The invention discloses an image encryption method based on interpretable block model driving. The method comprises the following steps: inputting a digital image to be processed; calculating the spatial information complexity of the image; calculating the statistical information complexity of the image; calculating the visual perception complexity of the image; establishing a weight optimization objective function; the optimal weight configuration is solved; calculating an image complexity evaluation value, and calculating the image complexity by using the optimal weight configuration; the maximum and minimum block size logarithms are set, and the self-adaptive block size is determined; setting parameters of the Chen's chaotic system; generating a pseudorandom sequence, and starting the Chen's chaotic system by using a set initial value to obtain a three-dimensional chaotic sequence; inputting a plaintext image, extracting basic parameters, image height and image width, and calculating the total number of image pixels; according to the method, a completely quantifiable block decision model is constructed, the mapping process from the image features to the block strategy is accurately described through a mathematical formula, and the problem that a traditional block method lacks a theoretical basis is fundamentally solved.
Owner:LIAONING TECHNICAL UNIVERSITY

Tree-structure-based object rendering method and apparatus, and electronic device, computer-readable storage medium and computer program product

Disclosed in the embodiments of the present application are a tree-structure-based object rendering method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: acquiring an object node evaluation list associated with a three-dimensional virtual object to be rendered in a virtual scene; acquiring an image complexity of a target texture map of a target node; on the basis of a target node bounding box to which the target node belongs, determining target size information of the target node bounding box, and on the basis of the target size information, a camera line-of-sight between the target node and a target camera, and the image complexity, performing node evaluation on the target node to obtain a node evaluation result; and if the node evaluation result indicates that the target node meets a rendering display condition and the target node is located within the field-of-view corresponding to the target camera, when the three-dimensional virtual object is rendered and displayed, rendering and displaying the target node, and performing node hiding on a child node of the target node.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Medical image intelligent labeling and auditing method based on deep learning

The invention provides a medical image intelligent labeling and auditing method based on deep learning, and relates to the technical field of medical image labeling, and the method comprises the steps: carrying out the preprocessing of medical image data, and obtaining the preprocessed image data; performing two-dimensional evaluation on the preprocessed image data based on image complexity and labeling task complexity to obtain a total complexity score; dividing the preprocessed image data into a simple level, a medium level and a complex level according to the total complexity score; carrying out labeling processing on the preprocessed image data by adopting a differential labeling strategy to obtain a labeling result; performing quantitative evaluation on the labeling result to obtain a quality score; extracting labeling process features, and dividing labeling results into high quality, medium quality and low quality; and determining an auditing strategy, auditing the annotation result, and outputting the annotation result which is audited to be qualified. According to the method, objective evaluation of labeling quality and reasonable configuration of auditing resources can be realized, and the consistency of auditing is guaranteed.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

Image processing method and device based on deep learning, equipment and medium

The application relates to a deep learning-based image processing method, device, equipment and medium. The method comprises the following steps: acquiring a to-be-processed dynamic video stream, wherein the to-be-processed dynamic video stream comprises to-be-processed images of each video frame; acquiring a hardware resource state parameter corresponding to a video frame of the to-be-processed image, and calculating a corresponding hardware state value of each to-be-processed image based on the hardware resource state parameter; calculating the image complexity and the interframe difference degree of the to-be-processed image; performing image processing deep learning model state judgment on the to-be-processed image based on the hardware state value, the image complexity and the interframe difference degree, and generating an image processing deep learning model state judgment result of the to-be-processed image. The method can realize hardware adaptation and dynamic model architecture fine-tuning, and significantly improve the model inference efficiency, image processing quality and hardware resource utilization rate of video stream processing.
Owner:LIAONING UNIVERSITY

Method and apparatus for dynamic codec configuration

The invention relates to a system, apparatus and method for dynamic encoder configuration. In one exemplary embodiment, a machine learning model uses pixel features and encoding features from previous stages of an image processing pipeline (IPP) to dynamically adjust bit rates. The machine learning model is trained to select bit rate adjustments for an encoder such that the expected image quality of the video stream remains at a selected quality level (e.g., SSIM, VMAF, VIF, HVS-PSNR, etc. Conventional dynamic coding solutions focus on one-time coding over-delivery (best effort) applications, the exemplary IPP being designed for real-time applications that may not benefit from actual subsequent coding quality analysis; instead, proxy data (pixel features and encoding features), which are representative approximations of image complexity, are used.
Owner:GOPRO INC

Self-adaptive reversible information hiding method and extraction method based on image complexity

The invention relates to the technical field of information hiding, and discloses a self-adaptive reversible information hiding method and extraction method based on image complexity. The method comprises the following steps: firstly, calculating a global Shannon entropy setting complexity threshold value, and recursively partitioning; discrete wavelet transformation is performed on the high-complexity sub-graph, secret information is embedded into a low-frequency sub-band through prediction error extension, and then inverse transformation is performed and the secret information is spliced with other sub-graphs to generate a secret-carrying graph. During extraction, information can be extracted and an original image can be recovered without loss by using the same threshold value and the same blocking rule and carrying out reverse operation on the secret-containing sub-image. According to the method, the embedding capacity is dynamically adjusted on the premise of ensuring the visual quality, the calculation overhead is low, the safety is high, and the method is suitable for scenes of digital watermarking, copyright protection, medical images, secret communication and the like.
Owner:HAINAN NORMAL UNIV

Adaptive image super-resolution reconstruction method based on compressed sensing

The present invention relates to the field of image processing technology, and in particular to an adaptive image super-resolution reconstruction method based on compressed sensing. The present invention designs a trainable compressed sensing sampling module for performing sparse observation compression on an input image. The module supports end-to-end joint training and has data-driven optimization capabilities. Subsequently, the content structure in the compressed image is modeled through an image prior analysis module, and the image complexity index is calculated as a reference for subsequent path selection. The present invention combines compressed sensing theory with an image structure complexity perception mechanism, and realizes a more efficient and quality-controlled super-resolution image reconstruction process through an adaptive path selection strategy.
Owner:SUZHOU GAIDE PHOTOELECTRIC TECH CO LTD

VGGT model accelerated reasoning method based on dynamic sparse calculation

The invention discloses a VGGT model accelerated reasoning method based on dynamic sparse calculation, and the method comprises the following steps: carrying out the normalization preprocessing of an input image sequence, and enabling the input image sequence to meet the input requirements of a VGGT model; and generating a dynamic sparse mask for subsequent sparse calculation according to the input image and the output of the model intermediate layer. And then, in the VGGT model, reasoning is carried out by adopting a sparse calculation method, so that the calculation amount is effectively reduced. According to image complexity and model performance feedback, the sparse rate is dynamically adjusted, and calculation efficiency and model precision are both considered. And finally, outputting final prediction results such as a camera pose code, a depth map and a point tracking result. According to the method, the VGGT model reasoning process is optimized through dynamic sparse calculation, certain precision is guaranteed while the calculation efficiency is improved, an efficient solution is provided for related image analysis and processing tasks, and the method can be widely applied to the field of computer vision such as automatic driving and intelligent monitoring scenes.
Owner:LINKER

An automatic recognition system for pipeline shallow section images based on deep learning

The present invention discloses a system and method for automatic recognition of pipeline shallow section images based on deep learning, which comprises an introduction module, a model improvement module, a data acquisition module, a training module and a recognition module. The model improvement module integrates an attention feature pyramid network, a target detection dynamic head structure and an optimization loss function into a basic model to form an initial recognition model. The attention feature pyramid network dynamically adjusts the feature map extraction weights; the target detection dynamic head structure adjusts parameters according to image complexity; and the optimization loss function adjusts the penalty terms of angle, scale and aspect ratio. The data acquisition module collects data to form a training set, a test set and a validation set. The training module trains the initial model based on the training set data, and adjusts the hyperparameters through the validation set until the model performance meets the preset conditions. The recognition module uses the optimization model to recognize the test set image to obtain the final result. The present invention can effectively improve the accuracy and efficiency of automatic recognition of pipeline shallow section images.
Owner:NINGBO SHANGHANG SURVEYING & MAPPING

A control system based on intelligent multi-screen cooperation

The application relates to the field of screen control, in particular to a control system based on intelligent multi-screen cooperation, which comprises a processing unit, an analysis unit, a load optimization unit and a display optimization unit. The processing unit comprises a plurality of sub-processing modules and is used for processing target videos respectively. The analysis unit is used for determining the display state of each target screen according to a picture misadjustment reference value and a display decline reference value, and determining an optimization mode as load distribution optimization or environment optimization according to a first display state proportion and a first display state aggregation degree. The load optimization unit is used for determining the screen category of the target screen according to an image complexity reference value and a dynamic reference value, and determining an optimization processing mode as task distribution optimization or task order optimization based on a balance coefficient. The display optimization unit is used for determining the environmental light influence state corresponding to each target screen according to an illumination intensity coefficient and an illumination distribution coefficient, and determining whether to adjust the brightness of the target screen according to the environmental light influence state. The application improves the display effect under the control of the multi-spliced screen.
Owner:BEIJING ZHONGYICHENG TECHNOLOGY CO LTD

A digital watermark image generation method and system

This invention discloses a method and system for generating digital watermarked images. The method includes: acquiring an original image to be watermarked; generating an image coding mask based on the image complexity of the original image; performing frequency domain decomposition on the original image to obtain a low-frequency component image; embedding an initial watermark image into the low-frequency component image according to the embedding strength values ​​corresponding to the masks in the image coding mask to obtain a low-frequency watermark image; different masks in the image coding mask correspond to different embedding strength values; and performing an inverse frequency domain transform on the low-frequency watermark image to obtain a synthesized watermark image. This invention can improve the success rate and accuracy of watermark extraction.
Owner:ZHOUPU DATA TECH NANJING CO LTD

Character recognition method, device, intelligent food collection cabinet, electronic device and storage medium

The present application provides a character recognition method, device, intelligent food collection cabinet, electronic device, and storage medium. The character recognition method includes: converting an image to be recognized into a standard format image; determining a character region within the standard format image; and identifying a character string within the character region based on a lightweight character recognition neural network. The present application simplifies the complexity and content of images requiring neural network processing. Therefore, a lightweight character recognition neural network can be used to quickly and accurately implement character recognition, achieving a faster response speed. Furthermore, the use of a lightweight character recognition neural network significantly reduces the computing performance requirements of the device.
Owner:BOE TECHNOLOGY GROUP CO LTD

Blood vessel recognition method, system and device based on image features and medium

The invention discloses a blood vessel recognition method, system and device based on image features and a medium, and relates to the technical field of data transmission processing, and the method comprises the steps: obtaining a target image, and dividing the target image into a plurality of sub-pixel regions; obtaining a gray gradient amplitude of each sub-pixel region, and obtaining a dynamic extension threshold of two adjacent sub-pixel regions; acquiring gray variation of two adjacent sub-pixel areas, and connecting the two adjacent sub-pixel areas with the gray variation smaller than a dynamic extension threshold to form a to-be-extended section; obtaining a preset angle threshold value, obtaining the number of the to-be-extended sections formed by each sub-pixel region, and obtaining a target angle threshold value according to the number of the to-be-extended sections formed by each sub-pixel region and the preset angle threshold value; if the included angle between the adjacent to-be-extended sections is larger than or equal to the target angle threshold value, the adjacent to-be-extended sections are connected, and a to-be-blood-vessel section is formed. The method has the advantages of image complexity perception, bifurcation and extension recognition and adaptive threshold adjustment.
Owner:SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV

Image anti-counterfeiting method based on vector quantization and phase mask

The invention discloses an image anti-counterfeiting method based on vector quantization and a phase mask, and the method comprises the steps: firstly, selecting a two-dimensional code which is simple in structure and only has transverse and longitudinal black and white change characteristics as an information medium, and simplifying a subsequent phase modulation structure while keeping the recognition characteristics clear; a vector quantization image compression technology is utilized to effectively code a high-definition image with rich content into a simplified graphic expression in which a two-dimensional code can be embedded, and effective mapping and compression from image complexity to a barcode structure are realized; and finally, a low-order pure-phase optical mask is generated based on the two-dimensional code, embedding and optical reconstruction of a high-definition image are realized under the condition that the phase quantization order is relatively low, and both the image identification degree and the mask preparation feasibility are considered. Therefore, the innovative image-level anti-counterfeiting coding method which integrates physical anti-counterfeiting and digital anti-counterfeiting advantages and has high information bearing capacity and strong physical anti-counterfeiting performance is constructed. The invention provides a feasible, efficient and low-cost new path for physical anti-counterfeiting of high-definition images.
Owner:DALIAN MARITIME UNIVERSITY

A printing timing control method, a printer, a program product, and a storage medium.

PendingCN122284934ASmooth generation of time-consuming fluctuationsAvoid read and write conflictsDigital dataGeneration process
This application provides a printing timing control method, a printer, a program product, and a storage medium, relating to the field of electronic digital data processing. Utilizing a double-buffered structure of a character buffer and a transmit character buffer, the generation process of print data is physically isolated from the hardware transmission process. This avoids the serial dependency of waiting to transmit and transmitting while waiting to generate, as found in related technologies. It allows the main control module to pre-generate subsequent data and store it in the first-level cache during the hardware transmission intervals. This smooths out fluctuations in data generation time caused by differences in image complexity and avoids hardware stalls caused by computational delays.
Owner:BEIJING SHUOFANG INFORMATION TECH CO LTD

A vehicle-mounted video encoding code rate preprocessing method, system, vehicle and medium

PendingCN122513564AVideo encodingIn vehicle
This invention relates to the field of vehicle-mounted video coding technology, and discloses a preprocessing method, system, vehicle, and medium for vehicle-mounted video coding bitrate. The method includes: performing multi-dimensional feature detection on the original image frame before coding, and generating a comprehensive complexity factor reflecting the bitrate fluctuation of vehicle-mounted video coding based on the detection results; detecting targets in the original image frame, and identifying core and background regions in the original image frame based on the targets; performing corresponding image preprocessing operations on the core and background regions according to the comprehensive complexity factor to obtain the target image frame. This method does not rely on the underlying modification of the hardware encoder, but only fuses the relevant elements in the vehicle scene that are most likely to cause bitrate fluctuations through multi-dimensional feature detection, so that the intensity of subsequent preprocessing is precisely matched with the image complexity, thereby effectively suppressing bitrate peaks and improving bitrate stability; and performing differentiated processing on different regions, ensuring the effectiveness of the core region while controlling the overall bitrate.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Remote sensing image subtitle generation method and system

The invention discloses a remote sensing image subtitle generation method and system. The method comprises the steps of obtaining a combined remote sensing image data set; for each remote sensing image in the combined remote sensing image data set, generating first annotation information of the remote sensing image based on a remote sensing image category, a target detection bounding box and a semantic segmentation mask, determining image complexity based on the first annotation information, and determining a subtitle generation mode according to the image complexity, the subtitle generation mode comprises generation of subtitles according to rules, generation of subtitles by a multi-modal large model and generation of subtitles by combining generation of subtitles according to rules and generation of subtitles by the multi-modal large model; generating subtitles of the remote sensing image based on a subtitle generation mode; all the image subtitle pairs form a remote sensing image-text pairing data set; and training a CLIP model based on the remote sensing image-text pairing data set. According to the method, the workload of manual labeling is effectively reduced, and the construction cost of the high-quality remote sensing image-text data set is reduced.
Owner:MILITARY INTELLIGENCE RES INST OF THE CHINESE PEOPLES LIBERATION ARMY ACAD OF MILITARY SCI

A method for image complexity evaluation based on convolutional neural network

This paper discloses a method for image complexity assessment based on convolutional neural networks. The method comprises the following steps: using a two-branch convolutional neural network to extract image detail features and semantic information respectively, and fusing these two features in the prediction phase; proposing a spatially distributed attention module specifically designed for image complexity, so that the extracted features can be adaptively optimized based on the spatial distribution of the features; in the prediction phase, the features are input into two prediction head branches, wherein the global complexity score prediction head predicts the global complexity of the input image using a fully connected neural network, and the local complexity heat map prediction head predicts the local complexity heat map of the input image using a convolutional neural network. The method demonstrates excellent complexity prediction performance on complexity assessment datasets, surpassing existing methods. For a given image, the method can accurately provide a global complexity score and a pixel-level complexity heat map.
Owner:NANKAI UNIV