Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

902 results about "Real image" patented technology

In optics, a real image is an image which is located in the plane of convergence for the light rays that originate from a given object. If a screen is placed in the plane of a real image the image will generally become visible on the screen. Examples of real images include the image seen on a cinema screen (the source being the projector), the image produced on a detector in the rear of a camera, and the image produced on an eyeball retina (the camera and eye focus light through an internal convex lens). In ray diagrams (such as the images on the right), real rays of light are always represented by full, solid lines; perceived or extrapolated rays of light are represented by dashed lines. A real image occurs where rays converge, whereas a virtual image occurs where rays only appear to diverge.

Deep forgery detection method based on visual language model

The invention discloses a deep forgery detection method based on a visual language model, and relates to the field of image forensics. The deep forgery detection method based on the visual language model aims to combine multi-source information to improve the discrimination capability of the model on a real image and a generated image. The method comprises the following steps: firstly, extracting image features through an image encoder of a pre-trained CLIP model; meanwhile, a frequency domain enhanced counterfeit perception adapter is embedded in the image encoder to mine potential anomalies of counterfeit images in the image domain and the frequency domain. Secondly, a manual feature extraction module is provided, discriminative low-dimensional features are extracted from the four aspects of the edge, the texture, the frequency and the symmetry of the image, and the discriminative low-dimensional features are used as auxiliary information input in the forgery detection process, so that the robustness and the interpretability of the model are improved; meanwhile, the text cue words are converted into feature vectors through a text encoder of a pre-training CLIP model; and finally, the model predicts a forgery score by calculating the cosine similarity between the image features and the text features so as to realize the discrimination of the authenticity of the image. According to the method, the problem that the detection capability of the model on the cross-dataset is insufficient is effectively improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Flow field measurement method based on event camera

The invention discloses a flow field measurement method based on an event camera, and the method comprises the steps: generating a PIV data set, each time sequence sample sequence comprising a plurality of frames of continuous particle images, a corresponding velocity vector field, and particle event data at all moments; establishing a flow field data acquisition device based on an event camera and a high-speed camera, acquiring real event data and real image data which are synchronous in time so as to adjust parameters of an event simulator, and verifying and updating particle event data in the PIV data set according to the adjusted event simulator so as to obtain a flow field data acquisition result; obtaining the updated PIV data set as a training data set; building an event camera optical flow method model, and training by adopting the training data set; and on the basis of the trained event camera optical flow method model, event sequences in two adjacent time periods are used as inputs to calculate a velocity vector field corresponding to a middle moment. According to the invention, the flow field velocity field at the required moment can be obtained based on the event data within a period of time.
Owner:ZHEJIANG UNIV

Road and bridge settlement displacement monitoring system and method based on image detection

The invention discloses a road and bridge settlement displacement monitoring system and method based on image detection, and the method comprises the steps: selecting a plurality of static background reference points in a stable background region of a monitoring scene, so as to construct a virtual and stable image internal reference system; furthermore, by accurately tracking image coordinate changes of the background reference points in the initial reference frame and the current frame, a transformation matrix capable of accurately describing disturbance of the camera from the initial pose to the current pose is reversely calculated, and the transformation matrix is applied to observation coordinates of a monitored target point; therefore, the virtual displacement component introduced by the camera disturbance is accurately stripped from the total displacement, and finally the real image displacement generated only by the motion of the structure is obtained. By means of the mode, the system can effectively resist interference of external factors such as environment vibration and temperature change, it is ensured that the height of the finally calculated physical displacement is close to the real settlement value of the structure, and therefore the accuracy and reliability of the monitoring result are greatly improved.
Owner:HEBEI JITONG ROAD&BRIDGE CONSTRUCT CO LTD

Sea temperature complementing method and system based on asynchronous diffusion Schrodinger bridge

The invention belongs to the technical field of sea temperature complementation, and discloses a sea temperature complementation method and system based on an asynchronous diffusion Schrodinger bridge, and the method comprises the steps: firstly, generating a weight anomal which reflects the abnormal degree of a pixel through a preprocessing step S1, so as to guide a subsequent diffusion process; s2, establishing a bidirectional diffusion path between the initial complementation image and the real image based on a diffusion Schrodinger bridge theory, and dynamically adjusting a diffusion coefficient according to anomay to generate an intermediate state xt; predicting a score function pred through a U-Net network S3, and reconstructing a current image x0 for updating a state or calculating loss; model parameters are optimized through iteration during training, multi-round denoising reconstruction is carried out during inference, and finally a complete high-quality sea surface temperature image SSTrecon is output. According to the invention, local details are fully reserved, and the accuracy of image completion is improved.
Owner:OCEAN UNIV OF CHINA

Titanium alloy microstructure prediction method and system based on conditional generative adversarial network and storage medium

The invention discloses a titanium alloy microscopic structure prediction method and system based on a conditional generative adversarial network and a storage medium, and belongs to the following steps: firstly, constructing a process-structure mapping model, and taking the output of the model as a rule constraint condition; inputting the random noise vector and the rule constraint condition into a conditional generative adversarial network to generate a prediction image; according to the generative adversarial network, thermal dynamic constraints based on physical quantities of microscopic structures are introduced in the training process, so that the interpretability of a prediction result is improved. And carrying out quantitative comparison on the predicted image and the real image, verifying the consistency of the statistical characteristics, and if the verification is passed, outputting a prediction result. According to the method, end-to-end prediction from process parameters to microscopic structure images is realized, the limitation that only symbolization or parameterization prediction can be carried out in a traditional method is broken through, and the intuition, the interpretability and the engineering application value of the method are remarkably enhanced.
Owner:SHANGHAI JIAOTONG UNIV

Face forgery detection algorithm for multi-view fusion processing based on style guidance

The invention discloses a face forgery detection algorithm based on style-guided multi-view fusion processing. The method comprises the following steps: firstly, carrying out standardized preprocessing on a face video sample, and extracting multi-scale image features based on an OfficientNet-B4 backbone network; by constructing a local texture map, an attention enhancement map and a style vector sequence, precise modeling and discrimination of a forged area are realized. The algorithm further utilizes a multi-branch sequence convolutional network to carry out time sequence modeling on fusion features, and outputs global style change representation for classification of forged and real images. The method comprehensively fuses the spatial texture, the semantic style and the time feature, has the advantages of high detection precision, strong generalization ability, good robustness and the like, and is suitable for complex and diverse depth forgery detection tasks.
Owner:NANJING TECH UNIV +1

Clear imaging method of mask under optical objective lens

The invention discloses a method for clearly imaging a mask under an optical objective, and relates to the technical field of optical detection and image processing, and the method comprises the following steps: S1, constructing a light intensity distribution model based on the boundary condition of the field of view of the optical objective in combination with the reflection characteristic and incident angle change rule of a metal edge material, and determining the coverage range of metal edge reflection crosstalk, generating a mask boundary interference prediction map; s2, according to the mask boundary interference prediction map, performing region division on the micro-reflectivity image, extracting brightness gradient characteristics of reflectivity lifting in an interference region, and generating a brightness gradient parameter set required by image filtering; the method is based on light intensity modeling, fusion direction filtering, pixel stripping, gradient reconstruction and credibility weighted fusion, realizes closed-loop control from interference prediction to pixel restoration, has high resolution and adaptivity, can accurately strip edge reflection artifacts and restore a real image structure, remarkably reduces misjudgment and rework rate, and is suitable for large-scale popularization and application. And the mask yield and the stability of the detection system are improved.
Owner:ZHONGKEZHUOXIN SEMICON TECH (SUZHOU) CO LTD

Nerve radiation field rendering method based on dynamic hash coding

The invention discloses a neural radiation field rendering method based on dynamic hash coding, and the method comprises the steps: employing the feature sequence data as the input, calculating the density value and color value of each sampling point through the forward propagation of a neural network, carrying out the volume rendering integral operation according to the ray tracing principle in the direction of a ray, and obtaining the feature sequence data; judging a final color output result of the current pixel point; according to an error value between the color output result and a real image, updating a network parameter weight through a back propagation algorithm, and if the error value is greater than a convergence threshold, continuing to iterate the training process to adjust a feature coding strategy to obtain an optimized neural radiation field model parameter; and after the rendering performance configuration parameters are obtained, optimizing a storage allocation strategy of feature data through a memory pool management mechanism, and if the current memory occupancy rate exceeds a safety threshold, starting a data compression algorithm to reduce the storage space requirement, and obtaining a real-time rendering output result. According to the invention, high-quality real-time rendering of the dynamic scene is realized.
Owner:ZHEJIANG UNIV OF TECH

Medical image cross-modal generation method and device based on wavelet high-frequency enhancement

The invention discloses a medical image cross-modal generation method and a medical image cross-modal generation device based on wavelet high-frequency enhancement, which have the following effects: common features and unique features of a multi-modal medical image are effectively learned by using multi-scale local-global features and high-frequency texture detail information, and an accurate and fine target modal image is obtained. The method can effectively deal with the limitation of medical conditions and reduce the cost of obtaining multi-modal images, and has great application value. A CMMB block of the global branch encoder aggregates global information and multi-scale local features, fully learns information of different anatomical structures and muscle textures, and generates fine edge textures and tissue details of a target modal image; an RSTB block of the high-frequency branch encoder extracts high-frequency features from an input mode and aggregates the high-frequency features with global branches, so that the texture fidelity of a generated image is promoted, and the image is close to a real image; the CAG mechanism of the decoder promotes feature interaction between the decoder and the encoder, redundant information is removed, and an efficient image is generated.
Owner:HUNAN UNIV

Deep forgery detection method and system, storage medium and computer equipment

The invention relates to the technical field of deep counterfeit image detection, and discloses a deep counterfeit detection method and system, a storage medium and computer equipment. The method comprises the following steps: firstly, constructing a reference data set containing a forged image and an original real image; secondly, through an integrated model, generating antagonistic samples for the reference data set, and integrating the successfully attacked antagonistic samples into an antagonistic sample set; and finally, merging the reference data set and the adversarial sample set, and constructing a robustness enhanced data set containing four types of samples. In the model training stage, multi-classification cross entropy loss and comparative learning loss are combined, and expression of the model in a feature space is optimized through comparative learning constraint, so that the model learns discriminative features with more compact intra-class features and more dispersed inter-class features. The model trained by the method not only can effectively defend against attack and improve robustness, but also surpasses original detection performance on clean samples, and has remarkable technical advantages and application value.
Owner:GUANGDONG UNIV OF TECH

Neural radiation field high-fidelity representation method and system based on geometric spectrum coupling

The invention discloses a geometric spectrum coupling-based neural radiation field high-fidelity representation method and system. The method comprises the following steps of: obtaining sparse multi-view image data and calculating internal and external parameters of a corresponding camera; generating light based on internal and external parameters of the camera, and performing point sampling on the space along the direction of the light to obtain sampling points; constructing a differentiable coding layer, predicting the density of sampling points through forward propagation of a neural network, and obtaining a principal curvature, a density gradient and an included angle between a normal vector of the sampling points and an observation direction based on a density prediction result; constructing a spectrum-geometric coupling function, and calculating to obtain a spectrum coupling coding vector; the spectrum coupling coding vector is used as an input to be transmitted into a NeRF main network, and NeRF volume rendering is carried out according to a density and color prediction result; and performing difference optimization training on the rendering result and the real image to obtain a high-fidelity composite image. According to the invention, the expression ability of the neural radiation field model in the high-frequency detail area can be enhanced.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Optical lens

The invention provides an optical lens, which comprises four lenses with focal power and sequentially comprises a first lens with negative focal power, a second lens with negative focal power, a third lens with positive focal power, a fourth lens with negative focal power, a fifth lens with negative focal power and a sixth lens with negative focal power from an object side to an imaging surface along an optical axis, the object side surface of the second lens is a convex surface, and the image side surface of the second lens is a concave surface; the object side surface of the third lens is a convex surface, and the image side surface of the third lens is a convex surface; the image side surface of the fourth lens is a convex surface; wherein the maximum field angle FOV of the optical lens and the real image height IH corresponding to the maximum field angle of the optical lens meet the following conditions: 35 degrees / mmlt; fOV / IHlt; 46 degrees / mm; the focal length f3 of the third lens and the focal length f4 of the fourth lens satisfy 0.5 lt; f3 / f4lt; and 0.8. According to the optical lens provided by the invention, through specific surface shape matching and reasonable focal power distribution, the lens has one or more advantages of an ultra-wide angle, a large aperture, high imaging quality, high collimation and the like.
Owner:JIANGXI LIANCHUANG ELECTRONICS CO LTD

Visual detection optimization control method, device and equipment based on digital twinning and storage medium

The invention discloses a visual detection optimization control method, device and equipment based on digital twinning and a storage medium, and relates to the technical field of visual detection, and the method comprises the steps: obtaining an analog image and a real image in a digital twinning environment, and carrying out the frequency domain feature transformation of the analog image and the real image, thereby obtaining an image spectrum feature; a frequency domain alignment model is established based on multi-band spectrum envelope guide residual mapping, and the structure of virtual and real image spectrum features is kept aligned; further extracting features through multi-scale convolution and channel dependence mapping to obtain virtual-real fusion features; quantifying channel similarity and establishing a covariance regularization constraint, performing channel correction on the cross-domain features, and eliminating feature drift to obtain second virtual-real fusion features; and finally, performing visual detection and micro defect identification based on the features. The problem that a virtual sample and a real sample are different in local texture structure and channel distribution is solved.
Owner:SUZHOU HENGZHI INTELLIGENT TECH CO LTD

Track bolt detection method and system

The invention belongs to the field of bolt defect detection, and particularly relates to a track bolt detection method and system. According to the obtained bolt real image of each track bolt to be detected on the production line, the geometric appearance size value of each bolt is obtained through bolt edge information obtained through edge detection; whether the size of each bolt is qualified or not is judged according to whether the geometric appearance size value of each bolt meets the set size qualification condition or not; bolts with unqualified sizes are judged as unqualified products; and respectively inputting the bolt real image of each to-be-detected bolt with the qualified size into the trained defect detection model, and according to a defect detection result output by the defect detection model, obtaining a judgment result containing whether each bolt with the qualified size has a defect or not and a defect position of the bolt with the defect according to the judgment result. The detection process almost does not need manual participation, and the method can guarantee the accuracy and reliability of the detection result on the basis of guaranteeing the high detection efficiency.
Owner:XINYANG AEROSPACE FASTENER FACTORY

Underwater visible light signal recovery method based on transfer learning and communication system

The invention relates to an underwater visible light signal recovery method based on transfer learning, and aims to solve the problems of signal noise, distortion and communication performance reduction caused by disturbance of water turbidity, flow velocity, ambient light intensity and the like on an underwater visible light communication (UVLC) system. In combination with transfer learning and a generative adversarial network (GAN), data are acquired through a self-developed hardware platform, background data are generated through MATLAB simulation, simulation and real image features are fused by using a conditional generative adversarial network (cGAN) to generate simulation data, signal recovery capability is trained by using a U-NetGAN model, and finally the system is deployed at a receiving end. Compared with the prior art, the scheme overcomes the problems that traditional modulation is poor in adaptability under disturbance of water turbidity, flow velocity and the like and is difficult to deal with a complex underwater environment; the defects that a deep learning method depends on large-scale data and cross-scene generalization is weak are overcome. Through the combination of transfer learning and GAN, the UVLC system communication performance is significantly optimized, the delay is low, the complexity is low, and actual deployment requirements are better met.
Owner:HUZHOU UNIVERSITY

Camera calibration method and device, electronic equipment and storage medium

The invention provides a camera calibration method and device, electronic equipment and a storage medium, and belongs to the technical field of data processing, and the method comprises the steps: obtaining laser radar point cloud data, inertial measurement unit data and reference camera image data of a vehicle, and obtaining initial calibration parameters of a to-be-calibrated camera; based on the laser radar point cloud data, the inertial measurement unit data and the reference camera image data, constructing a point cloud map containing color textures and a pose track of a vehicle, and converting the point cloud map into a surface grid model; according to the pose track and the initial calibration parameters, the surface grid model is rendered to a visual angle of the to-be-calibrated camera, and a rendered image and a rendered depth map corresponding to the to-be-calibrated camera are generated; based on the real image data, the rendered image and the rendered depth map, establishing a corresponding relationship between two-dimensional image features and three-dimensional space points in the real image data; and constructing a re-projection error objective function based on the corresponding relationship, and optimizing the external reference and the internal reference of the camera to be calibrated by minimizing the objective function.
Owner:IFLYTEK CO LTD

Three-dimensional reconstruction method, device and equipment based on 3D Gaussian

The invention discloses a three-dimensional reconstruction method, device and equipment based on 3D Gaussian. The method comprises the steps of performing data preprocessing based on a three-dimensional point cloud output by sparse reconstruction of a target scene, a camera internal reference and a camera pose; initializing and training a 3D Gaussian model based on a real image, the three-dimensional point cloud after candidate object mask screening, a camera pose, a depth information graph and a scale alignment parameter to obtain a coarse-grained 3D Gaussian model of the target scene; generating a dynamic target mask by taking a rendered image generated by rendering the target scene based on the coarse-grained 3D Gaussian model as a reference and combining the real image and the candidate object mask; and on the basis of the coarse-grained 3D Gaussian model, the dynamic target mask and all input data obtained by initializing and training the coarse-grained 3D Gaussian model, carrying out refined reconstruction and optimization on the coarse-grained 3D Gaussian model. According to the method, unification of high precision, high efficiency and high robustness is realized, and the increasing actual industrial application requirements are met.
Owner:SHENYANG MXNAVI CO LTD

Quartz surface defect detection method based on visual identification

The invention discloses a quartz surface defect detection method based on visual identification, and the method comprises the following steps: S1, carrying out the image collection of a quartz surface through a high-resolution camera, and carrying out the preprocessing of the collected image; s2, carrying out edge extraction on the image by adopting a Canny edge detection algorithm; s3, generating more defect images by using a hybrid generation model based on a variational auto-encoder and an energy guide mechanism; s4, combining the generated defect image and the real image, and performing defect classification through an image segmentation network in combination with a self-attention mechanism; s5, jointly training the improved variational auto-encoder and the image segmentation network; and S6, feeding back the category of the defect and the edge information of the defect to a production line control system in real time. According to the method, a high-resolution image acquisition technology, a variational auto-encoder and an image segmentation network are combined, and automatic detection and accurate classification of quartz surface defects are realized through fusion of deep learning and a traditional image processing method.
Owner:SUZHOU ANYI ROBOT TECHNOLOGY CO LTD

Dynamic scene reconstruction method, computer equipment and program product

The invention discloses a dynamic scene reconstruction method, computer equipment and a program product. According to the method, picture information and camera poses of the same scene are obtained, a standard field composed of three-dimensional Gaussian primitives is constructed in a three-dimensional space, a deformation field is constructed, the standard field and the deformation field at different times are subjected to joint rendering in a three-dimensional Gaussian splashing mode, and various parameters are trained and optimized according to the difference between a synthetic image and a real image. Performing motion perception according to the motion condition of the three-dimensional Gaussian primitives, taking the primitives with smaller motion as static primitives, and performing partition modeling on the primitives with larger motion in time to obtain a layered dynamic scene model; and calling the model for rendering according to the given time and the camera pose in the reasoning stage, and outputting a target image. According to the method, the reconstruction precision of the high-dynamic target can be improved while the rendering efficiency is ensured, and the continuity and visual quality of the dynamic scene in the time dimension are improved by introducing the time consistency constraint at the boundary of the time subinterval.
Owner:ZHEJIANG UNIV

Multi-feature 3D (three-dimensional) Gaussian reconstruction method based on laser vision

A multi-feature three-dimensional reconstruction 3D Gaussian method based on laser vision comprises the steps that laser radar point cloud and camera images are aligned through space-time calibration, and a unified coordinate system is established; extracting geometric features by using point cloud data acquired by the Lidar point cloud, and initializing a Gaussian ellipsoid according to the Lidar point cloud; optimizing the brightness, the contrast ratio and the structural similarity of the rendered image and the real image by combining the mean absolute error L1 and the structural similarity SSIM; the curvatures of Gaussian ellipsoids of K-nearest neighbors are forced to be consistent, and long and short axes and line and surface features of the Gaussian ellipsoids are aligned to reduce geometric distortion; the distribution density of 3D Gaussian is dynamically adjusted through line / surface features and visual structure information extracted by Lidar, and balance between geometric detail enhancement and calculation efficiency is achieved. According to the method, the position, the scale and the rotation parameters of Gaussian are uniformly optimized, and the details and the calculation efficiency of the model are balanced while the consistency of the model structure is improved.
Owner:CHINA UNIV OF MINING & TECH

Head reconstruction method based on Gaussian sputtering

The invention discloses a head reconstruction method based on Gaussian sputtering, and belongs to the field of computer vision. Performing alignment and clone splitting on the initial point cloud model by using the supervision image to obtain a head accurate point cloud of the supervision image, and taking the head accurate point cloud position as a Gaussian point cloud position; projecting the Gaussian point cloud to an image coordinate system, and performing feature extraction by taking the supervised image as convolutional neural network input to obtain a convolutional feature map; sampling the head precise point cloud and the convolution feature map of the image coordinate system to obtain a feature vector, and decoding the feature vector to obtain a Gaussian point cloud parameter so as to obtain a three-dimensional head Gaussian model; a head image of any viewpoint is quickly rendered through the three-dimensional head Gaussian model, pixel-by-pixel loss calculation is performed on the rendered image and a real image, and the three-dimensional head Gaussian model is optimized. According to the method, under the condition of sparse input, the convolutional neural network and Gaussian sputtering are fitted, and the obtained high-fidelity three-dimensional head Gaussian model can be rendered quickly in real time and has generalization.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Titanium alloy microstructure grain image generation method based on potential diffusion model

The invention discloses a titanium alloy microstructure grain image generation method based on a potential diffusion model, and belongs to the technical field of image generation, and the method comprises the following steps: obtaining seven mechanical properties and corresponding image data of titanium alloy microstructure grains; preprocessing the image data, and forming a data set corresponding to the mechanical properties and the image data; constructing a titanium alloy microstructure grain generation model, designing a continuous condition encoder, and taking mechanical properties as conditions to be embedded into the model for training; and training the generative model by using the data set, wherein the trained model can generate a corresponding microstructure grain image according to the input mechanical properties. According to the method, the potential diffusion model is combined with mechanical property condition input, high-quality generation of the titanium alloy microstructure grain image is achieved, the generated image has high similarity with a real image, the input mechanical property parameters can be effectively correlated, and the controllability of microstructure image generation and the generalization ability of the model are improved.
Owner:BEIJING INST OF TECH

System and method for railway foreign object detection

A computer-implemented system for foreign object detection in a scene. The system includes a memory-suppress diffusion network module adapted to reconstruct a reconstructed image from an encoded image, and a contrastive dissimilarity network adapted to combine the input image and the reconstructed image to predict an anomaly map for the input image. The encoded image is based on an input image, and the memory-suppress diffusion network module and the contrastive dissimilarity network are trained using only normal, real images. The system leverages only normal images in training and does not compromise the detection performance at the inference stage.
Owner:CITY UNIVERSITY OF HONG KONG

Projection lens

The invention provides a projection lens, which comprises five lenses in total, and sequentially comprises a first lens with negative focal power, a second lens with negative focal power, a third lens with negative focal power, a fourth lens with negative focal power, a fifth lens with negative focal power and a sixth lens with negative focal power from a projection surface to an image source surface along an optical axis, the surface of the projection side of the second lens is a convex surface, and the surface of the image source side of the second lens is a convex surface; the surface of the projection side of the third lens is a concave surface, and the surface of the image source side of the third lens is a concave surface; the surface of the projection side of the fourth lens is a convex surface, and the surface of the image source side of the fourth lens is a convex surface; the surface of the projection side of the fifth lens is a convex surface, and the surface of the image source side of the fifth lens is a convex surface; wherein the real image height IH corresponding to the maximum field angle of the projection lens and the effective focal length f of the projection lens meet the following conditions: 0.6 lt; iH / flt; and 0.8. According to the projection lens provided by the invention, the projection quality of the projection lens is improved through the reasonable configuration of the surface types of the lenses and the reasonable matching of the focal power.
Owner:HEFEI LIANCHUANG OPTICAL CO LTD

Prompt-to-prompt image editing with cross-attention control

Some implementations are directed to editing a source image, where the source image is one generated based on processing a source natural language (NL) prompt using a Large-scale language-image (LLI) model. Those implementations edit the source image based on user interface input that indicates an edit to the source NL prompt, and optionally independent of any user interface input that specifies a mask in the source image and / or independent of any other user interface input. Some implementations of the present disclosure are additionally or alternatively directed to applying prompt-to-prompt editing techniques to editing a source image that is one generated based on a real image, and that approximates the real image.
Owner:GOOGLE LLC

Pavement PBR material inversion method and system

The invention provides an inversion method and system for a pavement PBR material, and belongs to the field of computer vision and physical simulation, and the method comprises the steps: carrying out the space-time alignment of obtained IMU data, three-dimensional pavement data and two-dimensional pavement texture images, and carrying out the reconstruction of a vehicle track; generating parameter combination sample points in a multi-dimensional parameter space of the constructed initial simulation model; simulating the parameter combination sample points to calculate a loss value between a simulation image and a real image, and constructing a data set; a data set machine learning agent model is used for training, the trained model is used for recognition, and then sensitivity analysis is carried out to obtain a parameter subset with the strongest coupling effect; and rendering a simulation image by using the parameter subset, calculating the difference between the simulation image and a realistic image, synchronously updating the parameter subset, and outputting a pavement PBR material map and system parameters during convergence. Based on the method, the invention further provides an inversion system of the pavement PBR material, and high-fidelity pavement PBR material diagram and system parameter output are realized.
Owner:ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3

Image classification model generation method and device, equipment and storage medium

The invention belongs to the technical field of neural networks, and discloses an image classification model generation method and device, equipment and a storage medium. The method comprises the following steps: acquiring a training image and a real image category; predicting the training image by the plurality of teacher models to obtain a soft label; the student model predicts the training image according to different temperature parameters to obtain a soft prediction result and a hard prediction result; constructing a total loss function according to the real image category, the hard prediction result, the soft label and the soft prediction result, and optimizing the student model to obtain an image classification model; according to the invention, a plurality of teacher models are integrated to generate the soft label, low-temperature prediction and high-temperature prediction are carried out in combination with temperature adjustment, the model can be ensured to finally output accurate category judgment when the complex category relationship of the teacher soft label is focused on learning, and the student model which is light and rapid and can maintain high classification precision is trained. And the contradiction between the model precision and the efficiency is effectively solved.
Owner:WUHAN UNIV OF TECH

Industrial part small sample target detection method based on synthetic data

The invention discloses an industrial part small sample target detection method based on synthetic data, and the method consists of a synthetic data generation module and an enhanced target detection model named YOLO-DC, and comprises the steps: separating a training process of the target detection model from dependence on large-scale real labeled data; and training the YOLO-DC model only by using virtual data generated by the synthetic data generation module. The trained model can be directly deployed in a real physical environment to accurately detect industrial parts, so that a visual perception task is completed. According to the method, a user can quickly deploy a detection system adaptive to new parts without collecting and marking real images, so that the development and deployment period of an industrial visual system is shortened, the data cost is remarkably saved, and meanwhile, the flexibility and the intelligent level of a manufacturing system are greatly improved.
Owner:YANTAI ZHONGKELANDE CNC TECH CO LTD

Time sequence prediction method based on multi-modal contrast learning technology

According to the time series prediction method based on the multi-modal contrast learning technology, original multivariable time series data are converted into structured visual representation and language representation, multi-modal representation with consistent inner performance can be constructed without depending on external natural language or real image data, and the time series prediction method is high in practicability. The deep semantic understanding capability of the model on the complex operation state of the rail transit is effectively enhanced; a multi-modal contrast learning mechanism is introduced, visual and text modal representation is aligned in a shared embedding space, positive sample consistency is maximized through InfoNCE loss, negative sample interference is suppressed, and the robustness and generalization ability of time sequence features are remarkably improved; and the importance of each variable on a prediction task is dynamically evaluated by using the aligned multi-modal representation, and key variables are automatically screened, so that the redundant information interference is reduced, and the prediction precision and the calculation efficiency of the model in a high-dimensional multi-variable scene are also improved.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Simulation-to-real image migration method and device for unmanned system

The invention discloses a migration method and device from simulation to a real image for an unmanned system, and the method comprises the steps: constructing a double-flow architecture comprising a convolution encoder and a visual Mama encoder, wherein the convolution encoder is used for extracting the local texture information of an image, and the visual Mama encoder employs a visual state space model to capture the global context information of the image; and the two features are input into a decoder after channel dimension fusion, and a real style image after style migration is generated. In order to improve the unsupervised learning ability of the model, a cyclic consistency training framework is introduced, training can be completed under the condition of no paired image samples, and the image structure consistency is kept. According to the method, the sense of reality and the structural fidelity of a simulation image are effectively improved, and compared with an existing method, the method shows lower distortion and higher generalization ability in an image translation task. The method can be widely applied to simulation data field adaptation and training sample enhancement tasks in the fields of automatic driving, virtual simulation, robot vision and the like.
Owner:WUHAN UNIV +1