Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

105 results about "Computational vision" patented technology

Image review method fusing semantic comprehension and visual identification

The invention belongs to the technical field of machine room safety monitoring, and discloses an image review method fusing semantic understanding and visual recognition, which comprises the following steps: calculating visual / semantic feature dynamic credibility in real time through an exponential weighted moving average algorithm in combination with environment interference and equipment state parameters; resNet50 is adopted to extract visual features in global and local branches, and a BERT model is adopted to encode semantic features; correcting the feature correlation degree according to the scene, calculating a dynamic weight, carrying out heterogeneous calibration through an attention mechanism, and calling priority rules such as'physical security features are higher than behavior features' for arbitration during conflicts; the method is advantaged in that low-credibility feature interference fusion is avoided, abnormity identification accuracy in a complex scene of a machine room is improved, multi-modal data cooperation demands are adapted, scene labels are marked based on time, work orders and historical data dimensions, an exclusive sub-model is constructed for a high-density scene through DBSCAN density clustering, and low-frequency scene parameters are migrated to a similar model.
Owner:QINGYUN CLOUD COMPUTING (SHENZHEN) CO LTD

Robot dynamic grabbing control method based on visual sense and tactile sense depth fusion

The invention relates to a robot dynamic grabbing control method based on visual sense and tactile sense depth fusion. The method comprises the following steps that firstly, based on visual information, surface geometric features of a target object are extracted, grabbing adaptability scores are calculated, an optimal grabbing point is selected, and a collision-free approaching track is planned; secondly, in a contact establishment stage, detecting a contact event through a touch sensor, fusing vision and touch data to unify a coordinate system, calculating a vision-touch consistency score, and applying an initial grabbing force; and thirdly, in the stable holding stage, slip detection is conducted through wavelet packet energy entropy and pressure gradient, and the grabbing force and impedance parameters are dynamically adjusted in combination with self-adaptive impedance control. According to the method, intelligent control over the whole process from approaching to contact to stable holding is achieved, and the adaptability, stability and safety of robot grabbing in the dynamic environment are improved.
Owner:INEXBOT

Bidirectional visual semantic interaction enhancement method and system for generalized zero sample learning

The invention provides a bidirectional visual semantic interaction enhancement method and system for generalized zero sample learning, and the method comprises the steps: collecting visual samples containing visible classes and unvisible classes and corresponding cross-class semantic descriptions in a generalized zero sample learning scene; respectively generating a global visual feature vector and a semantic word vector through a visual feature extractor and a semantic encoder; fusing the two by using a multi-head attention mechanism to generate a semantic enhanced visual embedding representation; then visual-to-semantic and semantic-to-visual two-way comparative learning with expert attribute vectors as alignment targets and starting points is carried out, and visual semantic embedding and attribute vector alignment are promoted in combination with attribute regression loss; then training a bidirectional visual semantic interaction enhancement model based on multi-loss joint optimization; and finally, calculating visual semantic embedding for a test image by utilizing the trained model, and outputting a classification result of a visible class or a non-visible class by combining similarity calculation with a candidate class attribute vector with a self-calibration item.
Owner:WUHAN TEXTILE UNIV

Self-adaptive visual training method based on eye characteristics

The invention relates to a self-adaptive visual training method based on eye features, and belongs to the technical field of visual health and artificial intelligence. In order to solve the problems that traditional visual training equipment is heavy in structure, single in training mode and lack of personalized regulation and control and concentration state monitoring, the visual training equipment obtains eye images of a user in real time through an image acquisition module, adopts visual intelligent analysis, extracts double features of eye contours and pupil positions, calculates the visual concentration degree, and improves the visual training efficiency. And a concentration state is judged by combining a dynamic self-adaptive threshold value. When the concentration degree is insufficient, the control module triggers intervention mechanisms such as prompt or training pause and the like, and the remote interaction module uploads data to the cloud platform to generate a personalized training scheme. Intelligent and personalized regulation and control of the training process are achieved, the training effect and compliance are improved, the system is compact in structure and suitable for various scenes, and an efficient solution is provided for vision health.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Method for driving emotion interaction of intelligent device based on multi-modal understanding

The invention relates to the technical field of data processing, in particular to a method for driving emotion interaction of an intelligent device based on multi-modal understanding, and aims to eliminate illumination and noise interference and output a standardized face video stream, an effective voice segment and a touch thermodynamic diagram through an environment adaptive acquisition module. The feature extraction module extracts facial action optical flow features, voice Mel-frequency cepstral coefficient vectors and tactile pressure gradient parameters. The cross-modal correlation model adopts a tensor decomposition algorithm to calculate a space-time correlation matrix of visual and voice features, and the tactile feature weight is dynamically adjusted in combination with environmental parameters. According to the response strategy, an intervention scheme is retrieved based on a graph database, emotion confirmation statements, guide statements and behavior suggestions are fused to generate multi-mode response, and PID adjustment of the temperature control device and tactile pulse output of the vibration device are synchronously driven. And the feedback evaluation module verifies the emotion recognition consistency through a Pearson's correlation coefficient, triggers conflict sample separation storage and model increment training, and realizes closed-loop optimization.
Owner:BEIJING HAOXINQING MOBILE MEDICAL TECH CO LTD

Weakly supervised group behavior identification method based on dynamic prompt tuning

The invention discloses a weak supervision group behavior recognition method based on dynamic prompt tuning, and belongs to the field of video understanding. According to the method, a pre-trained vision-language model is adaptively expanded to a group behavior recognition task for challenges such as complex group behavior semantics, strong spatio-temporal context dependency and lack of individual labels in a video. According to the technical scheme, a dynamic prompt generation technology with visual conditions is provided, instance-level text prompts can be automatically generated according to input video content, the accuracy of vision-text semantic alignment is enhanced, meanwhile, a time sequence fusion module is integrated in a model, and the purpose is to fuse key time information in multiple frames through modeling inter-frame time sequence dependence. The model calculates a similarity score between visual and text features and takes the similarity score as a classification prediction basis, and finally, a cross entropy loss function is adopted to carry out end-to-end training on the network. The validity of the method is verified on a volleyball data set and an NBA data set.
Owner:BEIJING UNIV OF TECH

Sheep manure conveying blockage identification method and system based on machine vision

The invention belongs to the technical field of image recognition, and particularly relates to a sheep manure conveying blockage recognition method and system based on machine vision, and the method comprises the steps: firstly fusing the local brightness deviation and gradient features of pixels, and calculating a visual interference index to generate a stability mask; a visual interference area caused by highlight, shadow and the like on the surface of the material is shielded; then, space gating is carried out on an inter-frame difference result of the image by using the mask, a robust motion feature map is generated, and a static connected region is extracted from the robust motion feature map; and finally, calculating a blockage index by combining the area of the static region and the average static degree, and comparing with a decision threshold to judge a blockage event and generate a response signal. According to the method, the visual interference degree under the complex working condition can be reduced, and the accuracy and the automation degree of conveying system blockage event recognition are improved.
Owner:SHAANXI YATAI DAIRY CO LTD

Inertia / satellite / vision adaptive integrated navigation method and system

The invention provides an inertia / satellite / vision adaptive integrated navigation method and system. The method comprises the following steps: carrying out inertia measurement and navigation calculation to obtain inertia data; calculating satellite navigation observed quantity according to the satellite navigation data and the inertial data, and constructing a satellite observation model; calculating visual navigation observed quantity according to the visual navigation data and the inertial data, and constructing a visual observation model; aiming at a satellite and a visual observation model, respectively adopting an innovation-based adaptive covariance estimation method to obtain satellite and visual observation noise covariance updated values; constructing a satellite and visual navigation health degree, and calculating a satellite and visual fusion weight according to the satellite and visual navigation health degree; controlling updating of satellite and visual navigation observed quantity according to the satellite and visual fusion weight; and calculating a weighted equivalent observation matrix and a noise covariance, and carrying out filtering estimation. According to the method, an adaptive noise estimation and sensor health degree evaluation mechanism is introduced, dynamic weighted fusion of multi-sensor data is realized, and the robustness and precision of a navigation system in complex environments such as satellite signal lock losing and visual feature missing are improved.
Owner:BEIJING AUTOMATION CONTROL EQUIP INST

Short video content accurate recommendation system based on artificial intelligence image recognition

The invention discloses a short video content accurate recommendation system based on artificial intelligence image recognition, particularly relates to the technical field of short video personalized recommendation, and is used for solving the problem of matching of user interest and visual bearing capacity. The method comprises the following steps: extracting short video key frame spatial-temporal characteristics through a multi-scale convolutional neural network, constructing a cognitive load threshold curve in combination with a user micro gesture sequence, representing the tolerance range of a user to visual complexity in different time periods, and generating a visual complexity vector; calculating a visual load matching index by using an adaptive dynamic warping algorithm, and adjusting a candidate video sequence through a load penalty factor to generate an initial recommendation probability; modeling a user dynamic interest vector based on a gated loop unit network in combination with an attention mechanism; and finally, user interests and video semantics are fused through multi-dimensional features, a final recommendation list is generated through a multi-objective optimization algorithm under the constraint of visual load, and personalized recommendation of interest matching degree maximization and visual comfort optimization is realized.
Owner:ANHUI JUYUN ZHONGLIAN NETWORK TECHNOLOGY CO LTD

Robot vision calibration device and method

According to the robot vision calibration device and method, through collaborative design of the industrial personal computer, the vision system, the robot, the calibration jig and the calibration plate, the problems caused by complex structure, low calibration precision and poor adaptability in the prior art are solved. A floating mechanism of the calibration jig is combined with a displacement sensor to realize multi-degree-of-freedom dynamic displacement compensation, so that mechanical positioning errors are remarkably reduced; the conical surface guide structure and the reference hole layout of the calibration plate optimize the stability of the positioning reference, and ensure the accurate embedding of the calibration needle. The flexible installation of the visual system adapts to different scene requirements. According to the calibration method, the generation of a high-precision visual calibration matrix is realized through the processes of multi-reference hole coordinate recording, calibration ring coordinate calculation and visual image acquisition and matching. According to the whole scheme, through a dynamic compensation mechanism and intelligent process design, the calibration precision, efficiency and robustness are improved, meanwhile, the operation complexity is reduced, and reliable technical support is provided for an industrial automation scene.
Owner:伯朗特机器人股份有限公司

Subject entity labeling method and system fusing image recognition and knowledge graph

The invention discloses a subject entity labeling method and system fusing image recognition and a knowledge graph, and relates to the technical field of image recognition and natural language processing. The method comprises the following steps: carrying out preprocessing and image-text association on multi-source heterogeneous subject data; detecting a visual entity in the image through an improved YOLO model, and extracting and linking a text entity in combination with a subject dictionary and a knowledge graph; cross-modal collaborative disambiguation is realized by calculating the semantic similarity of visual candidate entities and text context vectors; multi-modal entities are combined, relation reasoning and enrichment labeling are carried out in a knowledge graph, and a deep labeling result containing the entities and a semantic relation network of the entities is generated; the problems of difficulty in multi-source data fusion, inaccurate professional entity recognition and difficulty in semantic ambiguity elimination are effectively solved, the depth and accuracy of subject knowledge semantic understanding are remarkably improved, and key technical support is provided for intelligent education application.
Owner:CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP

Multi-modal cross-domain small sample facial expression recognition method based on relation distillation self-paced learning

The invention discloses a multi-modal cross-domain small sample facial expression recognition method based on relational distillation self-paced learning, and relates to a computer vision technology. Constructing a multi-modal semantic enhancement module, generating semantic descriptions of expressions by using a large language model, performing CLIP coding, and performing alignment and fusion with image visual features in the multi-modal semantic enhancement module to construct a multi-modal prototype; a self-paced learning mechanism based on relational distillation is designed, visual and semantic structural errors are calculated, and progressive training from easy to difficult is realized through a soft and hard mixed sample selection strategy and a mixed sample selection mechanism regulated and controlled by a dynamic threshold value. And cross-domain migration of emotional knowledge from basic expressions to fine-grained composite expressions can be effectively realized. And under the condition that only a small number of labeled samples are provided, rapid adaptation and accurate recognition of new expressions can be realized. The method is remarkably superior to a traditional supervised learning method, has higher practicability and expansibility, and can better meet the requirement for efficient recognition of new expressions in practical application.
Owner:XIAMEN UNIV

Robot dynamic obstacle trajectory prediction navigation method based on visual model

The invention provides a robot dynamic obstacle trajectory prediction navigation method based on a visual model, and relates to the field of power equipment maintenance, the method comprises the following steps: collecting original data of dynamic scene multi-modal perception, carrying out space-time alignment, extracting thermal infrared dynamic feature vectors and visual dynamic feature vectors, and carrying out dynamic scene multi-modal perception on the thermal infrared dynamic feature vectors and the visual dynamic feature vectors; simultaneously calculating the fuzzy degree of the visual image and the dynamic complexity of the scene; calculating a fusion weight of the thermal infrared feature and the visual feature through a dynamic modal weight algorithm, and outputting a fusion feature vector; inputting the fusion feature vector into a visual model, and calculating the trajectory probability distribution of the dynamic obstacle; calculating comprehensive uncertainty and risk coefficients based on the trajectory probability distribution; and based on the fusion feature vector and the trajectory probability distribution, integrating the uncertainty and the risk coefficient, and planning an obstacle avoidance path and a decision report in combination with an A algorithm, so as to realize the dynamic navigation of the robot. According to the technical scheme, the reliability and intelligence of robot navigation can be improved, and the navigation requirement of a complex environment is met.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Machine room robot inspection system based on industrial vision

The invention relates to the technical field of robot inspection, in particular to a machine room robot inspection system based on industrial vision, which comprises a data acquisition module used for acquiring a visible light image, a current temperature and a current electromagnetic field intensity value of a target inspection point location, and a visual anomaly score output by a visual detection model; the visual complexity quantification module is used for calculating a visual complexity index based on the visible light image; the basic anomaly evaluation module is used for calculating a basic anomaly score based on the current temperature, the current electromagnetic field intensity value and preset historical statistical data; the dynamic fusion module is used for determining a dynamic weight based on the visual complexity index; the comprehensive decision-making module is used for generating a comprehensive abnormal score based on the dynamic weight, the visual abnormal score and the basic abnormal score, passive adaptation to the environment is converted into active quantification and dynamic compensation to the environment, and therefore the intelligent level and reliability of inspection decision making are improved.
Owner:RUNHE WORLD UNION DATA TECH CO LTD

Aviation landing gear cable teleoperation collaborative assembly system and method

The invention relates to the technical field of aeronautical manufacturing and robot assembly, and discloses an aeronautical landing gear cable teleoperation collaborative assembly system and method, and the system comprises a master end interaction subsystem, a slave end assembly subsystem and a central control processing subsystem. The method comprises the following steps: acquiring multi-modal data through an environment monitoring camera, a six-dimensional force sensor and a touch sensor, and realizing space-time registration and state estimation by using Kalman filtering; and planning an optimal path by using a genetic algorithm based on the digital twinborn model and providing augmented reality guidance. And the system calculates visual confidence in real time, and adaptively switches a vision dominant control mode according to the visual confidence: when the vision is limited, the rigidity of the mechanical arm is automatically reduced, the force feedback weight is enhanced, and an operator is assisted to complete compliant operation in a blind area. According to the invention, the problems of sensing shielding and precise control of flexible cable assembly in a narrow cabin are solved, and the safety and efficiency of assembly operation are improved.
Owner:河北工业职业技术大学

Film space position prediction method and system based on machine vision

The invention belongs to the technical field of image processing, and particularly relates to a film space position prediction method and system based on machine vision, and the method comprises the steps: collecting a film operation image, extracting a transverse projection value of an edge sub-pixel point, and synchronously obtaining physical field data such as linear speed, mechanical wear and tension; a dynamic drift index is determined by using differential operation, and a steady-state characteristic index is solved by combining fluid dynamics and a physical constant; the spatial displacement of the film in a visual processing hysteresis stage is accurately predicted by calculating the processing time consumption of a visual system based on dynamic drift, steady-state characteristics and a dynamic second-order compensation item, so that a transverse prediction projection value is obtained, and the transverse prediction projection value is further mapped back to a physical spatial position. According to the method, the phase lag problem of visual detection under the high-speed working condition is effectively solved, and the prediction precision and the operation stability of closed-loop control are improved.
Owner:WEINAN DADONG PRINTING PACKING MASCH CO LTD

Privacy computing visual sensing unit and output desensitization visualization method thereof

The invention discloses a privacy computing visual sensing unit and a desensitization visualization method for output of the privacy computing visual sensing unit. The sensing unit is formed by integrating an optical signal acquisition module, an edge calculation module and a data output module, and an event-type image sensor is adopted to asynchronously acquire pixel-level illumination change and output an original event stream. The edge calculation module locally converts an event stream into a privacy calculation visual data stream containing body part semantics, readable expression of the motion state is achieved by distributing distinguishable visual features to different parts, and image reconstruction or artificial intelligence recognition is not depended on. The whole system does not generate, store and transmit complete image frames in any form, and privacy protection is achieved from the source. The visual data stream can be stored in a computer readable medium or transmitted through a network, and is suitable for scenes with high privacy requirements, such as home-based care for the aged, medical monitoring and the like.
Owner:YUNNONG (SHENZHEN) TECHNOLOGY CO LTD

An immune cell state analysis system based on image processing technology

The application relates to the fields of biomedical image processing and intelligent control technology, in particular to an immune cell state analysis system based on an image processing technology; the system comprises feature extraction, state evaluation, decision generation and self-adaptive correction modules; the system extracts features by using a space-time graph neural network, the core of which is to calculate visual semantic entropy based on classification probability and feature response field, and to solve decision confidence weight by combining a cell motion diffusion index; accordingly, an AI regulation instruction and a conservative instruction based on a kinetic tolerance boundary are weighted and fused to generate a final instruction and to self-adaptively calibrate boundary parameters according to an observation error; by quantifying visual uncertainty and analyzing motion characteristics, the application effectively overcomes image blurring and non-biological interference, and significantly improves the recognition precision of active cells and the system robustness in a complex environment.
Owner:XI AN DONGAO BIOSCIENCES CO LTD

Intelligent park target identification method and system based on artificial intelligence

The invention belongs to the technical field of artificial intelligence, and particularly relates to a smart park target identification method and system based on artificial intelligence, and the method comprises the following steps: accessing a park monitoring video stream, analyzing a camera topological structure to determine a blind area channel, and carrying out the mapping and historical image analysis to obtain a target identification result; obtaining a physical path length and an illumination change intensity index of a blind area channel, and extracting a track point sequence of a target in a visible area; the residence time after the target enters the blind area is monitored, the confidence coefficient weight of the visual features is calculated through a preset feature effectiveness attenuation model according to the speed fluctuation condition before the target enters the blind area and the environment complex factors of the blind area channel, and the confidence coefficient weight is in nonlinear attenuation along with the increase of the residence time. According to the method, the problems of appearance failure and large motion prediction deviation caused by long-time retention are effectively solved, and the recognition accuracy is greatly improved.
Owner:ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1

Industrial vision-assisted warehouse automated storage and retrieval optimization method and system

The present application relates to the technical fields of industrial automation, artificial intelligence and computer vision application, in particular to a warehouse automatic access optimization method and system assisted by industrial vision, comprising: collecting original perception data stream, obtaining feature vector of abnormal data area; obtaining physical reality probability and optical false probability through perception classification model processing; calculating visual confidence entropy; in response to the visual confidence entropy being greater than a preset entropy activation threshold, performing illusion evolution deduction to determine a physical confirmation index; executing hierarchical decision logic, including: in response to the visual confidence entropy being not greater than the preset entropy activation threshold, outputting a first access optimization action; in response to the visual confidence entropy being greater than the preset entropy activation threshold, outputting a second access optimization action; obtaining a physical reality label, and correcting the perception classification model and the model used in the illusion evolution deduction; the present application realizes active verification of passive perception failure, and improves decision reliability.
Owner:KENTUO (TIANJIN) IND AUTOMATION TECH CO LTD +4

Kinematic parameter adaptive identification system for mechanical arm based on binocular vision guidance

The application relates to the technical field of mechanical arms and discloses a mechanical arm kinematic parameter self-adaptive identification system based on binocular vision guidance, which comprises the following modules: a modeling module obtains joint angles and temperatures to construct an augmented parameter set, calculates an end theoretical pose, a joint Jacobian matrix and a task Jacobian matrix; a vision module obtains measurement parameters through a binocular camera and calculates a visual confidence matrix; a monitoring module constructs a Hessian matrix based on the Jacobian matrix and the confidence matrix and sends a micro-activation instruction when a condition number exceeds a standard; a control module generates joint micro-offsets in combination with a tolerance band and triggers the binocular camera to collect actual measurement poses according to torque derivatives; and an updating module updates the augmented parameter set by using the residual errors of the actual measurement poses and the theoretical poses in combination with the confidence matrix. The application utilizes binocular vision guidance and stable torque triggering, overcomes the coupling of structural thermal deformation and error parameters, realizes high-precision mechanical arm kinematic parameter self-adaptive identification, and improves calculation stability.
Owner:CHENGDU AEROSPACE KAITE ELECTROMECHANICAL TECH CO LTD

Self-adaptive multi-scale dynamic fire detection method and system based on few-sample learning

The invention relates to the technical field of computer vision and public safety, in particular to a self-adaptive multi-scale dynamic fire detection method and system based on few-sample learning. The method comprises the following steps: constructing a multi-scale semantic attribute library; generating a multi-scale dynamic feature map fusing color, time domain motion and multi-scale optical flow information based on the video sequence; obtaining visual and text feature vectors, and performing cross-modal semantic alignment through comparative learning to construct a basic model; a visual prototype and a semantic prototype are calculated and fused to obtain a scene prototype, and the basic model is finely adjusted to generate a scene special model; and acquiring a real-time video sequence, calculating to obtain a real-time multi-scale dynamic feature map, and extracting cosine similarity between a real-time feature visual vector and a scene prototype, thereby performing real-time fire detection. Due to the fact that multi-scale optical flow information is extracted, full-scale dynamic characteristics from microscopic flickering to macroscopic spreading can be accurately captured, and fire disasters can be efficiently detected.
Owner:SHANGHAI TECHN INST OF ELECTRONICS & INFORMATION

A zero-shot anomaly image detection method based on learnable prompts

The application discloses a zero-shot abnormal image detection method based on a learnable prompt. A learnable prompt generation module based on context optimization is designed, which contains a learnable prompt and an image abnormal state prompt that can be optimized. A multi-level visual coding feature of a to-be-detected image is obtained by using an image coding network of a visual language large model, and a text feature of a learnable prompt embedding is obtained by using a text coding network. A multi-level cosine similarity between the visual coding feature and the text feature is calculated to construct an image abnormal area calculation module, so that an abnormal area of the to-be-detected image is obtained. The learnable prompt avoids the complexity and instability of manually designed prompts, improves the accuracy of image abnormal detection, guarantees the effectiveness and efficiency of zero-shot learning, and greatly reduces the cost of pre-training of a visual language large model to a downstream task.
Owner:COMPUTER INNOVATION TECH RES INST OF ZHEJIANG UNIV

Method for fusing information and video data of ship automatic identification system

The invention discloses a ship automatic identification system information and video data fusion method, which comprises the steps of preprocessing AIS information, extracting and tracking a ship target from a video by using a YOLOX algorithm in combination with a ByteTrack algorithm and Kalman filtering, and generating visual data of historical and predicted trajectories; performing space-time alignment on the AIS data through a constant-speed model to generate a corresponding pixel coordinate sequence; constructing a timing constraint residual mechanism to calculate a similarity matrix of the visual trajectory and the AIS data, and iteratively calculating the minimum matching cost by using a screening rule and a Hungary matching algorithm to obtain an optimal matching pair; aIS information with the most matching times of each ship is screened out through a sliding window voting system, and stable and accurate fusion of the AIS and video data is achieved; the fusion accuracy can be improved, noise can be effectively resisted, and the anti-interference capability is high; and the calculation complexity is reduced, and error accumulation caused by overlong time span is reduced.
Owner:DALIAN MARITIME UNIVERSITY +1

Multi-role rpa collaboration financial shared process intelligent processing method and system

The present application relates to the technical field of data processing, and more particularly to a multi-role RPA collaborative financial shared process intelligent processing method and system, which introduces data flow information entropy rate, quantifies data redundancy by combining upstream write rate and byte probability distribution, truly reflects the actual load pressure of data flow on downstream nodes, and avoids the evaluation distortion problem of traditional rate indicators. Through pixel-by-pixel two-norm difference calculation of visual rendering delay, the completion time of downstream RPA node interface rendering is accurately judged, the complete process time from data writing to result presentation is obtained, and all processing links of the RPA process are fully covered. An information-rendering coupling impedance index is proposed to quantify the nonlinear coupling effect between data flow load and visual rendering delay, which can identify potential congestion inflection points in advance and avoid sudden collapse caused by nonlinear growth.
Owner:JIANGSU ECOLOGICAL ENVIRONMENT BIG DATA CO LTD +1

A cross-modal retrieval method based on prototype regularization learning

The application provides a cross-modal retrieval method based on prototype regularization learning, and belongs to the field of artificial intelligence and multi-modal information processing, and comprises the following steps: extracting visual features and text features through a visual encoder and a text encoder and calculating a basic contrast loss; dividing visual clusters and text clusters through a clustering algorithm and calculating visual prototypes and text prototypes, selecting text features most similar to the visual prototypes as anchor points, and constructing cross-modal prototypes; calculating a prototype-level discriminant loss according to the visual prototypes and the text prototypes, calculating an instance-level discriminant loss according to the visual clusters and the text clusters, and calculating a prototype projection loss according to the cross-modal semantic prototypes; constructing a joint optimization objective function by combining all the losses, training the visual encoder and the text encoder, and performing cross-modal retrieval after the training is completed. The application solves the problems of excessive intra-class variance and excessively high inter-class similarity caused by appearance bias in the existing cross-modal retrieval method, and causes the problems of insufficient retrieval accuracy and robustness.
Owner:SHANGHAI EYE DISEASE PREVENTION & TREATMENT CENTER

Cosmetic product detail picture element extraction method and system based on multi-modal model

The invention relates to the technical field of deep learning, in particular to a makeup product detail picture element extraction method and system based on a multi-modal model, and the method comprises the steps: extracting a text region sequence and a visual region sequence in a detail picture, and obtaining the position coordinates of each text region and each visual region; calculating a visual weight according to the information entropy of each visual area, and carrying out weighted correction on the initial visual features of the visual areas to obtain enhanced visual features; calculating the semantic similarity between the text features and the enhanced visual features; constructing a logic correlation coefficient according to a vertical distance between the text region and the visual region, and correcting the semantic similarity to obtain a comprehensive similarity; and determining a visual area corresponding to the text area according to the comprehensive similarity, and outputting an element extraction result. Through the technical scheme of the invention, accurate extraction and alignment of the claim text, the experimental data, the component graph and other visual demonstration elements in the detail graph of the beauty makeup product are realized.
Owner:GUANGZHOU XINSHU INTELLIGENT TECH CO LTD

Visual enhancement method for color vision disorder, electronic equipment and medium

The embodiment of the invention provides a visual enhancement method for color vision disorder, electronic equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing visual enhancement on an original training image through an original visual enhancement model to obtain a visual enhancement image aiming at the performance of color vision disorder; performing color discrimination on the vision enhancement image through a color vision color simulation model so as to perform model training on the original vision enhancement model according to the calculated target loss values of the vision enhancement image, the original training image, the target color discrimination name and the real color name; and performing visual enhancement on the target image through the trained visual enhancement model. According to the embodiment of the invention, color name discrimination is carried out on the vision enhancement image through the color vision color simulation model so as to simulate a color naming process perceived by people with color vision disorder, and the vision enhancement model is trained in combination with the target loss value, so that the people with color vision disorder are effectively helped to calibrate the consistency of color naming and public language standards.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

A magnetic field vision fusion detection system and method of a UAV platform

This invention relates to a magnetic field visual fusion detection system and method for an unmanned aerial vehicle (UAV) platform, belonging to the field of airborne magnetic detection technology. It includes an UAV platform and its mounted magnetic detection module, camera equipment, data receiver, UAV remote controller, and connected ground control terminal. The ground control terminal is configured to receive flight data, detection data, and visual images. It performs interpolation and color mapping using an interpolation algorithm to generate a magnetic anomaly distribution map. Combining the field-of-view parameters of the camera equipment, it calculates the latitude and longitude coordinates corresponding to each pixel in the visual image, obtaining a visual image with latitude and longitude mapping. The magnetic anomaly distribution map and the visual image with latitude and longitude mapping are superimposed to generate a fused positioning map. This achieves synchronous acquisition and fusion processing of multi-source data, avoiding the system complexity and equipment redundancy caused by introducing an additional host computer, and improving the portability and on-site deployment efficiency of the magnetic detection module.
Owner:QINGDAO HAIYUEHUI TECH CO LTD