Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Visual matching" patented technology

Welding path planning method based on multi-modal fusion and deep learning

The invention provides a welding path planning method based on multi-modal fusion and deep learning, and relates to the technical field of welding automation, the method comprises the following steps: synchronously acquiring three-dimensional coordinate information of a welding workpiece through a laser ranging module and a CCD visual sensor, and fusing laser height parameters and visual plane feature points; an initial path is generated based on the dynamic welding adaptation coefficient and the welding environment complexity, and laser errors, visual matching deviation and environment interference intensity are quantified; laser and visual data are dynamically weighted by using multi-modal feature fusion weights, spatial-temporal features are extracted through a deep learning model, and a welding track is optimized; and the six-axis mechanical arm is controlled to execute welding operation, weld quality is monitored in real time, and path parameters are adjusted in a closed-loop mode. According to the method, the problem of insufficient precision of a single sensor is solved through a multi-modal data complementation mechanism, sudden obstacles and temperature fluctuation are dealt with in combination with dynamic parameter calculation and real-time path correction, and the local path optimization capability is improved by using the convolutional neural network and a self-attention mechanism.
Owner:SHAOXING UNIVERSITY

Underwater LED fish gathering lamp control method and system based on light field scattering model

PendingCN120475567AElectrical apparatusFishingVisual matchingUnderwater light field
The invention discloses an underwater LED fish gathering lamp control method based on a light field scattering model, and the method comprises the following steps: obtaining water optical characteristic parameters and environmental parameters of a target water area in real time, and building an underwater light field scattering model through employing a radiation transfer equation in combination with a Monte Carlo simulation method; with improvement of target fish school gathering efficiency and reduction of energy consumption as optimization objectives, environment parameters are taken as constraint conditions, illumination effects under different LED parameter combinations are simulated through a light field scattering model, and LED control parameters with optimal comprehensive performance are determined; a control instruction is generated according to the LED control parameters, and the working state of each LED unit is adjusted; meanwhile, fish school gathering data and current change values of water temperature, salinity and turbidity are monitored in real time and fed back to the light field scattering model for parameter correction, and closed-loop control is formed. The underwater LED fish gathering lamp control method solves the technical problems of low light field distribution prediction precision, insufficient fish school visual matching performance, poor environmental adaptability and unbalanced energy consumption optimization in the existing underwater LED fish gathering lamp control method.
Owner:ZHUHAI LEEDMART TECH CO LTD

Part free edge visual matching method and system based on model matching

The invention provides a part free edge visual matching method and system based on model matching, and the method comprises the steps: analyzing an assembly three-dimensional model or a cutting two-dimensional model of a ship part, and building a part model database; determining a visual matching template database according to the two-dimensional contour of the part and the part features; collecting an image of a to-be-detected part; determining part features of the to-be-detected part; matching the part features of the to-be-tested part with the visual matching template database, and determining a matching result; if the matching result is successful matching, determining pixel coordinates of the to-be-detected part in the image of the to-be-detected part according to a visual matching template corresponding to the image of the to-be-detected part; and converting the pixel coordinates of the to-be-detected part in the image of the to-be-detected part into actual coordinates of a physical space, and determining the free edge of the to-be-detected part on the machine tool of the ship. According to the method and the device, the free edges of various parts of the ship are positioned, and the efficiency and the quality of ship manufacturing are improved.
Owner:SHANGHAI JIAOTONG UNIV +1

Creative thinking auxiliary generation method and system based on AI

The invention discloses an AI-based creative thinking auxiliary generation method and system, relates to the technical field of artificial intelligence, and solves the problem of inaccurate sound and picture matching in a traditional method by performing timestamp alignment and feature extraction on audio data and a visual image frame and calculating the correlation between the audio data and the visual image frame by using a cross-modal attention mechanism. According to the method, user eye movement track data is introduced, real attention points of a user are mapped into a visual sequence, and an optimized weight matrix is generated by constructing attention masks and fusing model attention weights, so that a generation result is more in line with perception key points of the user. And meanwhile, a feedback mechanism is established based on the synchronization error score, and when the sound and the picture are detected to be asynchronous, the visual frame timestamp can be dynamically adjusted, so that the self-adaptive correction of the content is realized. On the whole, the method has remarkable advantages in the aspects of improving modal alignment precision, enhancing user perception consistency and optimizing generation result naturalness.
Owner:ZHEJIANG NORMAL UNIV

Unmanned aerial vehicle rear-end loopback detection method and device based on multi-sensor fusion extended Kalman filtering

The invention discloses an unmanned aerial vehicle rear-end loopback detection method and device based on multi-sensor fusion extended Kalman filtering. The method comprises the following steps: constructing a linearization state equation and a linearization observation equation of a target unmanned aerial vehicle; collecting visual feature points and laser radar feature points; performing target unmanned aerial vehicle rear-end loopback detection based on the visual feature points and the laser radar feature points to obtain a visual matching degree and a laser radar matching degree; calculating a visual proportion coefficient and a laser radar proportion coefficient in the loopback detection fusion process; taking the linearized state equation and the linearized observation equation as initial pose states of the target unmanned aerial vehicle, and iteratively updating a Kalman filtering gain by using a visual proportion coefficient and a laser radar proportion coefficient; the pose state of the target unmanned aerial vehicle is corrected and updated according to the Kalman filtering gain, and back-end loopback detection is carried out on the target unmanned aerial vehicle again based on the pose state of the target unmanned aerial vehicle; according to the invention, the error of back-end loopback detection in the mapping process of the unmanned aerial vehicle is reduced.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Visual mouth matching equipment for vacuum cup processing

The utility model discloses visual mouth matching equipment for vacuum cup processing, which comprises a rack, and a turntable mechanism, a cup mold ejecting mechanism, a mouth matching front detection support frame and a pressing support frame which are arranged on the rack, a mouth matching front detection mechanism is arranged on the mouth matching front detection support frame, and a pressing mechanism is arranged on the pressing support frame. The cup mold ejection mechanisms comprise cup molds used for fixing vacuum cups, the cup mold ejection mechanisms are correspondingly arranged below the opening matching front detection mechanism and the pressing mechanism, and the opening matching front detection mechanism comprises a front detection camera with a lens facing the corresponding cup mold ejection mechanism. The downward pressing mechanism comprises a downward pressing driving source arranged on the downward pressing supporting frame, a downward pressing telescopic rod driven by the downward pressing driving source to move up and down, a downward pressing guiding assembly arranged on the downward pressing telescopic rod and moving up and down along with the downward pressing telescopic rod, and a downward pressing buffering assembly arranged at the end of the downward pressing guiding assembly. The utility model solves the problems of low production efficiency, unstable rate of finished products and the like in the production process of the conventional opening matching procedure of the vacuum cup.
Owner:YONGKANG QIHUI ROBOT CO LTD

Scene fusion supply and demand panoramic visualization matching decision-making system based on big data

The invention discloses a scene fusion supply and demand panoramic visualization matching decision-making system based on big data, and relates to the technical field of information sharing. The scene fusion supply and demand panoramic visual matching decision-making system comprises a data acquisition module, a portrait construction module, a matching module, a scene fusion module, an information sharing platform, an advanced decision-making module and a visual display module. The portrait construction module is used for constructing each supply end portrait by using the collected supply data of the supply ends and constructing each demand end portrait by using the demand data of the demand ends; the supply end and the demand end are combined to generate a supply-demand group; the scene fusion module is used for analyzing the influence of different environmental data on all supply end portraits and demand end portraits in the market, and fusing the influence of different environmental data to construct a comprehensive influence model; and the advanced decision module is used for making adjustment decisions of all supply and demand groups in advance through the information in the information sharing platform.
Owner:QINGDAO CISCO WANFANG ECONOMIC INFORMATION CONSULTING CO LTD

Synchronous speed visual matching method and system, electronic equipment and storage medium

The invention relates to the technical field of stereoscopic vision, and discloses a synchronous speed visual matching method and system, electronic equipment and a storage medium, and the method comprises the steps: synchronously collecting a stereoscopic video sequence with a predefined frame rate; executing multi-target hybrid tracking and motion induction detection, and outputting target motion information including position and velocity vectors; extracting hierarchical motion features of the target from continuous multiple frames of the stereoscopic video sequence, and performing unified space-time coding; under geometric constraints of stereoscopic vision, scale cosine similarity, direction similarity and trajectory consistency measurement are calculated and serve as observation evidences to be input into the probabilistic reasoning model for fusion, and a posterior probability representing matching reliability is output; the weight distribution of the speed similarity and the direction similarity is adjusted according to the motion characteristics of the targets in the scene, and the stable corresponding matching relation between the left view target and the right view target is established. According to the method, high-time-resolution information can be utilized, motion features and geometric constraints can be effectively fused, and the method has self-adaptive capacity.
Owner:TIANXIANG RUIYI

Low slow small flight target detection method based on visual matching

The invention relates to the technical field of target detection, in particular to a low-slow-small flight target detection method based on visual matching, which comprises the following steps: running a real-time target detection thread and a periodic salient target detection thread in parallel, and performing low-slow-small target detection on each frame of visual image by the real-time target detection thread; the periodic salient target detection thread is executed once every a preset period, sky segmentation and saliency detection are carried out on the current frame of visual image, salient candidate targets in a sky area are extracted, the salient candidate targets are added into a candidate target queue, and priority ranking is carried out on the salient candidate targets; and if the significant candidate target with the highest priority enters a preset countering distance, detecting the confidence coefficient of the significant candidate target through a real-time target detection thread, and performing countering decision. According to the method, the problems of difficult identification, easy tracking loss, insufficient tail end precision and the like of the low-slow small flight target in a complex background are effectively solved.
Owner:长春长光博翔无人机有限公司

Unmanned aerial vehicle navigation deception detection method based on cross-view visual matching

The invention provides an unmanned aerial vehicle navigation deception detection method based on cross-view visual matching. The method comprises the following steps: acquiring a first aerial image collected by an unmanned aerial vehicle at the current moment and unmanned aerial vehicle positioning information; acquiring a first satellite image of the corresponding area based on the unmanned aerial vehicle positioning information; respectively inputting the first aerial image and the first satellite image into a pre-trained first twin network branch and a pre-trained second twin network branch, generating a first aerial image cross-view matching feature and a first satellite image cross-view matching feature, and sharing parameters of the first twin network branch and the second twin network branch, each of the ViT and the ViT comprises a ViT encoder and a cross-view feature mapping module; and determining whether a navigation spoofing attack exists based on the first aerial image cross-view matching feature and the first satellite image cross-view matching feature. By implementing the method, the accuracy and generalization ability of navigation deception detection can be effectively improved.
Owner:BEIHANG UNIV

Inertial vision fusion positioning method and system based on graph optimization in indoor cross-floor environment

The invention relates to an inertial vision fusion positioning method and system based on graph optimization in an indoor cross-floor environment, and the method comprises the steps: obtaining the short-distance relative pose estimation data of a pedestrian based on the multi-source motion parameters of the pedestrian; based on the priori three-dimensional feature map and the image of the current scene of the pedestrian, acquiring visual absolute pose data; and based on a pre-constructed factor graph model, fusing the relative pose estimation data and the visual absolute pose data, and obtaining the global optimal three-dimensional position and pose of the pedestrian through nonlinear optimization solution. According to the method, continuity of deep coupling inertial navigation and absolute precision of visual positioning are optimized through the factor graph, inertial navigation accumulated drift is effectively inhibited, the problem of visual matching failure is solved, decimeter-level positioning precision (the average error is 0.24 m) and 100% continuous positioning are realized in an indoor cross-floor complex scene, and positioning robustness is remarkably improved.
Owner:HUAIYIN TEACHERS COLLEGE

A method for cluttered scene object grasping based on visual-linguistic-action joint modeling

The application discloses a method for cluttered scene target object grasping based on visual-language-action joint modeling. The application uses object-centered representation to realize a method for cluttered scene target object grasping based on visual-language-action joint modeling, processes the object-centered representation through a pre-trained visual-language model and a grasping model, obtains visual-language features and grasping features of each bounding box, and uses a transformer to implement cross-attention mechanisms among visual-language-action multimodality, generates visual-language-action cross-attention features, and then generates decisions and executes, so that higher sample utilization is realized, and additional data collection and training in the simulation-physical migration process are avoided; compared with a two-stage strategy, visual attributes and planner screening rules for language-visual matching need not be artificially designed, so that more flexible language instructions can be adapted, and better task generalization is achieved.
Owner:ZHEJIANG UNIV

A Visual Localization Method Based on Point Cloud Map

The present invention provides a visual positioning method based on a point cloud map, including a point cloud map generation module, a visual inertial odometer construction module, and a visual matching and positioning module based on the point cloud map. Among them, the point cloud map module establishes a high-precision point cloud map by fusing laser, IMU, and GPS information; the visual inertial odometer construction module first extracts visual feature points, uses the optical flow method to track the feature points of the front and rear frames, and performs pre-integration on the IMU, and fuses with the visual feature points to construct a visual inertial odometer, outputs the initial pose of each frame, and restores the depth of the feature points; the visual matching and positioning module based on the existing map extracts a sub-map from the point cloud map according to the initial position, projects the feature points restored by vision into the 3D space map coordinate system, and queries the nearest points in the current sub-map. For the nearest points matched by vision and the map, the RANSAC algorithm based on dual quaternion is used to optimize the pose of the current frame.
Owner:NANJING UNIV +1

Same-view-field cross-lens real-time target tracking system and method based on visual matching

The invention provides a same-view-field cross-lens real-time target tracking system and method based on visual matching, and relates to the technical field of multi-target tracking, and the system comprises an acquisition module, a target detection module, a feature extraction module, a trajectory tracking module, a database management module and a cross-lens matching module. Wherein the trajectory tracking module can track a detected target in combination with target feature information, generate an identity label and output trajectory information, so as to ensure identity continuity and trajectory integrity of the target under the same lens; the database management module is used for uniformly storing the identity label, the feature information and the track information of the target and providing reliable data support for cross-lens tracking; and the cross-lens matching module is used for matching targets under different lenses based on the information in the database management module so as to realize cross-lens association of multiple lenses in the same view field. Through cooperation of multiple modules, multi-lens cross-lens real-time tracking under the same field of view can be effectively realized, and the accuracy and real-time performance of target matching are improved.
Owner:GUANGDONG UNIV OF TECH

Low-altitude visual matching navigation methods, devices, systems, and storage media

This invention discloses a low-altitude visual matching navigation method, device, system, and storage medium, comprising: in an offline phase, optimizing an aerial image sequence into a high-fidelity 3DGS map model; in an online phase, rapidly fusing coarse poses for rendering using inertial pre-integration and global descriptor retrieval; based on the coarse pose, using the 3DGS differentiable rasterization pipeline to synthesize a high-fidelity new perspective reference image in real time by jointly optimizing the pose increment and reference view fusion weights; obtaining a 2D-2D correspondence between the real-time image and the new perspective reference image through deep learning matching, and converting the depth map synchronously generated by the 3DGS model into a 2D-3D association; and calculating the high-precision visual pose of the aircraft through the PnP algorithm and nonlinear optimization. This invention can improve the problems of low quality of new perspective reference images, insufficient multi-sensor information fusion, and limited positioning accuracy in complex low-altitude scenarios.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-scene-oriented security robot intelligent inspection method and system

The invention discloses a multi-scene-oriented security robot intelligent inspection method and system, and relates to the field of intelligent security, and the method comprises the steps: collecting machine body vibration data generated by interaction between a robot and the ground in an environment with good visibility, and building a standard fingerprint band and a historical image; in a low-visibility environment, the robot collects a fuzzy visual image and vibration data in real time, generates a current inspection fingerprint and matches the current inspection fingerprint with a standard fingerprint band; and when deviation is detected, calling a visual memory library and determining the closest reference landmark by utilizing a deep learning visual matching network, so that the steering angle of the robot is adjusted, and the robot is enabled to be close to the central path of the fingerprint zone again. The method can be applied to industrial control software, is deployed on security robots of multiple models to realize process control of inspection, solves the problem that the security robots are unstable in inspection positioning in a low-visibility environment, and realizes path self-correction and intelligent inspection control based on vibration fingerprints and depth vision matching.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

Automatic driving vehicle-mounted sensor fusion positioning method, device, equipment and medium

The invention relates to an automatic driving vehicle-mounted sensor fusion positioning method, device and equipment and a medium, and the method comprises the steps: fusing a coarse positioning result corresponding to I MU data and a coarse positioning result corresponding to wheel speed meter data based on a preset Kalman filtering algorithm, recalculating the position deviation, and obtaining the position deviation of the I MU data; carrying out constraint calculation on the relative displacement so as to determine fusion positioning information of the autonomous vehicle; and determining a vehicle moving track of the autonomous vehicle according to the fused positioning information, and performing map matching or visual matching based on the vehicle moving track to complete map matching positioning of the autonomous vehicle. According to the invention, high-precision positioning of data fusion of the automatic driving vehicle-mounted sensor can be realized without the help of high-cost and high-precision inertial navigation equipment under the condition that the GNSS signal is unavailable.
Owner:CHINA NANHU ACAD OF ELECTRONICS & INFORMATION TECH

Unlisted non-motor vehicle driver identification method, system and program product

The invention belongs to the technical field of intelligent traffic, and particularly discloses an unlisted non-motor vehicle driver identification method and system and a program product, and the method comprises the steps: carrying out the time-space correlation matching of a front image and a back image of a driving non-motor vehicle and a driver, and an electronic license plate number of the non-motor vehicle; the method comprises the following steps: acquiring a front image and a back image of a non-motor vehicle, performing visual matching on the corresponding front image and back image, performing space-time and visual double-base judgment based on results of the two matching modes, determining a target front image corresponding to a target back image, finally performing face recognition on the target front image, and determining identity information of a corresponding non-motor vehicle driver without listing a tag. According to the method, the identity of the non-motor vehicle driver can be accurately and efficiently traced and identified, so that a complete evidence chain is formed, and an effective basis is provided for traffic management personnel to treat and punish non-motor vehicle non-listed behaviors.
Owner:BEIJING BOHONG KEYUAN INFORMATION TECH CO LTD

Ship robot double-wire welding process and welding equipment based on model driving and visual matching fusion

The invention discloses a ship robot double-wire welding process and welding equipment based on model driving and visual matching fusion. The ship robot double-wire welding process comprises the following steps that an adaptive robot double-wire welding process is selected; visual scanning, model introduction and fusion are completed by the structured light shooting camera; clicking the welding plan to generate a welding operation file; a welding seam information json file is imported, and a robot welding operation list is generated; gun cleaning operation is executed, and welding is conducted after gun cleaning is completed; a welding task is automatically executed according to the welding operation list, gun cleaning is automatically completed according to the length of the welding seam, and welding of the next welding seam is continued; through the model driving and visual matching fusion algorithm, collaborative operation of model importing and visual scanning is achieved, the success rate of one-time workpiece recognition is high, the recognition precision is high, the welding seam positioning time is shortened, the arcing rate of the robot is increased, and the overall welding capacity and the unit area output efficiency are improved.
Owner:SHIPBUILDING TECHNOLOGY RESEARCH INSITITUTE (NO 11 INSTITUTE OF CSSC)

GENERATE SYNCHRONIZED SOUND FROM VIDEOS

Method (200) for recognizing visually matching tones, wherein the method comprises: Receiving visual training data (105) at a visual coder (110) that has an initial machine learning (ML) model; Identifying data corresponding to a visual object in the visual training data (105) using the first ML model; Receiving audio training data (107) synchronized with the visual training data at an audio forwarding regulator (115) which has a second ML model, wherein the audio training data (107) has a visually matching tone and a visually mismatched tone, both of which are synchronized with one and the same frame in the visual training data (105) which contains the visual object, wherein the visually matching tone corresponds to the visual object, whereas the visually mismatched tone is generated by a sound source which is not visible in the same frame; Filtering data matching the visually appropriate tone from an output of the second ML model using an information bottleneck (120); and Training a third ML model following the first and second ML models (235) using the data corresponding to the visual object and data corresponding to the visually inappropriate tone.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Spatial positioning method and device, equipment and storage medium

The invention provides a spatial positioning method and device, equipment and a storage medium. The method comprises the following steps: constructing a neural radiation field based on three-dimensional point cloud data and a multi-view discrete image of a target scene; obtaining user behavior data which comprises a basic pose image shot by a user and / or editing data when a virtual object is edited, extracting an effective pose from the user behavior data, and generating a virtual shooting pose set; generating a rendered image set bound with a pose through the neural radiation field; and finally, receiving a positioning request, and determining a target pose in combination with the to-be-positioned image, the rendering image set and the three-dimensional point cloud data. The target pose is determined in combination with the multi-modal data, the precision limitation of a single-modal sensor in a complex scene is broken through, visual matching accumulative errors and dynamic environment interference are eliminated, the computing resource consumption is reduced while the positioning precision is improved, meanwhile, the effective pose is extracted in combination with the user behavior data for rendering, and the user experience is improved. And the problem of complex or inaccurate calculation caused by redundant rendering is avoided.
Owner:LINGBAN INTELLIGENT (HANGZHOU) INFORMATION TECHNOLOGY CO LTD

Authenticate a user before performing a sensitive operation associated with a UE in communication with a wireless telecommunication network

The system receives an indication of a sensitive operation. The system obtains a unique ID of a user's UE. Based on the unique ID of the UE, the system retrieves a visual authentication method including a visual ID. The system records the visual ID, and retrieves a corresponding stored visual ID. The system performs a liveness check associated with the visual ID, to determine whether the visual ID is a recording or a live version of the visual ID. Upon determining that the visual ID is the recording, the system refuses to authenticate the user. Upon determining that the visual ID is the live version of the visual ID, the system compares the visual ID and the corresponding stored visual ID to determine whether the visual ID and the corresponding stored visual ID match. Upon determining that the visual ID and the corresponding stored visual ID match, the system authenticates the user.
Owner:T MOBILE US INC

Leveraging audio matches to improve visual matching recall between video content items

Audio matching is performed between a first video content item and a second video content item to identify a matching audio segment. First temporal boundaries within the first video content item and second temporal boundaries within the second video content item corresponding to the identified matching audio segment are identified. A visual matching between the first video content item within the first temporal boundaries and the second video content item within the second temporal boundaries is performed using a modified visual similarity threshold that is lower than a baseline visual similarity threshold. Whether a match exists between the first and second video content items is determined based on the visual matching.
Owner:GOOGLE LLC

Multi-stage cognitive modeling and parameter optimization method based on visual pairing comparison task

The invention relates to a multi-stage cognitive modeling and parameter optimization method based on a visual pairing comparison task, and aims to solve the technical problems of incomplete cognitive process modeling, weak crowd distinguishing ability, task parameter empirical and the like in the existing VPC task evaluation technology. Precise evaluation of the cognitive function and optimization design of task parameters are achieved through a five-step method, wherein VPC task parameterization design and eye movement data collection are carried out; constructing a'familiarity accumulation-memory attenuation-novelty attention and utility 'three-stage cognitive model; performing backstepping on individual free parameters based on maximum likelihood estimation; carrying out multi-population cognitive parameter distribution modeling and subtype division; and task parameter closed-loop optimization based on simulation and efficiency discrimination. According to the method, quantitative distinguishing of learning efficiency, memory stability and novel preferences is achieved, cognitive parameter recovery precision and multi-crowd distinguishing efficiency are improved, a standardized and transplantable technical path is provided for early screening of cognitive impairment and crowd typing, and the method is suitable for cognitive neuroscience research and clinical cognitive evaluation scenes.
Owner:HANGZHOU DIANZI UNIV

Image automatic labeling method and system based on landscape element knowledge graph and visual matching

A method and system for automated image annotation based on a landscape element knowledge graph and visual matching is disclosed. The method includes: constructing a landscape element knowledge graph; performing structured association modeling of landscape nodes, landscape element nodes, image nodes, and cultural knowledge nodes; extracting relevant landscape nodes from the graph based on user-inputted images and shooting locations to limit the range of elements to be identified; extracting features from the input image using a zero-shot visual matching model and comparing them with image samples of candidate landscape elements to obtain identification results; extracting attributes and associated cultural knowledge content from the graph based on the identified landscape element nodes; semantically fusing the identification results with cultural knowledge according to preset rules to generate structured or natural language annotation text; and outputting the annotation text to a terminal for display. This invention combines geographic location, knowledge graph, and visual matching to achieve accurate, information-rich, and culturally profound automated annotation of landscape images.
Owner:ZHEJIANG UNIV OF TECH

Webpage data processing method, apparatus, device, and medium

The present disclosure provides a webpage data processing method and device, equipment and medium, relates to the technical field of artificial intelligence, in particular to the technical field of webpage development and deep learning. The method comprises: determining an interactive operation for a target webpage; obtaining a screenshot of the target webpage and attributes of a plurality of webpage elements in the target webpage, the attributes indicating functions possessed by the webpage elements; based on the attributes of the plurality of webpage elements, screening a plurality of candidate webpage elements related to the interactive operation from the plurality of webpage elements; determining the positions and sizes of the plurality of candidate webpage elements, and cutting the screenshot to obtain visual segments of the plurality of candidate webpage elements; determining the intention matching degrees of the attributes of the plurality of candidate webpage elements and the interactive operation; determining the visual matching degrees of the visual segments of the plurality of candidate webpage elements and the interactive operation; and based on the intention matching degrees and the visual matching degrees, determining a target webpage element from the plurality of candidate webpage elements and executing the interactive operation.
Owner:BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Robot task planning method based on multi-modal large model

PendingCN122347174AVisual matchingData set
The application relates to the technical field of robot control, and discloses a robot task planning method based on a multimodal large model. The application collects multimodal data of a robot motion scene, including voice information and visual information, pre-processes the information to obtain a task instruction sequence required to be executed by the robot and an entity data set in a scene where the robot is located, inputs the pre-processed data into a multimodal large model based on a large language model, calculates, based on the pre-processed data, a semantic evaluation coefficient, a visual matching coefficient, a feasibility coefficient and a priority coefficient corresponding to each task instruction in the task instruction sequence, and obtains a task comprehensive execution coefficient corresponding to each task instruction by weighted summation, reorders the task instructions of the task instruction sequence based on the task comprehensive execution coefficient, and obtains a final task planning sequence, thereby improving the rationality and efficiency of task execution of the robot.
Owner:KEYI COLLEGE OF ZHEJIANG SCI TECH UNIV

A multi-scale machine vision matching method

The present invention discloses a multi-scale machine vision matching method, belonging to the technical field of vision matching. It includes determining an original scale template image I1 based on a search image I0 input by a user, and obtaining a set of template feature points under each layer of the original scale pyramid; calculating a template scale step s; calculating the upper and lower limits of the template scale; calculating a set of scale factors S; traversing the set of scale factors S to obtain a set of feature points for each layer of the pyramid corresponding to different scale factors; performing feature correction on the set of feature points for each layer of the pyramid with a scale factor less than 1; calculating a set of pyramid image sequences of the search image I0 input by the user; performing top-layer template similarity calculation to obtain a top-layer matching result; performing neighborhood suppression on all top-layer matching results to obtain a set of top-layer matching results; and adopting a layer-by-layer approximation method to traverse all pyramid layers, and the obtained matching result is the final matching result. This application can effectively locate targets at multiple scales.
Owner:SHENZHEN RUIDA TECH CO LTD

Data transmission system and method for operation guidance in remote interventional surgery

This application discloses a data transmission system and method for operation guidance in remote interventional surgery, relating to the field of intelligent sensor technology. The method includes: acquiring a control dataset; calculating the video end-to-end latency, visual matching ambiguity, and operator hand tremor components based on the control dataset; processing the optimal pairing set corresponding to the control dataset and the visual matching ambiguity using a preset filter to obtain motion state estimates; performing temporal cross-correlation analysis on the video end-to-end latency and operator hand tremor components to calculate the delay-tremor coupling strength; constructing a future position prediction distribution model based on the motion state estimates and the delay-tremor coupling strength; superimposing the future position prediction distribution model with a preset anatomical structure model, and displaying the superimposed result on the current surgical video screen to guide the surgeon in remotely controlling the surgical operation from the control terminal. This application improves the safety of remote precision operations.
Owner:THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE